We're looking for a QA Engineer with experience testing AI-powered applications. The ideal candidate has hands-on experience evaluating non-deterministic AI systems (LLMs, AI agents, and RAG applications) and building or maintaining evaluation frameworks (eval engineering) to measure quality, detect regressions, and improve AI system performance.
What are the main responsibilities of this role?
- Create detailed, comprehensive, and well-structured test plans and test cases.
- Review client needs and technical specifications with the development team to provide timely and meaningful feedback.
- Estimate, prioritize, plan, and coordinate testing activities.
- Identify, record, document thoroughly, and track bugs.
- Design, develop, and execute automation scripts (including Playwright-based automation).
- Perform thorough regression testing when bugs are resolved.
- Track quality assurance metrics and reports.
- Design and maintain evaluation suites (evals) to validate AI-powered features and detect regressions in model behavior.
- Test non-deterministic AI systems, including LLM-based features, AI agents, and RAG applications, using qualitative and quantitative evaluation methods.
- Analyze AI outputs, identify failure patterns (hallucinations, instruction-following issues, grounding, etc.), and collaborate with engineers to improve system quality.
- Stay up to date with new testing tools, AI evaluation techniques, and quality assurance strategies while driving continuous improvement.
- Investigate the causes of software defects and collaborate with the development team to implement preventive solutions.
- Collaborate on internal initiatives.
Which skills and experience do we envision for this role?
- Background in Software Engineering, Computer Science, or a related field.
- 3+ years of experience as a QA Engineer.
- Advanced English skills.
- Hands-on experience with mobile testing (required).
- Experience with test automation, preferably using Playwright.
- Experience using AI tools (e.g., Claude or similar) to optimize testing processes, including test case creation, bug reporting, edge case identification, and automation support.
- Experience testing AI-powered applications or LLM-based systems.
- Understanding of the challenges of testing non-deterministic systems and validating AI-generated outputs.
- Experience with evaluation engineering, including building or maintaining evals, benchmark datasets, or automated AI quality evaluations.
Nice to have:
- Knowledge of CI/CD pipelines and automated test integration.
- Experience testing AI agents, RAG applications, or conversational AI systems.
- Familiarity with AI evaluation frameworks or LLM observability tools such as LangSmith, Braintrust, Arize Phoenix, Promptfoo, OpenEvals, or similar.
Join our team and enjoy:
- Access to Rootstrap University, Conferences/Certifications, and a Mentorship Program for your professional growth.
- Learning Bonus.
- Opportunities to organize cross-functional initiatives and receive 360 Feedback to improve your skills continuously.
- The flexibility to work remotely or from our offices in Montevideo, Buenos Aires, and Medellín, with a flexible time schedule and Workation program.
- Gym benefits, psychological counseling, and weekly lunch reimbursements with special foods in the offices.
- An Onboarding kit and access to cutting-edge technologies and tools to make your work easier.
- After offices, Prizes, and special occasions gifts to celebrate your achievements.
- People Care referent to support your well-being and personal development.
- We value your well-being and happiness. We grow together!
How we work:
- We offer a flexible and diverse work environment that fosters multicultural talent.
- We value autonomy and creativity, We want people to bring ideas to the table and to think as leaders.
- We put our focus on real human connections.
- We have a commitment to excellence.
At Rootstrap, we help companies launch, grow, and scale their digital products and services. With over 160 global team members, we’ve built solutions for industry leaders like MasterClass, Universal, and Google, as well as for innovative startups and inspired pioneers. Our team takes pride in making an impact that goes beyond the code, earning regular recognition as a Fastest Growing Company and Best B2B Service Provider.
We embrace an AI-first mindset, continuously exploring how artificial intelligence can enhance the way we build products, solve problems, and deliver value. Across all stages of the development process, we encourage our teams to actively leverage AI to boost productivity and drive innovation.
With offices in Uruguay, Argentina, and Colombia, we offer the flexibility to work remotely. Discover more about life at Rootstrap at https://www.rootstrap.com/careers.
All of our candidate searches are covered under Uruguayan Law 19691 (the purpose of this law is to promote the inclusion of people with disabilities).
As part of our Security and Compliance process, all seniority levels and roles must follow the company’s information security policies and secure development practices, and promptly report any suspected security incidents or vulnerabilities.