About Toloka
At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety.
About the Position
Led applied ML initiatives within the Delivery division, integrating LLM and AI-agent technologies into active client engagements with a focus on Anti-Fraud.
Optimize output quality and automate manual workflows to scale operational capacity and improve business margins.
Core Outcomes You’ll Drive
Automation lift: Achieve double-digit percentage improvements per account by replacing manual steps with robust LLM/agent pipelines.
Quality & reliability: Raise client acceptance rates and reduce rework through automated LLM checks and evaluation harnesses.
Throughput & margin: Shorten cycle times and expand task capacity by productizing reusable components across projects.
Safety & governance: Enforce strict guardrails and solution auditability through systematic red-teaming.
What you’ll do
Solution architecture: Design and deploy agentic workflows (tool use, planning, retrieval, critique loops) for data generation, evaluation, anti-fraud, and safety use cases.
Automated evaluation: Build judge models, auto-grading systems, and self-verification pipelines to codify acceptance criteria into repeatable checks.
Productization: Create reusable reference architectures, templates, and components to scale delivery across multiple client accounts.
AI-driven scaling: Integrate LLMs into contributor workflows (pre-label → verify → escalate) to eliminate operational toil and maximize expert leverage.
Observability: Implement production measurement frameworks (task KPIs, drift), online A/B tests, and cost/latency dashboards.
Customer collaboration: Partner with Delivery Leads, PMs, and client teams to translate business requirements into pragmatic ML designs.
Technical stewardship: Set engineering standards for prompts, agents, data pipelines, and CI/CD via active code and design reviews.
What we're looking for
Production experience: Over 5–8+ years of expertise in applied ML and LLM development, with a documented history of launching agentic workflows into production.
Domain expertise: Strong background in building anti-fraud platforms, trust & safety systems, or real-time anomaly detection frameworks.
Delivery focus: A pragmatic approach, capable of transforming vague project needs into executable solutions defined by strict timelines and measurable KPIs.
AI quality assurance: Extensive experience in building automated evaluation pipelines and grading systems to maintain consistency at scale.
Automation mindset: A natural drive for automation, with a track record of identifying and optimizing manual operations to enhance throughput and margins.
System architecture: Skill in pragmatic engineering, prioritizing simple, observable, and efficient designs that align with latency and cost goals.
Hands-on engineering: A technical profile featuring advanced Python proficiency and the ability to tune agents and prompts within live systems personally.
What we can offer
You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
Competitive compensation package including base salary, bonus, and ESOP
Paid PTO and benefits will vary depending on location
We offer a full remote or hybrid model (if you are based in NL or Serbia)
IT setup and home office allowances.
Published on: 8/8/2026

Toloka
At Toloka AI we create data that powers leading GenAI models and innovations.
Please let Toloka know you found this job on Wantapply.com. It helps us to get more jobs on our site. Thanks!