Senior AI Agent Evaluation Engineer
- Role
- QA
- Experience
- Senior
- Employment
- Part-time20h/week
Open to GB only. Set where you work from to check your eligibility.
No BS summary
Senior software engineer with 5+ years of experience in Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis. Must have experience writing functional and integration tests. Requires B2+ English proficiency. The role involves creating and evaluating AI coding agents by building realistic developer environments and designing challenging tasks and tests.
Core skills
Required skills
Required languages
What you'll do
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What they require
- 5+ years in software development
- English proficiency - B2+
Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners. Mindrift, powered by Toloka, connects top domain experts with cutting-edge AI initiatives.
What people say about this company
4.0/ 5