Skip to main content
Mindrift

AI Evaluation Engineer (Python, QA or Security)

RemoteJapan, South Korea, New Zealand +1 more only
Published
Role
QA
Experience
Senior
Employment
Contract
$50+/hr
Check eligibility

Open to JP, KR, NZ, SG only. Set where you work from to check your eligibility.

No BS summary

Project-based AI evaluation work for engineers with 5+ years in software development. Must know Python/FastAPI, JavaScript/TypeScript/React, Docker, Postgres, Kafka, Redis, testing, and English B2+. Hiring in Japan, South Korea, New Zealand, and Singapore.

Core skills

PythonTestingAI Evaluation

Required skills

FastAPIJavaScriptTypeScriptReactDockerPostgreSQLKafkaRedisFunctional testingIntegration testing

Required languages

English B2+

What you'll do

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

What they require

  • Please submit your CV in English and indicate your level of English proficiency.
  • Participation is project-based, not permanent employment.
  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+
  • You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.
  • Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
  • Tasks are estimated at ~20 hours each; you set your own schedule.

Benefits

  • You set your own schedule.

Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners. Mindrift, powered by Toloka, connects top domain experts with cutting-edge AI initiatives.

🇮🇳 IndiaAI
$50+/hr