Skip to main content
Anyone AI
Anyone AI

Human Data Evals Lead (Remote/US/LATAM)

RemoteLATAM+United States
Published
Role
AI / ML
Experience
Lead
Salary not disclosed
Check eligibility

Open to Anywhere in LATAM + US. Set where you work from to check your eligibility.

No BS summary

You will own Anyone AI’s data initiatives and proposals to AI labs, from the data proposal or responding to requests, through pilot delivery.

Required languages

English Fluent

Optional languages

Spanish

What you'll do

  • Proposals & requests. Study public benchmarks and eval targets, and turn them into proposals and sample packages that demonstrate capability and win the work. Respond to lab data requests and pilots.
  • Sample & benchmark development. Design and build the sample packages, working with subject-matter experts. Every package meets the bar of our current sample set:
  • Subject-matter experts. Recruit, brief, calibrate, and review a pool of experts across coding, agentic/tool-use, and STEM/reasoning. Raise their output to our standard and keep it there; be the arbiter of what "correct" and "frontier-difficulty" mean.
  • Lab relationships. Be a direct point of contact for lab partners on Slack and calls, with support from the CEO and the wider team. Keep senior lab contacts informed, surface what they actually need, and pull in the CEO and subject-matter experts when the conversation calls for it.
  • Pilot delivery. Own pilots end to end: scoping, SOW, staffing, production, QC, and delivery. Nothing ships before it's lab-ready, and nothing comes back rejected as "not frontier-level" without us already knowing why.

What they require

  • Originated data or benchmark proposals for AI labs, translated eval targets into sample tasks that demonstrate capability, and owned the engagement through delivery.
  • Deep evaluation and quality expertise: LLM benchmarking, with real strength in code-model evaluation.
  • Built QC processes and artifact standards that met enterprise or lab requirements, and set a quality bar a team of experts was held to.
  • 5+ years in technical delivery, quality, or program management, with recent experience in AI/ML data, model evaluation, or benchmarking.
  • Hands-on experience delivering data or evaluation work to AI labs or enterprise ML teams, scoping through delivery.

Anyone AI creates high-quality STEM training data for frontier AI models used by leading AI labs to improve model reasoning in scientific domains.

AIStartup
Salary not disclosed