Skip to main content

SENIOR AI INTERACTION EVALUATOR (CODEX / CLAUDE CODE)

RemoteNot specified. Estimate: United States · 66% confidence
Published
Role
AI / ML
Experience
Senior
Employment
Part-time10–20h/week
$100–$200/hr
Check eligibility

The listing doesn't say where it hires from. It may hire in United States (66% confidence). This is an estimate, not an eligibility rule; verify before applying.Signals: salary level.

No BS summary

We’re looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code. This is not a traditional engineering role. You won’t be writing production code.

Core skills

OpenAI CodexClaude Code

Required skills

TypeScript/JavaScript/PythonCursor

Optional skills

Cursorprompt design

What you'll do

  • Evaluate AI-generated coding interactions end-to-end
  • Judge whether outputs are: Useful, Correct (at a high level), Aligned with how a strong engineer would think
  • Assess the quality of explanations and reasoning, not just code
  • Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4)
  • Provide clear, opinionated feedback on: What worked, What didn’t, What felt “off” or misleading

What they require

  • Staff / Principal-level engineer (or equivalent experience)
  • Strong background in one of the below: TypeScript / JavaScript, Python
  • Hands-on experience using: OpenAI Codex, Claude Code, Cursor
  • Deep familiarity with modern AI-assisted dev workflows
  • Able to evaluate code without needing to fully execute or deeply review every line
$100–$200/hr