Skip to main content
Figma
Figma

Director, Research - AI Evals

RemoteUnited States only
Published
Role
Research
Experience
Lead
Employment
Full-time
Company size
Enterprise
$258k–$348k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Research/product leader with 10+ years in product, research, applied research, or related work, including 2+ years managing people. Must have hands-on experience evaluating AI/LLM-powered products and building human plus automated evaluation approaches. US-based only, from a Figma US hub or remote in the United States.

Core skills

AI evaluationLLM evaluation

Optional skills

BraintrustLangSmithDeepEval

What you'll do

  • Own AI evaluation methods and operations for Figma's AI-powered experiences, including defining quality dimensions, designing how to measure them, and turning results into decision-ready signal.
  • Build and maintain evaluation frameworks, rubrics, golden datasets, and quality bars, combining human evaluation with automated or model-based approaches such as LLM-as-judge where appropriate.
  • Partner with engineering to stand up repeatable, reproducible evaluation pipelines and regression testing so evaluation is a routine part of how AI features are built and shipped.
  • Produce clear readouts and dashboards that let stakeholders confidently make go/no-go and prioritization decisions.
  • Socialize a shared definition of quality so evaluation standards are adopted across teams rather than re-invented.
  • Advocate for evaluation as a strategic partner in the product process.
  • Manage a small team to execute AI evals in partnership with contractors, internal staff, and/or LLMs.

What they require

  • 10+ years of experience in product, research, applied research, or a closely related field, including 2+ years of management experience.
  • Direct, hands-on experience owning the evaluation of AI/LLM-powered products.
  • Expertise designing and running AI evaluation, including human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability, and judgment about when and how to apply automated/model-based approaches such as LLM-as-judge and their limitations.
  • Strength across both qualitative and quantitative methods.
  • Comfort with data and metrics.
  • Ability to reason about model behavior.
  • Demonstrated success in identifying the riskiest assumptions behind an ambiguous quality question, prioritizing them, and designing right-sized evaluation to build confidence.
  • Proven track record of gaining buy-in from executive and cross-disciplinary stakeholders, transcending methodology to articulate a larger user story and the "so what" to inspire action.
  • Preferred: Experience standing up a new function, practice, or discipline from scratch.
  • Preferred: 2+ years in product design, user-centric product management, data science, product development, and/or front-end engineering.
  • Preferred: Familiarity and depth of experience using Figma's products.
  • Candidates must keep cameras on during video interviews.
  • If hired, candidates are required to attend in-person onboarding.

Benefits

  • Equity for employees.
  • Health coverage.
  • Dental coverage.
  • Vision coverage.
  • Retirement benefits with company contributions.
  • Parental leave.
  • Reproductive or family planning support.
  • Mental health and wellness benefits.
  • Paid time off.
  • Paid sick leave.
  • Holidays.
  • Other leave benefits in compliance with applicable federal, state, and local laws.
  • Employer-provided paid flexible PTO for exempt employees.
  • Flexible paid sick leave for exempt employees.
  • Company recharge days may be included.
  • Cell phone reimbursements may be included.
  • Home internet reimbursements may be included.
  • Lifestyle spending accounts may be included.
  • Annual bonus plan for eligible non-sales roles.

online vector graphics editor and prototyping tool

🇺🇸 United StatesSoftwareMid-sizefigma.com

Details

Apply routeGreenhouse
$258k–$348k/yr