Skip to main content
Moonshot

Head of AI Safety

RemoteIreland, United Kingdom only
Published
Role
AI / ML
Experience
C-Level
Employment
Full-time
€88k–€102k/yr
Check eligibility

Open to IE, GB only. Set where you work from to check your eligibility.

No BS summary

Moonshot is seeking a Head of AI Safety to lead its AI Safety portfolio, focusing on preventing online harms like violence, extremism, and CSEA. The role requires hands-on experience in red teaming and adversarial evaluation of AI systems, strong project and team management skills, and the ability to build relationships with frontier AI companies, governments, and regulators. Candidates must have experience in trust & safety or related fields and be eligible to work in the UK.

Core skills

AI SafetyRed TeamingAdversarial Evaluation

Optional skills

model safetyadversarial evaluation of LLMsLLM architecturesafety toolingtrust & safety policychild safety evaluationteen-safety product workgrooming and CSEA detection

What you'll do

  • Lead and quality-assure Moonshot's applied AI safety work across harm categories including pathways to violence, extremism, CSEA, abuse and grooming, mental health and crisis, and risks affecting children and teens, using methods such as red teaming and adversarial evaluation of AI systems.
  • Advise frontier AI companies on how to improve the safety of their models, products, policies, and intervention systems.
  • Translate insights from psychologists, child-safety specialists, violence-prevention practitioners, safeguarding experts, and other subject-matter experts into clear, actionable guidance for model safety, policy, product, research, and engineering teams.
  • Set the methodological approach for the portfolio, translating violence-prevention, safeguarding, and behavioural-risk expertise into structured and testable evaluation frameworks.
  • Lead and participate directly in red teaming and adversarial evaluation, working in detail with test scenarios, model responses, scoring criteria, safety policies, and evaluation results.
  • Identify patterns, edge cases, and potential safety failures, and develop practical recommendations for improving model behaviour and user protections.
  • Maintain rigour and clear documentation across the team's technical deliverables, suitable for technical, government, and foundation audiences.
  • Ensure work is delivered within a clear ethical framework and in compliance with contractual, legal, data protection, and ethics obligations.
  • Identify, manage, and escalate operational, reputational, delivery, and partnership risks.
  • Serve as Moonshot's primary applied AI safety counterpart for frontier AI company partners, governments, regulators, and the wider ecosystem invested in AI safety.
  • Build trusted relationships with model, policy, trust and safety, product, research, and engineering teams.
  • Build and sustain relationships across the wider AI safety ecosystem, including governments, foundations, regulators, academics, researchers, civil society organisations, and specialist practitioners.
  • Represent Moonshot externally in meetings, briefings, workshops, and sector engagement, including with regulators and policymaker audiences.
  • Provide direct leadership, coaching, and management to Moonshot's AI safety team.
  • Foster a collaborative, accountable, and mission-driven team culture, with particular attention to wellbeing given the sensitive nature of the work.
  • Support workforce planning, performance management, and professional development across the team.
  • Ensure effective coordination with internal teams supporting the portfolio, including operations, finance, research, and technical teams.
  • Develop Moonshot's AI safety portfolio, identifying strategic opportunities, partnerships, and funding.
  • Lead proposal development, scoping, and renewals with technical credibility, using precise, defensible language suited to technical and government audiences.
  • Develop repeatable methodologies, service offerings, and partnerships that allow the portfolio to grow while maintaining methodological rigour and delivery quality.
  • Support external communications, publications, briefings, and thought leadership that establish Moonshot as a credible voice in applied AI safety.
  • Oversee project planning, staffing, budgeting, forecasting, and delivery timelines across the portfolio.

What they require

  • Experience in trust & safety, online harms, or a closely related field such as violence prevention, safeguarding, or public health, and the ability to adapt that knowledge to AI systems.
  • Curiosity about AI and the ability to build technical fluency quickly, enough to engage credibly with technical counterparts at AI companies.
  • Experience designing research, evaluation frameworks, or interventions for harm categories such as violent extremism, CSEA, self-harm and crisis, or targeted violence.
  • Demonstrated experience managing projects, teams, budgets, partners, and clients, with strong people management skills.
  • Excellent written communication, with experience producing credible (not promotional) material for government, foundation, or enterprise audiences.
  • Comfort and demonstrated resilience working with highly sensitive or graphic content (CSEA, extremist material, crisis content), with awareness of wellbeing practices for this kind of work.
  • Strong judgment and the ability to navigate ambiguity, competing priorities, and sensitive stakeholder environments, including representing organisations externally.
  • Willingness to travel and work outside regular hours where needed to accommodate clients or respond to incidents.
  • Highly trustworthy, with discretion and diplomacy, and willing to undertake relevant security clearance procedures.
  • Experience supporting business development, grant funding, or procurement.
  • Commitment to Moonshot's mission.
  • Eligibility to work in the UK and pass any relevant security clearance procedures per the needs of clients.
  • Desirable: Direct experience in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems.
  • Desirable: Understanding of LLM architecture, safety tooling, or trust & safety policy.
  • Desirable: Prior experience in child safety evaluation, teen-safety product work, or grooming and CSEA detection.
  • Desirable: Familiarity with government or regulatory engagement, such as briefing officials or supporting policy submissions.
  • Desirable: Experience with intervention or diversion programme design that can transfer to AI-mediated interventions.
  • Desirable: Academic or applied background in radicalisation studies, forensic psychology, or violence risk assessment.
  • Desirable: Familiarity with taxonomy or classifier development, including how testing data feeds a classifier.

Benefits

  • 30 days' paid annual leave, excluding public holidays.
  • Flexible public holiday policy with the option to work public holidays in exchange for a day off at another time.
  • Private healthcare package with access to specialist mental health cover, including coverage for partners and children.
  • Dental and Vision Insurance.
  • Life Insurance & Income Protection.
  • Employee Assistance Programme providing access to mental health support.
  • Generous maternity and paternity leave: 26 weeks paid maternity leave, 8 weeks paid paternity leave.
  • All permanent employees are granted share options upon employment.

Moonshot works to counter violent extremism and online harms, combining expertise in violence prevention, behavioural risk, and online harms with applied AI safety evaluation.

AI SafetyStartupmoonshot.audio/
€88k–€102k/yr