Moonshot
Senior Freelance Consultant, AI Safety
RemoteUnited Kingdom, Canada, United States +1 more only
- Experience
- Senior
- Employment
- Contract
Salary not disclosed
Check eligibility
Open to GB, CA, US, IE only. Set where you work from to check your eligibility.
No BS summary
Senior Freelance Consultant needed for AI Safety, focusing on red teaming, adversarial evaluation, and methodological development across harm categories. This is a technical, project-based role requiring experience in trust and safety or online harms, with a close to full-time commitment for approximately 6 weeks.
Core skills
AI SafetyRed TeamingAdversarial Evaluation
Optional skills
Model safetyadversarial evaluation of LLMs or other AI systemsLLM architecturesafety toolingtrust and safety policyChild safety evaluationteen safety product workgrooming and CSEA detection
What you'll do
- Red teaming and adversarial evaluation of AI systems against defined harm categories.
- Reviewing model responses against harm and risk criteria and providing expert judgement.
- Bringing subject matter expertise to a specific harm area, such as grooming and CSEA, radicalisation pathways, crisis signalling, or teen online safety.
- Supporting the design of evaluation frameworks that translate real world harm knowledge into structured, testable criteria.
- Contributing to the design of intervention logic that connects at risk users to appropriate support.
- Drafting methodology or findings suitable for technical and government audiences.
What they require
- Experience in trust and safety, online harms, or a closely related field such as violence prevention, safeguarding, or public health, with the ability to apply that knowledge to AI systems.
- Demonstrated experience designing research, evaluation frameworks, or interventions for harm categories such as violent extremism, CSEA, self-harm and crisis, or targeted violence.
- The ability to translate real world knowledge of how a harm works into a way of testing whether an AI system handles it safely.
- Comfort and demonstrated resilience working with highly sensitive or graphic content (violence, extremist material, crisis content), with awareness of wellbeing practices for this kind of work.
- Strong written communication, able to produce credible, non promotional material for technical and government audiences.
- Sound judgement working with ambiguity and sensitive material.
- Availability for a close to full time commitment over approximately 6 weeks.
- Willingness to undertake relevant security clearance procedures if required by the engagement.
Benefits
- Flexible working arrangements.
- Opportunity to work on diverse, impactful projects.
- Competitive consultancy rates.
- Remote working options available.
Moonshot works to counter violent extremism and online harms, combining expertise in violence prevention, behavioural risk, and online harms with applied AI safety evaluation.
Salary not disclosed