Adversarial Task Writer for AI Security RL Gyms
- Role
- Security
- Employment
- Full-time
Open to RS only. Set where you work from to check your eligibility.
No BS summary
Full-time Serbia-based adversarial prompt-injection task writer for AI security RL gyms. Needs YAML, Docker/CLI comfort, systematic model testing, and realism in at least one domain such as e-commerce, finance, HR, enterprise SaaS, healthcare, or travel. Pentesting, appsec, LLM security research, or red teaming is strongly preferred.
Core skills
Required skills
On behalf of SafetyTech Client #1, SD Solutions is looking for a talented Adversarial Task Writer for AI Security RL Gyms.
SD Solutions is a staffing company operating globally. Contact us to get more details about the benefits we offer.
Responsibilities:
You design prompt injection scenarios in YAML, run them against frontier models, validate success rates, and submit passing tasks. 5 high-quality tasks per week (full-time equivalent). Per-task compensation, paid on acceptance.
Requirements:
- Adversarial mindset: you think like an attacker and understand how to exploit an AI agent’s helpfulness, authority assumptions, or trust in its environment
- Prompt injection expertise: direct (role-play, encoding, context flooding) and indirect/environment-embedded (poisoned tool responses, malicious content in documents, cross-context leakage)
- Technical writing in YAML
- Comfortable with Docker, CLI tools, and running systematic tests against multiple models
- Domain realism in at least one vertical: e-commerce, finance, HR, enterprise SaaS, healthcare, travel
- Background in pentesting, appsec, LLM security research, or red teaming strongly preferred
The Task
You build adversarial prompt injection tasks for Alice’s RL Gym platform. Each task is a self-contained YAML scenario simulating a realistic AI agent deployment, testing whether the agent can be manipulated into violating its safety policies.
About the company:
A company building specialized evaluation infrastructure for AI safety and robustness testing. Their platform simulates adversarial conditions used by AI development teams to validate agent behavior before deployment. Currently expanding a freelance contributor pool for scenario and environment development.
By applying for this position, you agree to the terms outlined in our Privacy Policy. If you have any questions or concerns regarding our Privacy Policy, please feel free to contact us.
What you'll do
- Design prompt injection scenarios in YAML.
- Run prompt injection scenarios against frontier models.
- Validate success rates.
- Submit passing tasks.
- Produce 5 high-quality tasks per week full-time equivalent.
- Build adversarial prompt injection tasks for Alice’s RL Gym platform.
- Create self-contained YAML scenarios simulating realistic AI agent deployments.
- Test whether agents can be manipulated into violating safety policies.
What they require
- Adversarial mindset: think like an attacker and understand how to exploit an AI agent’s helpfulness, authority assumptions, or trust in its environment.
- Prompt injection expertise: direct techniques such as role-play, encoding, and context flooding, and indirect/environment-embedded techniques such as poisoned tool responses, malicious content in documents, and cross-context leakage.
- Technical writing in YAML.
- Comfortable with Docker, CLI tools, and running systematic tests against multiple models.
- Domain realism in at least one vertical: e-commerce, finance, HR, enterprise SaaS, healthcare, travel.
- Preferred: Background in pentesting, appsec, LLM security research, or red teaming.
Benefits
- Contact SD Solutions to get more details about the benefits they offer.
SD Solutions is a staffing company operating globally. On behalf of AdTech Client #1, it is hiring for a large-scale advertising technology company building audience intelligence and campaign automation tools for brands, agencies, and publishers across multiple digital channels.