Software Engineer, Infrastructure & Platform
- Role
- Backend
- Experience
- Mid
Open to US only. Set where you work from to check your eligibility.
No BS summary
US-based (remote) backend/infrastructure engineer with 3–5+ years building sandboxed, reproducible environments where AI models can execute code and use tools. Strong Python, Docker/Kubernetes, GCP/AWS, Linux, and infrastructure security. Prior AI experience is welcome but not required.
Core skills
Required skills
Optional skills
About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations, and intelligence collection enable engineering, safety, and security teams to stay ahead of evolving threats and deploy AI systems safely. Software Engineer, Infrastructure & Platform About the Role We are seeking a Software Engineer, Infrastructure & Platform to build the systems and infrastructure that power advanced AI evaluations, including evaluations focused on autonomous model behavior, agentic systems, and loss-of-control risks. This is a hands-on engineering role at the intersection of backend systems, infrastructure, and AI. You will build secure and reproducible environments where frontier models can interact with tools, execute code, complete complex tasks, and operate across realistic multi-step workflows, including machine learning research and engineering. The ideal candidate has strong backend and infrastructure fundamentals with attention to security, enjoys debugging complex distributed systems, and is excited to apply those skills to difficult problems in AI safety and evaluation. What You'll Do Design and build sandboxed evaluation environments where AI models can safely execute code, use tools, interact with services, and complete complex tasks. Build backend services and infrastructure supporting large-scale, repeatable AI and agentic evaluations. Develop agent scaffolding and evaluation harnesses, including tool-use loops, context management, retries, state management, token budgets, and multi-agent or subagent workflows. Build systems for provisioning and orchestrating isolated environments using technologies such as Docker, Kubernetes, VMs, and cloud infrastructure. Design secure approaches to networking, permissions, secrets, credentials, and resource isolation for model-driven environments. Develop APIs, internal tools, and automation that allow analysts, engineers, and subject-matter experts to create and run evaluations efficiently. Improve the reliability and reproducibility of evaluations through logging, observability, snapshotting, debugging tools, and automated testing. Build systems capable of running thousands of evaluation tasks reliably and capturing the artifacts and telemetry needed to understand model behavior. Partner with analysts, red teamers, and domain experts to translate complex evaluation ideas into robust technical systems. Investigate failures across the evaluation stack and distinguish between model limitations and infrastructure, harness, or environment failures. What We're Looking For 3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering. Strong programming skills in Python and experience building production-quality software. Experience designing and operating backend services or distributed systems. Hands-on experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies. Experience working with GCP, AWS, or similar cloud infrastructure. Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security. Experience with infrastructure-as-code or automation tools such as Terraform. Strong debugging skills and comfort diagnosing failures across application, infrastructure, and networking layers, especially in agentic loops. Ability to build systems that are reproducible, observable, scalable, and secure. Comfort working on ambiguous technical problems where the architecture and requirements may evolve quickly. Interest in AI systems, agentic workflows, AI security, or model evaluations. Prior professional AI experience is helpful but not required. Nice to Have Experience building developer platforms, CI/CD systems, test infrastructure, sandboxes, or ephemeral compute environments. Experience with agent frameworks, LLM APIs, tool-calling systems, or AI evaluation infrastructure. Experience designing secure execution environments for untrusted or semi-trusted code. Background in SRE, platform engineering, cloud infrastructure, cybersecurity, or developer tooling. Experience with distributed task execution, queues, workflow orchestration, or large-scale automated testing. Familiarity with AI safety, adversarial testing, model evaluations, or autonomous-agent systems. Familiarity with agentic AI fundamentals, including common harnesses, Model Context Protocol, agent benchmarks, and security risks to AI agents. Compensation & Benefits Salary Range: $110K–$160K, depending on experience and location Bonus: Performance-based annual bonus Professional Development: Support for conferences, continuing education, or leadership training Work Environment: Fully remote, U.S.-based Health Benefits: Comprehensive health, dental, and vision coverage Time Off: Generous PTO and paid holiday schedule Retirement: 401(k) plan
What you'll do
- Design and build sandboxed evaluation environments where AI models can safely execute code, use tools, and complete complex tasks.
- Build backend services and infrastructure supporting large-scale, repeatable AI and reinforcing-dependent autonomous-model research.
- Develop agent scaffolding and evaluation harnesses (tool-use loops, context management, retries, token budgets, multi-agent workflows).
- Provision and orchestrate isolated environments with Docker, Kubernetes, VMs, and cloud infrastructure.
- Design secure approaches to networking, permissions, secrets, credentials, and resource isolation for model-driven environments.
- Develop APIs, internal tools, and automation to create and run evaluations at scale.
- Improve reliability and reproducibility through logging, observability, snapshotting, debugging tools, and automated testing.
- Run thousands of evaluation tasks reliably and capture artifacts and telemetry.
- Investigate failures across the evaluation stack, distinguishing model limitations from infrastructure/harness failures.
What they require
- 3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering.
- Strong programming skills in Python and experience building production-quality software.
- Experience designing and operating backend services or distributed systems.
- Hands-on experience with Docker, Kubernetes, virtual machines, or other container/orchestration technologies.
- Experience working with GCP, AWS, or similar cloud infrastructure.
- Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security.
- Experience with infrastructure-as-code or automation tools such as Terraform.
- Strong debugging skills in agentic loops.
- Ability to build reproducible, observable, scalable, secure systems.
- Comfort with ambiguous technical problems that evolve quickly.
- Interest in AI systems, agentic workflows, AI security, or model evaluations.
Benefits
- Salary range $110K–$160K depending on experience and location.
- Performance-based annual bonus.
- Professional development support for conferences, continuing education, or leadership training.
- Fully remote work.
- Comprehensive health, dental, and vision coverage.
- Generous PTO and paid holidays.
- 401(k) plan.
10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Its adversarial red teaming, model evaluations, and intelligence collection enable engineering, safety, and security teams to stay ahead of evolving threats and deploy AI systems safely.