Forward Deployed AI Engineer
- Role
- AI / ML
- Experience
- Mid
Open to GB only. Set where you work from to check your eligibility.
No BS summary
Engineer to build production agentic LLM systems embedded with enterprise clients. Strong Python-based production LLM/agent engineering, RAG/retrieval and data-engineering skills required. Client-facing role; production cloud delivery, observability and security/compliance experience expected.
Core skills
Required skills
Optional skills
Required languages
gravity9 is a boutique IT consulting company headquartered in the UK with offices in the US, Canada, Poland, and Colombia. Our team has deep experience in engineering, experience design and product management. We enjoy a challenge and pride ourselves on working with our clients on their most complex problems, finding elegant and flexible solutions that help them transform their businesses.
Join Our Growing Team - Future Opportunities Await!
We are excited to share that our company is experiencing significant growth and expansion, and as a result, we're on the lookout for talented individuals to join our team.
We want to be upfront about our hiring process. At this moment, we are in the process of securing projects that align with our business goals and objectives. Our intention is to ensure that once we bring new team members on board, they will have a meaningful and impactful role to play from day one. While we might not be able to extend an offer immediately, we are highly interested in considering you for these future projects.
If you're open to the idea of potentially joining our team in the future, we encourage you to continue with our interview process. We value your time and effort and believe that getting to know you better will help us make an informed decision.
Once you have effectively concluded the entire recruitment procedure, we will retain your profile and reach out to you once a suitable opportunity emerges. Additional interviews will not be necessary at this stage, and we anticipate being able to extend an offer to you promptly.
What does the recruitment process look like?
1. Recruiter Screen
In the first call, our recruiter will learn more about you and your story to check a potential fit for gravity9. This is also an opportunity to ask your questions about the role and company. This step might take around 60 mins.
2. Tech 1 Interview with our AI Practice Lead
In this meeting your potential future teammate will take a deeper dive into your experience and what you could bring to the team. This step will take around 75 mins.
3. Tech 2 Interview with our software consultants
In this meeting your potential future teammate will take a deeper dive into your experience and what you could bring to the team. This step will take around 60 mins.
4. Final Interview
You made it to the very last stage! Here we already strive to cooperate with you and give you an opportunity to meet our leadership team during this informal talk. This step will take around 30 mins.
Thank you for considering the opportunity to be a part of our growing journey. We look forward to the possibility of working together and achieving great things.
Role Description
gravity9 is expanding its Forward Deployed Engineering team to build production agentic AI systems for enterprise clients, in partnership with leading frontier-model providers. We already support clients from different verticals to have agentic and RAG systems live in production across healthcare, financial services, retail and global logistics.
As a Forward Deployed AI Engineer you work embedded in the client's environment, from discovery, through architecture and build, to production and handover. This is end-to-end agent engineering, not proofs of concept and not advisory work. Two flavours of engagement:
- Internal enterprise use cases, such as reconciliation and KYC workflows in financial services, clinical and claims workflows in healthcare, operations and supply-chain workflows in logistics.
- Product- and customer-facing agents, more greenfield, often exploratory, built into the client's own product.
Engagements are typically a team of engineers over a few months, working shoulder-to-shoulder with the client's team including architects, DevOps and QA.
- Embed with the client. Work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers. You are visible to the client from day one.
- Design and build production agentic AI systems. Multi-agent orchestration, tool and function calling, retrieval, planning and routing, human-in-the-loop checkpoints, state and checkpointing, guardrails, failure handling and recovery.
- Do the unglamorous data work. A large share of every engagement is data engineering: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing. Models reason over data, bad data beats a good model every time.
- Own evaluation and accuracy. Establish a baseline eval dataset at the start of the engagement, automate grading, and track groundedness, faithfulness and retrieval quality per tool, not just at the agent level. Be ready to defend accuracy numbers to a sceptical enterprise stakeholder.
- Engineer for cost and latency. Model selection and routing (cheaper, faster models for non-reasoning steps; frontier models where reasoning genuinely earns it), prompt and context budgeting, caching. Cost is a non-negotiable metric on every engagement.
- Ship it properly. Observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review.
- Transfer knowledge deliberately. We don't run a long-term support business. Every engagement is designed so the client owns and can extend the system after we leave. You architect with their team in the room, pair with their engineers, and hand over working monitoring and documentation.
Technical skills
Essential
- Programming: Strong Python (async, typing, testing, packaging). Working competence in TypeScript / Node is a plus.
- LLM application engineering: Hands-on production experience with frontier models: prompt and context engineering, structured outputs, tool and function calling, streaming, token and context-window management.
- Agentic architecture: Built and shipped multi-agent or agentic workflows, planner/router patterns, supervisor and sub-agent designs, ReAct-style loops, state and checkpointing, retries and interrupts, human-in-the-loop review gates.
- Frameworks: At least one of: Anthropic Agent SDK, LangGraph, LlamaIndex, or an equivalent orchestration framework, plus the judgement to know when to use none of them.
- RAG and retrieval: Chunking strategy, embeddings, hybrid and semantic search, re-ranking, citation and provenance, natural-language-to-query translation.
- Data platforms: NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or a strong equivalent, plus SQL. Comfortable designing schemas for agent retrieval, not just for OLTP.
- Data engineering: Building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
- Evaluation: Building eval harnesses and eval datasets, LLM-as-judge with its limits understood, regression testing of prompts and agents, metric selection per use case.
- Cloud: Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
- Engineering discipline: Git, code review, testing, CI/CD, IaC (Terraform or equivalent), observability.
Strong advantage
- Anthropic certification (Claude Developer / Anthropic-issued credential), an explicit advantage at shortlisting.
- Anthropic Agent SDK production experience, a significant bonus.
- MCP (Model Context Protocol): building servers and clients, tool exposure, auth patterns.
- Claude Code as an autonomous SDLC agent: sub-agents, hooks, custom skills, agent marketplaces, context sharing across a team.
- LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar).
- Regulated-environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails and red-teaming.
- Graph or taxonomy-based knowledge representation alongside vector retrieval.
- Kafka / streaming, Databricks, or comparable large-scale data platform experience.
Nice to have
- Voice and multimodal agents; evaluation of non-text outputs.
- FinOps for AI workloads; unit-economics modelling for agent systems.
- Open-source contribution to the agent / LLM ecosystem, or conference speaking.
- Prior experience as an FDE, solutions architect, or delivery consultant at a frontier-model, data platform, or infrastructure vendor.
Soft skills
An FDE is an engineer who is safe in front of a client. We screen as hard on this section as on the technical one.
- Client presence and credibility. Can hold a technical conversation with a client architect and a business conversation with their CTO or head of operations in the same hour, and be trusted by both. Can present, whiteboard, and answer hard questions without deflecting.
- Translation. Turns a business problem described by non-technical people such as a clinician, campaign manager or supply-chain planner into a technical design, and explains the technical design back in their language.
- Ego-free collaboration in a supporting role. On some engagements the tech lead will come from the client or a partner, not from us. You need to contribute strongly, disagree well, and take direction without friction. Brilliant-but-territorial doesn't work here.
- Comfort with ambiguity and greenfield. Engagements often start before the requirements, the data access, or sometimes the problem statement are settled. You make progress anyway, and you make the ambiguity visible rather than hiding it.
- Bias to production. Instinctively asks "how does this get deployed, monitored and maintained?" rather than stopping at a working notebook.
- Teaching instinct. Actively enjoys upskilling the client's engineers, because self-sufficiency at handover is the definition of success, not follow-on billing.
- Ownership and pace. Short engagements, small teams, no layers to hide behind. You unblock yourself, chase the access request, and follow the thread to the answer.
- Commercial awareness. Understands that scope, cost and the client's willingness to pay are part of the engineering problem, and contributes to scoping and estimating honestly.
- Resilience and adaptability. New client, new domain, new stack every few months; occasional travel; occasionally a sceptical stakeholder who has been told AI is coming for their job. You stay steady and constructive.
- Written communication. Clear design docs, decision records, handover material and status updates in English. Much of this is asynchronous and cross-timezone.
- Genuine curiosity about the field. This ecosystem changes monthly. We want people who are already reading, building side projects and forming opinions — not waiting for training to be scheduled for them.
Experience profile
- FDE (mid): 3+ years software engineering, 1+ years hands-on LLM / agentic work with at least one system live in production. Client-facing exposure.
- Senior FDE: 5+ years engineering, 2+ years AI or agentic AI, has owned the architecture of at least one production agent system end-to-end and led a client conversation about it.
- Lead FDE: 7+ years, has led delivery teams of 3–7 people, can act as engagement tech lead, contributes to pre-sales and estimation, and can mentor a growing FDE bench.
What success looks like
- First 90 days — owning a meaningful component of a client-facing agent system, including its evals, and trusted in front of the client.
- 6 months — designing agent architectures independently, leading a workstream, and running the handover and knowledge-transfer track.
- 12 months — a reference point for others on the team, contributing to reusable assets, pre-sales and interviewing.
What we offer
- Support for vendor certifications and access to partner training programmes.
- Company-wide AI tooling as part of how we work day to day.
What you'll do
- Embed with the client and work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers.
- Design and build production agentic AI systems: multi-agent orchestration, tool and function calling, retrieval, planning and routing, human-in-the-loop checkpoints, state and checkpointing, guardrails, failure handling and recovery.
- Perform data engineering: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing.
- Own evaluation and accuracy: establish a baseline eval dataset, automate grading, and track groundedness, faithfulness and retrieval quality per tool.
- Engineer for cost and latency: model selection and routing, prompt and context budgeting, caching.
- Ship production concerns: observability and tracing, CI/CD, IaC, monitoring the client can operate, security and compliance review.
- Transfer knowledge deliberately so the client owns and can extend the system after handover; pair with client engineers and provide documentation and monitoring.
- Work across internal enterprise and product-facing agent engagements and adapt approach per use case.
- Work embedded in the client's environment, from discovery, through architecture and build, to production and handover.
- Build end-to-end agent engineering systems, not proofs of concept and not advisory work.
- Work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers.
- Be visible to the client from day one.
- Do the unglamorous data work: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing.
- Establish a baseline eval dataset at the start of the engagement.
- Automate grading.
- Track groundedness, faithfulness and retrieval quality per tool, not just at the agent level.
- Defend accuracy numbers to a sceptical enterprise stakeholder.
- Engineer for cost and latency.
- Handle model selection and routing, prompt and context budgeting, and caching.
- Ship observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review.
- Transfer knowledge deliberately.
- Architect with the client's team in the room.
- Pair with the client's engineers.
- Hand over working monitoring and documentation.
- Own a meaningful component of a client-facing agent system in the first 90 days, including its evals.
- Design agent architectures independently after 6 months.
- Lead a workstream after 6 months.
- Run the handover and knowledge-transfer track after 6 months.
- Contribute to reusable assets, pre-sales and interviewing after 12 months.
What they require
- Strong Python (async, typing, testing, packaging).
- Working competence in TypeScript / Node is a plus.
- Hands-on production experience with frontier models: prompt and context engineering, structured outputs, tool and function calling, streaming, token and context-window management.
- Built and shipped multi-agent or agentic workflows, planner/router patterns, supervisor and sub-agent designs, ReAct-style loops, state and checkpointing, retries and interrupts, human-in-the-loop review gates.
- Experience with at least one orchestration framework: Anthropic Agent SDK, LangGraph, LlamaIndex or equivalent.
- RAG and retrieval: chunking strategy, embeddings, hybrid and semantic search, re-ranking, citation and provenance, natural-language-to-query translation.
- NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or equivalent, plus SQL.
- Data engineering experience building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
- Building eval harnesses and eval datasets, understanding limits of LLM-as-judge, regression testing of prompts and agents, metric selection per use case.
- Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
- Engineering discipline: Git, code review, testing, CI/CD, IaC (Terraform or equivalent), observability.
- Client-facing exposure and ability to present and hold technical/business conversations with client stakeholders.
- Written communication: clear design docs, decision records, handover material and status updates in English.
- Experience profile (examples provided): FDE (mid): 3+ years software engineering, 1+ years hands-on LLM/agentic work with at least one system live in production. Senior FDE: 5+ years engineering, 2+ years AI or agentic AI and owned architecture of at least one production agent system. Lead FDE: 7+ years, led delivery teams of 3–7 people and engaged in pre-sales/estimation/mentoring.
- At least one of: Anthropic Agent SDK, LangGraph, LlamaIndex, or an equivalent orchestration framework, plus the judgement to know when to use none of them.
- Data platforms: NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or a strong equivalent, plus SQL.
- Comfortable designing schemas for agent retrieval, not just for OLTP.
- Data engineering: Building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
- Evaluation: Building eval harnesses and eval datasets, LLM-as-judge with its limits understood, regression testing of prompts and agents, metric selection per use case.
- Cloud: Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
- Preferred: Working competence in TypeScript / Node is a plus.
- Preferred: Anthropic certification (Claude Developer / Anthropic-issued credential), an explicit advantage at shortlisting.
- Preferred: Anthropic Agent SDK production experience, a significant bonus.
- Preferred: MCP (Model Context Protocol): building servers and clients, tool exposure, auth patterns.
- Preferred: Claude Code as an autonomous SDLC agent: sub-agents, hooks, custom skills, agent marketplaces, context sharing across a team.
- Preferred: LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar).
- Preferred: Regulated-environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails and red-teaming.
- Preferred: Graph or taxonomy-based knowledge representation alongside vector retrieval.
- Preferred: Kafka / streaming, Databricks, or comparable large-scale data platform experience.
- Preferred: Voice and multimodal agents; evaluation of non-text outputs.
- Preferred: FinOps for AI workloads; unit-economics modelling for agent systems.
- Preferred: Open-source contribution to the agent / LLM ecosystem, or conference speaking.
- Preferred: Prior experience as an FDE, solutions architect, or delivery consultant at a frontier-model, data platform, or infrastructure vendor.
- Client presence and credibility.
- Can hold a technical conversation with a client architect and a business conversation with their CTO or head of operations in the same hour, and be trusted by both.
- Can present, whiteboard, and answer hard questions without deflecting.
- Turns a business problem described by non-technical people such as a clinician, campaign manager or supply-chain planner into a technical design, and explains the technical design back in their language.
- Ego-free collaboration in a supporting role.
- Contribute strongly, disagree well, and take direction without friction.
- Comfort with ambiguity and greenfield.
- Make progress when requirements, data access, or the problem statement are not settled, and make the ambiguity visible rather than hiding it.
- Bias to production.
- Teaching instinct.
- Actively enjoys upskilling the client's engineers.
- Ownership and pace.
- Unblock yourself, chase the access request, and follow the thread to the answer.
- Commercial awareness.
- Understands that scope, cost and the client's willingness to pay are part of the engineering problem, and contributes to scoping and estimating honestly.
- Resilience and adaptability.
- Stay steady and constructive with a new client, new domain, new stack every few months, occasional travel, and sceptical stakeholders.
- Written communication.
- Clear design docs, decision records, handover material and status updates in English.
- Comfort with asynchronous and cross-timezone work.
- Genuine curiosity about the field.
- Already reading, building side projects and forming opinions.
- FDE (mid): 3+ years software engineering, 1+ years hands-on LLM / agentic work with at least one system live in production.
- Client-facing exposure.
- Senior FDE: 5+ years engineering, 2+ years AI or agentic AI, has owned the architecture of at least one production agent system end-to-end and led a client conversation about it.
- Lead FDE: 7+ years, has led delivery teams of 3–7 people, can act as engagement tech lead, contributes to pre-sales and estimation, and can mentor a growing FDE bench.
Benefits
- Support for vendor certifications and access to partner training programmes.
- Company-wide AI tooling as part of how we work day to day.
gravity9 is a boutique IT consulting company headquartered in the UK with offices in the US, Canada, Poland, and Colombia. Our team has deep experience in engineering, experience design and product management.