Skip to main content
gravity9

Forward Deployed AI Engineer

RemoteUnited Kingdom only
Published
Role
AI / ML
Experience
Mid
Salary not disclosed
Check eligibility

Open to GB only. Set where you work from to check your eligibility.

No BS summary

Engineer to build production agentic LLM systems embedded with enterprise clients. Strong Python-based production LLM/agent engineering, RAG/retrieval and data-engineering skills required. Client-facing role; production cloud delivery, observability and security/compliance experience expected.

Core skills

Agentic LLM systemsRAG / retrieval & embeddingsData engineering for LLMs

Required skills

PythonTypeScriptNode.jsPrompt engineeringTool and function callingStreamingToken and context-window managementAnthropic Agent SDK/LangGraph/LlamaIndexRetrieval-augmented generation (RAG)EmbeddingsSemantic searchVector searchRe-rankingCitation and provenanceNatural-language-to-query translationMongoDBSQLData ingestion and enrichmentDocument extraction (PDFs, unstructured sources)Evaluation harnesses and regression testingAWSAWS BedrockAzureGCPContainers (Docker)ServerlessNetworking basicsSecrets managementIAMCI/CDIaC (Terraform)Observability and tracingKafkaDatabricksLLM-as-judge evaluationasynctypingtestingpackagingLLM application engineeringcontext engineeringstructured outputstool callingfunction callingtoken managementcontext-window managementmulti-agent workflowsagentic workflowsplanner/router patternssupervisor and sub-agent designsReAct-style loopsstate and checkpointingretries and interruptshuman-in-the-loop review gatesorchestration frameworkRAGchunking strategyhybrid searchNoSQLaggregation pipelinesAtlas SearchAtlas Vector Searchschema designdata engineeringingestion pipelinesenrichment pipelinesPDFsobject storageeval harnesseseval datasetsLLM-as-judgeregression testingBedrockcontainersGitcode reviewIaCTerraformobservabilityAgentic AI

Optional skills

Anthropic certification (Claude Developer / Anthropic-issued credential)Anthropic Agent SDK production experienceModel Context Protocol (MCP)Claude Code autonomous SDLC patternsLLM observability tooling (LangFuse, LangSmith, Arize or similar)Regulated-environment delivery experience (HIPAA, GDPR, SOC 2, FCA/PRA)Graph or taxonomy-based knowledge representationVoice and multimodal agents

Required languages

English

What you'll do

  • Embed with the client and work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers.
  • Design and build production agentic AI systems: multi-agent orchestration, tool and function calling, retrieval, planning and routing, human-in-the-loop checkpoints, state and checkpointing, guardrails, failure handling and recovery.
  • Perform data engineering: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing.
  • Own evaluation and accuracy: establish a baseline eval dataset, automate grading, and track groundedness, faithfulness and retrieval quality per tool.
  • Engineer for cost and latency: model selection and routing, prompt and context budgeting, caching.
  • Ship production concerns: observability and tracing, CI/CD, IaC, monitoring the client can operate, security and compliance review.
  • Transfer knowledge deliberately so the client owns and can extend the system after handover; pair with client engineers and provide documentation and monitoring.
  • Work across internal enterprise and product-facing agent engagements and adapt approach per use case.
  • Work embedded in the client's environment, from discovery, through architecture and build, to production and handover.
  • Build end-to-end agent engineering systems, not proofs of concept and not advisory work.
  • Work hand-in-hand inside the client's environment, codebase and cloud tenancy, often in a hybrid team alongside their engineers.
  • Be visible to the client from day one.
  • Do the unglamorous data work: ingestion, flattening deeply nested structures, extracting content from unstructured documents, classification, summarisation, tagging, enrichment, indexing.
  • Establish a baseline eval dataset at the start of the engagement.
  • Automate grading.
  • Track groundedness, faithfulness and retrieval quality per tool, not just at the agent level.
  • Defend accuracy numbers to a sceptical enterprise stakeholder.
  • Engineer for cost and latency.
  • Handle model selection and routing, prompt and context budgeting, and caching.
  • Ship observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review.
  • Transfer knowledge deliberately.
  • Architect with the client's team in the room.
  • Pair with the client's engineers.
  • Hand over working monitoring and documentation.
  • Own a meaningful component of a client-facing agent system in the first 90 days, including its evals.
  • Design agent architectures independently after 6 months.
  • Lead a workstream after 6 months.
  • Run the handover and knowledge-transfer track after 6 months.
  • Contribute to reusable assets, pre-sales and interviewing after 12 months.

What they require

  • Strong Python (async, typing, testing, packaging).
  • Working competence in TypeScript / Node is a plus.
  • Hands-on production experience with frontier models: prompt and context engineering, structured outputs, tool and function calling, streaming, token and context-window management.
  • Built and shipped multi-agent or agentic workflows, planner/router patterns, supervisor and sub-agent designs, ReAct-style loops, state and checkpointing, retries and interrupts, human-in-the-loop review gates.
  • Experience with at least one orchestration framework: Anthropic Agent SDK, LangGraph, LlamaIndex or equivalent.
  • RAG and retrieval: chunking strategy, embeddings, hybrid and semantic search, re-ranking, citation and provenance, natural-language-to-query translation.
  • NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or equivalent, plus SQL.
  • Data engineering experience building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
  • Building eval harnesses and eval datasets, understanding limits of LLM-as-judge, regression testing of prompts and agents, metric selection per use case.
  • Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
  • Engineering discipline: Git, code review, testing, CI/CD, IaC (Terraform or equivalent), observability.
  • Client-facing exposure and ability to present and hold technical/business conversations with client stakeholders.
  • Written communication: clear design docs, decision records, handover material and status updates in English.
  • Experience profile (examples provided): FDE (mid): 3+ years software engineering, 1+ years hands-on LLM/agentic work with at least one system live in production. Senior FDE: 5+ years engineering, 2+ years AI or agentic AI and owned architecture of at least one production agent system. Lead FDE: 7+ years, led delivery teams of 3–7 people and engaged in pre-sales/estimation/mentoring.
  • At least one of: Anthropic Agent SDK, LangGraph, LlamaIndex, or an equivalent orchestration framework, plus the judgement to know when to use none of them.
  • Data platforms: NoSQL databases such as MongoDB (aggregation pipelines, Atlas Search, Atlas Vector Search) or a strong equivalent, plus SQL.
  • Comfortable designing schemas for agent retrieval, not just for OLTP.
  • Data engineering: Building ingestion and enrichment pipelines over messy structured and unstructured sources, documents, PDFs, object storage.
  • Evaluation: Building eval harnesses and eval datasets, LLM-as-judge with its limits understood, regression testing of prompts and agents, metric selection per use case.
  • Cloud: Production delivery on AWS (incl. Bedrock), Azure or GCP, containers, serverless, networking basics, secrets management, IAM.
  • Preferred: Working competence in TypeScript / Node is a plus.
  • Preferred: Anthropic certification (Claude Developer / Anthropic-issued credential), an explicit advantage at shortlisting.
  • Preferred: Anthropic Agent SDK production experience, a significant bonus.
  • Preferred: MCP (Model Context Protocol): building servers and clients, tool exposure, auth patterns.
  • Preferred: Claude Code as an autonomous SDLC agent: sub-agents, hooks, custom skills, agent marketplaces, context sharing across a team.
  • Preferred: LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar).
  • Preferred: Regulated-environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails and red-teaming.
  • Preferred: Graph or taxonomy-based knowledge representation alongside vector retrieval.
  • Preferred: Kafka / streaming, Databricks, or comparable large-scale data platform experience.
  • Preferred: Voice and multimodal agents; evaluation of non-text outputs.
  • Preferred: FinOps for AI workloads; unit-economics modelling for agent systems.
  • Preferred: Open-source contribution to the agent / LLM ecosystem, or conference speaking.
  • Preferred: Prior experience as an FDE, solutions architect, or delivery consultant at a frontier-model, data platform, or infrastructure vendor.
  • Client presence and credibility.
  • Can hold a technical conversation with a client architect and a business conversation with their CTO or head of operations in the same hour, and be trusted by both.
  • Can present, whiteboard, and answer hard questions without deflecting.
  • Turns a business problem described by non-technical people such as a clinician, campaign manager or supply-chain planner into a technical design, and explains the technical design back in their language.
  • Ego-free collaboration in a supporting role.
  • Contribute strongly, disagree well, and take direction without friction.
  • Comfort with ambiguity and greenfield.
  • Make progress when requirements, data access, or the problem statement are not settled, and make the ambiguity visible rather than hiding it.
  • Bias to production.
  • Teaching instinct.
  • Actively enjoys upskilling the client's engineers.
  • Ownership and pace.
  • Unblock yourself, chase the access request, and follow the thread to the answer.
  • Commercial awareness.
  • Understands that scope, cost and the client's willingness to pay are part of the engineering problem, and contributes to scoping and estimating honestly.
  • Resilience and adaptability.
  • Stay steady and constructive with a new client, new domain, new stack every few months, occasional travel, and sceptical stakeholders.
  • Written communication.
  • Clear design docs, decision records, handover material and status updates in English.
  • Comfort with asynchronous and cross-timezone work.
  • Genuine curiosity about the field.
  • Already reading, building side projects and forming opinions.
  • FDE (mid): 3+ years software engineering, 1+ years hands-on LLM / agentic work with at least one system live in production.
  • Client-facing exposure.
  • Senior FDE: 5+ years engineering, 2+ years AI or agentic AI, has owned the architecture of at least one production agent system end-to-end and led a client conversation about it.
  • Lead FDE: 7+ years, has led delivery teams of 3–7 people, can act as engagement tech lead, contributes to pre-sales and estimation, and can mentor a growing FDE bench.

Benefits

  • Support for vendor certifications and access to partner training programmes.
  • Company-wide AI tooling as part of how we work day to day.

gravity9 is a boutique IT consulting company headquartered in the UK with offices in the US, Canada, Poland, and Colombia. Our team has deep experience in engineering, experience design and product management.

🇬🇧 United KingdomIT Consulting
Salary not disclosed