Skip to main content
Mitratech

Senior Software Engineer - AI/ML

RemoteGermany only
Published
Role
AI / ML
Experience
Senior
Company size
Enterprise
Salary not disclosed
Check eligibility

Open to DE only. Set where you work from to check your eligibility.

No BS summary

Senior Software Engineer specializing in Generative AI and LLMs, with a focus on agentic systems, RAG, and AI evaluations. Must have production experience with multi-agent workflows, RAG pipelines, and evaluating generative AI outputs. Proficiency in Python, AWS ecosystem (Bedrock, SageMaker), and MLOps is required. Master's degree in ML or CS preferred.

Core skills

Generative AILarge Language Modelsagentic systems

Required skills

Agent Orchestrationmulti-agent systemstool usememory/state managementfault-tolerant routingLangChainLangGraphAutoGenRAG pipelineschunkingembedding modelsvector databasesretrieval tuninganswer synthesisgenerative AI evaluationautomated scoringhuman evaluation designhallucination mitigationdrift monitoringfoundation modelsGenAI providersAWS BedrockOpenAIAnthropicMetafine-tuninginstruction tuningprompt engineeringclassical NLP techniquesNERtext classificationintent detectiontopic modellingAmazon BedrockBedrock AgentsBedrock Knowledge BasesBedrock GuardrailsSageMakerLambdaECS/EKSS3OpenSearchIAMCloudWatchVPC networkingMLOpsCI/CD for MLfeature storescanary deploymentsmonitoringrollbackTerraformAWS CDKPythonpackagingtestingpytesttype hintsasync programmingML frameworksexperiment trackingLangfuseArizeLangsmith

Optional skills

LLM-as-judge patternsBedrock Model Evaluationtraditional ML frameworksfine-tuning workflows

What you'll do

  • Design, build, and operate multi-agent workflows and tool-enabled agents, implementing orchestration logic, state management, safety guardrails, and fallback strategies for resilient production pipelines.
  • Architect and maintain end-to-end RAG systems, covering document ingestion, chunking, embedding, vector retrieval, reranking, and answer synthesis with a focus on quality, attribution, and latency.
  • Evaluate and integrate LLMs and GenAI services across cost, performance, and privacy dimensions, selecting the right mix of managed and in-house models.
  • Develop, version, and optimise prompting strategies; implement automated prompt testing and regression tracking to maintain output quality and reliability.
  • Define and own evaluation frameworks for generative outputs, including automated metrics, LLM-as-judge approaches, human evaluation protocols, hallucination detection, and drift monitoring.
  • Apply classical NLP techniques where appropriate and maintain awareness of data distribution shifts that could impact model behaviour in production.
  • Build and operate scalable, secure AI infrastructure on AWS (Bedrock, SageMaker, Lambda, OpenSearch), following well architected principles and infrastructure-as-code practices.
  • Own the full deployment lifecycle: CI/CD for models and agents, testing strategies, observability, and rollback procedures.
  • Ensure data quality through rigorous validation and augmentation, and proactively source datasets for training, fine-tuning, and evaluation.

What they require

  • Production experience designing multi-agent systems with tool use, memory/state management, and fault-tolerant routing.
  • Hands-on experience building RAG pipelines end-to-end: chunking, embedding models, vector databases, retrieval tuning, and answer synthesis at production scale.
  • Strong experience defining and running evaluation pipelines for generative AI — automated scoring, human evaluation design, hallucination mitigation, and drift monitoring.
  • Demonstrated experience with foundation models and GenAI providers (AWS Bedrock, OpenAI, Anthropic, Meta).
  • Comfortable with fine-tuning, instruction tuning, and prompt engineering at scale.
  • Solid grounding in classical NLP techniques (NER, text classification, intent detection, topic modelling) and good judgement on when to apply them alongside or instead of LLMs.
  • Hands-on experience with Amazon Bedrock: foundation model APIs, Bedrock Agents, Knowledge Bases, and Guardrails.
  • Proficiency with SageMaker, Lambda, ECS/EKS, S3, OpenSearch, IAM, CloudWatch, and VPC networking.
  • Familiarity with model registries, CI/CD for ML, feature stores, canary deployments, monitoring, and rollback.
  • Experience with Terraform or AWS CDK for reproducible infrastructure provisioning.
  • Production-quality Python: packaging, testing (pytest), type hints, async programming, and clean ML pipeline abstractions.
  • Experience with Langfuse, Arize or Langsmith, or equivalent for tracking runs, metrics, and artefacts.
  • Ability to translate ambiguous business goals into concrete technical solutions and communicate tradeoffs to non-technical stakeholders.
  • Strong collaborative instincts — comfortable working across engineering, product, and data teams.
  • A rigorous, evidence-driven mindset: you ship with confidence because you measure, test, and monitor thoroughly.
  • A Master’s degree in Machine Learning, Computer Science with a preference for specialization in the NLP domain.

Benefits

  • We are an equal-opportunity employer that values diversity at all levels. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity, disability, or veteran status.

multinational technology company supporting corporate legal, compliance, and human resources

🇩🇪 GermanySoftwareEnterprisemitratech.com/
Salary not disclosed