Skip to main content
Pragmatike

AI Infrastructure Engineer (GPU)

RemoteEMEA· UTC-1…UTC+2
Published
Role
AI / ML
Experience
Senior
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to Anywhere in EMEA · UTC-1…UTC+2. Set where you work from to check your eligibility.

No BS summary

Seeking an AI Infrastructure Engineer with 4+ years of experience in ML Ops, Platform Engineering, or SRE, focusing on production-grade model serving and AI infrastructure. Must have hands-on experience with model serving frameworks (vLLM, TGI, Triton), container orchestration, GPU workloads, and MLOps tooling. Proficiency in Python and infrastructure-as-code tools is required. Must have an ownership mindset and be able to operate independently in a remote-first environment.

Core skills

GPUmodel servingML inference

Required skills

vLLMTGITritonPythonTerraformHelm

Optional skills

KubeflowMLflowKubeAICUDAROCm

Required languages

English Fluent

What you'll do

  • Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent
  • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models
  • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers
  • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance
  • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health
  • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments
  • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities
  • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment

What they require

  • 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems
  • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent
  • Strong background in container orchestration and operating GPU-based workloads in production
  • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines
  • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar)
  • Strong understanding of distributed systems, performance tuning, and production reliability engineering
  • Ability to effectively use AI coding assistants to accelerate development and debugging workflows
  • Ownership mindset with the ability to operate independently in a remote-first environment

Benefits

  • Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform
  • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup
  • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture
  • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems
  • Influence core engineering decisions and define best practices that will scale with the company.

Pragmatike is recruiting on behalf of a global leader in mobile payment and digital content monetization. The client operates across 60+ countries, connecting telecom operators, merchants, content providers, media companies and brands, powering Direct Carrier Billing, Mobile Money, Local Payment Methods, content distribution, interactivity programs and multi-channel activations.

FintechStartup
Salary not disclosed