Skip to main content

Platform Architect (AI/ML Infrastructure, GCP-focused)

RemoteLATAM
Published
Role
DevOps
Experience
Senior
Employment
Contract
Salary not disclosed
Check eligibility

Open to Anywhere in LATAM. Set where you work from to check your eligibility.

No BS summary

Senior platform architect for AI/ML infrastructure on Google Cloud. Requires 5+ years in platform engineering/SRE/MLOps with hands-on production ML serving and pipelines, deep Terraform, Kubernetes/GKE and GitOps (ArgoCD/Flux). Remote in LATAM; working hours EST; payment in USD.

Core skills

TerraformArgoCD/FluxGKE

Required skills

KubernetesVPCCompute EngineIAMCloud StorageBigQueryDataflow/Pub/Sub/DataprocGitHub Actions/Cloud Build/GitLab CI

Optional skills

GKE node auto-provisioning / GPU schedulingVertex AIGemini APIArgo WorkflowsKubeflowCloud ComposerAirflowVertex AI Pipelines

What you'll do

  • Build and operate model and inference serving infrastructure managing latency, throughput, autoscaling, and reliability for real-time and batch inference across multiple tenants.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies (canary, shadow, A/B), and safe rollback.
  • Operate agentic and LLM workloads in production, managing providers/gateways, quota/throttling, guardrails, prompt/version management, and graceful degradation.
  • Build reproducible, automated ML pipelines as code with lineage and reproducibility; extend IaC patterns to ML systems using Terraform and multi-project design.
  • Run ML workloads on multi-tenant Kubernetes (GKE), managing GPU/accelerator scheduling, workload placement, tenant isolation, observability, SLOs and cost efficiency.

What they require

  • 5+ years in platform engineering, SRE, MLOps, or infrastructure with meaningful time operating production systems at scale.
  • Hands-on experience deploying and operating ML or AI workloads in production (serving, inference, or training infrastructure).
  • Deep Terraform expertise managing complex state, reusable modules, and multi-project configurations with CI-driven plan/apply workflows.
  • Strong GitOps background (ArgoCD or Flux) and experience operating declarative infrastructure management in production.
  • Strong GCP background including VPC networking, Compute Engine, IAM, Cloud Storage and hands-on BigQuery experience (partitioning, clustering, query cost/performance tuning).

Benefits

  • Payment in USD
Machine Learning
Salary not disclosed