Skip to main content
Mirantis

AI Infrastructure & Platform Operations Engineer

RemoteEU
Published
Role
DevOps
Experience
Mid
Employment
Full-time
$60k–$67k/yr
Check eligibility

Open to Anywhere in EU. Set where you work from to check your eligibility.

No BS summary

Platform/infrastructure operations engineer with 3+ years in infrastructure, SRE, cloud, network, datacenter, or related ops. Must have strong Linux troubleshooting, production Kubernetes knowledge, networking fundamentals, and be able to work shifts. Remote role hiring in the EU.

Core skills

LinuxKubernetesNVIDIA GPU

Required skills

NVIDIA GPU infrastructure

Optional skills

InfiniBandNVIDIA UFMGrafanaPrometheusELKOpenTelemetryInfrastructure as CodeNVIDIA GPU infrastructure

What you'll do

  • Monitor, operate, and support production AI infrastructure platforms.
  • Investigate and resolve infrastructure, networking, hardware, and platform-related incidents.
  • Support NVIDIA GPU infrastructure and associated platform services.
  • Monitor and troubleshoot Kubernetes-based environments.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, datacenter personnel, and service delivery teams to resolve technical issues.
  • Participate in incident response, root cause analysis, and operational improvement activities.
  • Contribute to improvements in monitoring, observability, automation, and operational processes.
  • Maintain operational documentation, runbooks, and knowledge articles.

What they require

  • 3+ years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or related technical roles.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of networking concepts and experience diagnosing infrastructure-related issues.
  • Working knowledge of Kubernetes in production environments.
  • Experience supporting production infrastructure and services.
  • Strong analytical and problem-solving skills.
  • Experience working within structured operational and incident management processes.
  • Excellent communication and collaboration skills.
  • Ability to work within a shift-based operational environment.
  • Preferred: Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.
  • Preferred: InfiniBand networking and NVIDIA UFM.
  • Preferred: Kubernetes platform operations.
  • Preferred: AI infrastructure or HPC environments.
  • Preferred: Site Reliability Engineering (SRE) or Platform Engineering.
  • Preferred: Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Preferred: Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Preferred: Large-scale distributed systems and production platforms.
  • Preferred: Experience in one or more of the following areas is highly desirable: InfiniBand networking and NVIDIA UFM.
  • Preferred: Experience in one or more of the following areas is highly desirable: Kubernetes platform operations.
  • Preferred: Experience in one or more of the following areas is highly desirable: AI infrastructure or HPC environments.
  • Preferred: Experience in one or more of the following areas is highly desirable: Site Reliability Engineering (SRE) or Platform Engineering.
  • Preferred: Experience in one or more of the following areas is highly desirable: Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Preferred: Experience in one or more of the following areas is highly desirable: Infrastructure automation technologies and Infrastructure-as-Code practices.
  • Preferred: Experience in one or more of the following areas is highly desirable: Large-scale distributed systems and production platforms.
  • Preferred: Experience in NVIDIA GPU infrastructure and accelerated computing platforms.
  • Preferred: Experience in InfiniBand networking and NVIDIA UFM.
  • Preferred: Experience in Kubernetes platform operations.
  • Preferred: Experience in AI infrastructure or HPC environments.
  • Preferred: Experience in Site Reliability Engineering (SRE) or Platform Engineering.
  • Preferred: Experience in observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
  • Preferred: Experience in infrastructure automation technologies and Infrastructure-as-Code practices.
  • Preferred: Experience in large-scale distributed systems and production platforms.

Benefits

  • Work with some of the most advanced AI infrastructure environments in production today.
  • Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
  • Help define how next-generation AI infrastructure is operated and supported.
  • Be part of a team shaping the future of AI-powered operations through k0rdent AI.
  • Join a growing organisation investing heavily in AI infrastructure and platform services.

cloud computing software and services company

🇺🇸 United StatesCloud InfrastructureMid-sizemirantis.com/

What people say about this company

3.2/ 5

$60k–$67k/yr