Skip to main content
Vultr

Infrastructure Production Engineer

RemoteUnited States only
Published
Role
DevOps
Employment
Full-time
$60k–$80k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Associate Infrastructure Production Engineer for GPU hardware validation in production cloud infrastructure. Needs strong Linux, Python/shell scripting, infrastructure diagnostics, and data center networking basics. Remote United States only.

Core skills

LinuxPythonAnsible

Required skills

Shell scriptingDHCPIPv6ICMP

What you'll do

  • Design, develop, and maintain automated diagnostic, validation, and remediation frameworks for production GPU hardware (NVIDIA and AMD) using vendor tooling, Python, and infrastructure automation technologies.
  • Engineer and support Python-based agents, services, APIs, and Ansible automation (playbooks, roles, pipelines) that orchestrate hardware provisioning, telemetry collection, health monitoring, and production onboarding workflows.
  • Analyze workload performance, utilization, thermals, and diagnostic output to identify hardware and system issues, enhance validation methodologies, and improve infrastructure readiness standards.
  • Execute and evolve testing and verification processes for production onboarding and Return Material Authorization (RMA), developing automation enhancements to improve reliability, scalability, and coverage.
  • Contribute to production stability by building tooling that gates hardware deployment, enforces quality standards, and reduces systemic infrastructure risk.
  • Document system designs, automation logic, validation methodologies, and operational guidance; maintain accurate Jira records reflecting engineering activities, findings, and outcomes.
  • Identify gaps or inefficiencies within testing, validation, or automation processes and design technical solutions to address them in collaboration with engineering teams.

What they require

  • Strong analytical skills and attention to detail with the ability to evaluate, enhance, and optimize validation and testing methodologies.
  • Ability to clearly document work performed, including commands executed, observations, and outcomes
  • Strong Linux proficiency, and experience working within Linux production environments in command-line system contexts.
  • Ability to design, write, and modify Python and shell scripts to support infrastructure diagnostics, automation, and validation workflows (required after initial training).
  • Strong written and verbal communication skills and the ability to collaborate effectively with technical teams
  • Understanding of data center networking concepts and protocols such as DHCP, IPv6, and ICMP
  • Ability to collaborate with engineering teams to deliver automation solutions, reliability improvements, and production-ready tooling.
  • Preferred: Familiarity with infrastructure automation or configuration management frameworks such as Ansible is strongly preferred.

Benefits

  • 100% company-paid insurance premiums for employee medical, dental and vision plans.
  • 401(k) plan that matches 100% up to 4%, with immediate vesting
  • Professional Development Reimbursement of $2,500 each year
  • 11 Holidays + Paid Time Off Accrual + Rollover Plan
  • Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year
  • $500 stipend for remote office setup in first year + $400 each following year
  • Internet reimbursement up to $75 per month
  • Gym membership reimbursement up to $50 per month
  • Company paid Wellable subscription

Vultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions.

Cloud InfrastructureEnterprise
$60k–$80k/yr