Skip to main content
Uvation

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

RemoteRomania only
Published
Role
Unknown
Experience
Senior
Employment
Contract
Salary not disclosed
Check eligibility

Open to RO only. Set where you work from to check your eligibility.

No BS summary

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms.

Core skills

LinuxBare Metal as a Service (BMaaS)GPU infrastructure

Required skills

UbuntuRed HatSUSEBIOS/UEFIRAID controllersFirmware managementiLOiDRACIPMINICsSmartNICsHBA cardsHardware diagnosticsGPUNVIDIA A100NVIDIA H100NVIDIA H200NVIDIA B200NVIDIA DGXGPU provisioningGPU monitoringCUDANCCLGPUDirect StorageNVIDIA Fabric ManagerLVMXFSEXT4NFSiSCSIFibre Channel SANMultipath I/OCephRBDCephFSRGWWEKAVAST DataDell PowerScalePure Storage FlashBladeNetAppNVMe-over-Fabrics (NVMe-oF)RDMABondingVLANsRoutingMTU optimizationDNSDHCP100G Ethernet200G Ethernet400G EthernetRoCESpine-Leaf architecturesNVIDIA Spectrum-XMellanox/NVIDIA ConnectXLayer 2Layer 3BashPythonLinux & Bare Metal Infrastructure/AI Factory & GPU Infrastructure/Enterprise Storage & Data Platforms/Networking & Infrastructure/Operations & Reliability

Optional skills

KubernetesKVMVMwareOpenShift VirtualizationAnsibleNVIDIA Base Command ManagerSlurmPrometheus

What you'll do

  • Designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure
  • Building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms
  • Deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • GPU provisioning and lifecycle management
  • GPU monitoring and performance optimization
  • Supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Advanced Linux storage administration
  • Experience with high-performance AI storage platforms
  • Strong networking knowledge
  • Experience with high availability, clustering, and disaster recovery
  • Troubleshooting across Linux operating systems, hardware platforms, GPU infrastructure, networking, and enterprise storage
  • Supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Creating operational documentation, runbooks, and infrastructure standards

What they require

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
  • Strong understanding of server hardware, including BIOS/UEFI, RAID controllers, firmware management, iLO/iDRAC/IPMI, NICs and SmartNICs, HBA cards, hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including A100, H100, H200, B200, or equivalent GPU platforms, NVIDIA DGX and OEM GPU servers, GPU provisioning and lifecycle management, GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Understanding of GPU resource allocation and scheduling, multi-GPU systems, GPU networking requirements, high-bandwidth, low-latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as CUDA, NCCL, GPUDirect Storage, NVIDIA Fabric Manager, NVIDIA Base Command (preferred)
  • Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, Multipath I/O
  • Strong hands-on experience with Ceph, including cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW, capacity planning, performance tuning, failure recovery
  • Experience with high-performance AI storage platforms such as WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, NetApp
  • Understanding of NVMe-over-Fabrics (NVMe-oF), RDMA, GPUDirect Storage, parallel file systems, AI data pipelines
  • Strong networking knowledge: bonding, VLANs, routing, MTU optimization, DNS, DHCP
  • Experience with high-performance data center networking: 100G/200G/400G Ethernet, RoCE, RDMA, Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting
  • Experience with high availability, clustering, and disaster recovery
  • Strong troubleshooting skills across Linux operating systems, hardware platforms, GPU infrastructure, networking, enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards

Uvation delivers technology solutions and maintains a high-quality product marketplace.

IT Services
Salary not disclosed