Skip to main content
DDN

Platform Support Architect

RemoteUnited Kingdom only
Published
Role
AI / ML
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to GB only. Set where you work from to check your eligibility.

No BS summary

Lead supportability and enablement for AI solutions, focusing on NVIDIA AI Enterprise services, vector databases, RAG/agentic workflows, and high-performance storage/networking. Act as a technical advisor, combining solutions architecture with L3 support, to ensure cohesive AI platform operations.

Core skills

KubernetesNVIDIA GPUsMilvus/Qdrant/Pinecone/pgVector/OpenSearch/Elasticsearch

Required skills

LinuxDockercontainerdHelmCUDANVIDIA GPU OperatorEXAScaler/Lustre/GPFS/Ceph/distributed object storage/enterprise NAS/SANRDMAEthernetInfiniBandNVIDIA NIMNVIDIA NeMoTriton Inference ServerTensorRTTensorRT-LLM

Optional skills

PrometheusGrafanaLokiELKNetQ

What you'll do

  • Act as the primary NVIDIA AI Enterprise and vector database solutions expert for HyperPOD customer environments, guiding diagnosis, optimization, and solution design.
  • Own complex end‑to‑end triage across GPU, NVAIE services, vector DB, Kubernetes, Docker, high‑speed networking, and Infinia storage, distinguishing product defects from environmental and integration issues.
  • Diagnose and resolve performance bottlenecks in RAG and agentic AI workflows, from model selection and prompt/RAG configuration through to vector search, GPU utilization, and data access patterns.
  • Author and maintain support triage runbooks and checklists for HyperPOD covering NVAIE services, Milvus/vector DB, GPU stack, Docker, Kubernetes resources, and their interaction with Infinia and the network fabric.
  • Build hands‑on labs and PoCs that mirror customer RAG and agentic AI use cases on HyperPOD, validating supportability and capturing “known good” configurations and troubleshooting patterns.

What they require

  • 5+ years in Linux-based infrastructure roles (SRE, MLOps, platform engineering, or L2/L3 support) supporting production systems; 8+ years total technical experience preferred.
  • Strong hands-on experience with containers and Kubernetes (Docker/containerd, Helm, Operators; debugging pods, DaemonSets, CSI, CNI, and ingress/load balancers).
  • Demonstrated experience operating GPU-accelerated workloads in production, including NVIDIA GPUs, drivers, CUDA concepts, GPU utilization/perf triage, and NVIDIA GPU Operator.
  • Practical experience with AI storage and networking for HPC/AI clusters, including high-performance storage systems and RDMA-accelerated/high-speed networking.
  • Experience with one or more vector databases (Milvus, Qdrant, Pinecone, pgVector, OpenSearch/Elasticsearch vectors, etc.), including schema design, ingestion, and operations.

DDN

DDN is positioned as NVIDIA’s storage and data intelligence partner for AI factories and the NVIDIA AI Data Platform.

Data Storageddnet.org/
Salary not disclosed