Skip to main content
Lightning AI

Senior Network Engineer

RemoteUnited States only
Published
Role
DevOps
Experience
Senior
Company size
Startup
$150k–$190k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior network engineer with 5+ years in large-scale data center networking, hands-on Cumulus Linux/NOS, and experience with spine-leaf L3 fabrics, BGP, EVPN, and VXLAN. Must be based in the continental U.S.; no visa sponsorship. AI/HPC or GPU-dense infrastructure experience is required.

Core skills

Cumulus LinuxBGPEVPN/VXLAN

Required skills

Cumulus NOSSONiCJunosVPCNFVDirect ConnectCloud ConnectEVPNVXLANPython/Ansible/TerraformNetwork observability toolingTelemetry pipelines

Optional skills

NVIDIA SpectrumNVIDIA QuantumNVIDIA BlueFieldRDMARoCEInfiniBandBare-metal provisioning systems

What you'll do

  • Design and deploy scalable spine/leaf network architectures for AI data centers
  • Engineer high-performance Ethernet fabrics supporting GPU clusters and AI workloads
  • Build and maintain EVPN/VXLAN, BGP, and high-speed routing environments
  • Optimize east-west traffic flows for AI training and inference operations
  • Support RoCE/RDMA networking and low-latency transport technologies
  • Support backbone, DCI, WAN, and edge connectivity solutions
  • Collaborate with compute, storage, AI platform, and operations teams to deliver integrated infrastructure solutions
  • Develop automation and Infrastructure-as-Code (IaC) solutions for network provisioning and operations
  • Troubleshoot complex network, performance, and congestion issues across distributed environments
  • Improve network observability, telemetry, and operational visibility

What they require

  • Hands-on Cumulus Linux expertise
  • Hands-on experience working with SONiC and Junos
  • Experience with cloud networking technologies including VPC’s, NFV, Direct Connect, Cloud Connect
  • You enjoy working with a small group of friendly, highly motivated, high-execution colleagues
  • You’re comfortable with a high degree of autonomy, can independently prioritize your work and understand how it maps to the overall needs and goals of the company
  • You’re knowledgeable in your domain but also enjoy wearing multiple hats and venturing outside of your comfort zone when the need arises
  • You value the ability to write well and understand the importance of good documentation
  • Experience with Cumulus NOS
  • 5+ years of experience in large-scale data center networking
  • Experience in spine-leaf architectures and L3 fabrics
  • Experience with BGP, EVPN, VXLAN
  • Experience operating high-performance computing (HPC) or GPU-dense environments
  • Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure
  • Experience in automation with (Python, Ansible, Terraform, or similar)
  • Experience with network observability tooling and telemetry pipelines
  • Proven ability to design systems that scale to thousands of nodes
  • Strong documentation and communication skills
  • Preferred: Familiarity with NVIDIA networking (Spectrum, Quantum, BlueField, etc.)
  • Preferred: Familiarity with RDMA, RoCE, or InfiniBand fabrics
  • Preferred: Experience with multi-region backbone design
  • Preferred: Exposure to bare-metal provisioning systems
  • Preferred: Experience working in high-growth infrastructure startups
  • Fully remote out of the continental U.S. with occasional team and company offsites
  • We are not able to provide visa sponsorship for this role at this time

Benefits

  • Discretionary bonus
  • Meaningful equity component
  • Comprehensive medical, dental and vision coverage (U.S.)
  • Private medical and dental insurance (U.K.)
  • Retirement and financial wellness support (U.S.)
  • Pension contribution (U.K.)
  • Generous paid time off, plus holidays
  • Paid parental leave
  • Professional development support
  • Wellness and work-from-home stipends
  • Flexible work environment
  • Medical, dental, and vision coverage for employees and eligible dependents
  • RSUs that give employees a stake in the company's long-term success
  • 401(k) matching (U.S.) and pension contributions (U.K.)
  • Unlimited PTO, company holidays, and floating holidays to support work-life balance
  • Two weeks of company closure each winter to disconnect and recharge
  • Paid leave to support you and your family through life's important moments
  • Annual learning and development allowance to support your professional growth
  • Wellness and work-from-home stipends to support your physical and mental well-being
  • Four weeks of paid sabbatical leave after four years of service
  • Flexible schedules and a hybrid work model for our office-based teams
  • Complimentary meals at our office hubs

Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction. Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute.

AI InfrastructureStartup

Details

Visa sponsorshipNo
$150k–$190k/yr