Skip to main content
Skylo

Staff Network Reliability Engineer, Cloud Operations

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Mid-size
$155k–$165k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level cloud infrastructure/SRE for hybrid Kubernetes (GKE + on-prem) production environments. Needs 8+ years SRE/infrastructure experience, deep Kubernetes, Postgres/Redis, Prometheus/VictoriaMetrics observability; US-remote role. Owns 24x7 platform health, SLOs, runbooks and incident RCA.

Core skills

KubernetesPrometheusPostgreSQL

Required skills

GKEPersistent Volume ClaimCSIVictoriaMetricsGrafanaOpenTelemetryRedisArgoCDHelmTerraformAnsiblePub/SubLinux internalscontainer runtime debuggingRBACnetwork policiesCephRook

Optional skills

KubeVirtHarvesterGoPythonBGPVXLANEVPNFinOps

What you'll do

  • Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, PVC availability, network policies, and multi-cluster federation.
  • Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).
  • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation.
  • Execute and own Cloud Infra runbooks for P2–P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation.
  • Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks reflecting cluster topology.
  • Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage, OpenTelemetry collector health, and alert routing via Pub/Sub.
  • Maintain database reliability: PostgreSQL streaming replication, backup/restore, failover testing, query performance monitoring; Redis cluster operations and persistence configuration.
  • Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, and query performance.
  • Serve as the L3 escalation authority for Cloud Infra incidents: diagnose at Kubernetes, storage, network, and database layers and deliver resolution or decision-grade root cause.
  • Lead Cloud Infra troubleshooting bridges and participate in global 24x7 on-call rotation as Cloud Infra domain escalation tier.
  • Define and maintain SLOs for Cloud Infra components; own error budget tracking and drive toil reduction and automation.
  • Own Cloud Infra RCA end-to-end and deliver initial RCA documentation within SLA windows.
  • Author, own, and maintain Cloud Infra runbooks and SOPs; validate and sign off on operational readiness for infrastructure changes.
  • Own operational oversight of GitOps tooling in production: ArgoCD sync health, Helm chart management, and drift detection.
  • Partner with security, NI, Platform Engineering, Core NRE and RAN NRE on infrastructure readiness, hardening, and operational context.

What they require

  • 8–10+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment with direct on-call ownership for Kubernetes-at-scale environments.
  • Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures.
  • Hybrid cloud operations experience: public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms).
  • Observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipelines.
  • Database reliability experience: PostgreSQL streaming replication, backup/restore, failover; Redis cluster operations.
  • GitOps in production: ArgoCD or Flux CD; Helm chart authorship; Terraform or Ansible for provisioning.
  • SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design.
  • Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.
  • Runbook authorship capability and strong written and verbal communication for RCA and MNO-facing summaries.

Benefits

  • Competitive compensation packages including a stock option-based equity program
  • Comprehensive benefits including medical, dental, vision, and retirement plan
  • Monthly allowances for wellness and education reimbursement
  • A generous time-off policy, holidays, and the opportunity to temporarily work abroad
  • Once-in-a-career opportunity to operate the world's first commercial, live direct-to-device satellite network
  • Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization
  • Open, transparent, inclusive culture blending Silicon Valley, Nordic, and South Asia characteristics

Skylo has pioneered a standards-based approach to satellite connectivity, connecting smartphones and IoT devices directly to satellites through a commercial NTN vRAN platform that bridges terrestrial and satellite networks.

🇺🇸 United StatesTelecommunicationsStartupskylo.tech/
$155k–$165k/yr