Skip to main content
Skylo

Senior Network Reliability Engineer, Incident Management

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
$125k–$135k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior NRE/telecom operations engineer with 5-10+ years in 24x7 production incident management. Must be US-remote and strong in 5G/RAN/Core, observability tools, Kubernetes triage, ticketing, and on-call workflows.

Core skills

5GGrafana/Prometheus/LokiKubernetes

Required skills

RANCUSMeCPRIPTPSyncE5G CoreAMFSMFUPFCloudOSSkubectlJira/ServiceNowPagerDuty

Optional skills

OSS/BSSFCAPSSNMPEMS3GPP NASNG-APRRCIMSI

What you'll do

  • Serve as the central command point during network degradations, service disruptions, and subscriber-impacting events from first alert through full-service restoration.
  • Own end-to-end incident lifecycle management for Sev 1-4 incidents.
  • Initiate bridge calls, identify cause and impacted domain, page the correct on-call NRE, maintain bridge discipline, and drive incidents to restoration.
  • Prioritize incidents according to urgency and business impact using alarm signatures, subscriber impact data, and domain KPI telemetry from OSS systems.
  • Escalate to subject matter experts in Operations and Engineering when critical or time-sensitive resolution is required.
  • Provide technical context, structured problem statements, and documented timelines during escalations.
  • Engage and interface with vendor support teams for RAN, Core, and cloud infrastructure incidents.
  • Track vendor SLA response and escalate vendor delays to the domain NRE.
  • Support hypercare operations during major network launches, high-risk change windows, and special events.
  • Act as first responder for degradation during hypercare windows.
  • Ensure trouble tickets are created promptly in Jira/ServiceNow with complete technical details, troubleshooting steps, MOPs followed, and outcome documentation.
  • Produce a structured incident timeline artifact within two hours of closure.
  • Manage the open incident backlog by tracking ageing tickets, escalating stalled items, and ensuring documented resolution paths or justified deferrals.
  • Coordinate post-incident review scheduling.
  • Compile incident records and gather logs from in-house observability tools and other relevant sources.
  • Deliver structured problem statements to domain NREs owning root cause analysis.
  • Handle internal, external, and MNO partner incident escalations and follow-ups.
  • Interface with Market Operations, OEM contacts, and partner NOCs for joint incident resolution.
  • Ensure external-facing communications are approved before transmission.
  • Assure Skylo's operated network meets agreed availability KPIs and MNO partner SLA commitments.
  • Track availability metrics and flag degradation trends before they breach SLA thresholds.
  • Track recurring issues and feed continuous improvement inputs.
  • Document repeat-incident patterns and identify operational gaps such as missing runbooks, stale thresholds, or absent alerts.
  • Route operational findings to the appropriate domain NRE for action.
  • Contribute to weekly and monthly Network Performance Reports covering incident counts, MTTR trends, top issues, SLA compliance, and KPI deviation analysis.
  • Drive proactive measures for network issue detection and isolation.
  • Participate in the Service Assurance and Automation domain and provide operational input for closed-loop automation requirements.
  • Participate in the global follow-the-sun on-call rotation as the incident coordination tier.
  • Maintain shift handoff hygiene with written end-of-shift summaries.
  • Manage on-call paging workflow by acknowledging alerts within SLA target windows and escalating to L2 within defined thresholds.
  • Ensure no alert goes unacknowledged across shift boundaries.
  • Support planned maintenance and change windows.
  • Validate pre-change observability coverage.
  • Confirm rollback readiness with the NI team.
  • Execute rollback runbooks if a deployment causes service degradation.
  • Collaborate across RAN NRE, Core NRE, Cloud Infrastructure NRE, BOSS, Network Implementation, Product Engineering, and Market teams during incidents.
  • Maintain a single source of truth on the bridge during cross-functional incidents.
  • Engage appropriate stakeholders based on incident signature, subscriber impact, and domain ownership.
  • Avoid over-escalation and under-escalation through disciplined severity classification.
  • Interface with MNO partner NOC teams during shared-impact events.
  • Relay technical status updates to MNO partner NOC teams.
  • Manage partner communication cadence.
  • Escalate partner requests through the correct internal channel.
  • Surface repetitive manual incident steps to the service assurance automation team.
  • Document manual incident step, frequency, and toil cost as input to the automation backlog.
  • Maintain working knowledge of RAN architecture relevant to incident triage, including CUSM, eCPRI/CPRI, PTP, SyncE, Netconf, and NTN-specific RAN alarm patterns.
  • Understand 5G Core network function roles enough to classify alarms, identify subscriber impact, and escalate with technical context.
  • Operate virtualized NTN infrastructure observability tools.
  • Navigate Grafana dashboards and use custom in-house tooling.
  • Run kubectl commands to assess pod health.
  • Correlate OSS alarms with underlying infrastructure events.
  • Build domain depth progressively across the first 12 months.

What they require

  • 5-10+ years of experience in telecom/wireless operations, network operations, or NRE in a production 24x7 environment.
  • Demonstrated ability to independently manage incident bridge calls, open the war room, maintain bridge discipline, drive to resolution, and produce a structured incident record.
  • Strong understanding of telecom network environments and 5G functional components.
  • Sufficient knowledge of RAN, including CUSM, eCPRI, PTP/SyncE.
  • Sufficient knowledge of 5G Core, including AMF, SMF, UPF.
  • Sufficient knowledge of Cloud/OSS to triage intelligently and escalate with context.
  • Incident and outage management expertise.
  • Ability to prioritize by urgency and impact.
  • Ability to manage multiple simultaneous events.
  • Ability to operate effectively under high-pressure 24x7 conditions.
  • Hands-on experience with at least one observability platform: Grafana dashboards, Prometheus alerting, Loki log queries, or equivalent.
  • Ability to independently navigate to relevant signals during an active incident.
  • Kubernetes operational literacy sufficient to run kubectl get pods, describe a failing pod, read container logs, and identify health issues for platform-layer incident triage and escalation.
  • Structured written communication skills for clear incident timelines, executive stakeholder updates, and post-incident summaries under time pressure.
  • Ticketing system proficiency with Jira, ServiceNow, or equivalent for incident lifecycle management, escalation workflows, and backlog hygiene.
  • On-call tooling experience with PagerDuty or equivalent for alert acknowledgement, escalation policy management, and on-call scheduling.
  • Ability to span departments and build strong working relationships with RAN, Core, Cloud, Engineering, and external partner teams to drive joint incident resolution.
  • Preferred: Prior experience in a satellite, NTN, or space-to-ground connectivity operational environment.
  • Preferred: Experience with OSS/BSS platforms, including FCAPS alarm management, event correlation, SNMP trap handling, or EMS integration.
  • Preferred: Familiarity with 3GPP NAS/NG-AP signaling flows, RRC state machine, or IMSI lifecycle procedures.
  • Preferred: ITIL Foundation certification or demonstrated practical application of ITIL incident, problem, and change management processes.
  • Preferred: Scripting ability in Python or Bash sufficient to automate repetitive operational tasks, parse log output, or build quick diagnostic utilities.
  • Preferred: Experience supporting hypercare operations, major network launches, or high-risk change windows.
  • Preferred: Exposure to closed-loop automation frameworks or service assurance platforms used in network operations.

Benefits

  • Competitive compensation packages including a stock option-based equity program.
  • Comprehensive benefits including medical, dental, vision, and retirement plan.
  • Monthly allowances for wellness and education reimbursement.
  • Generous time-off policy.
  • Holidays.
  • Opportunity to temporarily work abroad.
  • Opportunity to operate the world's first commercial, live direct-to-device satellite network.
  • Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization.
  • Open, transparent, inclusive culture that blends Silicon Valley, Nordic, and South Asia characteristics.

Skylo has pioneered a standards-based approach to satellite connectivity, connecting smartphones and IoT devices directly to satellites through a commercial NTN vRAN platform that bridges terrestrial and satellite networks.

🇺🇸 United StatesTelecommunicationsStartupskylo.tech/

Details

Apply routeDom
$125k–$135k/yr