Skip to main content
Honeycomb

Staff Field Reliability Engineer

RemoteUnited States only
Published
Role
SRE
Experience
Staff
Employment
Contract
Company size
Mid-size
$200k–$240k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level reliability engineer, 9+ yrs in SRE/infra/DevOps. Deep Kubernetes (EKS) and AWS expertise, observability and OpenTelemetry leadership, senior incident command. Remote in the US; no visa sponsorship.

Core skills

KubernetesAWSOpenTelemetry

Required skills

TerraformHelmChefAnsibleGo/Python/Java/TypeScript/Node.js/.NET

Optional skills

OTel CollectorsRefinerytail-based sampling

What you'll do

  • Define architecture and operational standards for Refinery as a Service (RaaS) and Honeycomb Private Cloud (HnyPC) across multiple AWS accounts and regions
  • Architect Terraform modules, Helm charts, and deployment automation that other FREs build on and extend
  • Set technical direction for how Honeycomb instruments, monitors, and operates its own managed infrastructure
  • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments
  • Serve as final technical escalation point for the most novel, highest-stakes customer situations
  • Resolve deep infrastructure and observability issues spanning distributed systems, Kubernetes clusters, AWS networking, and service meshes
  • Provide senior incident command for managed services (RaaS, HnyPC)
  • Build playbooks, diagnostic frameworks, and tooling that let IC3/IC4 engineers operate independently
  • Shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and represent Honeycomb in OTel SIGs
  • Build reference architectures and integration guides for common customer environments (Kubernetes, ECS, serverless)
  • Lead feature and architecture contributions to Honeycomb's open source projects (Refinery, Honeycomb Collector Distro)
  • Be the final technical authority Solutions Architects call when a deal goes deeper than any existing playbook
  • Lead architecture reviews, SLO workshops, and instrumentation deep-dives for complex customer environments
  • Step into high-stakes customer-facing POCs and pilots as technical lead
  • Build internal tools and UIs that change how the FRE function operates
  • Drive alignment across Solutions Architecture, Customer Success, Support, Product, and Engineering
  • Mentor IC3/IC4 engineers on career development
  • Communicate at the executive level with Honeycomb leadership and customer C-suite

What they require

  • 9+ years in engineering, SRE, infrastructure, DevOps, or equivalent, with Staff-level technical scope and impact
  • Deep hands-on experience with Kubernetes (EKS strongly preferred) — deployed, scaled, and operated production clusters at scale
  • Strong AWS expertise across core services (EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, Route53), multi-account architecture, service quotas, and cost optimization
  • Track record of senior incident command — owning incident response, triage, and postmortem process improvements at the function level
  • Infrastructure as Code mastery (Terraform, Helm, Chef, Ansible) — architected modules and standards others build on
  • Deep observability expertise: structured logging, distributed tracing, metrics, SLOs/SLIs, and the full instrumentation lifecycle
  • Strong command of OpenTelemetry (SDK, Collector architecture, processors, exporters, semantic conventions) with public community contribution and leadership
  • Proficiency in at least two of: Go, Python, Java, TypeScript/Node.js, .NET
  • Excellent executive communication skills — credible with a customer's staff SRE and their VP or C-suite
  • Demonstrated ability to set direction in ambiguous, high-pressure situations with no precedent
  • Background leading customer-facing engineering functions — solutions architecture, field engineering, or technical consulting
  • Track record of building platforms and organizational systems that let a team scale without proportional headcount
  • Public recognition or leadership role in the CNCF/OpenTelemetry ecosystem — maintainer status, SIG leadership, or conference speaking
  • Familiarity with Honeycomb or event-based observability approaches
  • Experience operating telemetry pipelines at scale (OTel Collectors, Refinery/sampling, tail-based sampling, pipeline reliability)
  • Experience with managed SaaS deployments, private cloud offerings, or multi-tenant infrastructure operations
  • Prior experience at an observability, monitoring, or developer tools vendor
  • Experience mentoring senior or staff-track engineers on scope and career progression

Benefits

  • Generous equity with employee-friendly stock program
  • Transparent pay based on levels relative to experience
  • Unlimited PTO
  • Distributed-first mindset and culture
  • Home office, co-working, and internet stipend
  • Full benefits coverage for employees, with additional coverage available for dependents
  • Up to 16 weeks of paid parental leave, regardless of path to parenthood
  • Annual development allowance

Honeycomb is a service defining observability and raising expectations of what developer tools can do. It works with companies including HelloFresh, Slack, LaunchDarkly, and Vanguard.

Software DevelopmentMid-size

Details

Visa sponsorshipNo
$200k–$240k/yr