Skip to main content
DDN

Staff Software Engineer

RemoteUnited States only
Published
Role
Backend
Experience
Staff
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff backend engineer with 8+ years of experience, deep Go proficiency, and production observability experience. Must know Prometheus, VictoriaMetrics, OpenTelemetry, Linux systems, and telemetry pipelines for clustered/distributed/cloud-native systems. Remote role in North Carolina.

Core skills

GoPrometheusOpenTelemetry

Required skills

VictoriaMetricsLinux

What you'll do

  • Execute core projects within the observability domain.
  • Ensure seamless integration between proprietary Go infrastructure and open-source tools.
  • Help design and refine how OpenTelemetry, Prometheus, and VictoriaMetrics handle massive metrics volume generated by the storage cluster.
  • Act as a key technical contributor to the Software Defined Storage control plane.
  • Write clean, performant Go code.
  • Provide rigorous code reviews.
  • Participate in the Scrum lifecycle from initial design and coding to automated testing, usability reviews, and release.
  • Document and standardize telemetry frameworks.
  • Make it easy for other engineering teams to instrument their components properly.
  • Contribute to the global team on-call rotation.
  • Use observability tools to provide high-level technical support for the distributed footprint.

What they require

  • 8+ years of backend development experience, with a target of matching senior engineering standards.
  • Deep proficiency in Go for building high-performance, low-overhead system components.
  • Hands-on experience implementing and operating telemetry pipelines for software-defined clustered, distributed, or cloud-native solutions.
  • Strong technical knowledge of Prometheus operators, alerting rules, and scraping mechanics.
  • Strong technical knowledge of the VictoriaMetrics stack for long-term, high-cardinality storage.
  • Practical experience with the OpenTelemetry ecosystem, including custom OTel collector configurations, instrumentation SDKs, and data processing.
  • Solid understanding of Linux networking, filesystems, and how clustered storage applications behave under heavy I/O workloads.
  • Proven ability to work effectively across geographically distributed teams.
  • Ability to drive technical clarity through code reviews and clear documentation.
  • Proven track record of developing or extending proprietary Go components to efficiently ingest, process, and forward massive streams of metrics, logs, and traces.
  • Experience managing CPU and memory footprint of monitoring agents so they do not compete with core storage data paths.
  • Dedication to writing robust unit and integration tests to ensure telemetry components remain stable during live cluster upgrades.
  • Proven ability to take ownership of complex technical initiatives and independently make progress in a fast-paced environment.
  • Ability to turn abstract cluster state data into logical, well-structured telemetry frameworks that bring visibility to complex system scenarios.

DDN

DDN is positioned as NVIDIA’s storage and data intelligence partner for AI factories and the NVIDIA AI Data Platform.

Data Storageddnet.org/

Details

Apply routeDom
Salary not disclosed