Перейти к основному содержимому
Diné Development Corporation (DDC)

Site Reliability Engineer

УдалённоUnited States толькоВ архиве
Опубликовано
Роль
SRE
Опыт
Синьор
Занятость
Полная занятость
Зарплата не указана
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Senior SRE/SME needed to support the GEOMAP platform's reliability, scalability, and operational resilience in secure cloud environments for the U.S. Air Force. Requires 8+ years of experience with AWS, Linux, Kubernetes, and CI/CD in a DoD environment. Active Secret clearance is mandatory.

Ключевые навыки

AWSKubernetesSite Reliability Engineering

Обязательные навыки

Linux administrationscriptingCI/CD pipelinesrelease automationinfrastructure-as-codemonitoring toolslogging toolsalerting tools

Желательные навыки

AWS Cloud Onesecure federal cloud environmentsgeospatial platformsEsri-based platformsArcGIS EnterpriseSLIsSLOserror budgets

Чем предстоит заниматься

  • Provide senior-level engineering support to improve reliability, availability, performance, and maintainability of GEOMAP cloud-hosted systems and services.
  • Analyze production issues, recurring incidents, and operational trends to identify root causes and recommend durable corrective actions.
  • Support the design and implementation of monitoring, alerting, logging, and observability solutions across applications, infrastructure, and containerized services.
  • Develop and recommend automation approaches that reduce manual effort, improve deployment consistency, and increase system resilience.
  • Partner with software engineers, DevSecOps engineers, Kubernetes engineers, database engineers, and production support personnel to improve service health and release readiness.
  • Support incident response, problem management, service restoration, and post-incident reviews for high-priority operational issues.
  • Evaluate system performance, capacity, and scalability needs and provide recommendations for optimization and operational risk reduction.
  • Assist in defining service reliability objectives, operational metrics, and support models for sustained mission operations.
  • Contribute to infrastructure and platform engineering efforts involving cloud environments, CI/CD pipelines, container orchestration, and secure deployment patterns.
  • Support architecture reviews, technical assessments, and engineering analyses related to reliability, recoverability, and production operations.
  • Develop or refine runbooks, standard operating procedures, reliability engineering practices, and technical documentation.
  • Provide reach-back support for surge requirements, complex production investigations, and priority modernization or stabilization efforts as directed.
  • Performs other related duties as assigned.

Что требуется

  • Active Secret clearance required.
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; Master’s degree preferred.
  • Minimum of 8 years of experience supporting enterprise systems, cloud platforms, site reliability engineering, production engineering, systems engineering, or related technical roles.
  • Experience supporting AWS environments, including monitoring, performance tuning, troubleshooting, incident response, and operational sustainment.
  • Experience with Linux administration, scripting, and troubleshooting distributed applications in production environments.
  • Experience with containerized systems and orchestration platforms such as Kubernetes.
  • Experience supporting CI/CD pipelines, release automation, infrastructure-as-code, and operational reliability in Agile or DevSecOps environments.
  • Experience with monitoring, logging, and alerting tools used to support enterprise application performance and infrastructure visibility.
  • Strong analytical, troubleshooting, documentation, and communication skills, with the ability to translate operational issues into engineering improvements.
  • Ability to work effectively across cross-functional teams in a mission-focused DoD environment.
  • Preferred Experience supporting AWS Cloud One or other secure federal cloud environments.
  • Experience supporting geospatial or Esri-based platforms, including ArcGIS Enterprise or related technologies.
  • Familiarity with service reliability practices such as SLIs, SLOs, error budgets, incident postmortems, and capacity planning.
  • Experience with Risk Management Framework (RMF), STIG compliance, vulnerability remediation, and secure system hardening practices.
  • AWS, Kubernetes, or other relevant cloud or reliability engineering certifications.
  • Experience supporting technical refresh, platform modernization, or high-availability design initiatives in enterprise environments.

Преимущества

  • Eligible full-time employees receive a comprehensive benefits package, including medical, dental, vision, life and disability coverage, retirement savings with company match, paid time off, voluntary supplemental benefits, and access to an employee assistance program.
  • The package also includes educational assistance, with tuition reimbursement.

Diné Development Corporation (DDC) is a Navajo Nation–owned family of companies that provides government agencies and commercial organizations with high-quality IT, professional, environmental, and research and development services. DDC is dedicated to empowering the Navajo Nation and the communities we serve.

Government ITСредняя
Зарплата не указана