Перейти к основному содержимому
Tech Holding

Lead Site Reliability Engineer (Performance & Scalability)

УдалённоUnited States только
Опубликовано
Роль
SRE
Опыт
Лид
Занятость
Contract
Зарплата не указана
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Hands-on Lead SRE for a contract role, remote in the US. Must have deep experience in performance engineering, distributed systems, observability, capacity planning, and SLOs. You'll define baselines, lead load testing, and drive scalability improvements. Must be authorized to work for any employer in the U.S.

Ключевые навыки

Site Reliability EngineeringPerformance EngineeringCapacity Planning

Обязательные навыки

Platform EngineeringDistributed SystemsObservabilityPerformance AnalysisReliability EngineeringCloud InfrastructureDatabasesNetworkingCachingQueueingComputeStorageSLOsSLIsError BudgetsLoad TestingStress TestingSoak TestingScalability TestingResilience TestingProfilingBottleneck DiagnosisArchitecture TuningGraceful DegradationDependency Failure PlanningIncident ManagementRoot Cause Analysis

Желательные навыки

High-scale SaaSIdentityDNSRegistryInfrastructureCapacity-Cost ModelsForecasting Infrastructure RequirementsCI/CD Performance Gates

Чем предстоит заниматься

  • Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows
  • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds
  • Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services
  • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams
  • Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments
  • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events
  • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements
  • Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates
  • Own technical readiness assessments for major pilots, partnerships, and production launches
  • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures
  • Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work
  • Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership
  • Recommend capacity and reliability investments before they become production constraints

Что требуется

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering
  • Strong hands-on experience with cloud infrastructure and production distributed systems
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics
  • Hands-on experience performing load, stress, soak, scalability, and resilience testing
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios
  • Strong incident management and root-cause analysis experience
  • Ability to translate technical performance and reliability risks into clear business implications for senior leadership
  • Strong judgment around when systems genuinely require optimization versus when additional complexity is premature

Преимущества

  • Remote work
  • Equal Opportunity Employer
  • Diverse and inclusive workplace
Consulting
Также опубликовано ещё в 1 канале
Зарплата не указана