Senior Site Reliability Engineer
- Role
- SRE
- Experience
- Senior
- Employment
- Full-time
Open to US only. Set where you work from to check your eligibility.
No BS summary
Senior SRE needed to build and scale critical infrastructure across global data centers, cloud, and on-premise systems. Focus on automation-first solutions, GitOps, self-healing infrastructure, and cluster autoscaling. Requires 4+ years of AWS experience, Kubernetes expertise, and strong Go/Python automation skills.
Core skills
Required skills
Optional skills
At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together. The Crown Is Yours As a Senior Site Reliability Engineer, you’ll build and scale the critical Kubernetes infrastructure that powers our platforms and services. You’ll solve complex reliability challenges across public cloud and on-premise environments, designing automation-first solutions that strengthen performance and simplify operations. You’ll help shape architectural decisions, advance stability at scale, and build tools that give our teams the confidence to move quickly and deliver reliably. What you’ll do as a Senior Site Reliability Engineer Drive stability, performance, and scalability across our global compute platform spanning multiple public clouds and on-premise environments. Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams. Operate and evolve our GitOps delivery model, using Rancher Fleet, Flux, and Helm to deploy core Kubernetes services and application workloads consistently and reliably. Own Kubernetes scaling and capacity strategies using technologies including Karpenter, Horizontal Pod Autoscaler (HPA), Kubernetes Event-Driven Autoscaling (KEDA), and predictive scaling based on event and calendar data. Define and monitor service-level objectives and reliability metrics for platform components using Datadog and our logging pipeline. Strengthen our engineering practices by sharing knowledge, contributing to architectural and design discussions, and participating in an on-call rotation. What you’ll bring A Bachelor’s Degree in Computer Science or a related field, or equivalent education, experience, and training. At least 4 years of experience managing distributed cloud and on-premise environments at scale, including strong hands-on experience with Amazon Web Services; experience with Google Cloud Platform, vSphere, or Nutanix is a plus. Deep expertise in Kubernetes and container orchestration, with experience designing, scaling, and troubleshooting complex workloads. Strong software development experience using languages such as Go and Python to build automation and infrastructure tooling. Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting. Experience with Infrastructure as Code and configuration management tools to build scalable, consistent, and repeatable infrastructure. Join Our Team We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment. Don’t worry, we’ll guide you through the process if this is relevant to your role. The US base salary range for this full-time position is 128,000.00 USD - 160,000.00 USD, plus bonus, equity, and benefits as applicable. Our ranges are determined by role, level, and location. The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific pay range and how that was determined during the hiring process. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability. DraftKings Inc. (Nasdaq: DKNG) is a digital sports entertainment and gaming company. It’s simple, at DraftKings, we believe life’s more fun with skin in the game. For that reason, we’re committed to responsibly creating the world’s favorite games and betting experiences. Headquartered in Boston, with offices around the globe, we believe we can continue to define what it means to be a technology company in sports entertainment. We love what we do, and think you will too.
What you'll do
- Drive stability and scalability across our global compute platform spanning numerous data centers, multiple public clouds, and on-premise environments, serving as the foundation for every product.
- Operate and evolve our GitOps delivery model, using Rancher Fleet and Flux with Helm to deploy core cluster services and application workloads declaratively and repeatably.
- Build self-healing, fault-tolerant infrastructure and internal tooling that eliminates repetitive operational work and reduces toil for both platform and application teams.
- Own cluster autoscaling and capacity strategy, including Karpenter, HPA and KEDA, and predictive scaling driven by event and calendar data.
- Define SLOs and reliability metrics for platform components, using Datadog and our logging pipeline to surface cluster and workload health.
- Support technical growth by sharing knowledge, participating in design discussions, and contributing to a collaborative team culture, including on-call rotation.
What they require
- Bachelor's degree in Computer Science or relevant education, experience, and training.
- At least 4 years managing distributed cloud and on-premise environments at scale, with strong hands-on AWS experience.
- Deep expertise in container orchestration with Kubernetes, including the ability to design, scale, and troubleshoot complex workloads.
- Strong experience developing software for automation and infrastructure tooling such as Go and Python.
- Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting.
- Experience with Infrastructure as Code (IaC) and configuration management tools to ensure scalable and repeatable infrastructure provisioning.
Benefits
- plus bonus, equity, and benefits as applicable
We’re a publicly traded (NASDAQ: DKNG) technology company headquartered in Boston. As a regulated gaming company, you may be required to obtain a gaming license issued by the appropriate state agency as a condition of employment.
What people say about this company
3.3/ 5
- Employees enjoy the dynamic and fast-paced work environment.
- Many appreciate the opportunities for career growth and development.
- Some reviews mention issues with management and communication.