Site Reliability Engineer (SRE)
- Role
- SRE
- Experience
- Mid
- Employment
- Contract
Open to AR, CL, CO only. Set where you work from to check your eligibility.
No BS summary
Our engineering team is looking to add a skilled and proactive Site Reliability Engineer (SRE). This position serves as the connective link between software development and systems operations. Software engineering principles will be applied to automate operations, scale infrastructure, and keep systems highly available, resilient, and performant.
Core skills
Required skills
Required languages
Our engineering team is looking to add a skilled and proactive Site Reliability Engineer (SRE) . This position serves as the connective link between software development and systems operations. Software engineering principles will be applied to automate operations, scale infrastructure, and keep systems highly available, resilient, and performant. The core mission involves building, running, and safeguarding the production environments that power our applications, keeping downtime to a minimum while enabling fast, safe software deployment. Responsibilities Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormation Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs) Respond to production incidents and lead troubleshooting efforts to restore services Conduct blameless post-mortems to identify root causes and prevent recurrence Partner with software developers to optimize system performance and plan capacity Ensure services can scale to handle growth and traffic spikes Requirements 2+ years of experience in systems administration, DevOps, or systems-focused software development Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust Experience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and Kubernetes Understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS Passion for automation, eliminating toil, and building resilient systems that fail gracefully Advanced proficiency in English (C1+)
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
What you'll do
- Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormation
- Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
- Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog
- Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Respond to production incidents and lead troubleshooting efforts to restore services
What they require
- 2+ years of experience in systems administration, DevOps, or systems-focused software development
- Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
- Experience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and Kubernetes
- Understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS
- Advanced proficiency in English (C1+)
Benefits
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
Our Workforce Experience team helps clients realize digital and IT transformation by identifying and bridging gaps in employee knowledge, behaviors, and mindset. We design, develop, and implement a variety of solutions for online and in-person learning, including videos, articles, podcasts, interactive eLearning, coaching, experiential learning, instructor-led training, and assessments, using the latest insights from the learning sciences. In other words, we go beyond "training" to create learning experiences that really work.