Skip to main content
Arista Networks

Senior Site Reliability Engineer - CloudVision

RemoteIreland only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Enterprise
Salary not disclosed
Check eligibility

Open to IE only. Set where you work from to check your eligibility.

No BS summary

Senior Site Reliability Engineer with 5+ years in infrastructure or systems roles, remote from Ireland. Needs Go, Python, or Bash plus Linux/UNIX, production systems at scale, infrastructure-as-code, server provisioning, and incident response.

Core skills

Linux/UNIXInfrastructure-as-codeObservability

Required skills

Go/Python/Bash

Optional skills

KubernetesDockerVirtualizationPrometheusGrafanaGitLabSpinnakerTerraform

What you'll do

  • Design, build, and deploy production systems focused on scalability, reliability, observability, performance, and security
  • Develop and maintain automation solutions to eliminate toil and improve operational efficiency across production environments
  • Monitor production systems, establish alerting strategies, and implement automated incident response mechanisms
  • Create and maintain incident response runbooks
  • Conduct postmortem analyses after incidents to identify root causes and prevent recurrence
  • Collaborate with software engineering teams to identify and resolve infrastructure bottlenecks
  • Design solutions that improve product deployment workflows
  • Manage and optimise monitoring infrastructure using industry-standard tools
  • Plan, communicate, and execute maintenance windows on production systems with minimal service disruption
  • Triage platform and infrastructure issues
  • Engage with third-party vendors and support teams as required
  • Deploy new systems and updates in a staged, risk-managed manner
  • Survey and adopt best practices in infrastructure and platform management
  • Study design and implementation details of open-source systems to improve troubleshooting and issue resolution
  • Communicate system status, planned maintenance, and infrastructure improvements to stakeholders
  • Design, build, and deploy production systems with a focus on scalability, reliability, observability, and performance, ensuring systems meet stringent security standards
  • Develop and maintain comprehensive automation solutions to eliminate toil and streamline operational efficiency across production environments
  • Proactively monitor production systems, establish intelligent alerting strategies, and implement automated incident response mechanisms to minimise downtime
  • Create and maintain detailed incident response runbooks; conduct thorough postmortem analyses following incidents to identify root causes and prevent recurrence
  • Collaborate with software engineering teams to identify and resolve infrastructural bottlenecks, designing innovative solutions that enhance product deployment workflows
  • Manage and optimise monitoring infrastructure using industry-standard tools, ensuring comprehensive visibility across all systems
  • Plan, communicate, and execute maintenance windows on production systems with minimal disruption to service availability
  • Triage platform and infrastructural issues with decisiveness and analytical rigour; engage with third-party vendors and support teams as required
  • Deploy new systems and updates in a staged, risk-managed manner, ensuring safe and incremental rollouts
  • Survey and adopt best practices in infrastructure and platform management to maintain secure, scalable, and fault-tolerant systems
  • Study the design and implementation details of open-source systems to enhance troubleshooting capabilities and accelerate issue resolution
  • Work transparently with stakeholders to communicate system status, planned maintenance, and infrastructure improvements

What they require

  • Bachelor's degree in Computer Science, Engineering, or equivalent professional experience
  • 5+ years in a related infrastructure or systems role
  • Ability to implement medium-complexity automation workflows
  • Strong knowledge of Linux or UNIX from both administration and debugging perspectives
  • Hands-on experience operating software systems, infrastructure, and complex applications at scale in production environments
  • Demonstrated expertise in infrastructure-as-code principles and practices
  • Strong problem-solving and software troubleshooting skills with a methodical, analytical approach
  • Experience with server provisioning, particularly from storage and networking perspectives
  • Proven ability to work collaboratively within cross-functional teams and communicate technical concepts clearly
  • Experience with incident response, postmortem analysis, and continuous improvement methodologies
  • Preferred: Understanding of distributed systems architecture and principles
  • Preferred: Experience with performance tuning and system optimisation
  • Preferred: Knowledge of security best practices in infrastructure and systems design
  • Preferred: On-call support experience and comfort with incident response responsibilities
  • Bachelor's degree in Computer Science, Engineering, or equivalent professional experience (5+ years in a related infrastructure or systems role)
  • Proficiency in one or more programming languages: Go, Python, or bash shell scripting , with the ability to implement medium-complexity automation workflows
  • Desirable Skills and Experience: Experience with container orchestration platforms, particularly Kubernetes
  • Desirable Skills and Experience: Hands-on experience with Docker and virtualisation technologies
  • Desirable Skills and Experience: Proficiency in managing monitoring stacks, including Prometheus and Grafana
  • Desirable Skills and Experience: Experience with CI/CD systems such as GitLab tools or Spinnaker
  • Desirable Skills and Experience: Knowledge of infrastructure-as-code frameworks, particularly Terraform
  • Desirable Skills and Experience: Experience managing databases such as PostgreSQL or equivalent relational database management systems
  • Desirable Skills and Experience: Experience with artifact repositories and Docker registries
  • Desirable Skills and Experience: Familiarity with cloud platforms (Google Cloud Platform, Amazon Web Services, or Microsoft Azure)
  • Desirable Skills and Experience: Understanding of distributed systems architecture and principles
  • Desirable Skills and Experience: Experience with performance tuning and system optimisation
  • Desirable Skills and Experience: Knowledge of security best practices in infrastructure and systems design
  • Desirable Skills and Experience: On-call support experience and comfort with incident response responsibilities

Benefits

  • Engineers have complete ownership of their projects
  • Flat and streamlined management structure
  • Software engineering is led by people who understand software engineering
  • Engineers have access to every part of the company
  • Opportunities to work across various domains
  • All R&D centers are considered equal in stature
  • Culture that values invention, quality, respect, and fun
  • Arista stands out as an engineering-centric company. Our leadership, including founders and engineering managers, are all engineers who understand sound software engineering principles and the importance of doing things right.
  • We hire globally into our diverse team.
  • At Arista, engineers have complete ownership of their projects.
  • Our management structure is flat and streamlined, and software engineering is led by those who understand it best.
  • We prioritize the development and utilization of test automation tools.
  • Our engineers have access to every part of the company, providing opportunities to work across various domains.

Arista Networks is an industry leader in data-driven, client-to-cloud networking for large data center, campus and routing environments, using cloud computing, artificial intelligence, and software-defined networking.

🇺🇸 United StatesNetworkingEnterprisearista.com/

What people say about this company

4.1/ 5

  • Employees appreciate the innovative and cutting-edge technology used at Arista.
  • The company culture is often described as collaborative and supportive.
  • Many employees highlight opportunities for professional growth and development.
Salary not disclosed