Skip to main content
Arctic Wolf

Senior Reliability Developer

RemoteUnited States only
Published
Role
DevOps
Experience
Senior
Employment
Full-time
Company size
Enterprise
$145k–$180k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior reliability/devops engineer for cloud infrastructure and monitoring. Needs strong Terraform, AWS, Kubernetes, Prometheus/Grafana, and scripting (Python/Bash/JS). US-based candidates only (telecommute permissible from anywhere in the US).

Core skills

TerraformKubernetesPrometheus

Required skills

PuppetChefDockerECSEKSPythonBashJavaScriptGroovyPowerShellGrafanaZabbixAlertManagerPagerDutyAWS (EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, CloudWatch)OpenStackJenkinsGitLabGitHubBitbucketKeeperEBS

Optional skills

GaiaCylance AWS Amazon cloud platformLucidChartBackstageConfluenceJira

What you'll do

  • Responsible for changes and updates to public/private cloud infrastructure.
  • Owns all monitors that support AWN business services and our customers.
  • Implement effective automation using programming languages such as Bash, Python, and JavaScript.
  • Ensure full infrastructure stack is resilient with day-to-day care and feeding.
  • Write and manage Terraform modules and configuration trees for deployments into AWS and OpenStack.
  • Manage automated patching solutions for instances deployed across multiple regions to ensure compliance with strict security requirements.
  • Monitor internal and external TLS certificates for expiration, renewals, and re-deployment.
  • Create and improve service monitoring solutions using Prometheus, Grafana, Zabbix, and various Prometheus exporters.
  • Migrate monitoring for legacy production services to a new monitoring system and recreate existing monitoring within Prometheus + Grafana.
  • Collaborate frequently with development teams to implement new features, fixes, or troubleshoot services in lab/production environments.
  • Troubleshoot problems with microservices running within containers and administer containers in Kubernetes clusters.
  • Manage cloud components deployed in multiple AWS regions and maintain CI/CD pipelines.
  • Manage Kubernetes cluster upgrades, module upgrades, logging configurations, and resizing of EBS volumes and instances.
  • Troubleshoot issues between interconnected services and resolve API/service unavailability.
  • Create, improve, and maintain documentation for MOPs, service architecture, runbooks, and monitoring configurations.
  • Plan and execute service decommissioning in lab and production environments.
  • Participate in on-call rotation to support business-critical services and conduct root cause analysis during incident reviews.

What they require

  • Bachelor’s degree or foreign degree equivalent in Computer Information Systems, or related field and five (5) years of progressive, postbaccalaureate experience in a Technology related role or job offered or related role.
  • Experience utilizing Infrastructure-as-Code (IaC) tools including Terraform and configuration management tools including Puppet and Chef to design, implement, and manage multi-region cloud infrastructure deployments across AWS and OpenStack environments.
  • Experience utilizing containerization and orchestration technologies including Docker, Kubernetes, ECS, and EKS to deploy, manage, and troubleshoot microservices architectures in production environments.
  • Experience utilizing Python, Bash, JavaScript, Groovy, and PowerShell scripting languages for automation of infrastructure provisioning, certificate lifecycle management, and automated patching solutions.
  • Experience utilizing monitoring and observability tools including Prometheus, Grafana, Zabbix, AlertManager, and PagerDuty to create dashboards, configure alerts, and implement service health monitoring solutions.
  • Experience utilizing AWS cloud services including EC2, ECS, EKS, ELB, S3, RDS, IAM, Lambda, CloudFormation, VPC, Route53, and CloudWatch.
  • Experience utilizing CI/CD pipeline tools including Jenkins, GitLab, GitHub, Bitbucket, and Gaia.
  • Experience managing Kubernetes cluster operations including upgrades, module upgrades, logging configurations, node scaling, EBS volume management, and pod orchestration across multiple AWS regions.
  • Experience utilizing certificate management and security tools including Keeper for secret management, TLS certificate monitoring, renewal coordination, and automated deployment.
  • Experience troubleshooting distributed systems and microservices communication issues including API failures, container networking, and deployment failures.
  • Experience creating and maintaining technical documentation using Atlassian tools including Jira, Confluence, Backstage, and LucidChart.
  • Background checks are required for this position.
  • May require authorization to receive software or technology controlled under U.S. export control laws and regulations (EAR).

Benefits

  • Equity for all employees
  • Flexible time off and paid volunteer days
  • RRSP and 401k match
  • Training and career development programs
  • Comprehensive private benefits plan including medical, mental health, dental, disability, life and AD&D, and value-added services
  • Robust Employee Assistance Program (EAP) with mental health services
  • Fertility support and paid parental leave
  • Variable incentive compensation and new hire equity grants

At Arctic Wolf, we're not just navigating the cybersecurity landscape - we're redefining it. Our global team of dedicated Pack members is driving innovation and setting new industry standards every day.

🇺🇸 United StatesCybersecurityEnterprisearcticwolf.com
$145k–$180k/yr