Skip to main content
ConsumerAffairs

AI Native, Tech Ops Engineer

RemoteUnited States only
Published
Role
DevOps
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Tech Ops Engineer with 5+ years in system administration/Linux, AWS, CI/CD, scripting and monitoring. Must be strong in Kubernetes/EKS, Terraform, AWS infrastructure, observability, logging, security remediation, and automation. United States role.

Core skills

AWSKubernetesTerraform

Required skills

LinuxCI/CDScriptingMonitoringEC2VPCEFSS3EKSArgoCDGitOpsPythonJavaScriptJenkins/ConcourseDatadogPrometheusOpenSearchVectorKafkaPostgreSQLAurora RDSRedis

What you'll do

  • Monitor and maintain the organization’s infrastructure, including servers, networks, storage systems, and applications.
  • Perform routine system checks and preventive maintenance to ensure optimal performance and uptime.
  • Respond to system alerts and incidents, diagnosing and resolving issues promptly to minimize downtime.
  • Provide technical support to resolve infrastructure-related issues, working closely with other technical teams.
  • Troubleshoot and resolve hardware, software, and network issues, escalating to higher-level support when necessary.
  • Maintain detailed documentation of issues, solutions, and processes to improve the team’s knowledge base.
  • Plan and execute system upgrades, patches, and configuration changes, ensuring minimal disruption to business operations.
  • Test and validate updates in development environments before deploying them to production.
  • Ensure that all systems comply with security standards and best practices.
  • Identify opportunities to automate routine tasks and processes, improving operational efficiency and reducing manual workload.
  • Implement scripts, automation tools, and AI skills to streamline system management and monitoring.
  • Continuously evaluate and optimize infrastructure performance, capacity, and resource utilization.
  • Support the development and execution of disaster recovery plans to ensure business continuity in case of system failures.
  • Manage backup and restore processes for critical systems and data, ensuring data integrity and availability.
  • Participate in regular disaster recovery testing and drills.
  • Plan and execute decommissioning of legacy infrastructure, including EC2 instances, VPCs, and load balancers, coordinating Terraform state cleanup and DNS cutover.
  • Work closely with development, network, and security teams to ensure alignment and effective communication on infrastructure projects.
  • Provide input on infrastructure design and architecture to support new projects and initiatives.
  • Communicate effectively with non-technical stakeholders, providing updates on system status and issues.
  • Design and maintain log ingestion pipelines (e.g., Vector → OpenSearch, Vector → Kafka), including index retention, document shape optimization, and failure recovery.
  • Triage and remediate security vulnerabilities (CVEs) across infrastructure components, including container base images, OS packages, and third-party services.

What they require

  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent work experience.
  • 5+ years of experience in system administration, or a similar role.
  • 5+ years of professional experience in Linux administration, managing AWS resources, developing CI/CD and server orchestration pipelines, scripting and monitoring.
  • Expert in cloud-based production systems at scale.
  • Expert in Amazon Web Services (EC2, VPC, EFS, S3, EKS etc.).
  • Expert in production experience running workloads in Kubernetes (EKS), including ArgoCD GitOps deployments.
  • Expert in Infrastructure as Code tools, primarily Terraform.
  • Experience working in a Python and JavaScript-centric codebase and familiarity with their related best-practices.
  • Experience creating CI/CD pipelines with Jenkins, Concourse or other CI/CD implementation.
  • Experience with monitoring tools, like Datadog or Prometheus.
  • Experience scripting for server side automation, auditing, and monitoring.
  • Experience maintaining logging, monitoring, and alerting capabilities using OpenSearch, Vector log pipelines, Prometheus, and Kafka.
  • Experience configuring and managing data sources like PostgreSQL/Aurora RDS, OpenSearch, Redis, and message streaming platforms like Kafka.
  • Preferred: A working knowledge of modern software practices and technologies such as Agile methodologies.
  • Preferred: Promoting and establishing development standard methodologies for AWS infrastructure-as-code.
  • Preferred: Experience with AWS Well-Architected principles.
  • Preferred: Experience with High Availability implementations.
  • Preferred: Experience around Security and Compliance.
  • Preferred: Exceptional analytical and problem-solving skills.
  • Intellectual curiosity, a willingness to learn new skills and the ability to contribute new ideas.
  • Detail-oriented with a focus on maintaining high standards of operational reliability.
  • Adaptable and flexible, able to manage multiple tasks and prioritize effectively in a fast-paced environment.
  • Strong communicator with the ability to collaborate across teams and provide clear, concise technical support.
  • Proactive and self-motivated, with a passion for continuous learning and improvement in technology operations.
  • Learns quickly and using whatever resources to solve new problems.
  • Obsessed with ensuring an exceptional customer experience- for both internal and external customers.
  • Stands up for decisions, takes responsibility for results, and shares both good and bad outcomes transparently.
  • Demonstrates a relentless focus on results with a commitment to deliver; Takes decisive action, and confidently changes course if unsuccessful.
  • Displays a growth mindset to continually improve; encourages everyone around them to be tenacious and never settle.
  • Constantly seeks feedback to improve; Focuses on solving issues through teamwork, and collaboration.
  • Acts with urgency; delivers top results in hours and days instead of weeks and months.
  • Relentless in their pursuit of success and possessing the willpower to embrace challenges as opportunities.

Benefits

  • At ConsumerAffairs, your voice matters.
  • We foster a collaborative environment where you’re encouraged to take initiative, experiment boldly, and grow professionally.
  • We're committed to work-life harmony, career development, and celebrating wins together.
  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Paid Time Off (Vacation, Sick & Public Holidays)
  • Family Leave (Maternity, Paternity)
  • Short Term & Long Term Disability
  • Training & Development

ConsumerAffairs is an AI-forward company.

Salary not disclosed