Skip to main content
SS&C

Site Reliability Engineer/L3 Support

RemoteNot specified
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Enterprise
$110k–$130k/yr
Check eligibility

The listing doesn't say where it hires from. Check the description or the employer's site before applying.

No BS summary

Site Reliability Engineer (SRE) / L3 Support Engineer needed for a FedRAMP High cloud platform. Requires 3-6 years of experience in SRE, Production Engineering, DevOps, or senior production support. Must be a U.S. Citizen eligible to work on FedRAMP High systems. Experience with Kubernetes, AWS, and scripting is essential.

Core skills

KubernetesAWSPython

Required skills

LinuxBashPowerShellGoPrometheusGrafanaCloudWatchDatadogSplunkOpenTelemetry

Optional skills

FedRAMP HighDoD IL5/IL6EKSRDSIAMRoute 53VPC networkingAWS Backup

Required languages

English

What you'll do

  • Monitor the health, availability, performance, and security of production services.
  • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing.
  • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers.
  • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support.
  • Participate in an on-call rotation for critical production incidents.
  • Lead incident response activities, including coordination, communication, and post-incident reviews.
  • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence.
  • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures.
  • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators.
  • Work closely with software engineering teams to improve service reliability, scalability, and resilience.
  • Identify opportunities to automate operational tasks and eliminate repetitive manual work.
  • Support production deployments, infrastructure changes, and maintenance activities.
  • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews.
  • Ensure operational activities comply with FedRAMP High security and compliance requirements.
  • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.

What they require

  • U.S. Citizenship (required).
  • 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support role.
  • Experience supporting mission-critical cloud-based production systems.
  • Strong understanding of Linux operating systems and networking fundamentals.
  • Experience troubleshooting distributed applications running in Kubernetes.
  • Experience with public cloud platforms, preferably AWS.
  • Experience with infrastructure as code and configuration management.
  • Strong scripting or programming skills (e.g. Python, Bash, PowerShell, Go, or similar).
  • Experience using monitoring and observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry.
  • Experience analysing application logs, metrics, and traces to diagnose production issues.
  • Understanding of incident management, problem management, and root cause analysis.
  • Strong analytical and troubleshooting skills.
  • Excellent written and verbal communication skills.
  • Preferred Qualifications Experience supporting systems operating under FedRAMP High, DoD IL5/IL6, or similar regulated environments.
  • Experience with Kubernetes in production.
  • Experience with AWS services including EKS, RDS, IAM, CloudWatch, Route 53, VPC networking, and AWS Backup.
  • Experience with CI/CD pipelines and deployment automation.
  • Knowledge of service mesh technologies such as Istio.
  • Familiarity with security best practices including IAM, least privilege, vulnerability management, and compliance monitoring.
  • Experience with PagerDuty, Jira Service Management, or similar incident management platforms.
  • AWS certification (Associate or Professional) is desirable.

Benefits

  • Hybrid Work Model & a Business Casual Dress Code, including jeans
  • 401k Matching Program
  • Professional Development Reimbursement
  • Flexible Personal/Vacation Time Off
  • Sick Leave
  • Paid Holidays
  • Medical, Dental, Vision
  • Employee Assistance Program
  • Parental Leave
  • Discounts on fitness clubs, travel and more!
  • Comprehensive total rewards package designed to support your wellbeing, growth, and future.
  • Medical, dental, and vision coverage
  • 401(k) plan with company match
  • Paid time off, holidays, and parental leave
  • Professional development reimbursement opportunity

an American multinational financial technology company

🇺🇸 United StatesFintechEnterprisessctech.com/

Details

Visa sponsorshipNo
$110k–$130k/yr