Skip to main content
Feeld

Senior Site Reliability Engineer (SRE)

RemoteBrazil only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to BR only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with a strong backend engineering background in Node.js and TypeScript. Needs hands-on AWS, Cloudflare and CloudWatch experience plus production incident response. Remote in Brazil, must be comfortable with out-of-hours on-call.

Core skills

Node.jsTypeScriptAWS

Required skills

CloudflareCloudWatch

Optional skills

SentryReact Native

What you'll do

  • Own observability for critical product and user journeys within your squad.
  • Define, build and maintain meaningful metrics, dashboards and alerts.
  • Define and maintain SLIs/SLOs for key services and product-level metrics.
  • Improve monitoring, logging, tracing and alerting across the squad's systems.
  • Act as the first responder for critical P0/P1 production incidents, including out-of-hours incidents.
  • Investigate production signals, identify potential root causes and begin mitigating issues independently.
  • Coordinate with other engineers when broader support or escalation is required.
  • Participate in incident triage, mitigation and postmortems.
  • Identify recurring reliability issues and drive improvements to infrastructure, tooling and incident-response processes.
  • Work closely with backend and product engineers in a distributed, autonomous squad.

What they require

  • Strong previous experience as a Backend / Software Engineer, with senior-level hands-on experience in Node.js and TypeScript.
  • Hands-on experience working in an SRE, Production Engineering or similar reliability-focused role.
  • Strong production experience with AWS.
  • Experience with Cloudflare and CloudWatch.
  • Experience with monitoring and observability across metrics, logging, tracing and alerting.
  • Practical experience responding to production incidents, including triage, mitigation and postmortems.
  • Ability to interpret monitoring signals and independently investigate and begin resolving production issues.
  • Understanding of both the application and infrastructure layers rather than infrastructure-only experience.
  • Strong communication skills and the ability to work autonomously within a distributed engineering team.
  • Comfortable participating in out-of-hours incident response as part of the team's coverage model.

Benefits

  • Health insurance.
  • Wellbeing budget.
  • Sport coverage.
  • Learning budget.
  • 18 business days of paid vacation days per year.
  • Paid sick leaves.
  • All public holidays are paid days off.

location-based social discovery service application for iOS and Android

DatingStartupfeeld.co
Salary not disclosed