Skip to main content
Yuno

Staff Incident Manager

RemoteUnited States only
Published
Role
SRE
Experience
Staff
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level incident manager with 7+ years of experience running major production incidents. Must know distributed systems, cloud-native operations, observability/incident tooling, on-call programs, and advanced English. Payments/PCI-DSS and Spanish are preferred.

Core skills

DatadogIncident ManagementObservability

Required skills

PagerDuty/OpsGenie

Optional skills

PCI-DSSSRESLO

Required languages

English Advanced proficiency required.

Optional languages

Spanish Business-level proficiency preferred.

What you'll do

  • Act as incident commander on major and critical incidents, owning coordination, decision-making cadence, and escalation from detection through resolution.
  • Treat merchant transaction impact, PSP and acquirer dependency failures, settlement and reconciliation knock-on effects, and PCI-DSS scope as first-class concerns in every response.
  • Drive down time to detect, time to engage, and time to recover.
  • Own MTTR as a headline metric and the operational discipline behind Yuno's 99.99% uptime target.
  • Run rotations, escalation policies, paging hygiene, and alert quality.
  • Reduce noise so on-call engineers trust their pages.
  • Own internal stakeholder alignment during an event and drive clear, accurate merchant-facing updates, including the status page, in coordination with Support and account teams.
  • Run blameless postmortems, hold the room to a no-blame standard, and make sure action items are concrete, owned, and tracked to closure.
  • Translate recurring incident patterns into reliability work, partnering with engineering teams to turn postmortem findings into roadmap items, not orphaned tickets.
  • Define and maintain incident severity levels, response runbooks, and the operating model for declaring and managing incidents.
  • Report on incident trends, reliability posture, and SLA and SLO performance to engineering leadership.

What they require

  • 7+ Years of Experience.
  • Individual Contributor.
  • Proven experience running major incident response in a production environment that real customers depend on, ideally as an incident commander or in a dedicated incident management function.
  • Strong working knowledge of modern distributed systems and cloud-native operations — comfortable holding your own on a bridge call with senior engineers during an outage, following the technical thread, and keeping the response moving without needing every detail spoon-fed.
  • Experience defining or maturing an incident management practice from the ground up: severity frameworks, operating models, runbooks, and on-call programs built to last, not just inherited.
  • Fluency with observability and incident tooling — metrics, logging, tracing, alerting, and paging platforms.
  • Familiarity with a paging platform (PagerDuty, OpsGenie, or similar) and a status page tool is expected.
  • A track record of running postmortems that change behavior, and of closing the loop between incidents and engineering work.
  • Excellent written and verbal communication — able to write a clear merchant-facing status update and a precise internal escalation under pressure.
  • Calm, decisive judgment during high-pressure events — you hold structure when others are reacting.
  • Comfort being on-call as a regular part of the role, including for critical incidents outside business hours.
  • English — advanced proficiency required.
  • Preferred: Familiarity with PCI-DSS and the operational obligations that come with handling payment flows at scale.
  • Preferred: Prior experience in payments, fintech, or another high-availability, regulated domain.
  • Preferred: Exposure to SRE practices, error budgets, and SLO-driven prioritization.
  • Preferred: Experience managing incidents across multi-timezone, remote-first engineering organizations.

Benefits

  • Competitive Compensation.
  • Remote Work – You can work from everywhere!
  • Home Office Bonus – A one-time allowance to help you create your ideal home office.
  • Work Equipment.
  • Stock Options.
  • Health Plan wherever you are.
  • Flexible Days Off.
  • Language, Professional, and Personal Growth courses.

Yuno is the AI-native operating system of global commerce, powering the financial infrastructure of enterprise merchants, banks, and wallets. With a single API, Yuno connects them to pay-ins, payouts, fraud prevention, KYC/KYB, and stablecoins globally, so they can operate everywhere with a single integration.

Fintech
Salary not disclosed