Перейти к основному содержимому
LinkedIn

Principal Staff Software Engineer, Systems Infrastructure

УдалённоUnited States только
Опубликовано
Роль
SRE
Опыт
Принципал
Занятость
Полная занятость
Размер компании
Крупная
$226k–$369k/yr
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Principal/staff-level reliability engineering leader. Deep distributed-systems and SRE expertise (SLOs/SLIs, HA, failover), 10+ years of engineering with 5+ years in principal/architect/technical leadership. Must be able to operate at company-wide scale; role is Mountain View (CA) / US-eligible or remote/hybrid per LinkedIn policy.

Ключевые навыки

observabilitydistributed systems reliabilitySLO/SLI design

Обязательные навыки

SLO designSLI designmonitoringalertingcapacity planningdisaster recoverymulti-region failoverJavaGoC++Pythondistributed systemsresiliency engineering

Желательные навыки

LLM and agent evaluationAI observabilityautonomous remediationself-healing systems

Чем предстоит заниматься

  • Define and drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering
  • Lead adoption and evolution of service criticality models that set reliability expectations based on business impact and blast radius
  • Serve as a technical authority for architecture decisions related to reliability, resiliency, availability, and failure handling
  • Partner with infrastructure and product engineering teams to improve system design, reduce incident risk, and strengthen operational readiness
  • Identify high-risk systems and drive cross-organizational initiatives to improve reliability of critical services
  • Establish and evolve reliability standards including SLOs, SLIs, uptime expectations, redundancy, monitoring, alerting, and failover patterns
  • Influence engineering culture by promoting reliability-focused design, incident review rigor, and postmortem-driven improvements
  • Provide architectural guidance and mentorship to senior engineers and technical leaders across teams
  • Balance technical strategy, hands-on engineering judgment, and cross-functional influence to drive measurable improvements in site stability
  • Help shape how LinkedIn builds and operates resilient systems as the platform continues to scale
  • Drive the strategy for applying LLMs to alert triage, root cause analysis, and incident summarization at scale

Что требуется

  • BA/BS in CS or related field or equivalent practical experience
  • 10+ years of experience in software engineering, infrastructure engineering, distributed systems, SRE, production engineering, or reliability engineering
  • 5+ years of experience in a technical leadership, architect, or principal-level engineering role
  • Experience designing, building, or operating large-scale distributed systems
  • Experience defining or driving reliability standards such as SLOs, SLIs, uptime targets, incident reduction, or operational readiness frameworks
  • Understanding of high availability, redundancy, fault tolerance, failure modes, and resiliency patterns
  • Experience influencing architecture and engineering practices across multiple teams or organizations
  • Software engineering experience in one or more languages such as Java, Go, C++, Python

Преимущества

  • Annual performance bonus, stock, benefits and/or other applicable incentive compensation plans
  • LinkedIn benefits (see https://careers.linkedin.com/benefits)
  • Reasonable accommodations for applicants with disabilities

LinkedIn is the world’s largest professional network, built to create economic opportunity for every member of the global workforce.

🇺🇸 Соединенные ШтатыSocial NetworkКрупная
$226k–$369k/yr