Перейти к основному содержимому
CloudLinux
CloudLinux

Lead Site Reliability Engineer - Imunify Reliability Platform

УдалённоPoland, Romania, Bulgaria +2 more только
Опубликовано
Роль
SRE
Опыт
Лид
Занятость
Полная занятость
Зарплата не указана
Проверьте доступность

Доступно для: PL, RO, BG, RS, AM only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Lead SRE with substantial production-engineering experience, including defining SLO frameworks. Must have strong Python and be comfortable reading/modifying Go or Rust. Requires deep practical grip on time-series and event telemetry at scale, distributed systems debugging on bare metal, and configuration management/CI at production scale. This is a remote-only role.

Ключевые навыки

PythonGoPrometheus

Обязательные навыки

RustOpenMetricsGrafanaAlertmanagerClickHouseAnsibleGitLab CIJenkins

Желательные навыки

OpenTelemetryeBPFSentryKubernetes

Чем предстоит заниматься

  • Define what "working" means for ~70 components
  • Run SLI definition with squad leads and senior engineers
  • Facilitate and hold the standard; the owning squad signs the SLI
  • Build the taxonomy this product actually needs, which is broader than availability and latency: Service SLIs, Fleet SLIs, Control-efficacy SLIs, Delivery SLIs, Pipeline SLIs
  • Enforce one non-negotiable design rule: an SLI must be measurable from outside the gate of the thing it measures
  • Attach an SLO, an error budget and an owning squad to each
  • Build the collection system
  • Design and build the pipeline that gets these indicators off the fleet and into a queryable store: push-based, sampled, privacy-constrained, and with a cardinality budget you set and defend
  • Extend agent-side and service-side instrumentation where the signal does not exist yet, in Python, Go and Rust, working with the owning squads
  • Consolidate the current sprawl of dashboards, ad-hoc queries and reporting paths into a defensible set of instruments, and retire what does not earn its keep
  • Build alerting and alert management
  • Symptom-based, SLO-anchored alerting with multi-window burn-rate semantics
  • A three-tier taxonomy — page / ticket / dashboard — with an explicit rule for what is allowed to page a human at 03:00
  • Every alert ships with an owner, a runbook and a documented failure mode, or it does not ship
  • Alert hygiene as a standing practice: quarterly review, deletion counted as a win, actionable-rate tracked
  • Build escalation
  • Component → owning squad ownership map, kept current, machine-readable, and wired into routing so an alert reaches the right seven people rather than a shared channel
  • Severity matrix, acknowledgement SLAs, follow-the-sun rota design across UTC−5 … UTC+8, and clean handoff protocol
  • Incident command practice and blameless postmortems within 24 hours
  • Design the escalation system so that squads carry their own pagers
  • Build and operate the platform and coach on the practice

Что требуется

  • Substantial production-engineering or SRE experience, including at least one environment where you defined the SLO framework rather than inherited it
  • We will ask you to walk through SLIs you personally wrote and how you negotiated them with resistant teams
  • Strong Python
  • Comfortable reading and modifying Go or Rust — our agents are written in them and instrumentation lands there
  • Deep practical grip on time-series and event telemetry at scale: Prometheus/OpenMetrics, Grafana, an Alertmanager-class routing layer, and a columnar store for high-cardinality fleet data (ClickHouse or equivalent)
  • Distributed systems debugging on bare metal and long-lived hosts
  • Configuration management and CI at production scale — Ansible, GitLab CI, Jenkins or close equivalents
  • The judgement to design measurement for machines you do not own and cannot scrape: push telemetry, sampling, clock skew, partial reporting, and the privacy constraints that come with running on a customer's server
  • Written communication that holds up async
  • Security product background — WAF, EDR, AV, vulnerability management — and the instinct that a security control's SLI is about enforcement, not uptime (Valuable)
  • Monitoring under audit: SOC 2 CC7.x, ISO 27001 A.8.16, NIST SP 800-137 continuous monitoring (Valuable)
  • Cost- and cardinality-aware telemetry design (Valuable)
  • Fluency with agentic development tooling — we run a Cursor/Claude-first SDLC with internal and third-party MCP servers, and engineers here are assessed on how well they work with it (Valuable)

Преимущества

  • A strong focus on professional development with opportunities for learning and growth: Interesting and challenging projects, Mentor and other knowledge-exchange programs
  • Fully remote work with flexible working hours, that allows you to schedule your day and work from any location worldwide
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves to ensure you maintain a healthy work-life balance
  • Compensation for private medical insurance
  • Co-working and gym/sports reimbursement
  • The opportunity to receive a reward for the most innovative idea that the company can patent, fostering a culture of creativity and innovation

CloudLinux provides tools and solutions for web hosting providers, agencies, developers, and site owners to improve Linux server stability, security, performance, and website reliability. Its portfolio includes CloudLinux, Imunify security products, KernelCare, WordPress optimization tools, and lifecycle support offerings for hosting infrastructure.

🇺🇸 Соединенные ШтатыInformation Technologycloudlinux.com
Зарплата не указана