Skip to main content
Software Mind

Senior Site Reliability Engineer (SRE) – Kubernetes

RemotePoland only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to PL only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 5+ years of experience, specializing in Kubernetes production operations, observability, and troubleshooting Node.js and JVM services. Must have strong Linux and networking fundamentals, and experience with CI/CD tools like Helm and GitOps.

Core skills

KubernetesSplunkPrometheus

Required skills

GrafanaHelmGitOpsArgoCD/FluxLinuxDNSTCPHTTPHTTP/2Node.jsJVMJavamTLSJWT

Optional skills

Web ComponentsLitKEDA

Required languages

English Very good spoken and written

What you'll do

  • Support the deployment, operation, and reliability of production services running on Kubernetes.
  • Monitor service health and investigate production incidents across distributed applications.
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.
  • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring.
  • Work within a client-directed backlog and established priorities.

What they require

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering , or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.
  • 3+ years of hands-on production Kubernetes experience strongly preferred.
  • Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment.
  • Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication
  • Very good spoken and written English.
  • Preferred: Web Components / Lit experience, to perform first-level debugging of UI-related issues
  • Preferred: Server-side rendering or isomorphic runtime experience
  • Preferred: Canary rollout / multi-version production operations
  • Preferred: Distributed tracing and request-context correlation
  • Preferred: KEDA or event-driven autoscaling
  • Preferred: Experience with enterprise platform integration layers

Benefits

  • Flexible employment and remote work
  • International projects with leading global clients
  • International business trips
  • Non-corporate atmosphere
  • Language classes
  • Internal & external training
  • Private healthcare and insurance
  • Multisport card
  • Well-being initiatives

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects!

🇺🇸 United StatesSoftware Services
Salary not disclosed