Skip to main content
Skyflow

Senior Site Reliability Engineer/Cloud Platform Engineer

RemoteIndia only
Published
Role
DevOps
Experience
Senior
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to IN only. Set where you work from to check your eligibility.

No BS summary

Senior platform/SRE engineer with 8+ years experience, strong Go/Python skills, deep Kubernetes and Terraform expertise, and experience with AWS/GCP/Azure. Must be comfortable in B2B enterprise environments with compliance and security requirements. Lead and mentor junior engineers.

Core skills

GoKubernetesTerraform/Pulumi/OpenTofu

Required skills

PythonAWS/GCP/Azure

Optional skills

IstioEnvoyArgoCDFluxOPAAerospikePostgreSQLCassandra

What you'll do

  • Design and build automation that provisions and manages cloud infrastructure end-to-end — new environment onboarding, upgrades, scaling, and decommissioning — across multiple cloud providers (AWS, GCP) and multiple deployment models (multi-tenant and dedicated/BYOC).
  • Write production-grade Go and Python services and CLIs that turn infrastructure operations into self-service platform capabilities for other engineering teams, rather than one-off scripts or manual runbooks.
  • Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) that defines every environment, and drive migrations across the fleet (Kubernetes version upgrades, node pool migrations, service mesh changes) with minimal customer impact.
  • Operate and scale core platform services — Kubernetes clusters, service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads — with a focus on capacity planning, cost efficiency, and right-sizing.
  • Build and improve observability and alerting (metrics, logs, synthetic monitoring) so that failures are caught before customers notice, and drive the automation that turns repeat incidents into permanent fixes.

What they require

  • 8+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles, with real ownership of production cloud infrastructure.
  • Strong software engineering skills in Go and/or Python — you write tested, maintainable code and think of infrastructure automation as software, not scripting.
  • Deep hands-on experience with Kubernetes in production (workload scheduling, networking, autoscaling, upgrades) and Infrastructure-as-Code tooling such as Terraform, Pulumi, or OpenTofu.
  • Solid experience with at least one major public cloud (AWS or GCP or Azure); experience operating in all the three is a strong plus given our multi-cloud footprint.
  • Track record of building tools or platforms that other engineers use — internal CLIs, provisioning frameworks, self-service portals, or automation pipelines — not just maintaining existing infrastructure.

Benefits

  • Work from home expense
  • Excellent Health Insurance Options
  • Very generous PTO
  • Flexible Hours
  • Generous Equity

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive data, enable safe AI deployment, and unlock full value from their applications, data platforms, and AI systems.

🇺🇸 United StatesCybersecurityMid-sizeskyflowergame.com/
Salary not disclosed