Skip to main content
SentinelOne

Staff AI Platform Engineer, Infrastructure Services

RemoteUnited States only
Published
Role
DevOps
Experience
Staff
Employment
Full-time
$156k–$215k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level platform/infrastructure engineer with 8+ years in platform, infrastructure, or DevOps. Must own AI Gateway infrastructure on Kong, Kubernetes/GitOps, CI/CD, AWS/EKS, Terraform, LLM inference stacks, and production reliability. U.S.-based remote role.

Core skills

KubernetesKong AI GatewayLLM inference

Required skills

Kong/Envoy/ApigeeGitOpsArgoCDJenkinsGitHub ActionsArtifactoryXrayGitHub EnterpriseTerraformAWSEKSvLLM/NVIDIA Triton/NVIDIA NIM/TGI/Ollama

Optional skills

AI/LLM gateway patternsClaude CodeCopilotOktaOIDCLinearBQodoVector databases

What you'll do

  • Work on the AI Gateway platform: architect, harden, and scale our Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS), including auth (Okta/OIDC), consumer tiers and budgets, rate limiting, semantic caching, and observability.
  • Lead reliability and incident response: drive root-cause analysis and remediation for gateway issues (timeouts, latency, capacity, failover) and build the monitoring/alerting needed to catch them before users do.
  • Design across the platform, not just the gateway: work fluently with our CI/CD (Jenkins, JPAAS), GitOps and Kubernetes deployment tooling (ArgoCD across dev/gov/prod), artifact management (Artifactory/Xray), GitHub Enterprise administration, and GitHub Actions runner fleet, so that AI infrastructure decisions account for how the rest of the platform actually works.
  • Evaluate and roll out AI developer tooling: run structured pilots and adoption efforts for tools like AI-assisted PR review (Qodo) and engineering metrics platforms (LinearB), and make clear build-vs-buy recommendations.
  • Set technical direction and mentor: define architecture and standards for AI infrastructure, review designs across the team, and raise the bar for other engineers working in this space.
  • Partner cross-functionally: work directly with security, DevEx, and product engineering teams consuming the gateway to translate their needs into platform capabilities.
  • Host and serve local models: stand up and operate self-hosted/open-weight model serving infrastructure (e.g. vLLM, NVIDIA Triton/NIM, TGI, Ollama) for workloads where routing to an external provider isn't the right fit, including GPU capacity planning, autoscaling, and cost/performance tuning.
  • Support the broader model lifecycle: help build LLMOps practices such as model versioning, evaluation, and safe rollout, plus supporting infrastructure for retrieval-augmented generation (vector stores, embedding pipelines) as use cases mature.
  • Track usage and cost: build observability into token usage, latency, and spend across both API-based and self-hosted models so the business can see what AI infrastructure actually costs.

What they require

  • 8 or more years of experience in platform, infrastructure, or DevOps engineering, with a track record of owning systems end-to-end in production.
  • Hands-on experience with API gateway technologies (Kong, Envoy, Apigee, or similar); direct experience with AI/LLM gateway patterns (rate limiting, semantic caching, prompt/response observability) is a strong plus.
  • Strong Kubernetes and GitOps experience (ArgoCD or comparable), and comfort operating across multiple environments (dev, gov, prod).
  • Solid CI/CD background: Jenkins pipeline design and administration, build infrastructure, and runner/agent fleet management (GitHub Actions runners or equivalent).
  • Experience with artifact and package management systems (Artifactory, Xray, or similar) and source control platform administration (GitHub Enterprise).
  • Working knowledge of infrastructure-as-code (Terraform) and cloud platforms (AWS/EKS).
  • Experience deploying and operating self-hosted LLM inference stacks (vLLM, NVIDIA Triton/NIM, TGI, Ollama, or similar) and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
  • Familiarity with LLMOps practices: model versioning, evaluation harnesses, and usage/cost observability across API-based and self-hosted models.
  • Track record of setting technical direction, driving cross-team initiatives, and mentoring other engineers; this role has significant scope and minimal day-to-day oversight.
  • Clear, proactive communicator who can explain infrastructure trade-offs to both engineers and non-technical stakeholders.
  • Preferred: Experience operating LLM/AI-assisted developer tooling at scale (Claude Code, Copilot, or similar) inside an enterprise.
  • Preferred: Familiarity with Okta/OIDC and enterprise auth patterns for internal platforms.
  • Preferred: Experience with engineering productivity metrics tooling (LinearB or similar) and AI-based code review tooling (Qodo or similar).
  • Preferred: Experience with vector databases and RAG pipelines (e.g. Milvus, Pinecone, pgvector, or similar) in a production setting.
  • Preferred: Exposure to model fine-tuning or lightweight training pipelines (LoRA/QLoRA or similar) for domain-specific model adaptation.

Benefits

  • Restricted Stock Units (RSUs)
  • Employee Stock Purchase Plan (ESPP)
  • Flexible time off
  • Paid company holidays and paid sick time
  • Gender-neutral parental leave
  • Grandparent leave
  • Medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Life and disability insurance
  • Health and dependent care FSA
  • Voluntary benefits (hospital, accident, critical illness)
  • Employee Assistance Program (EAP)
  • ARAG pre-paid legal
  • Nationwide pet insurance
  • Cancer Care program
  • Global business travel medical insurance
  • Home office allowance
  • Mobile phone reimbursement
  • Wellness coach
  • Wellness/gym reimbursement
  • Fertility coverage
  • Adoption & surrogacy reimbursement

SentinelOne is a company at the intersection of AI and security, pioneering a new operating model for cybersecurity. Its AI-native platform unifies protection across endpoint, cloud, identity, data, and AI systems to deliver autonomous detection and response.

CybersecurityEnterprisesentinelone.com/

What people say about this company

3.3/ 5

  • Employees appreciate the innovative technology and products offered by SentinelOne.
  • The company has a collaborative work culture that fosters teamwork.
  • Some employees report issues with management and communication.

Details

Apply routeGreenhouse
$156k–$215k/yr