Senior Technical Product Manager - Compute API & Experience
- Role
- Product
- Experience
- Senior
- Employment
- Full-time
- Company size
- Mid-size
Open to NL only. Set where you work from to check your eligibility.
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The role The AI Compute Platform runs the GPU infrastructure frontier AI labs train and serve their models on. The Compute API is how customers — and a dozen internal product teams — actually use it. As the TPM for Compute API & Experience, you own that API end-to-end and the surface customers touch — console, CLI / SDK, instance metadata, self-service and the reliability signals they trust us on. It's a gatekeeper seat with real authority: you set the contract other teams build on, and own the VM lifecycle and scheduling logic that decide how customer VMs are placed, run and recovered — on one of the largest GPU fleets in Europe. What you own:
Strategy & roadmap for the Compute API & Experience area — discovery, delivery, adoption. The Compute API as a product — contract, versioning, idempotency, error semantics, deprecation; predictable reconciliation / desired-state behaviour. VM / instance lifecycle as a first-class API — start / stop / restart / recovery, maintenance behaviour — the behaviour customers actually experience. Scheduling & placement - capacity-aware placement, NER handling, operation priority and fairness — and how that logic surfaces to customers. Integration and interaction with adjacent platform services — storage, networking, IAM and others The customer-facing surface end-to-end — console, CLI / SDK, instance metadata, self-service, notifications. New compute platforms into the product — instance types, presets and platform capabilities — with the hardware and platform teams. A multi-region product experience — region-aware placement and a project model that works across regions the way customers expect.
Responsibilities:
Own end-to-end product responsibility for your area — strategy, roadmap, discovery, delivery, adoption, and measurable customer & platform outcomes. Design and govern platform contracts at hyperscaler quality. Co-design the compute control plane with engineering as a technical peer (scheduling, allocation, reconciliation) — not just word an API. Manage stakeholders and drive cross-team execution across engineering, networking, storage, product and sales / CX. Define success metrics for the Compute API and the customer experience, and be the escalation point for product decisions on your surface.
We expect you to have:
6+ years in Product / Platform / Infrastructure PM — or an SRE / Engineering Lead moving to product — shipping technically complex platform products with measurable impact. Owned a public cloud / compute / platform API as a product . A declarative / desired-state or control-plane API (Kubernetes CRDs / operators, or a cloud control plane) is a strong plus. Hands-on cloud experience — the VM / instance lifecycle, its API and its console at a public or large-scale cloud (AWS / GCP / Azure or another large-scale cloud), with a real under-the-hood understanding of how scheduling, allocation and the virtualization layer work (control-plane vs data-plane — how VMs are placed, scheduled and run). GPU / AI-cloud experience is a big plus, but not required. A systems understanding of how a cloud fits together — how compute, storage, networking, IAM and related services connect and interact, and how it all works under the hood — enough to design the Compute API coherently and reason about it with those teams. Gatekeeper craft — other teams have shipped features through an API you governed: design review, breaking-change control, pushback with a migration path; debates trade-offs with engineering as a peer. Strong analytical skills — comfort defining and instrumenting product metrics, working with telemetry, and building data-informed roadmaps. Experience leading discovery-heavy work — structured customer interviews (we build top-tier infrastructure and work directly with frontier AI labs), usage analytics — turning insights into shipped product. Strong communication and the ability to align engineering, SRE, customer-facing teams and exec stakeholders. High ownership , a bias to ship, and focus on outcomes and customer value.
Nice to have: API & developer-surface depth
System design of compute / platform services — contr
What you'll do
- Own end-to-end product responsibility for your area — strategy, roadmap, discovery, delivery, adoption, and measurable customer & platform outcomes.
- Design and govern platform contracts at hyperscaler quality.
- Co-design the compute control plane with engineering as a technical peer (scheduling, allocation, reconciliation) — not just word an API.
- Manage stakeholders and drive cross-team execution across engineering, networking, storage, product and sales / CX.
- Define success metrics for the Compute API and the customer experience, and be the escalation point for product decisions on your surface.
What they require
- 6+ years in Product / Platform / Infrastructure PM — or an SRE / Engineering Lead moving to product — shipping technically complex platform products with measurable impact.
- Owned a public cloud / compute / platform API as a product . A declarative / desired-state or control-plane API (Kubernetes CRDs / operators, or a cloud control plane) is a strong plus.
- Hands-on cloud experience — the VM / instance lifecycle, its API and its console at a public or large-scale cloud (AWS / GCP / Azure or another large-scale cloud), with a real under-the-hood understanding of how scheduling, allocation and the virtualization layer work (control-plane vs data-plane — how VMs are placed, scheduled and run).
- A systems understanding of how a cloud fits together — how compute, storage, networking, IAM and related services connect and interact, and how it all works under the hood — enough to design the Compute API coherently and reason about it with those teams.
- Gatekeeper craft — other teams have shipped features through an API you governed: design review, breaking-change control, pushback with a migration path; debates trade-offs with engineering as a peer.
Dutch company developing a portfolio of AI-related technology assets