Skip to main content
WEKA

Principal Product Manager, Augmented Memory Grid (AMG)

RemoteUnited States only
Published
Role
Product
Experience
Principal
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Principal Product Manager needed for Augmented Memory Grid (AMG) within the NeuralMesh platform. Requires deep technical expertise in AI inference infrastructure, high-performance networking, and enterprise storage. Must have 10+ years of PM experience, strong leadership, strategic vision, and excellent communication skills. Experience with LLM inference serving and enterprise customer engagement is essential.

Core skills

AI inference infrastructurehigh-performance networkingenterprise storage

Required skills

LLM inference servingvLLMNVIDIA TritonTensorRT-LLMNIMRay ServeKV-cacheprefix cachingcontext cachingquantizationbatching strategiescontext windowsmulti-tenant servingGPU schedulingRDMAGPUDirect StorageNVMe-oFrequirements gatheringproduction deploymentsSLAs

Optional skills

GPU cloudinference platformAI infrastructure startupparallel file systemsdistributed file systemsobject storageNVIDIA

What you'll do

  • Own the AMG product roadmap: KV-cache/prefix-cache offload, memory tiering, and integration with inference engines and orchestration layers (vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Kubernetes-based serving).
  • Partner with engineering to define architecture trade-offs across GPU memory, networking (RDMA, GPUDirect, NVMe-oF), and distributed storage — translating inference performance bottlenecks (time-to-first-token, throughput, context length) into product requirements.
  • Work directly with enterprise customers and GPU cloud partners: Nebius, CoreWeave, TogetherAI, etc., running production inference workloads to gather requirements, validate benchmarks, and prioritize features that reduce cost-per-token and improve SLAs at scale.
  • Partner with NVIDIA and other silicon/inference-stack partners on joint roadmap and certification work.
  • Define and track benchmarks (TTFT, throughput, cache hit rate) that demonstrate AMG's value versus standard GPU-memory-only inference.
  • Support sales and field teams with technical positioning, competitive differentiation, and enterprise deal support.

What they require

  • 10+ years of product management experience, ideally with some portion in infrastructure, ML platforms, or developer-facing technical products.
  • Strong leadership skills with a history of successfully leading cross-functional teams.
  • Product Managers are expected to inspire and motivate team members to achieve ambitious goals while maintaining a collaborative and positive working environment.
  • You understand how to influence without authority, and your recall of meaningful details supports verbal and written agility.
  • You are a strategic thinker who can develop and execute product strategies that align with market trends and customer needs, as well as think critically about existing strategies.
  • You have a proven ability to translate strategic goals into actionable plans and deliver results.
  • You have excellent communication and interpersonal skills, with the ability to articulate complex technical concepts to both technical and non-technical stakeholders.
  • You are comfortable presenting product strategies and roadmaps to internal teams and external customers.
  • Inference ecosystem depth: hands-on product or engineering experience with LLM inference serving — vLLM, NVIDIA Triton/TensorRT-LLM/NIM, Ray Serve, or comparable — and fluency in concepts like KV-cache, prefix/context caching, quantization, and batching strategies.
  • Model & systems familiarity: working knowledge of how modern LLMs are served in production (context windows, multi-tenant serving, GPU scheduling) well enough to translate model-level constraints into infrastructure requirements.
  • Networking/infrastructure fluency: comfort with the fundamentals of high-performance networking and distributed systems — RDMA, GPUDirect Storage, NVMe-oF, or equivalent — and how they affect inference performance.
  • Enterprise customer experience: track record working directly with large enterprise accounts — requirements gathering, production deployments, SLAs — not solely self-serve/PLG products.
  • Prior experience at a GPU cloud, inference platform, or AI infrastructure startup.
  • Familiarity with storage systems (parallel/distributed file systems, object storage) in AI/ML pipelines.
  • Experience partnering directly with NVIDIA or other accelerator/silicon vendors.
  • Studies have shown that women and people of color may be less likely to apply for jobs if they don’t meet every qualification specified. At WEKA, we are committed to building a diverse, inclusive and authentic workplace. If you are excited about this position but are concerned that your past work experience doesn’t match up perfectly with the job description, we encourage you to apply anyway – you may be just the right candidate for this or other roles at WEKA.
  • WEKA is an equal opportunity employer that prohibits discrimination and harassment of any kind. We provide equal opportunities to all employees and applicants for employment without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.

Benefits

  • We are Accountable : We take full ownership, always–even when things don’t go as planned. We lead with integrity, show up with responsibility & ownership, and hold ourselves and each other to the highest standards.
  • We are Brave : We question the status quo, push boundaries, and take smart risks when needed. We welcome challenges and embrace debates as opportunities for growth, turning courage into fuel for innovation.
  • We are Collaborative : True collaboration isn’t only about working together. It’s about lifting one another up to succeed collectively. We are team-oriented and communicate with empathy and respect. We challenge each other and conduct positive conflict resolution. We are being transparent about our goals and results. And together, we’re unstoppable.
  • We are Customer Centric : Our customers are at the heart of everything we do. We actively listen and prioritize the success of our customers, and every decision we make is driven by how we can better serve, support, and empower them to succeed. When our customers win, we win.

suite of machine learning software written in Java

AI InfrastructureStartup
Salary not disclosed