Skip to main content
Supabase

AI Platform Engineer

RemoteWorldwide
Published
Role
AI / ML
Employment
Full-time
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to Worldwide. Set where you work from to check your eligibility.

No BS summary

AI platform engineer who has shipped production LLM agent systems, designs evaluation gates, and can own Python/GCP infrastructure end to end. Must have deep API work experience and have authored MCP servers. Role is fully remote and hires globally.

Core skills

MCPPythonLLM agents

Required skills

GCPInfrastructure as CodeCI

Optional skills

LLM observabilityCost instrumentation

What you'll do

  • Ship the agent platform to production, including an event-triggered queue, headless model-agnostic runtime, durable state, human review gate, atomic rollback, and complete run logging in the warehouse.
  • Choose and close the open runtime architecture decision.
  • Own the evaluation layer and enable the gate that depends on it.
  • Build golden suites with behavioral assertions, judge criteria with written rubrics, safety cases for every run, and a CI gate that blocks regressions from merging.
  • Build and register reporting, drafting, linting, triage, and question-answering agents across executive, team-lead, and individual-contributor layers.
  • Build a meta layer that observes the platform and improves it.
  • Enforce governance in code through risk tiers, least-privilege credentials per agent, tool-permission gates, decision audit logs, and autonomy classes with no code path for dangerous actions.
  • Ensure nothing runs without a registered owner, tier, tool grant, and human gate.
  • Design how the system contacts people, including interruption budgets, message bundling, and useful context before requests.
  • Own platform tooling, including the compiler, validator, inventory integrity, and distribution paths for context and capabilities into repositories and chat surfaces.
  • Compute operating measures from production data, from raw system activity to a computed maturity grade per team.
  • Build defensible queries so teams can dispute maturity grades and receive data-backed answers.
  • Instrument the platform's own return with a ledger that logs work absorbed by each agent and computes the monthly figure.
  • Build generators and platform infrastructure rather than one-off artifacts.
  • Build the machine that decides whether every future agent is allowed to ship.
  • Design from failure modes backward, making forbidden agent actions structurally impossible.
  • Use agentic tools on real work in files and repositories with production rigor.

What they require

  • Have shipped production LLM agent systems that other people depended on, with operational history, real users, and at least one incident you can talk through.
  • Prompt engineering alone is not sufficient.
  • Design evaluations rather than spot checks.
  • Build golden sets, write behavioral assertions, define judge rubrics, set pass thresholds, and gate CI on evaluation results.
  • Can explain why manually reviewing outputs is not evaluation.
  • Have done deep API work against systems where work actually lives.
  • Have authored MCP servers.
  • Know the specific failure modes of those APIs.
  • Own infrastructure end to end in Python on GCP, with a cloud warehouse and infrastructure as code.
  • Can provision, deploy, monitor, and roll back your own systems.
  • Can close architecture decisions rather than routing them onward.
  • Have taste about how software contacts humans and treat every notification as spending limited trust.
  • Have your own agentic tools setup and can describe what triggers it, what it can touch, where humans approve, what it logs, and how you found out when it went wrong.
  • Preferred: Public work in this space, such as an open-source agent framework, an MCP server, an evaluation harness, or writing on agent reliability that other practitioners cite.
  • Preferred: Experience tracing agent runs, attributing spend per run, and building queries that turn raw logs into actionable reports.
  • Preferred: Have built an internal platform that non-engineers adopted voluntarily and can describe what changed after observing users.

Benefits

  • Fully remote work.
  • Global hiring and ability to work from anywhere.
  • No Supabase offices.
  • WeWork membership or co-working allowance usable anywhere in the world.
  • ESOP equity ownership for every team member.
  • Tech allowance for laptop, monitor, headphones, or other work setup needs.
  • Supabase covers 100% of employee health insurance and 80% for dependents, wherever you are.
  • Annual company off-site in a new city.
  • Flexible asynchronous work and trust to manage your own time.
  • Annual professional development education allowance for courses, books, conferences, or other learning.

open source backend platform for app development

TechnologyMid-sizesupabase.com

Details

Apply routeDom
Salary not disclosed