Skip to main content
Stratus

Senior Data Architect (Hands on)

RemoteUnited States only
Published
Role
Unknown
Experience
Senior
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

The Senior Data Architect owns the canonical data architecture — the schema, contracts, tenancy, and governance that every product and every AI/ML workload builds on. You are the single owner of the canonical data model: one normalized definition of the core business objects shared across our products, and the standard the rest of engineering builds against.

Core skills

MongoDB/Atlas M40+/aggregation framework/indexing/change streams/sharding/replica sets

Required skills

vector search/Atlas Vector Search/pgvectorAzure/AWS/GCPClaude Code/Copilot/Cursor

Optional skills

Knowledge-graphontologysemantic-layerCDCcross-engine syncMongoDB Change StreamsDebeziumDatabricks

What you'll do

  • Architect the data layer so AI/ML workloads — vector search, embeddings pipelines, RAG-grounded retrieval, model training — run on a clean, governed substrate.
  • Make production data AI-ready: well-modeled, contract-enforced, lineage-tracked, and drift-detectable.
  • Design the data-side integration patterns these workloads depend on, such as feature-store and vector-store patterns across document, relational, and embedding data.
  • Own the canonical data model — the normalized definition of the core business objects shared across our products — and decide what is canonical versus tenant-specific.
  • Establish data architecture standards, data contracts, and schema discipline the rest of engineering builds against, enforced in-repo.
  • Exercise strong polyglot-persistence judgment: what belongs in document vs. relational vs. vector stores, and how to migrate between them without big-bang rewrites.
  • Define the multi-tenant data architecture: tenancy isolation, data residency posture, and per-tenant cost attribution across storage and compute.
  • Lead staged modernization toward the right mix of stores and patterns for transactional, analytical, and AI/ML use cases — improving scalability, governance, and usability while minimizing disruption.
  • Own the architectural direction of the data pipeline and lake / lakehouse layer: ingestion, transformation, orchestration, and storage tiers.
  • Lead the move from homegrown pipelines to proven, industry-standard platforms, balancing build-vs-buy and total cost of ownership.
  • Modernize legacy data-access patterns via incremental, strangler-fig migrations that keep production stable.
  • Drive hands-on prototypes, reference implementations, and in-repo guardrails.
  • Define the data, storage, and retrieval patterns the rest of engineering builds against.
  • Establish data quality, testing, lineage, and observability standards for pipelines and AI/ML serving.
  • Mentor engineers on schema discipline, modern data practices, and AI/ML-readiness patterns.
  • Make canonical decisions that are time-boxed, written, and defensible; hold disagree-and-commit rather than letting schema debate become a standing committee.
  • Use AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for schema design, query tuning, and migration scripting.
  • Partner with database engineering on production data health while owning long-term architectural direction.
  • Partner with ML and application engineering on their data needs — structuring and governing data so it is retrieval-ready and safe to build on.
  • Partner with platform / infrastructure on reliability, disaster recovery, residency, and the multi-tenant operational posture.

What they require

  • 8+ years in data architecture, data engineering, database administration, or analytics engineering, with 3+ years in senior / lead roles.
  • Demonstrated ownership of a canonical or enterprise data model / cross-product schema — the model and contracts other teams built against.
  • Hands-on MongoDB at production scale (Atlas M40+ ideal): document modeling, aggregation framework, indexing, change streams, sharding, replica sets — and the judgment to recognize the Mongo-as-RDBMS anti-pattern.
  • Strong polyglot-persistence judgment: deciding what belongs in documents vs. relational vs. a vector store, and migrating between them incrementally.
  • Hands-on relational depth: schema design, indexing strategy, and query tuning, plus familiarity with vector search (Atlas Vector Search, pgvector, or equivalent).
  • Production experience making data AI/ML-ready: data architecture supporting RAG, semantic search, embeddings / vector pipelines, or agentic workloads.
  • Multi-tenant architecture experience: data residency and per-tenant cost attribution.
  • Pipeline / ELT / lake / lakehouse design at scale, with incremental migration strategies that minimize disruption.
  • Cloud-native data services (Azure, AWS, or GCP).
  • Strong grasp of data quality, testing, lineage, and monitoring — including observability for pipelines and AI/ML serving.
  • Comfortable modeling a complex, specialized domain. MEP / AEC / construction experience is a plus; appetite to learn the domain is required.

Benefits

  • Comprehensive and competitive health benefits plan
  • Matching 401k contributions
  • 20 days annual PTO
  • Primarily remote work with occasional annual team onsites

Stratus, is the leading cloud-based platform for MEP contractors. We’re on a mission to revolutionize the construction industry by providing innovative data driven solutions that seamlessly layer across a contractor’s entire workflow from design, to fabrication, to installation.

ConstructionStartupstratusband.com/
Salary not disclosed