Skip to main content
Harbor Compliance

Staff Data Engineer

RemoteUnited States only
Published
Role
Data Engineering
Experience
Staff
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level Data Engineer, 7+ years experience, U.S.-based remote role. Must have strong hands-on experience with streaming/CDC (Kafka/Kinesis/Flink/Debezium) and vector databases/embeddings; advanced SQL and Python required. Build a real-time, AI-ready data platform and vector/embedding infrastructure for executive/analytics use cases.

Core skills

KafkaDebeziumPinecone/Weaviate/pgvector/Milvus/Zilliz

Required skills

KinesisFlinkSQLPythonSnowflake/BigQuery/DatabricksdbtFivetranAirbyteLookerTableauPower BIClaudeCopilot

What you'll do

  • Design, build, and own near real-time data pipelines (CDC, streaming ingestion, event-driven architectures).
  • Evaluate, implement, and maintain vector database infrastructure and embedding pipelines to support semantic search, RAG, and AI agents.
  • Design and build ELT/ETL pipelines ingesting data from platform, financial systems, CRM (HubSpot) and HRIS for real-time and batch use cases; architect the underlying warehouse/lakehouse.
  • Own pipeline reliability and observability — monitoring, automated failure alerting, and lineage tracking across streaming and batch pipelines.
  • Build foundation for self-service and AI-powered reporting; implement data governance, documentation standards and access controls; partner cross-functionally with BI, Finance, Marketing and Ops.

What they require

  • 7+ years of hands-on data engineering experience, with meaningful depth in streaming/event-driven systems.
  • Proven experience designing and building near real-time pipelines from scratch (e.g., Kafka, Kinesis, Flink, Debezium/CDC) in production.
  • Hands-on production experience with vector databases and embeddings (e.g., Zilliz, Pinecone, Weaviate, pgvector, Milvus).
  • Advanced proficiency in SQL and Python; working knowledge of a cloud warehouse/lakehouse platform (Snowflake, BigQuery, or Databricks) and dbt.
  • Proven experience building or materially contributing to an end-to-end production data environment, ideally as an early or founding data hire; familiarity with B2B SaaS recurring revenue data models.

Benefits

  • Health benefits
  • Flexible paid time off
  • Parental leave
  • Fertility and adoption assistance
  • 401(k)

Harbor Compliance is a leading technology platform for entity compliance, helping more than 80,000 businesses and nonprofits manage licensing, tax registration, and legal entity requirements nationwide. Founded in 2012 and recognized repeatedly by the Inc. 5000 and Deloitte Technology Fast 500, we've grown through five strategic acquisitions — and are now backed by a 2026 majority growth investment from Bregal Sagemount to accelerate product, AI, and customer experience. We're a passionate team making compliance simpler and smarter for every organization we serve.

🇺🇸 United StatesFintechMid-size
Salary not disclosed