Skip to main content
Syllo

Staff Software Engineer, Search & Retrieval Infrastructure

RemoteUnited States onlyArchived
Published
Role
Backend
Experience
Staff
$190k–$230k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff/Principal-level software engineer with 8+ years scaling distributed, high-throughput systems for petabyte-level data. Must have deep production expertise in Lucene-based search, vector indexing, distributed pipelines, cloud infrastructure, Kubernetes, IaC, and backend/systems languages. US remote role.

Core skills

Elasticsearch/SolrLuceneVector indexing

Required skills

Kafka/KinesisGCPKubernetesInfrastructure-as-codeGo/Rust/Python/Java/C++

What you'll do

  • Scale the Retrieval Stack: Lead the optimization and architectural evolution of our existing hybrid search infrastructure, maximizing the throughput and efficiency of both lexical search (e.g., Elasticsearch, Lucene) and dense vector databases.
  • Advanced Data Tiering & Scanning: Design and implement intelligent, cost-effective tiering strategies across hot, warm, and cold data states. Evolve our distributed pipelines to efficiently execute asynchronous, massive-scale scans of petabytes of data in varying states of availability.
  • Relentless Optimization: Drive down latency and cost-to-serve. Deeply analyze system bottlenecks, tune indexing and querying algorithms, and optimize cloud infrastructure (compute, storage, and networking) for maximum efficiency at extreme scale.
  • Technical Leadership: Act as the domain expert and owner of the indexing and search ecosystem. Set the long-term technical vision for data storage and retrieval, guiding engineering teams on best practices for high-volume data modeling and performance tuning.
  • Resiliency at Scale: Ensure fault-tolerant, highly available operations during massive parallel ingest events and complex, concurrent querying across millions of documents.

What they require

  • Extreme Scale Experience: 8+ years of software engineering experience, with a proven track record operating at the Staff/Principal level optimizing and scaling highly distributed, high-throughput systems to handle petabyte-level data.
  • Search & Vector Mastery: Deep, production-level expertise tuning and scaling Lucene-based search engines (Elasticsearch, Solr) and modern vector indexing infrastructure. You deeply understand index internals, chunking strategies, and embedding retrieval optimization.
  • Cost-Aware Architecture: A strong history of managing the compute vs. storage trade-off. You know how to design sophisticated cold-storage scanning solutions and hot-index architectures that are highly performant but fundamentally cost-effective.
  • Distributed Systems: Extensive experience managing complex data pipelines, high-throughput event streaming (Kafka, Kinesis), and distributed compute architectures handling billions of records.
  • Cloud Infrastructure: Expert command of cloud primitives (GCP preferred), Kubernetes, and infrastructure-as-code.
  • Languages: Expert-level proficiency in systems-level and backend languages (Go, Rust, Python, or Java/C++).

Benefits

  • health insurance
  • equity

Syllo is defining the Litigation AI category. We are the first unified platform designed to autonomously manage the entire litigation lifecycle end-to-end.

LegalTechStartup
$190k–$230k/yr