Skip to main content
DEFCON AI

Data & ML Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Senior
Employment
Full-time
$150k–$200k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Data/ML engineer with 5+ years in data engineering, ML engineering, data architecture, applied ML, or production analytics engineering. Must have strong Python and SQL, US citizenship, active US Secret clearance, CAC eligibility, and willingness to travel up to 25%. Work is remote in the USA inside a controlled government cloud environment.

Core skills

PythonSQLMachine Learning

Optional skills

PostgreSQLpgvectorscikit-learnXGBoostPyTorchRetrieval-Augmented GenerationAWS GlueAirflow

What you'll do

  • Build the data and model layer behind an AI-enabled decision-support system operating inside an accredited environment.
  • Handle ingestion from many source systems, resolution of incoming records against a shared data model, relevance scoring, and generation of explanations a user can act on and defend.
  • Focus on new capability rather than maintenance: record matching, calibrated scoring, and grounded generation, hardened for the target environment.
  • Design and maintain the graph of entities, records, and the typed relationships between them.
  • Implement probabilistic matching, including blocking, candidate generation, pairwise scoring, clustering, and threshold policy.
  • Build deduplication and known-record suppression.
  • Establish provenance so that every node and edge traces to the source that asserted it.
  • Produce interface and data-flow design documentation detailed enough to serve as an implementation reference for other engineers.
  • Develop relevance and priority models over large, imperfect record sets.
  • Own calibration and threshold design, establishing what a score means rather than only how it ranks.
  • Design abstention policy that routes uncertain and high-risk cases to a person rather than returning a confident answer.
  • Perform feature engineering, establish baselines before introducing complex models, and conduct error analysis that accounts for the differing cost of false positives and false negatives.
  • Implement embeddings, vector storage, and retrieval across a large provenance-tracked evidence base.
  • Integrate language models through an approved managed service, and maintain a self-hosted or open-weight alternative within the same boundary.
  • Design prompts and output schemas.
  • Bind generated text to cited source records, and treat "insufficient evidence" as a valid system response rather than forcing a conclusion.
  • Own model packaging, serving, versioning, and rollback.
  • Build secure ingestion, transformation, validation, and publishing across structured, semi-structured, and unstructured sources.
  • Implement quality checks, schema validation, lineage capture, and audit logging.
  • Establish source drift detection so that degradation is surfaced rather than carried into the analysis.
  • Generate statistically representative synthetic data so that development can proceed ahead of live data access.
  • Work to the data model and standards set by the Data Lead, who approves designs and owns them through customer review.
  • Document assumptions, caveats, transformation logic, and known limitations, since deliverables are formally reviewed.
  • Instrument telemetry so that measurement does not require manual reconstruction.
  • Maintain the audit trail covering recommendations, human overrides, and model versions.
  • Submit model and pipeline changes through a gated release process rather than deploying in place.

What they require

  • 5+ years of experience in data engineering, data architecture, applied machine learning, ML engineering, or production analytics engineering.
  • Strong Python and SQL, with demonstrated experience working with large, imperfect operational data.
  • Experience delivering systems for sustained operational use rather than exploratory analysis alone.
  • Routine use of AI-assisted development, with informed judgment about where it adds value and where its output requires verification.
  • Ability to explain a technical decision to a stakeholder who must defend that decision without understanding its internals.
  • US Citizenship Required.
  • Active US Secret clearance.
  • The work is performed in a controlled government cloud environment and requires a favorable investigation and CAC eligibility from the start.
  • Elevated personnel security requirements apply to portions of this work and are discussed during screening.
  • Willingness to travel up to 25% to customer sites, DEFCON AI HQ, and vendor facilities as required.
  • Preferred: Clearance: active Top Secret.
  • Preferred: Matching: direct experience applying probabilistic matching to inconsistent identity data, including names, dates, addresses, and identifiers, and familiarity with the failure modes of each.
  • Preferred: Record linkage, master data management, or identity management.
  • Preferred: Graph data modeling.
  • Preferred: Graph algorithms applied in production.
  • Preferred: Modeling: model calibration and threshold design.
  • Preferred: Cost-sensitive learning where error types carry unequal consequences.
  • Preferred: Retrieval-augmented generation in production.
  • Preferred: Prompt and output-schema design.
  • Preferred: Establishing that generated output remains grounded in its sources, and testing to confirm it.
  • Preferred: Self-hosted or open-weight model operation.
  • Preferred: Fine-tuning, adapters, or custom embeddings.
  • Preferred: Pipelines: unstructured and semi-structured document ingestion.
  • Preferred: Synthetic or representative test data generation.
  • Preferred: Environment: federal DevSecOps, RMF, ATO, or DoW cloud environments.
  • Preferred: Hardened base images.
  • Preferred: Experience advancing a pipeline from development through accreditation and deployment.
  • Preferred: Domain: sensitive federal or defense data, and work performed under privacy or comparable handling constraints.
  • Preferred: Responsible AI: documentation, model cards, fairness testing, and model monitoring.
  • A data model that the rest of the team builds on without needing to redesign it.
  • Matching decisions that can be explained and defended to a non-technical reviewer.
  • Models whose miss rate is characterized, not only their overall accuracy.
  • Generated explanations that assert no more than the sources support, with the citation path intact.
  • Pipelines that surface problems early and trace them to a specific source.
  • Consistent development progress, including during periods when live data is not yet available.

Benefits

  • A fully remote, results-based environment.
  • Competitive salary, bonus, and equity package.
  • 100% employer paid, comprehensive health insurance including medical, dental, and vision for you and your family.
  • Unlimited PTO, with your manager’s approval.
  • Flexible work environment where you manage your work day.
  • 14 weeks of fully-paid parental leave.
  • Career track opportunity with potential for rapid advancement with strong performance as the firm grows.
  • 100% employer paid, comprehensive health care including medical, dental, and vision for you and your family.
  • Paid maternity and paternity for 14 weeks at employees' normal pay.
  • Unlimited PTO, with management approval.
  • Opportunities for professional development and continued learning.
  • Optional 401K, FSA, and equity incentives available.
  • Mental health benefits are available through Tara Mind.
  • Cost effective GLP-1 solutions available through Crux.

DEFCON AI is an insights company that leverages artificial intelligence, mathematical optimization, data analytics, and software engineering for resilient optimization of complex systems.

Defense TechStartup

Details

Apply routeGreenhouse
$150k–$200k/yr