Skip to main content
Apheris

Forward-Deployed Cheminformatician

RemoteGermany only
Published
Role
Data Science
Experience
Senior
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to DE only. Set where you work from to check your eligibility.

No BS summary

Cheminformatician needed to prepare binding data for AI models in drug discovery. Requires 3+ years of experience with biological assay data, Python, and RDKit. Must be able to work with pharma partners and turn ad-hoc cleaning into repeatable protocols.

Core skills

PythonRDKit

Required skills

SMILES normalizationtautomer handlingionization handlingstereochemistry handlingscaffold extractionquantitative binding assay data curationHTS data curationversion controltested modular scriptsvalidators

Optional skills

ChEMBLBindingDBPubChemBioAssayLLM tooling (Claude, Codex, Cursor)federated data networksmulti-party data collaborationpublication recordopen-source contributions

Required languages

English

What you'll do

  • Define and own the binding-data preparation protocol - data schema, small-molecule standardization, assay metadata model, value handling (KD, Ki, IC50, pIC50), qualifier and censored-value handling, duplicate and replicate aggregation.
  • Build the tooling that runs it - modular scripts, validators with actionable errors, and reusable pipelines that survive different pharma upstream systems (Dotmatics, Spot fire, in-house registries).
  • Work forward-deployed with pharma. Sit with their biologists and medicinal chemists, walk them through the protocol, sense-check what an assay column actually measures, and unblock retrieval.
  • Maintain the small-molecule representation pipeline - RDK it standardization, tautomer and ionization handling, stereochemistry preservation, and PAINS / frequent-hitter filtering.
  • Curate the public binding-data foundation - ChEMBL,BindingDB, PubChemBioAssay - prepared to the same standard, so our models train on the strongest public baseline anyone can assemble.
  • Hand the productized pipeline cleanly to engineering for scaling, and partner with ML to keep the data contract valid as models and networks evolve.

What they require

  • BSc, MSc, PhD or equivalent in cheminformatics, computational chemistry, or a related field, plus 3+ years preparing biological assay data in a discovery setting.
  • You are fluent in Python andRDKit. SMILES normalization, tautomer / ionization / stereochemistry handling, and scaffold extraction are second nature, and you understand why each matters for activity cliffs and model training.
  • You have hands-on experience curating quantitative binding assay data (KD, Ki, IC50, pIC50) and HTS data - censored values, qualifiers, duplicates, replicate aggregation, and assay metadata interpretation.
  • You write good engineering code - version control, tested modular scripts, validators that return useful errors.
  • You are comfortable forward-deployed with pharma medicinal chemists and biologists. You can sit in a sense-check meeting, pull out what is actually meant by a column label, and encode that back into the protocol.
  • You enjoy turning a messy ad-hoc cleaning job into a repeatable protocol others can run.
  • Bonus points if: You have practical familiarity with public binding-data sources (ChEMBL,BindingDB, PubChemBioAssay) and the gotchas in each.
  • Bonus points if: You have applied LLM tooling (Claude, Codex, Cursor) to accelerate data cleaning or metadata harmonization.
  • Bonus points if: You have worked across institutional data boundaries - federated, multi-party, or otherwise — where the data-preparation contract has to hold under partial visibility.
  • Bonus points if: You have a publication record or open-source contributions in cheminformatics or quantitative pharmacology.

Benefits

  • Industry-competitive compensation, including early-stage virtual share options
  • Remote-first work - work where you work best
  • Wellbeing budget, mental health support, work-from-home budget, co-working stipend, and learning budget
  • Generous holiday allowance
  • Office Days at our Berlin HQ or a different European location (3x per year)
  • A high-calibre, execution-focused team with experience from leading organizations

At Apheris, we are building the future of how AI is applied in pharmaceutical R&D. We enable leading pharmaceutical teams to discover and develop drugs faster. We host the industry’s largest federated data networks for drug discovery AI, spanning co-folding, ADMET, and antibody developability. Across these networks, models are trained on proprietary industry datasets to achieve higher performance and broader applicability while keeping data control and IP protected. We deliver these superior models through drug discovery applications that enable teams to run them at scale, further customize them, and integrate them into existing R&D workflows. AI Structural Biology (AISB) Network: Pharmaceutical companies collaborate in the field of co-folding, structure-based binding affinity predictions and antibody design.ADMET Network: Pharmaceutical and biotech companies collaborate to improve small-molecule property prediction and expand into further drug modalities.Antibody developability Network:Pharma partners collaborate to federate historical and purpose-built antibody developability data sets for secure ML training, without data leaving each partner’s environment.

BiotechStartupapheris.com/
Salary not disclosed