Overview
Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a computational data scientist who can define what a trustworthy spectral dataset looks like, and build the schema, QC gates, and annotation process that gets us there. This is a data and cheminformatics role, not a wet-lab role: though you'll work closely with analytical chemistry collaborators who run the instruments.
The Role
Developing a deep understanding of Novogaia's compound library, spectral data, and how both feed into our models
Working closely with our machine learning team to assess model training/validation leakage, de-duplicate against public datasets, and help select representative subsets for benchmarking
Working with analytical chemist collaborators to route ambiguous or high-value spectra for expert review, so results can be compared systematically against model predictions
Building and evaluating predictive models to infer molecular properties from molecular structure
Validating processing workflows for raw spectra arriving from analytical partners, including validating metadata, batch tracking, versioned releases, maintaining provenance, and licensing tags at the record level
In your first year, you'll build the data foundation everything else depends on: a documented schema, a QC process, and a dataset our modeling and evaluation teams can trust.
What We Require
Background in analytical mass spectrometry or cheminformatics (PhD or equivalent industry experience), ideally with exposure to natural products or small-molecule drug discovery
Deep familiarity with structural representation methods such as SMILES, SMARTS, SAFE, as well molecular fingerprinting and structural embedding
Familiarity with statistics and ML concepts, for close collaboration with the rest of the team
Hands-on experience working with LC-MS/MS data and standard formats and open-source tools (e.g. mzML, MSConvert) and spectral databases (e.g. GNPS, MassBank, MoNA)
Scripting ability in Python for developing algorithms and Nextflow/Snakemake for building pipelines
Familiarity with utilizing relational database schemas and ontologies to host and structure the variety of datatypes and datasets you will encounter
What We Value
Ability to intuitively interpret and assess mass spectrometry data and corroborate automated QC checks
Strong scientific judgment and a willingness to flag data that isn't ready, even under deadline pressure
Ability to turn "make this dataset AI-ready" into a concrete schema, checklist, and pipeline
Motivation to build data infrastructure other people will confidently rely on
Experience with natural product dereplication and compound classification
Curiosity, low ego, and a willingness to get close to the modeling and evaluation side of the work, even if it's outside your original training