Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a machine learning research engineer who excels at designing rigorous tests for molecular AI systems, and who can turn "does this model actually work" into a concrete, defensible answer. You will own the design of our internal benchmarks and evaluation pipelines; build the adversarial checks that catch shortcut learning and leakage. You will work closely with our modeling team to translate evaluation results into research priorities.
Develop a deep understanding of Novogaia's models, data, and evaluation needs
Design benchmark tasks that reflect real discovery problems: de novo molecular structure generation, molecular and spectral retrieval, mass spectrum simulation, molecular formula prediction, and compound property/class prediction
Build the datasets and controls that make a benchmark trustworthy: hard negatives, leakage-safe splits, and null baselines that catch a model exploiting shortcuts instead of genuine signal
Communicate evaluation results as clear findings for the modeling team, and as documentation and data cards that others can trust and reproduce
Work across the team to scope and lead evaluation work, including:
Defining what "good" looks like for a given model or task, and choosing the right test for it
Analyzing model behavior and interpreting results for researchers and non-technical stakeholders alike
Working with engineers to turn one-off analyses into repeatable, reproducible evaluation pipelines
Translate lessons from evaluation work into research priorities, data requirements, and R&D direction
In your first year, you'll take the lead on how Novogaia measures model quality, from benchmark design through adversarial testing and reporting. The tests you build will decide how much weight anyone can put on our models' outputs.
Research or applied experience in machine learning, with direct experience building or rigorously evaluating ML benchmarks
Exceptional technical communication skills, including the ability to explain evaluation findings clearly to both researchers and non-technical stakeholders
Ability to analyze model behavior and interpret computational results critically
Strong proficiency in Python, and comfort with reproducible, containerized pipelines
Familiarity with cheminformatics representations (SMILES, InChIKey, molecular fingerprints), or willingness to pick these up quickly
Familiarity with computational mass spectrometry or eagerness to learn
Ability to dive deep into a result until you know whether it's real or an artifact
Strong scientific judgment and a willingness to question the benchmark's own assumptions, not just the model's
Motivation to build infrastructure other people can confidently rely and build upon
Comfortable being the person who tells the team a result doesn't hold up
Curiosity, low ego, and a willingness to learn the chemistry side quickly, even if it's outside your original training