About Calibre
AI systems are moving from proof of concepts into the infrastructure of everyday life. That creates a fundamental technical problem: how do you know an increasingly autonomous AI system is secure, reliable and behaving as intended? And who do you trust to answer that question?
The institutions and processes we rely on to provide that confidence were built for a slower, more predictable world. AI is advancing faster than the mechanisms we use to trust it.
Calibre is building trust for the AI era as the first AI-native certification body. We develop core technologies to understand AI and cybersecurity risk, and use them to independently assess companies against leading standards.
Founded by two ex-Palantir technical co-founders, and a founding team spanning audit and AI research, we are pushing the frontier of what trust looks like for the AI economy. We’ve raised $3.3M from top tier SF based VCs.
The Role
We are hiring a genuinely exceptional Product Engineer to join the founding team and build a new class of products to supercharge our AI-native certification process.
This person is language-agnostic, obsesses over first principles, and has already integrated LLMs into their daily workflow as a true force multiplier.
What You Will Do
Own products across the whole stack. Build and ship high-quality, high-velocity code across the full stack (front-end, back-end, and AI-agent infrastructure)
Build products end-to-end. Design systems from discovery through to production.
Engineer reliable agent systems. Design harnesses, evals and safeguards that make quality non-compromisable.
What we’re looking for
Demonstrable AI engineering depth. You can show real systems where you made substantive decisions about agent harnesses, evals, hallucination containment, performance, agent sandboxing or MCP/tool design.
Strong product engineering and technical judgment. You ship across interface, backend, data and deployment, reason from first principles and choose tools deliberately.
Product instinct and clear collaboration. You ask good questions, form a view, test quickly, communicate clearly and care whether the work changes outcomes.
Evidence of exceptional output. Show us professional work, open-source contributions, personal products, research or hackathon results. Roughly 2+ years of experience is useful, but evidence matters more than tenure.
The Ideal Candidate
Exceptional technical talent with more than 1 year of professional experience. Your GitHub, personal projects, or hackathon wins speak for themselves.
A first-principles problem solver. You have a fundamentally strong computer science background (e.g., from a top university or equivalent experience) and are language-agnostic, picking the right tool for the job.
AI-Native: You are already a sophisticated user of LLM tools (e.g., Claude Code, Cursor, etc.) for coding. You understand their strengths, weaknesses, and how to build systems that leverage them effectively.
High-Leverage & High-Output: You thrive in a high-pressure, high-growth environment. You understand that in an early-stage startup, precision and speed are paramount.
Compensation Range: £60K - £120K