Research
verify
Our research goes into Lemma, our reasoning engine, today, and into our own models next.
Verifiable reasoning
Models that break problems into claims and prove each one, so conclusions can be checked.
Honest uncertainty
Step-level confidence that tells people where a model is likely to be wrong.
AI for science
Research agents that design experiments to falsify hypotheses, not confirm them.
Measurement
Evaluations that grade the reasoning, not just the final answer.
our own models
Where each stage stands today.
- In production
Now
The Lemma engine
Our verification layer on leading foundation models, running in our product and in enterprise agents.
- –Step-level confidence and review routing
- –Step Validity evaluation on real work
- –Full reasoning traces for every answer
- In development
Next
Our first model
A model trained on verified reasoning traces, built to check steps more reliably and more cheaply than a general model.
- –Trained only on data we have the right to use
- –Measured on Step Validity before it ships
- –Published with its evaluation results
- Planned
Then
Our models under the engine
Theoratix models do most of Lemma's work, with foundation models wherever they are still the better choice.
- –Same product, lower cost per answer
- –Private deployment inside client networks
- –Research published as we go
- Goal
Goal
The frontier of reasoning you can check
The most reliable reasoning systems for decisions that have to be explained, measured in public.
- –Leading results on public reasoning evaluations
- –A safety framework applied to every release
Oct 2026
Lemma engine
Our reasoning engine, in production in our product and first agents.
Our own evaluations, and the public benchmarks we report.
Step Validity
TheoratixScores each reasoning step for correctness and for whether later steps depend on it.
Faithfulness probes
TheoratixTests whether the reasoning a model shows is the reasoning that produced its answer.
Calibration under shift
TheoratixMeasures whether stated confidence still matches accuracy on unfamiliar problems.
Public benchmarks
ExternalMATH-500, GPQA Diamond and SWE-bench Verified, run with published settings.