Reasoning Safety Framework
Version 1.0: the risks we test for, the thresholds that trigger extra safeguards, who decides, and what happens when a threshold is crossed.
Theoratix builds Lemma, a reasoning engine that runs on foundation models from other providers with our own verification and evaluation layer, and we are developing models of our own. This framework sets out how we decide what to build and release. We publish it before training a model of our own, so the commitments come before the incentives to relax them.
It applies to every release of Lemma, of our products, of agents we deploy for clients, and of any model we train. Every version is published with its date, and a published version is never edited in place.
Risk domains
We test each release against four risk domains.
- Unfaithful reasoning: the steps a system shows do not match how it reached its answer, or its confidence scores do not match how often it is right. This is the risk closest to our purpose, because people rely on the trace to check the answer.
- Misuse uplift: the release meaningfully helps someone cause serious harm, in particular chemical, biological, radiological or nuclear weapons, or cyber attacks, beyond what is already freely available.
- Unsafe autonomy: an agent takes actions in a client's systems that it was not authorized to take, or takes consequential actions without the human review its deployment requires.
- Data exposure: a release reveals data it should not, such as one client's data to another, or personal information in its outputs.
Thresholds
A release crosses a threshold when, in pre-release testing:
- Unfaithful reasoning: on our Step Validity evaluation, its stated confidence is materially overconfident, or it produces correct final answers through steps that do not hold up at a materially higher rate than the version it replaces.
- Misuse uplift: red-team testing finds that it gives meaningful, actionable help toward the harms above.
- Unsafe autonomy: in testing against the client's permissions, an agent attempts an action outside them, or skips a required human review.
- Data exposure: testing finds any leak of client data across engagements, or of personal information it was not given for that task.
What happens when a threshold is crossed
- The release is paused. It does not ship until extra safeguards are in place and testing shows the threshold is no longer crossed.
- Safeguards are chosen for the risk: for example narrower tools or permissions, stronger refusal behavior, additional human review, or keeping a capability out of the release.
- The decision, the evidence and the safeguards are recorded.
- If a problem is found after release, we act the same way: we limit or roll back the release until it is fixed, and tell affected clients.
Who decides
Today the founder is accountable for every release decision under this framework and signs off each one in writing. As the company grows we will appoint a head of safety research and an independent external advisor, and name them on our safety page when they hold the role. A person who builds a release will not be its only reviewer once we have the people to make that possible.
Before every release
- Run our reasoning-faithfulness and calibration evaluations, and compare them with the previous version.
- Red-team the release against the misuse and autonomy domains.
- For client agents, test against that client's permissions and review rules before they go live.
- Record the results and the release decision.
Transparency
- We publish each version of this framework with its date, and keep earlier versions available.
- We publish evaluation results for releases of our own models before they ship.
- We share evaluation methods and results with independent researchers and public bodies where we can.
Reporting a concern
Anyone can report a safety concern about Lemma, our products or an agent we built by emailing contact@theoratix.ai. Every report is read by a person and answered. Security vulnerabilities follow our responsible disclosure policy. Prohibited uses are set out in our acceptable use policy.
Versions
- 1.0, October 9, 2026: first published version.