Research
Mohammad Imtiaz
A correct final answer can hide broken reasoning. Step Validity scores each step on its own.
Benchmarks usually grade only the final answer. That rewards models that reach the right answer by the wrong route.
Step Validity grades each intermediate step for correctness and for whether later steps depend on it. A model can only score well by reasoning well, not by guessing the answer and filling in plausible steps.
It is the evaluation our safety framework uses to test reasoning faithfulness before each release. We will release the evaluation set and grading code so others can reproduce our results, and will link them here when they are public.