GPT-5.2 wins this run.
Correct and easiest to learn from. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.
Give a correct proof, anticipate one misconception, and adapt the explanation for a first-year student.
Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.
Every new block adds another half; the total never settles.
Correct and easiest to learn from. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.
Scores follow the artifact. Every judgment stays attached to a visible state or recorded validator.
The winning output completed the primary journey without recovery help in three repeated runs.
Founders chose the fastest usable path; specialists rewarded control and evidence depth.
The lowest-scoring output hid one critical action and lost state after the second step.
Inspect the full task library or compare this winner against another model.