OUTCOME INDEXES

Rank the work,
not the model name.

Each index is scoped to a comparable job, locked conditions, current versions, and the people who need the output.

Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.

IndexLeaderPersona splitCoverageScore
Website builderBuild a conversion-ready ceramics studio siteClaude Opus 431%3 models · 9 runs92Lifecycle emailRecover a stalled free-trial userGPT-5.244%3 models · 9 runs89App builderPrototype a calm household budget appGemini 3 Pro36%3 models · 9 runs91Video + animationCompare motion where continuity breaksField evidence52%3 models · 9 runs95Playable gameWatch whether the first minute earns a replayField evidence57%3 models · 9 runs96Research agentAssess a new Toronto retail marketGemini 3 Pro28%3 models · 9 runs93Structured extractionExtract a 500-page product catalogFirecrawl19%3 models · 9 runs90Teaching answerExplain why the harmonic series divergesGPT-5.222%3 models · 9 runs94

A score without scope is advertising.

Correctness and coverage can gate a result. Persona preference cannot rescue a broken answer. Every row links back to the artifacts and failures that produced it.