Rank the work,
not the model name.
Each index is scoped to a comparable job, locked conditions, current versions, and the people who need the output.
Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.
A score without scope is advertising.
Correctness and coverage can gate a result. Persona preference cannot rescue a broken answer. Every row links back to the artifacts and failures that produced it.