METHODOLOGY / VERSION 1.0

Evidence first.
Preference second.

A good-looking artifact cannot outrank a broken one. A high score cannot escape its task, persona, conditions, or date.

Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.

01

Identity

Record provider, exact model version, alias, date, region, tools, and price.

02

Locked task

Freeze brief, conditions, success criteria, retries, and publication rules before execution.

03

Repeated runs

Execute comparable runs and preserve outputs, errors, latency, cost, and environment.

04

Objective gates

Correctness, schema validity, coverage, accessibility, and citation validity can disqualify outputs.

05

Persona judgment

Persistent disclosed personas evaluate fit without pretending to be customers, voters, or demand.

06

Confidence

Coverage, run variance, evaluator agreement, and freshness determine the confidence label.

What this does not prove

Synthetic evaluations do not establish real customer demand, market adoption, revenue, retention, or a statistically representative human preference. They are repeatable decision support attached to inspectable evidence.