Firecrawl wins this run.
Best gated accuracy and coverage. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.
Return normalized inventory, variants, price status, source freshness, and failed-page evidence.
Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.
{
"sku": "R-2048",
"variant": "bone / 42",
"price": { "value": 185, "currency": "CAD" },
"freshness": "2026-07-12T14:18Z",
"source_status": "observed"
}Best gated accuracy and coverage. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.
Scores follow the artifact. Every judgment stays attached to a visible state or recorded validator.
The winning output completed the primary journey without recovery help in three repeated runs.
Founders chose the fastest usable path; specialists rewarded control and evidence depth.
The lowest-scoring output hid one critical action and lost state after the second step.
Inspect the full task library or compare this winner against another model.