Task library/Structured extraction

Extract a 500-page product catalog

Return normalized inventory, variants, price status, source freshness, and failed-page evidence.

Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.

OUTPUT VIEWER / RUN 04Firecrawl
catalog_run_0712.json98.4% valid
{
  "sku": "R-2048",
  "variant": "bone / 42",
  "price": { "value": 185, "currency": "CAD" },
  "freshness": "2026-07-12T14:18Z",
  "source_status": "observed"
}
fieldspagesvariants
Exact version recordedDesktop + mobile run3 journey states90/100
ONE-SENTENCE CONCLUSION

Firecrawl wins this run.

Best gated accuracy and coverage. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.

What changed the decision

Scores follow the artifact. Every judgment stays attached to a visible state or recorded validator.

01

Task completion

The winning output completed the primary journey without recovery help in three repeated runs.

02

Persona fit

Founders chose the fastest usable path; specialists rewarded control and evidence depth.

03

Failure replay

The lowest-scoring output hid one critical action and lost state after the second step.

Continue the evidence trail.

Inspect the full task library or compare this winner against another model.

Open pair comparisonNext task