Task library/Research agent

Assess a new Toronto retail market

Produce a source-led decision brief with claims, contradictions, limitations, and a recommendation.

Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.

OUTPUT VIEWER / RUN 04Gemini 3 Pro
DECISION BRIEF / JULY 12Conditional go

Toronto's premium repair market is dense, but trust, not demand, is the entry constraint.

3 signals

Search intent rose 18%; four incumbent shops added waitlists; warranty language remains inconsistent.

2 contradictions

High household income does not predict repair adoption. Neighbourhood results diverge sharply.

12 sources9 claims verified2 limitations
Exact version recordedDesktop + mobile run3 journey states93/100
ONE-SENTENCE CONCLUSION

Gemini 3 Pro wins this run.

Highest citation validity. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.

What changed the decision

Scores follow the artifact. Every judgment stays attached to a visible state or recorded validator.

01

Task completion

The winning output completed the primary journey without recovery help in three repeated runs.

02

Persona fit

Founders chose the fastest usable path; specialists rewarded control and evidence depth.

03

Failure replay

The lowest-scoring output hid one critical action and lost state after the second step.

Continue the evidence trail.

Inspect the full task library or compare this winner against another model.

Open pair comparisonNext task