Task library/App builder

Prototype a calm household budget app

Build onboarding, transaction review, category editing, and a weekly decision summary.

Synthetic evaluation using demonstration outputs. Not real customer demand, adoption, or revenue.

OUTPUT VIEWER / RUN 04Gemini 3 Pro
still•••
WEEKLY PULSE

$2,184

Safe to spend after bills

72%plan confidence
Exact version recordedDesktop + mobile run3 journey states91/100
ONE-SENTENCE CONCLUSION

Gemini 3 Pro wins this run.

Fastest successful first session. The verdict changes when the persona prioritizes speed over editability, so the disagreement remains visible.

What changed the decision

Scores follow the artifact. Every judgment stays attached to a visible state or recorded validator.

01

Task completion

The winning output completed the primary journey without recovery help in three repeated runs.

02

Persona fit

Founders chose the fastest usable path; specialists rewarded control and evidence depth.

03

Failure replay

The lowest-scoring output hid one critical action and lost state after the second step.

Continue the evidence trail.

Inspect the full task library or compare this winner against another model.

Open pair comparisonNext task