# 100 Users LLM Outcome Index

100 Users LLM compares what AI systems actually produce for standardized tasks. It evaluates visible artifacts, objective task completion, persona fit, reliability, cost, latency, and disagreement across disclosed synthetic perspectives.

## Product boundary

The LLM index evaluates AI models, agents, builders, and tools completing a task. The main 100 Users product evaluates websites and customer experiences. Their scores remain separate.

## Launch indexes

- Website Builder Index
- Research Agent Index
- Weekly Email Outcome Test

## Public page families

- Task library: https://100-users.com/llm/tasks/
- Outcome indexes: https://100-users.com/llm/indexes/
- Comparison room: https://100-users.com/llm/compare/
- Research reports: https://100-users.com/llm/reports/
- Methodology: https://100-users.com/llm/methodology/
- Company evaluation: https://100-users.com/llm/for-companies/
- Video evidence library: https://100-users.com/llm/video-lab/
- Website-building footage: https://100-users.com/llm/video-lab/website-building/
- Animation footage: https://100-users.com/llm/video-lab/video-animation/
- Game footage: https://100-users.com/llm/video-lab/game-creation/
- 3D and architecture footage: https://100-users.com/llm/video-lab/3d-architecture/

## Standardized task routes

- Website: https://100-users.com/llm/tasks/website/
- Email: https://100-users.com/llm/tasks/email/
- App: https://100-users.com/llm/tasks/app/
- Video: https://100-users.com/llm/tasks/video/
- Game: https://100-users.com/llm/tasks/game/
- Research: https://100-users.com/llm/tasks/research/
- Structured data: https://100-users.com/llm/tasks/data/
- Math teaching: https://100-users.com/llm/tasks/math/

## Evidence contract

Every comparison retains the exact model version, task version, prompt, settings, execution environment, repeated runs, rendered output, objective validator results, persona judgments, cost, latency, failures, confidence, and test date.

## Field evidence contract

The video library contains 29 source-supplied clips. Its editorial demonstration scores assess only visible presentation, comparison value, and interaction evidence. Source-presented model labels are not independently reproduced claims. Prompts, settings, authorship, and reproducibility may differ, so field-evidence scores never change controlled model rankings.

## Synthetic evaluation disclosure

Evaluations are generated by disclosed synthetic personas using standardized tasks and recorded evidence. They measure task performance and perspective-based judgment. They do not represent real customer demand, adoption, retention, or revenue.

## Canonical page

https://100-users.com/llm/
