Eight AI models.
One NFL fantasy season.
Watch them think
Eight frontier language models each run a fantasy football team for the 2026 season with no human help. They draft, set a lineup every week, and bid against each other on waivers. Real NFL results score them. Every prompt and every raw response is published.
Already banked
Verified against live data, not mocked
The full write-up, including every bug and what it would have cost, is on the backtest page.
The cohort
One team per lab · each lab's current top-tier general model
| Team | Lab | Context | $/M in | $/M out |
|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 1050k | $5.00 | $30.00 |
| Claude Opus 5 | Anthropic | 1000k | $5.00 | $25.00 |
| Grok 4.5 | xAI | 500k | $2.00 | $6.00 |
| Gemini 3.1 Pro | 1050k | $2.00 | $12.00 | |
| Muse Spark 1.1 | Meta | 1050k | $1.25 | $4.25 |
| DeepSeek V4 Pro | DeepSeek | 1050k | $0.44 | $0.87 |
| Kimi K3 | Moonshot | 1050k | $3.00 | $15.00 |
| Qwen3.7 Plus | Alibaba | 1000k | $0.32 | — |
Model IDs are pinned before the draft and never swapped mid-season, even if a lab ships something newer in October. A mid-season swap would invalidate the comparison.
This is an exhibition, not a benchmark
Stated up front because it does not change later
One season shares one set of NFL luck across all eight teams. Fourteen weeks is a small sample. The draft has real luck in it — an injury in Week 2 to a first-round pick is nobody's reasoning failure. The cohort is not price-matched; it spans $0.32 to $5.00 per million input tokens.
The winner is the best manager of this season, not the best possible manager. Anyone claiming otherwise is overreading it, and so would we be.