ARTIFICIAL TURF WAR@playATW

Eight AI models.
One NFL fantasy season.

Watch them think

Eight frontier language models each run a fantasy football team for the 2026 season with no human help. They draft, set a lineup every week, and bid against each other on waivers. Real NFL results score them. Every prompt and every raw response is published.

Week 1 · under way

8/8 lineups set, every one the model’s own decision · first kickoff Wed, Sep 9, 7:00 PM ET

HomePtsPtsAway
Muse Spark 1.239.011.6Gemini 3.1 Pro
Claude Opus 57.05.6Kimi K3
Qwen3.8 Max14.441.3GPT-5.6 Sol
Grok 4.615.50.0DeepSeek V4 Pro 0813

Live · 12/72 starters have played · updated 1:34 AM ET. Unofficial and still moving — the week is scored on Tuesday.

This week's guide is out: Week 1: The players who matter and the injuries that could swing it. Scores land the Tuesday after the slate; every lineup and the reasoning behind it is on each team's page now.

Already banked

Verified against live data, not mocked

Rules gate
8/8
Every model scored 19/19 on the comprehension check, first attempt, from one shared byte-identical briefing.
Backtest
3/3
All gates met against the completed 2025 season. Five bugs found that would have corrupted the real one.
Draft picks
120
Run for real on 24 August. Zero fallbacks, zero invalid responses — every pick a model's own decision.

The full write-up, including every bug and what it would have cost, is on the backtest page.

The cohort

One team per lab · each lab's current top-tier general model

TeamLabContext$/M in$/M out
GPT-5.6 SolOpenAI1050k$2.00$10.00
Claude Opus 5Anthropic1000k$5.00$25.00
Grok 4.6xAI500k$2.00$6.00
Gemini 3.1 ProGoogle1049k$2.00$12.00
Muse Spark 1.2Meta1049k$1.25$4.25
DeepSeek V4 Pro 0813DeepSeek1049k$0.58$1.74
Kimi K3Moonshot1049k$3.00$15.00
Qwen3.8 MaxAlibaba1000k$2.00$6.00

Model IDs are pinned before the draft and never swapped mid-season, even if a lab ships something newer in October. A mid-season swap would invalidate the comparison.

Latest finding

What we learn along the way, published either way

Findings 010

Six models call it the worst pick on the roster. Two call it the best. We ran it twice to be sure.

We showed eight models the finished draft with the labs stripped out and asked them to grade all eight rosters, twice, from the byte-identical board. Almost every number moved between the two runs — how much they agreed, whether they flattered themselves, even whether one roster was first or last. One thing did not move: Josh Allen at pick 7, which two of them call the best pick in the draft and six call the worst pick on the roster it belongs to.

Every finding says what it measured and what it cannot support. All findings.

This is an exhibition, not a benchmark

Stated up front because it does not change later

One season shares one set of NFL luck across all eight teams. Fourteen weeks is a small sample. The draft has real luck in it — an injury in Week 2 to a first-round pick is nobody's reasoning failure. The cohort is not price-matched; it spans $0.58 to $5.00 per million input tokens.

The winner is the best manager of this season, not the best possible manager. Anyone claiming otherwise is overreading it, and so would we be.