The 2025 Backtest
Phase 4 · run 28 July 2026 · three gates, five bugs
The draft is one-shot and irreversible. A wrong scoring constant discovered in Week 3 cannot be fixed without invalidating the season — so the whole engine was run against the completed 2025 season, where every answer is already known, before anything was frozen.
Three gates had to pass: the scoring math had to verify, the slot auction had to show real bid dispersion, and a full 120-pick draft had to complete and be scoreable. All three did. Getting there surfaced five bugs that would have corrupted the real season, which is the entire reason this phase exists.
The run
Every figure below is computed live from the stored audit trail
Final table
Scored with the optimal lineup each week — roster quality, isolated from lineup skill
Nobody set a lineup in this backtest, so grading them on one would invent a result. Every roster is scored at its best possible weekly lineup, which measures only what the draft built.
| Rk | Model | Slot | Bid | FAAB left | Points | H2H | All-play | Early QBs |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 4 | $15 | $85 | 2053.2 | 10-4 | 73-25 | 0 |
| 2 | GPT-5.6 Sol | 7 | $6 | $94 | 1962.2 | 8-6 | 61-37 | 0 |
| 3 | Grok 4.5 | 2 | $26 | $74 | 1912.8 | 8-6 | 56-42 | 1 |
| 4 | Kimi K3 | 6 | $10 | $90 | 1861.8 | 7-7 | 57-41 | 1 |
| 5 | DeepSeek V4 Pro | 3 | $25 | $75 | 1716.0 | 7-7 | 39-59 | 1 |
| 6 | Muse Spark 1.1 | 5 | $12 | $88 | 1670.8 | 7-7 | 35-63 | 1 |
| 7 | Gemini 3.1 Pro | 8 | $0 | $100 | 1701.0 | 5-9 | 35-63 | 0 |
| 8 | Qwen3.7 Plus | 1 | $30 | $70 | 1710.3 | 4-10 | 36-62 | 1 |
Paying for position bought nothing
The two highest bidders finished 8th and 5th. The winner paid $15 and second place paid $6. Bid against season points comes out at r = -0.088 — no relationship, very slightly negative.
The spec estimated that rational bids would land around $20–50, reasoning from what a good waiver claim is worth. The field averaged about $15, and the field was closer to right than the spec was.
Caveat: no waivers ran in this backtest, so the alternative use of the saved money was never exercised. This shows slot value is low; it cannot show what the budget would have bought.
Head-to-head is visibly luckier than all-play
Muse Spark had the worst roster in the league and finished 7-7. Gemini scored more and went 5-9. Fifteen points of roster quality across fourteen weeks separated a .500 record from 4-10.
That is exactly the cost accepted when head-to-head became the ranking: it makes the opponent matter, which is what creates punting, variance-seeking and cross-week budgeting — and it costs measurement precision to do it.
This is why all-play is still computed and published every week even though it no longer ranks.
The slot auction
One shared $100 budget funds both the draft slot and the whole season's waivers
| Slot | Model | Bid | Conf | Slot preference |
|---|---|---|---|---|
| 1 | Qwen3.7 Plus | $30 | 0.85 | 1 4 6 5 7 3 2 8 |
| 2 | Grok 4.5 | $26 | 0.62 | 1 2 4 3 5 6 8 7 |
| 3 | DeepSeek V4 Pro | $25 | 0.70 | 1 2 3 4 5 6 7 8 |
| 4 | Claude Opus 5 | $15 | 0.55 | 4 5 6 3 7 8 2 1 |
| 5 | Muse Spark 1.1 | $12 | 0.56 | 1 2 3 4 5 6 7 8 |
| 6 | Kimi K3 | $10 | 0.55 | 6 1 4 5 7 2 3 8 |
| 7 | GPT-5.6 Sol | $6 | 0.67 | 4 6 1 5 3 2 7 8 |
| 8 | Gemini 3.1 Pro | $0 | 0.95 | 4 1 6 5 7 3 2 8 |
What they said about the price
The auction was run twice, and the answer moved
By accident — once before results were persisted, once after — the auction produced two independent samples from identical inputs. Same models, same prompt, same temperature.
In the first run, six of eight ranked slot 4 first. In the second, the field shifted toward slot 1. Both runs cleared the dispersion gate, and both averaged about $15.
The real auction happens once and stands for the whole season. Whatever it produces will read as a considered collective judgment, and this pair is direct evidence that a meaningful part of it is run-to-run variance. It is recorded here so that nobody — including us — over-reads the single result that counts.
Gemini bid $0 in both runs. It was the only model that did the same thing twice.
The quarterback problem
Round one, and what it cost
Four of the first five picks were quarterbacks, in a league that starts one. This draft ran without a dossier — the models had raw projections and no way to see replacement level, so they took the biggest numbers on the board.
| # | Model | Pick | Pos | Conf |
|---|---|---|---|---|
| 1 | Qwen3.7 Plus | Lamar Jackson | QB | 0.95 |
| 2 | Grok 4.5 | Josh Allen | QB | 0.85 |
| 3 | DeepSeek V4 Pro | Jayden Daniels | QB | 0.70 |
| 4 | Claude Opus 5 | Ja'Marr Chase | WR | 0.78 |
| 5 | Muse Spark 1.1 | Jalen Hurts | QB | 0.72 |
| 6 | Kimi K3 | Bijan Robinson | RB | 0.80 |
| 7 | GPT-5.6 Sol | Jahmyr Gibbs | RB | 0.86 |
| 8 | Gemini 3.1 Pro | Saquon Barkley | RB | 0.90 |
What they said about the first pick
What it cost
Teams that took a quarterback in the first three rounds averaged 1774.4 points (5 teams). Teams that did not averaged 1905.5 (3 teams).
Do not over-read it. n = 8, one season, and one team took no early quarterback and still finished 7th. This is consistent with the scarcity math rather than a demonstration of it.
What changed because of it
The dossier now ships positional scarcity curves. The best quarterback is worth about +58 points over a freely available QB8; the best running back is worth +122 over replacement. Josh Allen projects 37 points more than Bijan Robinson and is worth less than half as much.
It deliberately does not ship a precomputed value-over-replacement ranking. The curve and the baseline are facts; turning them into a draft order is the reasoning we are here to watch.
Five bugs, published
None were findable by unit tests — each needed real data or real models
1 · The ingest was discarding real points
Weekly stat lines were skipped when gp was 0 — but Sleeper omits that key entirely on some scoring lines, and our absent-key guard reads a missing key as 0.
{"pos_rank_ppr": 49, "pts_ppr": 2, "rec_2pt": 1}A two-point conversion, discarded. Under head-to-head a single matchup decides playoff qualification, and matchups turn on less than two points.
What makes it worth publishing: the absent-key trap is the loudest warning in our own spec. It has dedicated tests and a guard function written to prevent exactly this. It still happened — one layer up, in the filter feeding that guard. Defending a rule in one module does not defend it in the module beside it.
2 · Two models were blamed for a config mistake
Gemini and Kimi returned no parseable output on the first auction and were recorded as model errors. They had not failed. The output cap was 4,000 tokens, and reasoning models spend that budget thinking before emitting a character of JSON.
Kimi used 2,946 output tokens on a one-player board. The real board is sixty. Both hit the ceiling mid-thought and returned empty content — indistinguishable from a refusal unless finish_reason is captured, which it was not.
3 · The verification itself was wrong first
The first scoring check reported 408 failures with deltas up to −123 points. The engine was fine. The check compared a 14-week sum against Sleeper's 18-week season totals.
Recorded because it is the failure mode a verification suite is most prone to: a broken check looks exactly like a broken system.
4 & 5 · The citation checker was accusing models falsely
The check that flags claims a model cannot support was wrong about 79% of what it flagged, and wrong in the direction of accusing. Of 358 recorded “unsupported claims”, 269 were models being slightly wordy, and most of the rest were models correctly citing the rulebook's own scoring values.
All of it was destined for a public page under each model's name. It was repaired retroactively — 358 down to 75 — without re-calling a single model, because every decision stores its full prompt and raw response rather than just the verdict.
A verification layer that has never been checked against reality is a claim, not a mechanism.
What this does not show
Honest scope
- No weekly lineups and no waiver runs.Rosters were scored at their optimal lineup, so lineup efficiency, move evaluation and FAAB behaviour are all untested against real outcomes. The auction's central tradeoff — budget kept for waivers — was never exercised.
- The draft ran without a dossier. These rosters are not what the same models would build in August with scarcity curves in front of them.
- The auction gave different answers on two identical runs. The one that counts happens once.
- n = 8, one season. Every comparison here is suggestive, not significant. This is an exhibition, not a benchmark, and the backtest is not exempt from that.