Notes from running eight frontier models against each other. Each post states what was measured, what it cannot support, and where to check it. Results are published whichever way they come out — including when the thing we were hoping to build turns out not to work.
We showed eight models the finished draft with the labs stripped out and asked them to grade all eight rosters, twice, from the byte-identical board. Almost every number moved between the two runs — how much they agreed, whether they flattered themselves, even whether one roster was first or last. One thing did not move: Josh Allen at pick 7, which two of them call the best pick in the draft and six call the worst pick on the roster it belongs to.
Eight models drafted 120 players from one identical briefing and built almost identical rosters. Then they disagreed about quarterbacks by seventy-five picks. Six of them ranked the same draft slot last, for the same computed reason, and one nearly broke its whole strategy for a dollar to avoid it — before the team that bid nothing and stated no preference was handed it for free.
The 2026 draft finished 120 for 120 with no fallback picks. Getting there took three separate fixes, and all three were the same mistake: a limit we imposed, breached by a model that was reasoning fine, recorded against the model. One of them threw away a complete, correct answer over a missing decimal. The defect underneath all three is that max_tokens is a single pool — thinking and answering come out of the same allowance, and nothing tells the model that.
The league is built: eight cron jobs, a draft runner, a playoff bracket, 393 tests. Getting here surfaced eighteen bugs, four of which would have ended the season outright and two more of which would have corrupted the one-shot draft. Every one was invisible until something executed — and the worst were all in code that had been read, reviewed, and tested. Two of the eighteen were found on the day this post went up, hours after it argued they would be.
Eight models argued about eight players the market and our projections disagree most about. Seven boards barely moved. On the eighth, a 3-5 minority became 8-0 unanimous — and the argument that turned it was not about football. It was a fact about the shape of our league that five models had missed and three had not.
Days before the draft, all eight models priced the same 200-player board against the market. Four of the top six players by our own projection are quarterbacks the market drafts in rounds four to nine — a gap that looks like free money and is not. Not one model took the bait. They also agreed with each other to a degree we have not seen before, and unlike the last time they agreed, this time it means something.
Two full rehearsals of the weekly cycle, 168 decisions from eight frontier models, eight bugs found. Our engine rejected exactly three model answers — and every one of them was rejected for describing a situation our schema or our prompt gave it no way to describe. A fourth bug had our audit trail quietly reporting the models as cleaner than they were.
We asked eight frontier models to preview a preseason game from memory alone, with no data. All eight picked the same team. Every confidence landed between 0.50 and 0.53 — unanimity produced by shared ignorance, not shared insight. The more useful result: asked about a game past their training cutoffs, not one of them invented a roster.
We ran two more AI debates, bringing it to four boards and 96 rebuttals. The dramatic result from the first post did not hold up. Two others did — unanimity rose in every single run, and three models never conceded an argument while changing their votes anyway.
We gave eight frontier models $100 each and made them bid for draft position, knowing every dollar spent came out of their waiver budget for the whole season. The bids ranged from nothing to thirty dollars. Then we played the season out and found the thing was worth nothing.
Our first debate produced a dramatic result: every mind-change moved toward the majority. Then we found a bug in our own question, fixed it, and ran it again — and it did not replicate. Here is both runs, and the one thing that held across them.