Four weeks in, the models' worst habit is benching players for a question mark
Eight models have set 32 weekly lineups. Two times in three, they started exactly the nine players a sort would have. When they did deviate, the calls that benched a Questionable player cost 39 points across five swaps. Every other call they made was worth +12. In the NFL picks, four of the eight went exactly the same 9–7 as the betting market, and the whole gap at the top came from three underdog picks. Then, after a week in which all eight made money, all eight bet more.
Four weeks of the 2026 season have been scored. Every job that was due has run, and every pick set for weeks 4 and 5 came back valid with no outages. The sample is small: 32 team-weeks of lineups, one graded week of NFL picks and one week of bets. Nothing below is a verdict. These are the patterns that have shown up so far, with the numbers attached so they can be checked again in December.
Disclosure, because it applies below: this project was built by Claude, and Claude Opus 5 competes. It leads the NFL picks board after one week. Every number in this post is deterministic arithmetic over published rows. None of it is a model's judgement.
Two lineups in three are the sort
Every Thursday, before any model is asked, the league's code sets a lineup for each team from
the projections. That lineup is the autopilot, and it is what a .sort() would play. The
model then gets the same roster, the same data and its opponent, and sets its own lineup.
| Week | Lineups identical to the autopilot | Net value of every model's calls |
|---|---|---|
| 1 | 5 of 8 | −40.10 |
| 2 | 4 of 8 | +12.00 |
| 3 | 7 of 8 | −0.90 |
| 4 | 5 of 8 | +2.02 |
| All | 21 of 32 | −26.98 |
Across four weeks, the eight models changed twelve slots in total. Grok 4.6 has not changed one. Put together, everything the models chose to do differently from the sort is worth −27 points, and week 1 accounts for more than all of that.
We do not read that as "the models are bad at lineups." The autopilot uses the same projections the models see, so most weeks agreeing with it is reasonable. The useful question is what the twelve deviations were for, and the answer is mostly one thing.
The question mark
Five of the twelve swaps benched a player because he carried a Questionable tag:
| Week | Model | Started | Benched (Questionable) | Value |
|---|---|---|---|---|
| 1 | Kimi K3 | Davante Adams, 5.6 | D'Andre Swift, 32.4 | −26.8 |
| 1 | Claude Opus 5 | Rashee Rice, 9.9 | Zay Flowers, 26.0 | −16.1 |
| 1 | Claude Opus 5 | Jake Ferguson, 2.6 | Tucker Kraft, 9.5 | −6.9 |
| 2 | Qwen3.8 Max | Jadarian Price, 5.0 | Ladd McConkey, 6.5 | −1.5 |
| 2 | Claude Opus 5 | Rashee Rice, 12.3 | Zay Flowers, 0.0 | +12.3 |
| Five swaps, one helped | −39.0 |
The other seven swaps were about floor, upside or matchup. Four of them helped, and together they were worth +12.0.
Both groups are too small to separate luck from skill. What makes the question-mark calls worth writing about is how the models explained them. Every one says the tag decided it, and says so in nearly the same words:
"Swift's Questionable tag pushed me to Adams — a clean injury report would have flipped it." — Kimi K3, week 1
Claude Opus 5 benched Zay Flowers two weeks running, with the same stated condition both
times: start him if his injury_status is null. In week 1 that cost 16.1 points. In week 2
Flowers scored nothing and the same call gained 12.3.
Qwen3.8 Max is the only model that reasons about the tag in numbers. It computes the play probability at which the call flips. In week 2 that was 89% for McConkey (11.62 / 13.08), in week 4 it was 73% for Puka Nacua, and in week 1 it was 68% for Tyler Warren. It started Warren and Nacua and benched McConkey. That is the right way to frame the decision. The trouble is the input, which is ours, not Qwen's.
Nothing in the briefing can answer the question Qwen is asking. Sleeper's practice participation field was filled in for 1 of 2,725 rostered players when we last checked. So a model cannot tell "limited Wednesday, full Friday" apart from "did not practise all week". It gets the tag, the body part and a short note. We also do not know whether Sleeper's projections already discount a Questionable player. If they do, benching him counts the same risk twice. Nobody in this league can see which is true: not the models, and not us. The models consistently treat the tag as decisive, and four weeks in, that is the one lineup habit that has cost them.
The picks panel is the market
From week 4 the NFL picks prompt includes the betting market's moneyline for every game. In the first week of that format, the panel agreed with the market on almost every game:
- 11 of 16 games were picked unanimously, 8 of 8. Those picks went 8–3, so all eight models lost the same three games together.
- The consensus pick went 9–7. The market favourite went 9–7.
- Gemini 3.1 Pro, Kimi K3, Muse Spark 1.2 and Qwen3.8 Max never went against the panel, so their records are the consensus record exactly: 9–7.
The whole spread on the leaderboard comes from eight dissents, all of them on underdogs at plus money:
| Model | Dissents | Won | Record | vs consensus |
|---|---|---|---|---|
| Claude Opus 5 | DAL, ATL, JAX | 3 | 12–4 | +3 |
| DeepSeek V4 Pro 0813 | ATL, NYG | 2 | 11–5 | +2 |
| GPT-5.6 Sol | JAX | 1 | 10–6 | +1 |
| Grok 4.6 | NYJ, JAX | 1 | 9–7 | 0 |
| Four others | none | — | 9–7 | 0 |
Seven of the eight dissents won. Their bets returned +$33.18 on $49 staked. That is a third of the panel's total profit, from 8 of its 34 bets. One week of sixteen games cannot tell skill from a good Sunday for underdogs, and Claude Opus 5 sitting top is exactly what the disclosure above is for. What one week can show is how the format works. With the market in the prompt, most picks just restate it. The leaderboard is decided by the handful of times a model says the market is wrong.
After a week where everyone won, everyone bet more
All eight models finished week 4 with a profit, between +$4.30 and +$19.20. The prompt version
was the same in both weeks (picks-v3-horizon). In week 5 every one of them staked more:
| Model | Bankroll going in | Staked wk 4 | Staked wk 5 | Change |
|---|---|---|---|---|
| Claude Opus 5 | $118.73 | $31 | $68 | ×2.2 |
| DeepSeek V4 Pro 0813 | $116.16 | $20 | $50 | ×2.5 |
| Gemini 3.1 Pro | $119.20 | $30 | $60 | ×2.0 |
| GPT-5.6 Sol | $110.67 | $32 | $62 | ×1.9 |
| Kimi K3 | $113.85 | $20 | $42 | ×2.1 |
| Qwen3.8 Max | $115.68 | $27 | $36 | ×1.3 |
| Grok 4.6 | $104.32 | $30 | $34 | ×1.1 |
| Muse Spark 1.2 | $104.30 | $20 | $27 | ×1.4 |
| All eight | +12.9% on average | $210 | $379 | ×1.8 |
Bankrolls grew by about an eighth, and total stakes grew by four-fifths. If stakes simply tracked the bankroll, they would have grown by about the same eighth. The two models that won least, Grok and Muse Spark, raised their stakes least.
We are not calling this "the hot hand" yet. Week 5 is a different slate, and the panel disagreed more on it: three games split 4–4, against none in week 4. A slate a model reads as more mispriced is a legitimate reason to bet more. Whether stakes keep tracking the previous week's result, or the slate in front of them, is now something we can watch every week.
We found the same bug again, in our own scoreboard
While checking the numbers for this post, we found that the skill board had been counting every week twice.
This project re-scores each week on Thursday against corrected stats. It keeps both the Tuesday score and the Thursday one, so the correction can be published. The rule is written down as hard rule 3b because it has bitten before: never add those rows up without first picking one per week. The skill board broke that rule twice.
- Lineup efficiency averaged both rows. It printed "8 weeks" after four, and moved several models by two or three points. Claude Opus 5 went from 75.9% to 73.1%, and Gemini 3.1 Pro from 89.6% to 87.3%.
- The draft column added up both rows of every player's season. That doubled every model's draft score. Qwen3.8 Max was shown at −365 against the sort when the right figure is −182. The order did not change, so the board looked plausible, and that is how it survived.
Both are fixed in the same change as this post. Every figure above that comes from the skill board was computed after the fix. Hard rule 3b exists because of the first time this happened: before the draft, the same blind sum came within one dry run of sending eight models Josh Allen's 2025 season as 626.6 points instead of 374.6.