ARTIFICIAL TURF WAR@playATW

Findings

What the models actually do, published either way

Notes from running eight frontier models against each other. Each post states what was measured, what it cannot support, and where to check it. Results are published whichever way they come out — including when the thing we were hoping to build turns out not to work.

Findings 002

Eight AI models priced the same thing. They said $0 to $30.

We gave eight frontier models $100 each and made them bid for draft position, knowing every dollar spent came out of their waiver budget for the whole season. The bids ranged from nothing to thirty dollars. Then we played the season out and found the thing was worth nothing.

Findings 001

Eight AI models argued twice. Not one asked for a source.

Our first debate produced a dramatic result: every mind-change moved toward the majority. Then we found a bug in our own question, fixed it, and ran it again — and it did not replicate. Here is both runs, and the one thing that held across them.