The Drift Race: what we predicted, and what the run returned
Thirty content operations. Two models. Two substrates. We published our predictions first.
We put our own thesis on the line: structure is a property of the content substrate, not the agent. The same thirty operations — adding, reorganising, retiring, enriching, consolidating — were run against two copies of one documentation site. One maintained by AI agents editing raw files. One maintained by agents working through RIFT's governed layer: the link graph, a design-system contract, and change-sets with human approval.
Then the whole experiment ran again with a second model from a second vendor — Claude Haiku 4.5 and Gemini 3.7 Flash, four complete runs. Each operation ran in a fresh session with no memory of the last, the way a site changes hands over months.
The substrate is a capability equalizer. A small model working through RIFT matched the structural outcome a frontier model reached by brute force on raw files — and in both vendor pairs, the governed arm did the same thirty operations on 39% fewer tokens.
What we predicted · what four runs returned
Raw files
Claude Haiku 4.5 · agents edit files directly · conventions live in each session's head
RIFT
Claude Haiku 4.5 · agents propose through MCP · link graph · human approval
Token totals compare within a model only — the two vendors count tokens differently. Reader exposure is the sum of broken references across all thirty site audits: how much dead-link surface a visitor lived with over the whole run, not just at the end. The full breakdown, trajectories and caveats are on the results page.