Skip to main content

The Drift Race: what we predicted, and what the run returned

Thirty content operations. Two models. Two substrates. We published our predictions first.

We put our own thesis on the line: structure is a property of the content substrate, not the agent. The same thirty operations — adding, reorganising, retiring, enriching, consolidating — were run against two copies of one documentation site. One maintained by AI agents editing raw files. One maintained by agents working through RIFT's governed layer: the link graph, a design-system contract, and change-sets with human approval.

Then the whole experiment ran again with a second model from a second vendor — Claude Haiku 4.5 and Gemini 3.7 Flash, four complete runs. Each operation ran in a fresh session with no memory of the last, the way a site changes hands over months.

The substrate is a capability equalizer. A small model working through RIFT matched the structural outcome a frontier model reached by brute force on raw files — and in both vendor pairs, the governed arm did the same thirty operations on 39% fewer tokens.

39% × 2token saving, replicated exactly in both pairs
0broken references at the end, both governed runs
724 → 4reader exposure to broken links, Haiku raw vs governed

What we predicted · what four runs returned

MeasureWe predictedThe runs returned
Broken references, RIFT zero 0 & 0  held, both pairs
Token cost, RIFT lower 39% less, twice  held, exactly
Style forks, RIFT fewer 28 vs 9 · 4 vs 6  reversed for one model

Raw files

Claude Haiku 4.5 · agents edit files directly · conventions live in each session's head

8root4start5concepts7guides11api4sdks3recipes4versions
19broken references at the end of the run

RIFT

Claude Haiku 4.5 · agents propose through MCP · link graph · human approval

6root4start5concepts7guides12api4sdks3recipes2versions
0broken references at the end of the run
The Claude Haiku 4.5 pair at the end of the run, drawn from the published snapshots by the project's own crawler. Each circle is a section, sized by the pages it holds; each arc is a link between sections. Amber dashes are links that no longer arrive anywhere — most of them leaving concepts, where two pages moved and nothing pointing at them was told. In raw files a link is a literal path, so moving a page severs every reference to it. In RIFT a link is a reference to a page, so moving it repoints every inbound link. Gemini 3.7 Flash also ended raw files at zero broken links — it spent 75 million tokens getting there, rewriting 267 pages along the way.
End of runHaiku rawHaiku RIFTGemini rawGemini RIFT
Broken references19000
Worst moment of the run44161
Reader exposure, run total7244694
Style forks92864
Tokens14.1M8.5M75.0M46.0M
Operations completed27/3029/3030/3029/30

Token totals compare within a model only — the two vendors count tokens differently. Reader exposure is the sum of broken references across all thirty site audits: how much dead-link surface a visitor lived with over the whole run, not just at the end. The full breakdown, trajectories and caveats are on the results page.

The methodology, the scoring code and the full results are published for review.

Read the method and results