Orchestrated Multi-Model AI System

August 14, 2026

Status labels as a discipline

A trilobite fossil — cast and imprint in pale stone.

Every quantitative claim in this project carries one of six labels. This post is the argument that the labels are the most useful thing here, ahead of any of the results they decorate, and the argument is entirely about my own past behaviour.

The six

Theorem: proved from stated premises, with the proof checked. Benchmark: computed, reproducible, but not derived from anything deeper, and carrying no derivational authority. Conjecture: stated because it is testable, with its falsification route named. Preregistration: a rule set written before data, with no result claimed. Retracted: wrong, archived rather than deleted, with the failure mode documented. And proxy: a number that stands in for something not yet computed, labelled so nobody mistakes it for the thing.

The labels exist because of specific failures. The proxy label exists because a code-distance argument got treated as physics for two paper generations. The benchmark label exists because a trace ratio that turned out to be a standard result was cited as a framework theorem. The retracted label exists because five papers died and the corpses were more useful in public than in a drawer.

What the labels actually do

The mechanics are less interesting than the effect. A label is a promise about which conversations a number is allowed to join. A benchmark can be used for illustration and consistency checks; it cannot be load-bearing in a derivation. A conjecture can motivate a test; it cannot be cited as support. The dependency rule the programme now maintains makes this enforced rather than aspirational: every paper lists which prior results it consumes and their labels, and building on a benchmark requires saying so in the text.

The deeper effect is on writing speed. Before the labels, every draft was a negotiation with my own enthusiasm: is this result solid enough to phrase as a derivation, probably, move on. Now the negotiation happens once, at audit time, and the draft inherits a verdict. That sounds bureaucratic. In practice it is the difference between a correction being a one-line banner and a correction being a rewrite, and after fourteen corrections, one line is a gift.

The part that generalises

You do not need my ontology to use this. Anyone doing theory of any kind, economics, psychology, engineering, has claims that are proved, claims that are computed, claims that are hoped, and claims that are placeholders. The failure mode is not having wrong claims. It is having claims whose strength drifts upward as they get repeated, a computed number becoming established becoming derived in the telling, with nobody able to say when the drift happened. Labels freeze the strength at the moment of audit, and freezing is the whole trick. The rest of this method series is the machinery: what no-gos are for, how preregistration works, and how a second model earns its keep as the thing that catches your drift.

The next post is the case for the doors that shut, and for why three of the strongest results here are no-gos: what a no-go buys.

DPHmethod

← All writing · All topics