Orchestrated Multi-Model AI System

August 17, 2026

Auditing your own theory with a second model

The Brookhaven Cosmotron, an early particle accelerator, its ring of machinery filling the frame.

The largest single improvement to this programme’s error rate did not come from a new result or a better formalism. It came from pointing a second AI model at the existing corpus with one instruction: recompute every number from the printed formulas, and trust nothing that is merely asserted. This post is about what that found, and about the strange division of labour it revealed.

The setup

The rule is simple and worth copying verbatim: the auditor may read, but must not inherit. Any number it uses must be recomputed from the formulas in the text, from stated constants, or from its own derivations. If a paper says the value is X, the auditor’s job is to discover whether X follows from the paper’s own machinery, not to learn that X is the value. The output format is fixed: a list of confirmed results, a list of discrepancies, and a list of claims that could not be verified either way.

The model used for the second pass was deliberately not the one involved in drafting, for the same reason you do not proofread your own manuscript: shared blind spots. A different model has different reflexes about what looks right, which is exactly the property being purchased.

What it found

Three tiers of findings, in decreasing order of comfort. The comfortable tier: the mechanical spine survived. The anomaly cancellations, the dimension count, the congruence, the logistic saturation, the plateau arithmetic all confirmed to the digits, and the confirmed list is now the load-bearing part of the framework’s ledger.

The uncomfortable tier: discrepancies with a signature. A back-solved coefficient in the neutrino paper, its exact value reproducible only by starting from the answer. A phase chosen from a menu after the fit. An exponent written as if computed that was asserted. Each had the same shape: locally plausible, globally circular. No single paragraph exposes that shape; only the instruction to trace where each number came from does, because the trace terminates at the target instead of at the inputs.

The third tier was the most valuable and the most humbling: five claims that were wrong in ways no reader had noticed because the arithmetic was never done. A five-term sum copied from the literature instead of added. A formula that did not evaluate to its own printed value, off by sixty-six per cent. A decomposition that summed to the wrong total by five orders of magnitude. Each was checkable in minutes. Each had survived multiple drafts, multiple documents, and in one case a peer-review-shaped process, because everyone read the prose and nobody recomputed the line.

The division of labour

What the exercise taught me about humans and models: the models are tireless and literal, which makes them excellent at exactly the checks humans are worst at, the recomputation of a line that has been read too many times. They are also susceptible to the corpus’s own confidence, which is why the no-inheritance rule is the whole game; an auditor that has internalised your conclusions will confirm them. The human, meanwhile, is the only party who can judge which discrepancies matter, which retractions are acceptable, and what the corrected claim should say.

The honest cost: two of the second model’s own findings were wrong, one a misreading, one a recomputation error, and both were caught only because the programme’s rule is that the auditor’s claims get audited too. Audit is a loop, not an event. It ends when the discrepancies converge, not when the author is satisfied, and the corpus now carries the convergence history for anyone who wants to see how messy the middle of that process is.

The next post applies description length to theory selection, which is where Occam’s razor finally gets a definition: MDL as theory selection.

DPHmethodAI

← All writing · All topics