The plateau theorem
The plateau theorem is the result the whole computational side of the framework stands on, and this post gives its proof in full, with no physics and no assumptions beyond arithmetic. It is the rare theorem in this archive that would survive the framework being entirely wrong, and it is worth reading slowly for that reason alone.
Setup
Fix a model class M: a set of models, each describable in some formal language, with description length L of m for model m. Fix data D. The two-part description length of D under model m is the model cost, L of m, plus the residual cost, which is the length of encoding D’s deviation from m’s predictions, call it R of D given m. The MDL-optimal model is the minimiser of the sum. Define the novelty delivered by adopting model m over the previous model as the drop in total description length. The question: as you enrich M, does the total keep dropping?
The proof
The argument has three steps, each a few lines.
Step one, monotonicity: enlarging the class cannot increase the minimum. The new class contains the old one, so the minimum over the bigger set is at most the minimum over the smaller. Discovery, in the drop measure, is therefore non-negative as the class grows. So far, unlimited progress seems possible.
Step two, the convergence: the total description length is bounded below by the data’s essential content, which no encoding can go below. Formally, the sum L of m plus R of D given m is at least K of D, the shortest possible description, since the two parts together are a description of D. A monotone non-increasing sequence bounded below converges. Therefore the total description length converges, and the drops, which are the sequence’s decrements, form a summable series. The total future discovery budget is finite. That is the plateau.
Step three, the floor: if the true generating process of D is not in M or any enrichment you will use, then the residuals at the optimum are the true noise plus the model class’s structural error, and the structural error is incompressible within M. The minimum total length sits strictly above the floor set by the true structure, and no amount of work within M moves it. The rate of discovery goes to zero either way, but the terminal position differs, and that difference is measurable: it is the residual cost at the floor, and the walkthrough series showed how to compute it on a concrete example.
What the theorem does and does not say
It does not say science ends. It says science within a fixed description language ends, and the theorem’s escape routes are exactly two: enlarge the class, which is inventing new theory, or get new data whose structure the class cannot express, which is inventing new instruments. Both reset the sequence. Both are expensive, and the theorem’s corollary, that resets are the only engine of discovery, is the framework’s answer to where future novelty comes from.
It also does not say when the plateau arrives. The convergence proof gives no rate; the rate is empirical, and Varney’s law, derived in its own walkthrough, is the candidate answer for the domains where data is combinatorial. What the theorem gives is the shape’s inevitability, and the derivation gives the shape’s form, and the two together are the strongest result in this archive that does not depend on any contested physics.