Minimum description length as a theory-selection principle
Occam’s razor says prefer the simpler theory. It has been said for seven hundred years without anyone agreeing on what simpler means, and the definition matters, because under most definitions the razor is false: an arbitrarily simple theory can be arbitrarily wrong. Minimum description length is the version of the razor that is actually true, because it prices both sides of the transaction. This post is the method-series companion to the concrete walkthrough, which built the smallest example; this one is about using the principle on theories rather than data.
The two-part invoice
The principle: a theory’s total cost is the cost of stating the theory plus the cost of the data given the theory. Both terms are in bits, both are mandatory, and the winner is the minimum of the sum. The first term is the invoice for structure: every postulate, parameter, and mechanism is a line item. The second term is what you were trying to explain all along: the residual surprise in the data once the theory has done its work.
The failure modes of ordinary theorising are each a disease of one term. A baroque theory with twenty parameters that fits beautifully has bought its fit with a huge first term and hid the invoice in footnotes. An elegant one-parameter theory that misses the data has refused to pay the second term. Numerology, the disease this archive knows best, is a special case: it drives the second term to zero by construction, having tuned the first term until the residual vanished, and then reports only the residual.
What it changes in practice
Three comparisons from the programme, priced under the invoice. The power-law Yukawa ansatz against the modular-form mechanism: the ansatz has fewer effective parameters, so its invoice is smaller, but its residuals are enormous, missing the observed hierarchy by a factor of sixty, which is many bits of surprise per observable. The modular mechanism costs more to state, weights, a modulus, a level, and pays it back by compressing the ratios to within a per cent. The modular mechanism wins the sum.
The horizon-numerology claim against the derived capacity: the numerology cost nothing to state, one exponent, but its residual was fatal once the decomposition was summed honestly, five orders of magnitude of unexplained remainder. The derived capacity, S over pi squared, costs a derivation to state and compresses everything the numerology touched, plus several things it could not reach. The derived result wins, and the numerology is archived.
The unification document against a stack of independent hypotheses: this is the live comparison, and the honest entry is that the stack currently wins on the first term, the framework costs more to state, and wins only where its shared structure compresses genuinely different measurements. Where it does not, the document is obliged to say so, which is what the status labels are for.
The limit of the principle
MDL is relative to a description language, and the language is doing quiet work in every comparison above. Change the encoding and the sums shift; the framework’s papers therefore state their accounting, which parameters cost what, before comparing, so that disagreement about the verdict becomes disagreement about the code, which is a productive argument instead of a vague one. And the principle selects among stated theories; it does not generate them. It is a judge, not a muse. What it buys is exactly what a research programme needs mid-flight: a rule for killing theories that cannot be argued with, because the arithmetic is on the table, and the invoice, unlike the rhetoric, always adds up.