Orchestrated Multi-Model AI System

September 6, 2026

Novelty is a compression measurement

Margaret Hamilton standing beside a stack of Apollo guidance-computer code listings as tall as herself.

Before the plateau theorems can mean anything, the word novelty has to be rebuilt from the ground up, because the everyday meaning, something new, is too loose to converge. This post is the definition post, and the definition is one sentence: novelty is compression. Everything surprising in the computational series follows from taking that sentence literally.

The definition

Consider a data stream and a description language. Before you understand the data, describing it costs roughly its raw length. After you find the pattern, the cost drops: state the pattern, then the deviations. The drop is the structure you have learned, made into a number. Define novelty as exactly that drop: the reduction in description length won by absorbing new structure into your model class.

This definition has three properties that the everyday word lacks. It is quantitative: drops are measured in bits, comparable across fields and across time. It is operational: two observers can disagree about whether something was novel only by disagreeing about the encoding, which makes the disagreement explicit and auditable rather than aesthetic. And it is directional: a random string admits no compression, so randomness has zero novelty no matter how surprising it feels. Surprise is not novelty. A lottery result is surprising and compresses nothing. A new symmetry in a physical law is surprising and compresses enormously, because it reorganises everything downstream of it.

What the definition buys immediately

First, the Discovery Plateau becomes a theorem instead of a lament. If novelty is compression, and compression has a floor, the structure of any fixed data stream under any fixed language is finite, and its mining is a convergent series. The plateau posts walk the proof; the definition is what makes a proof possible at all.

Second, the distinction between novelty and data becomes sharp. Data without structure compresses nothing: a bigger telescope staring at a blank sky produces a longer raw string and zero novelty. Data with new structure compresses enormously. This is why the multiverse programme is obsessed with heterogeneous universes rather than merely more universes: by this definition, identical copies of our branch are the blank sky, infinite raw length, zero drops.

Third, the value of instruments becomes computable. An instrument is valuable exactly to the degree that the data it produces has structure outside the current model class, because that is the only data whose absorption yields drops. The framework’s instrument programme, the interferometers, the code-patch experiments, is chosen by this criterion, not by sensitivity league tables.

What the definition costs

Honesty requires the price. Compression-novelty is language-relative: the same observation can be a huge drop in one framework’s encoding and nothing in another’s, and the definition does not select which language is right. It is also retroactive: a drop is only measurable after the structure is absorbed, so the definition quantifies discovery but cannot rank proposals before they are made, which is where judgement still lives. And it is silent about significance: a compression of one bit that unlocks a thousand downstream drops is worth more than a hundred-bit drop that closes no doors, and the definition counts the bits, not the doors.

The framework’s answer to these costs is not to patch the definition but to pair it: the drop measure for accounting, human judgement for selection, and the status-label discipline for keeping the two from blurring. The definition’s job is not to replace taste. It is to make sure that when taste claims a discovery, a number exists that the claim has to move.

DPHcomputation

← All writing · All topics