3. The minimal model: factorization from bounded prediction
3.1 Design
The claim of §2 — that bounded capacity and a temporal-prediction objective jointly suffice to produce a factorization — is the kind of claim a minimal model can settle in the affirmative and can never settle in the general negative. What follows demonstrates the mechanism in one environment across a capacity sweep; §7 states plainly what a single environment cannot license. The demonstration is worth making concretely because the manner in which capacity produces factorization turns out to carry the paper's load: it is not that a starved system sees a blurred version of everything, but that it sees a coherent part of the world and is blind to the rest, and which part it keeps is, at the margin, not determined by the world at all.
The environment is a dot moving in a two-dimensional box with reflecting walls. Its latent causal state is four scalars — two positions (x,y) and two velocities (vx,vy) — and the environment is symmetric under exchange of the two spatial axes by construction: nothing in the dynamics, the rendering, or the noise distinguishes x from y. The system never observes this state. It observes a 16×16 pixel rendering with additive noise, and its only objective is to predict future frames' dot positions at offsets {5,10,20} from a short history of past frames. A GRU with hidden width h is the bottleneck; h is swept over {2,4,8,16}, and a 3-slot condition h=3 is added for the tie-break test of §3.4. After training, linear probes regress each true latent scalar on the final hidden state, on held-out data. Linear probes are a deliberate choice, not a convenience: the registered claims are about linearly decodable structure, the minimum standard for saying a variable is represented rather than merely recoverable in principle. A nonlinear probe would recover more and prove less.
Two registered runs stand behind the results. The first (seeds 0–19, the full sweep) tested emergence, an axis-based operationalization of structured blindness, and a diminishing-abstraction prediction. The second (seeds 20–39, h∈{2,3}) tested a revised, type-based operationalization of structured blindness on fresh seeds after the first run falsified the axis version. Both preregistrations, with committed thresholds and nulls, are in the supplementary materials; the sequence is reported honestly in §3.5 because the correction is part of the result, not an embarrassment to be smoothed over.
3.2 Emergence [R within the model]
A factorization appears, unsupervised. At h=8 the hidden state linearly encodes all four latent variables well enough to be called a representation of them — median across seeds R2=0.94 for each position and 0.63–0.68 for the velocities — and by h=16 the recovery is clean on every variable (medians 0.76–0.98). The system was told only to predict pixels; it built, inside its hidden state, the position–velocity coordinate system a physicist would have chosen. This is the affirmative half of §2's claim: temporal prediction under a bottleneck is sufficient to make the environment's causal variables exist, as linearly addressable quantities, for a system that was never given them.
Positions emerge first and everywhere. Even at h=2 — one scalar of capacity per spatial dimension, nowhere near enough for the full state — the positions are recovered at median R2≈0.86 while the velocities collapse to near zero. Capacity does not buy a uniformly degraded version of the whole state; it buys the whole of some variables and none of others. That observation is the subject of §3.3.
3.3 Structured blindness [R within the model]
The central result is about the shape of failure under starvation, and it is where the first registered prediction failed and taught us something.
The first run operationalized "structured blindness" as spatial-axis asymmetry: a starved system, we predicted, would keep one spatial axis and drop the other, so that the pilot run's apparent axis-collapse would recur in most seeds. It did not. Only 3 of 20 seeds showed axis asymmetry; the registered prediction failed as stated, and by the preregistration's own rules the pilot was thereby reclassified as a run that had landed in a minority outcome. But the failure was not uniform degradation — the committed null. Under a different cut, the blindness was total and structured in every seed. The cut is not spatial axis but variable type: position versus velocity. In 17 of those 20 seeds the starved system keeps both positions and discards both velocities.
The second run registered this type-based claim and tested it on seeds 20–39. It held. Type-structured blindness in 17/20; and — the load-bearing result — no uniform blur in any seed, across both runs combined, 40 of 40: in these runs, capacity starvation did not once degrade the four variables evenly. It sacrificed a coherent causal subspace whole. What the starved predictor becomes is a static observer: it represents where the dot is and has almost no representation of where it is going, and its prediction of the future is, in effect, the present held still. Velocities are the lower-value variables per unit of hidden capacity under a short-horizon prediction loss — the position at the next few offsets is mostly given by the position now — and so they are the subspace that goes. This is not a perceptual limitation in the ordinary sense. The system is not seeing a dim version of velocity; it has no velocity coordinate at all.
Which subspace is sacrificed is a function of the objective, and the short horizon is doing visible work here. Under a long-horizon objective, where the future has drifted far from the present, velocity becomes the higher-value variable and the blindness should invert — the system would keep the derivative and lose the instantaneous map. What the result claims survives across objectives is the shape of the failure, whole-subspace sacrifice rather than uniform blur; which particular subspace goes is objective-dependent, and §8 records this as a limit on the generality of the specific finding.
The variety gap of Paper VI is this result at institutional scale. An institution at its representational limit does not perceive a faint, evenly-attenuated copy of its environment; it maintains a coherent partial model — the variables that most reduce its prediction error — and is structurally blind to the rest, not blurred across it. The institutional analogue is a ministry, agency, or metric regime that retains the variables its reporting environment most rewards while losing the ones that would reveal motion, instability, or delayed consequence. The starved predictor keeps the map and loses the derivative: it knows the state of the world and not its motion.
A minority basin persists and is worth naming rather than hiding. The axis-mode outcome — keep one spatial axis with its velocity, drop the other axis entirely — recurred in exactly 3/20 seeds in both runs. A reproducible ∼15% minority across independent seed batches is a second attractor, not sampling noise. The two solutions are loss-ranked (the static-observer solution attains strictly lower validation loss than any axis-mode solution in the first run's data), which is why most initializations reach it and a stable minority do not. That two qualitatively different factorizations of the same environment are reachable under identical constraints, separated by a small loss gap, is the first appearance in this paper of the non-uniqueness that §4 treats in general — arriving here spontaneously, inside single training runs.
3.4 Discrete, symmetry-broken selection [R within the model]
If a starved system keeps positions and drops velocities, the natural next question is what happens at the margin — when capacity is increased by roughly one scalar above the position-only regime. Does the system add half of each velocity, or one velocity whole? The h=3 condition tests this, and the answer is sharp. It adds one velocity whole. Across all 20 seeds, positions are recovered (medians R2=0.91) and one velocity is recovered at R2≈0.6–0.7 while the other sits at essentially zero — a bimodal split with nothing in between. Capacity at the margin is not spread; it is committed.
And which velocity is committed to appears to be decided by nothing in the environment. The favored velocity split 11 to 9 between vx and vy across seeds — a coin flip. Because the environment is axis-symmetric by construction, and the ensemble split shows no directional bias, there is no evidence that the world selects vx over vy; the choice is made by the initialization and the training trajectory, and it is made discretely. This is symmetry breaking in the precise sense: a symmetric problem, an asymmetric solution, and an ensemble that restores the symmetry only in aggregate. The non-uniqueness of §4 is not merely that many factorizations could be chosen; it is that the choice is forced, sharp, and — where the world is symmetric — arbitrary.
The minimal model therefore gives three results, and they are the three ingredients the rest of the paper needs. First, prediction under a bottleneck produces the environment's latent causal variables without supervision — emergence. Second, when capacity is too small, failure is structured: the system loses whole variables rather than a little of everything — structured blindness. Third, at the margin the retained variable can be selected arbitrarily among symmetric alternatives — non-unique factorization. Emergence is the subject of §2's sufficiency claim; structured blindness grounds the variety gap; non-uniqueness is what §4 and §5 develop into the claim that coordination is selection within a privileged class rather than discovery of a unique truth.
3.5 What the sequence shows about method
The first registered prediction for §3.3 failed, and the paper is stronger for reporting it rather than for having guessed right. A single pilot run had shown axis-collapse; had we published it as the result, we would have reported a 15% attractor as the phenomenon. The multi-seed distribution corrected that, revealed the actual (type) structure, and a second registered run confirmed the correction on fresh seeds. This is the series' distributions-not-trajectories discipline (Paper IX) doing exactly what it exists to do, and it mirrors Paper XVIII's arc, where a registered early-warning index failed and forced a revision. The confirmed claims of this section — emergence, whole-subspace blindness, discrete symmetry-broken selection — rest on 40 seeds across two preregistrations. The one prediction that failed (diminishing abstraction: that capacity beyond the task quotient would buy pixel detail but not cleaner latent structure) is not carried into the paper's claims; h=16 improved velocity recovery over h=8, so the task quotient was not saturated at h=8 and the plateau claim is simply unsupported at this scale.