The Optimum That Depended on the Ruler

A multilevel-governance allocation hypothesis, and how twenty simulations took it apart

Björn Kenneth Holmström August 2026 20 min read

Most of what the Governance-as-Engineering series records is what survived scrutiny. This is a note about a long chain of hypotheses that mostly didn't — a line of computational exploration that started from a metaphor in a multi-model dialogue, produced a series of increasingly specific claims, and ended with its most attractive result dissolving under a robustness check I ran only because the discipline of the series demanded I try. I am writing it up for the same reason the retained-variety note exists: a framework that only reports its confirmations is not doing the thing it claims to value, and the shape of the failure here is more instructive than a clean win would have been.

The idea, and why it was attractive

The starting metaphor was that Westphalian sovereignty, centralized authority, and GGF-style subsidiarity are not three competing ideologies but three different engineering strategies for managing uncertainty: separation, aggregation, and appropriate scale. That is a governance-theory claim dressed as an engineering one, and the obvious move — the one the series' own method insists on — was to stop asserting it and try to break it computationally instead.

The first fifteen or so iterations of that attempt were organized around fear. Political institutions, the reasoning went, are technologies for managing uncertainty and threat perception; a fear equation driven by uncertainty, stakes, uncontrollability, and agency should differentiate the three architectures cleanly, and something like Paper VI's variety gap should fall out as a side effect. This produced a real result almost immediately, and a worrying one: a subsidiarity-style architecture reached near-zero fear while resolving fewer problems than the more anxious, boundary-obsessed Westphalian one. That looked less like evidence for subsidiarity and more like a controller that had learned to stop watching the fire.

The build-up, part one: fear

Four iterations were spent trying to save the fear framework by fixing what looked like a bookkeeping problem rather than a structural one. Agency was decoupled from free recovery and tied to actual resolution; objective and perceived risk were split into separate state variables; a false-security metric — fear low and risk high, simultaneously — was added explicitly rather than left to be noticed by eye. Once measured honestly, the false-security rate for the subsidiarity architecture reached 0.34–0.47 under sustained hostility or deception, on par with the failure mode the theory should have worried about in the first place.

The natural next question, in the series' own idiom, was whether this was a controller problem or a structural one. I built an "epistemic immune system" — trust that responds to evidence, with delayed discovery of deception so that betrayal costs something once it is found out — and it barely moved the number. Then I ran the decisive check: four qualitatively different calibration laws (a fast exponential update, a slow-integrating one, a bistable thresholded one, and a Bayesian evidence-gated one), holding everything else in the model fixed. False security under sustained hostility ranged only 0.434–0.469 across all four. [R within model] That was the falsification that mattered: the failure mode was not a badly tuned dial. It was upstream of calibration entirely, which meant fear and trust were not going to be the right primitives to keep building on.

The build-up, part two: variety

So I discarded them and rebuilt around Ashby's Law instead. The environment became a set of independent disturbance "channels"; the three architectures were rebuilt to differ only in topology — same agents, same capability, same coordination bandwidth where a coordination layer existed at all — never in raw power. Westphalian: full local parallelism, but a single joint decision slot for anything cross-boundary, regardless of how many distinct cross-boundary problems arrive in the same step. Centralized: local handling below a threshold, everything else through one hub with a fixed bandwidth. Subsidiarity: the same local parallelism as Westphalian, plus a bandwidth-limited coordination layer for whatever crosses a causal-jurisdiction threshold.

Two real implementation bugs were caught and fixed before I trusted anything downstream of them — one in how multi-agent events were counted toward a resolution rate, one in an escalation formula that diluted toward zero as variety rose, making the subsidiarity architecture's coordination layer untestable by construction rather than genuinely robust. After both fixes, a two-dimensional sweep across environmental variety and cross-boundary coupling (κ) located a real coordination-overload boundary for every architecture, including subsidiarity's — it simply sat much further out than the other two. [R within model]

An audit followed, because the result looked too convenient: was subsidiarity's advantage real regulation, or were unresolved local consequences quietly leaving the scoreboard? I added a leakage metric — independent of which tier processed an event, what fraction of its magnitude went entirely unaddressed — computed the same way for all three architectures. Subsidiarity's leakage stayed below the other two even at its own failure corner. [R within model] A twelve-hundred-step persistence test at that corner found no delayed reckoning either: agency collapsed as a leading indicator, well before — and mostly without — matching material harm ever showing up. That decoupling of institutional confidence from material regulation is, I think, the single most interesting and least expected finding in the whole arc, and it survived every check I threw at it.

The narrowing: what didn't help

Four more iterations spent trying to complete an intuitive causal chain — environment, observation, alignment, allocation, coordination, action — mostly produced honest nulls, and I want to preserve them explicitly rather than let them get absorbed into the surviving story:

  • Observer diversity alone did nothing. Correlating disturbance channels via a verified copula (audited the same way a controller's own noise floor would be) moved effective variety by a factor of four and moved no outcome at all, because none of the architectures had any mechanism for recognizing that recurring channels were the same underlying problem. [R within model]
  • A parallel observer/detection layer, built with the same discipline and equalized across all three architectures, also did nothing across a twenty-fold range of effective observer diversity. [R within model]
  • Alignment — whether a detection signal correctly identified what response was needed, as distinct from whether anything was detected at all — mattered a great deal, with a clean monotonic effect on objective risk for all three architectures. Observer diversity could be converted into that alignment, but only where a genuine multi-input integration hub existed to do the converting; an architecture with no such hub (Westphalian, by its defining structural feature) got exactly zero benefit from more redundant observers, which is what the control condition was for. [R within model]
  • Integration topology — flat, all-at-once aggregation versus cheaper, two-stage clustered aggregation, under a shared and genuinely finite budget — mattered only for the architecture whose coordination hub was already congested. For subsidiarity, whose whole structural point is to keep that hub uncongested, the choice of aggregation topology was close to irrelevant. [R within model]

None of these are embarrassing. Paper X's observer-correlation result and Paper XXVII's requisite-alignment result predict something close to this pattern; what the simulations added was a demonstration that quantity and relevance are dissociable in an architecture that has no mechanism to exploit either — which is a sharper, more falsifiable version of "more information doesn't automatically help" than the slogan usually gets.

Where it looked like a paper

The next several iterations built an adaptive allocator: instead of a fixed magnitude threshold deciding whether a cross-boundary disturbance escalates, a real-time comparison of expected cost (resolution probability times magnitude, plus an explicit administrative overhead for escalating at all) decided case by case. The first sweep looked unpromising — adaptive tracked almost exactly on top of "always handle it locally," and the fixed-threshold rule never beat that extreme either. A single-variable Monte Carlo, stripped down to just the two response curves and the routing inequality they imply, showed that an interior regime — selective escalation genuinely beating both pure extremes — does exist mathematically for the shape of curves the full simulation actually used. The reconciliation was that the full simulation's typical disturbance magnitude, produced by splitting a fixed budget across many active channels, rarely reached the region where the inequality favored escalation at all. That pointed at a genuinely new variable: not variety as channel count, but magnitude concentration — the same total disturbance mass, distributed across few large events or many small ones.

A preregistered sweep confirmed this in the full model, cleanly, on both predictions I committed to in advance: the fraction of disturbance mass an adaptive allocator chose to escalate fell monotonically as variety rose (holding total mass fixed), and adaptive's advantage over "always local" was roughly ten times larger at the concentrated end than at the diffuse end I had been testing all along. [R within model] A genuine interior regime appeared — not at either extreme, but in a bounded band of intermediate concentration. That was, by a wide margin, the most interesting and best-supported result in the allocation branch.

The turn

The obvious next question, and the one the series' method requires asking of any result this convenient, was whether the effect tracked concentration itself or was still secretly riding on channel count. I built an orthogonalizing factorial: a Dirichlet shape parameter that lets the evenness of a fixed budget's split be varied independently of how many channels are nominally active, audited against the standard concentration index to confirm it did what it claimed. Holding channel count fixed at the diffuse end and sweeping concentration across a wide range reproduced none of the interior regime. Holding concentration fixed and sweeping channel count reproduced it exactly as before. [R within model] The mechanism, once I looked for it directly, was mundane: coordination bandwidth in every version of the model up to that point was consumed per event, not per unit of magnitude, so what mattered was how many things were competing for a fixed number of slots, not how the total disturbance mass happened to be distributed among them.

That was survivable — a narrower, less exotic claim ("coordination burden scales with the count of concurrently active distinctions") is still a claim. But it rested on one modeling choice — event-count-based bandwidth — that I had made in the first variety-based iteration and never examined since. So I built the equally defensible alternative: a coordination layer with a fixed magnitude budget instead of a fixed event count, calibrated to admit comparable average workload at the reference point where the interior regime had been most robust. I reran the exact orthogonalizing factorial under both accounting rules.

Under magnitude-based accounting, the interior regime did not narrow. It disappeared completely — every cell, both ways of varying the parameters, at both coupling levels tested. [R within model] The entire allocation-frontier result, the best and most carefully checked finding of the second half of this work, turned out to be conditional on a specific and previously unexamined convention for what "coordination capacity" means. Change the ruler, and the optimum it had located was no longer there.

What was actually there

Removing the headline claim did not empty the file. Several findings survived every adversarial test I ran against them, including the one that eventually broke the allocation result:

Subsidiarity reduces coordination-layer burden, but the benefit is conditional, not structural. A four-cell factorial — filtering on or off, crossed with genuine local capacity or degraded catch-all treatment for whatever is filtered — showed that filtering without capacity at the receiving level (0.496) was worse than no filtering at all (0.415). Only filtering with real capacity (0.216) improved on the baseline. [R within model] This is close to a formal argument against symbolic or unfunded devolution, and it is the one result in this whole arc I would be most comfortable calling a design principle rather than a curiosity: delegation without capacity transfer is not a milder version of good subsidiarity, it is actively worse than not delegating.

Cross-boundary coupling, not raw variety, is what eventually overwhelms every architecture — and the three fail in qualitatively different ways under the same stress. Westphalian retains full agency while its effective regulation quietly fails (it still feels sovereign). Centralized's agency and material regulation collapse together, a clean overload signature. Subsidiarity's agency collapses first, as a leading indicator, largely decoupled from whether unresolved harm ever follows at the tested horizon. [R within model] None of these are better or worse in the abstract; they are different failure signatures, which is a more useful thing to know than a ranking.

The match-rate identity is a piece of arithmetic, not an emergent property, and worth stating because it disciplines every claim built on top of it: the fraction of coordination-relevant events that get a differentiated response is provably min(1, bandwidth / load) for every architecture that uses this admission mechanism, verified to machine precision. [R within model] Any apparent difference in "coordination intelligence" between architectures in this model is actually a difference in load reaching that stage, not in what happens once it arrives.

The boundary, stated honestly

It would be easy to read the allocation branch's collapse as evidence that adaptive, scale-sensitive allocation is not a real phenomenon. That is not what the model shows, and I want to be as careful about overshooting here as the retained-variety note was about its own negative result. What the magnitude-accounting test demonstrates is that this specific interior regime, in this specific model, depended on an unexamined convention. It does not show that no principled account of coordination capacity would recover something like it — only that "count-based" and "magnitude-based" are both defensible on their face and give qualitatively different answers, which means the modeler is obligated to say which one they mean and why, rather than treating "coordination capacity" as a self-evident quantity the way I initially did. A next attempt would need to motivate the capacity model from something outside the routing question itself — decision-maker attention, administrative throughput, genuine communication bandwidth — rather than choosing whichever accounting is convenient to code.

The surviving findings carry their own scope limits too. All of them are [R within model]: exact for the stated equations, seeds, and parameter envelopes, and claiming nothing directly about real institutions without the further, unavoidably interpretive step the series tags [IP]. None of this was preregistered in the disciplined sense Papers XIX–XXVII use — committed thresholds, independently regenerated model instances, a blind adjudication of what counts as confirmation. Several of the later sweeps in this arc did preregister a specific prediction before running, which is a meaningfully different and better practice than the earlier iterations, but it is still a single author's single model family, not the multi-observer protocol the series' own Paper X result says is necessary before trusting agreement.

Why this is worth recording

What twenty iterations actually taught me is not a law. It is a list of things that turned out not to be the same thing, each time I was tempted to treat them as interchangeable: fear is not risk; trust is not calibration; variety is not load; observer count is not usable information; information is not alignment; delegation is not capacity; and — the one that cost the headline result — coordination capacity is not a uniquely defined resource. That last one is the sharpest, because "capacity" sounds like the kind of quantity a model just has, not a choice the modeler makes. It isn't. What a governance architecture experiences as overload depends on what its capacity is assumed to be charged against, and two equally reasonable accountings gave opposite verdicts about whether an adaptive allocation policy was worth building at all.

The Westphalian/centralized/subsidiarity comparison that started this is, in the end, still standing in a narrower form than I began with: subsidiarity can genuinely reduce the variety that reaches a higher coordination layer, but only when the layer it's diverted to actually has something to meet it with, and the three architectures fail differently under stress rather than one simply failing less. Everything more specific than that — the shape of an interior allocation optimum, the role of magnitude concentration, the value of integration topology — turned out to be either absent, conditional, or dependent on a modeling choice I hadn't examined closely enough to defend. Locating exactly which parts of an attractive result survive contact with their own model's alternative specifications, rather than assuming the first specification that compiles is the right one, is the discipline this line of work leaves behind, whether or not it ever becomes a numbered paper.

Share this

GitHub Discord E-post RSS Feed

Built with open source and respect for your privacy. No trackers. This is my personal hub for organizing work I hope will outlive me. All frameworks and writings are offered to the commons under open licenses.

© 2026 Björn Kenneth Holmström. Content licensed under CC BY-SA 4.0, code under MIT.