Every measurement this estate signs becomes a card under 3 KB — score, n, interval, provenance. Cards supersede; the newest card is the current truth; the older ones stay auditable. SOVOS is the surface that reads that timeline from the top — 13 greenfield axes, the router's measured routes, and the algebra that guarantees the read always converges to the same state, in any order, on any machine.
Gold border = quotable (n ≥ 30 with interval). Cyan = measured bank, runs pending or interval not yet quotable. Red = correctly silent — an empty lane published as empty. Scores are macro-F1 with Wilson 95% intervals where carried.
"Route each task to the family that owns it" stopped being a slogan when the second route was measured. The small sovereign classifier owns injection detection; the frontier owns hard reasoning. The router picks neither favourite — it reads the card and sends the task to the measured winner.
| task family | sovereign small | frontier comparator | separation | route → |
|---|---|---|---|---|
| prompt-injection detection | recall 1.000 [0.938, 1.000] 200 KB TF-IDF · 0.065 ms/text · CPU |
gpt-4o-mini 0.793 [0.672, 0.877] | disjoint intervals McNemar p = 4.9×10⁻⁴ (12–0) |
sovereign small (n = 80) |
| hard reasoning (arithmetic) | 0/11 — no arithmetic ability | 11/11 = 1.000 [0.741, 1.000] | total separation | frontier (n = 11) |
The injection separation holds on one corpus (deepset prompt-injection, held out; 0 of 546 training rows overlap). Wider corpora are the next measurement, not an assumption. Toxicity, semantic-novelty and containment lanes are published as empty — no figure without a run.
Tested against the real 12-card board, not argued from analogy. The card knowledge-base is a CRDT-like commutative monoid: the same formal property distributed databases use to guarantee consistency without a coordinator — which is exactly what a sovereign, federated estate needs.
The boundary that keeps this honest: the algebra governs how cards combine — it cannot derive a measurement. You cannot algebra your way to a model's score; you still run the harness. Math applies to structure, convergence and embedding — never to generating the numbers.