Glossary
Appendix F of the corrigibility papers, which both carry it verbatim so the terminology is identical across the pair. Definitions are the papers’ own; both are CC0, so quote freely and cite the paper rather than this page.
- Terms
- 44
- Licence
- CC0
A
- Action Boundary
-
The deterministic envelope that wraps stochastic inference in an agent system. Inference outputs must pass through a hard-coded validation layer before execution. The action boundary is the architectural locus of GOVERN in agentic systems.
- Action Boundary Protocol
-
The five architectural requirements — deterministic validation, context-window isolation, exception handling, machine-readable specifications, binding constraints — that an agent system's action boundary must satisfy to support corrigibility. Independent of any specific external standard.
- Agent System
-
The tuple of model, orchestration harness, action boundary, action specifications, tool definitions, deployment-time configuration and persistent memory. Behaviour is a property of the system, not the model; each component bears distinct governance requirements and can capture or release sovereignty independently.
- AUDIT
-
Third parties can verify that the system's behaviour matches its disclosed specification without requiring operator authorisation — through tamper-evident logs, statistical bounds, or drift monitoring.
Mechanism
Mechanism
Concept
Test 3 · Independent verification
Test
C
- Citizen-Constraint Registry
-
A channel through which recognised subject-class representatives lodge protective constraints that the validator must evaluate before executing determinations in the certified class. The citizen-side counterpart of the recognised court's halt flag: structural standing at the boundary, not only recourse after the fact.
- CODE
-
Inspectability of the system's execution. For deterministic infrastructure this is source-code disclosure; for learned systems it is the LWD-R requirement.
- Collective-Standing Precondition
-
Ostrom's seventh design principle applied to GOVERN: the affected population's capacity to organise, aggregate claims and fund representation must be legally protected and historically exercised at the relevant stratum. Where this fails, GOVERN fails at that stratum.
- Compute Capture
-
A FORK failure mode specific to learned systems. Occurs when the compute cost to retrain a functionally equivalent model exceeds the compute accessible to non-operator actors by more than a policy-determined multiplier.
- Corrigibility
-
The structural capacity of those affected by a system to detect error, signal harm, and trigger correction without incurring material loss or irreversible consequence. Not a moral preference but an architectural stability requirement.
Mechanism
Test 2 · Inspectability
Test
Concept
Failure mode
Concept
D
- Data Capture
-
A FORK failure mode parallel to compute capture: the data required to train a functionally equivalent model exceeds what non-operator actors can access. Encompasses synthetic data, distilled outputs, RLHF traces, reward models and proprietary evaluation sets.
- Deployment Variables
-
Where the accountable authority sits, how many agents compose the system, how deeply orchestration nests, at what scale it runs, and what species of actor exercises each corrective function. The tests bind over all of them, fixing chain termini, checker independence and grant freshness rather than geometry.
- Digital Public Infrastructure
-
Shared digital systems — ledgers, registries, payment rails, data exchanges — intended to deliver public services at population scale. The subject of the deterministic-systems paper in the pair.
- Dual Exercise of the Tests
-
The two exercises each test admits: inward, by parties at or inside the operator's institutional perimeter, and outward, by the subjects the system decides about. A test discharged only inward has been verified for the operator, not for the governed.
Failure mode
Concept
DPI
Concept
Test
E
- Effect Surface
-
The resource that receives an action's effect, where an untyped action becomes typed again. Where the action channel is untyped, boundary enforcement migrates to the effect surface: some deterministic, specification-checkable gate must stand between inference and irreversible effect.
- Epistemic Public Infrastructure
-
Public infrastructure whose behaviour is generated by learned parameters rather than explicit logic. Includes AI-based identity, eligibility, triage and orchestration systems. The subject of the learned-systems paper in the pair.
- EXIT
-
Affected parties can refuse or leave the system without prohibitive penalty, or verified Functional Exit Equivalence is demonstrated where literal exit is infeasible. For agentic deployments the test decomposes into five layers — subject, memory, workflow, operator, bystander — and failing any layer fails EXIT.
Concept
EPI
Concept
Test 1 · Reversibility of participation
Test
F
- FORK
-
Parties outside operator control hold the legal permission, the public artifacts and the user-state portability needed to recreate a functionally equivalent instance. Constructed legal barriers disqualify; natural economic friction does not.
- Functional Exit Equivalence
-
Architectural guarantees that recreate the error-signal strength of literal exit when literal exit is impossible, as with monopoly identity systems. Achieved through multi-issuer mandates, credential acceptance diversity and legally protected statutory fallbacks; verified FEE discharges the EXIT test for essential systems.
Test 5 · Independent reproduction
Test
FEE
Mechanism
G
- GOVERN
-
Affected populations have binding authority over the system, not merely advisory input. For deterministic infrastructure this is RFC or charter governance; for learned systems it is the action boundary protocol combined with citizen-side mechanisms — juridical, class-action, deliberative.
- Governance Denial of Service
-
Failure mode in agentic systems where action throughput outruns governance correction velocity. Without explicit termination criteria and bounded action specifications, agents loop indefinitely or escalate beyond legitimate authority.
- Gradient Quantifier
-
The evaluation rule that takes pass/fail determinations at the least-resourced stratum of the affected population — the argmin of detection and correction probability — rather than at the mean. An audit that samples only median-resource users has not assessed whether the loop is closed.
Test 4 · Constitutive constraint
Test
GDoS
Failure mode
Test
I
- Inference vs. Training Forkability
-
Systems that publish weights but withhold training data or code achieve only inference forkability: the model can be run and fine-tuned, not independently reproduced or corrected. Training forkability — third-party reproduction of the training run itself — is what satisfies FORK for learned systems.
- Injunction Hook
-
An API in the action boundary built to accept rapid, cryptographically signed halt or rollback flags issued by recognised courts, executed without waiting for operator compliance.
Concept
Mechanism
L
- Latent Forkability
-
The equilibrium the framework requires for coordination goods such as identity, money and law: the credible threat of reproduction disciplines the operator without routine divergence. Where a good's value derives from uniqueness, an exercised fork is mutually assured destruction, so the threat must remain latent to remain usable.
- LWD-R
-
Four-layer transparency requirement for learned systems: logic (architecture and inference code), weights (trained parameters), data (training corpus with provenance), and representation (the operative categorical schema of the deployed system).
Concept
Logic, weights, data, representation
Concept
M
- Machine-Verifiable Legitimacy
-
The requirement that authority and attribution be record-borne acts rather than inferences from artifacts: machine-interpretable authority credentials, observable revocation propagation, tamper-evident execution traces, incident-triggered mitigation, and origination marking.
Mechanism
N
- Nested Delegation
-
Composition of agent systems across delegation depth, where one system's harness contains another's. The tests compose across the nesting: authority attenuates, audit records reconstruct across levels, EXIT propagates down the chain, and the deployment's corrigibility is the minimum over its depth.
Concept
O
- Ontological Capture
-
The condition in which a learned system's operative categories become non-contestable administrative facts, surrendering interpretive sovereignty regardless of weight publication. The mitigation is R-layer change control over an operative schema treated as a commons rather than a vendor asset.
- Open-washing
-
The invocation of openness, interoperability or digital sovereignty as reputational signals without satisfying reproducibility or governance conditions. As an umbrella failure it spans symbolic-openness, coerced-legitimacy and substantive-substitution washing; as a specific tactic it names the release of peripheral SDKs while core execution logic stays proprietary.
- Operative Representation
-
The categorical schema that emerges during inference under specific deployment conditions — quantisation, pruning, routing, output filtering, the retrieval pipeline — as distinguished from the latent space of the trained artifact in isolation. Disclosure must describe operative representation, not nominal.
- Origination Marking
-
The record states whether a human or an agent operated an action. Absence of a marker is never evidence of human operation; consulting a marker may only narrow authority, never enlarge it.
Failure mode
Failure mode
Concept
Mechanism
P
- Provenance Chain
-
The sequence of generators producing training data for downstream models. Transparency does not survive an opaque link in this chain.
Concept
R
- R-Layer Change Control
-
The temporal stability condition applied to operative representation. Schema changes above a declared materiality threshold require pre-deployment notice, suspensive effect for objections from recognised affected-class representatives, and versioned diffs so drift is itself auditable. A schema that shifts faster than the affected community can contest it fails GOVERN.
- Representation Capture Risk
-
The risk that a learned system's representational categories — eligibility, risk, fraud — become uncontestable administrative facts. Operates upstream of the action boundary.
- Requisite Variety
-
Ashby's Law: a controller must possess at least as many states as the system it regulates. Applied to public infrastructure, governance mechanisms must match the population's diversity of conditions to remain corrective.
- Rule of the Ledger / Rule of the Workflow
-
Successive regimes in which authority is exerted through synchronised, self-executing artifacts. The rule of the ledger describes deterministic record-based authority; the rule of the workflow describes orchestration-mediated authority in agentic infrastructure.
Mechanism
RCR
Failure mode
Concept
Concept
S
- Scale as a Governed Variable
-
Agentic deployment multiplies effective scale: environmental variety tracks principals times orchestration fan-out times action rate, so the threshold at which governance variety falls short is crossed far earlier. Fan-out limits, action-rate limits and blast-radius caps are first-class GOVERN instruments, not performance tuning.
- Strict Interpretability Firewall
-
Eligibility rule for the high-stakes tier: a learned system whose operative representation cannot be disclosed to the operative-R standard is ineligible for adoption in rights-affecting determinations. The failure sits at CODE; nothing in the other four tests compensates for it.
- Subject Receipt
-
A signed, structured record emitted for each rights-affecting determination — action class, validator decision, timestamp, inclusion proof against the audit log — delivered through a channel independent of the operator's logging infrastructure, independently verifiable, and machine-readable for juridical and class-action aggregation.
- Switching Cost
-
The aggregate friction of leaving a system: data export, service downtime, credential re-establishment, network loss and legal barriers. When the minimum across all plausible alternatives exceeds the calibrated threshold, EXIT and FORK functionally fail regardless of formal rights. The implication is one-way — a low switching cost licenses no pass on its own.
Concept
Mechanism
Mechanism
Failure mode
T
- Temporal Stability Condition
-
A system is corrigible only if the rate of effective correction exceeds the rate of error accumulation. Correction velocity carries both a lower bound — harm otherwise accumulates — and an upper bound, since correction faster than verification capacity lets manipulated corrections bypass review.
Concept
V
- Variety Drift
-
The growing gap between a frozen model's variety and the evolving environmental variety it is meant to regulate. Drives corrigibility from a one-shot certification into a persistence condition.
Temporal variety gap
Failure mode
W
- Workflow Capture Coefficient
-
A diagnostic quantifying epistemic delegation risk as a function of reproduction cost, ontological substitutability and model portability. Computed as one minus the harmonic mean of the per-dimension sovereignty scores, so a single weak dimension dominates and cannot be hidden by strength in others.
WCC
Concept
The papers
Definitions here are compressed. The papers carry the derivations, the evaluation and the cases — both released under CC0, no permission needed.