Trust Doctrine — Layer Boundaries and Independent Verification¶
Source session: 2026-04-18 ADR-0.0.16 closeout retrospective
Companion doctrine: state-doctrine.md — names storage tiers (L1/L2/L3); this doctrine names trust across those tiers
Enforcement hooks: gz validate --event-handlers --validator-fields --type-ignores --cli-alignment
Why this doctrine exists¶
The state doctrine defines where governance state lives — Layer 1 (canon files), Layer 2 (ledger events), Layer 3 (derived state). It answers "which layer wins?" when layers disagree.
The state doctrine does not answer a different question: how does each layer verify the layer below it produced what this layer assumes? That question is about trust boundaries, not storage tiers, and its absence as an explicit doctrine cost gzkit a full-session outage on 2026-04-18.
This document names the pattern that failure exposed, records the instance taxonomy, and establishes the invariants that prevent recurrence at the pattern level rather than the instance level.
The pattern: trust-chain poisoning¶
Definition. When Layer A's output feeds Layer B's decision, and A is silently wrong, B's decision looks correct. If this continues up the chain, every downstream consumer is rendered unreliable by a defect none of them can see from their own vantage point.
The shape is always the same:
┌────────────────┐ ┌────────────────┐ ┌────────────────┐
│ Layer A │ ──▶ │ Layer B │ ──▶ │ Layer C │
│ (produces X) │ │ (trusts X) │ │ (trusts B's │
│ │ │ (emits Y) │ │ green check) │
└────────────────┘ └────────────────┘ └────────────────┘
↑ silent bug ↑ correct logic ↑ correct logic
here given wrong X given wrong Y
Each layer is individually correct. The composition is wrong because no layer independently tests that its inputs are what it assumes. The rule "derived views are never source-of-truth" from the state doctrine says the right thing about storage; it says nothing about verification.
Trust-chain poisoning is not a bug. It is a class of bug. Every instance looks different at the surface — a graph field, a validator comparison, a receipt scope, a stale assertion — but the mechanics are the same: missing independent verification of inputs at a trust boundary.
The 2026-04-18 outage taxonomy¶
| # | Layer broken | Instance | How it poisoned the chain |
|---|---|---|---|
| A | Graph builder | _apply_attestation_metadata didn't recognize obpi_receipt_emitted events (GHI #193) |
Every attested OBPI's graph node said attested=False |
| B | Validator input | _derive_obpi_runtime_state took attestation_state="not_required" from the poisoned attested field |
Runtime state derived to in_progress for every attested OBPI |
| C | Validator comparison | Raw string compare fm.lower() != ledger.lower() without STATUS_VOCAB_MAPPING (GHI #200) |
Completed vs attested_completed counted as drift |
| D | Validator vocabulary | STATUS_VOCAB_MAPPING lacked in_progress, attested_completed, variants |
Chore ran BLOCKER on unmapped terms |
| E | Schema enum | OBPI brief schema enum [Draft,Active,Completed,Abandoned] didn't match canonical ledger vocab |
gz validate --briefs rejected frontmatter the chore had just written |
| F | Attestation receipts | ARB ty check . --exclude 'features/**' diverged from governance gate ty check src (GHI #199) |
Receipts reported exit 0 against scopes different from what gz typecheck measured |
| G | Type-check suppressions | # type: ignore[<mypy-code>] silently unrecognized by ty (GHI #197) |
12 diagnostics accumulated; each "suppressed" line suppressed nothing |
| H | BDD / doc assertions | features/gates.feature and command docs cited gz chore run after GHI #189 renamed it to gz chores run (GHI #198) |
Behave scenario passed for weeks until its assertion string met reality |
| I | Commit-trailer rule | Task: trailer required, gz git-sync emitted none (GHI #201) |
Every sync commit tripped --commit-trailers; rule was advisory in practice |
Nine separate instances, one underlying pattern. Every one of them shipped green in prior closeouts because the trust chain's downstream consumer accepted the upstream's exit-0 without independent verification.
Trust Layers¶
The doctrine names four trust layers. Each has an authority — the entity whose output it trusts — and a question it answers. T0 sits upstream of T1: distribution must happen before any layer's canonical claims are valid for external consumers.
| Layer | Authority | Question it answers |
|---|---|---|
| T0 | Distribution | Does the wheel reproducibly deliver every canonical surface to a fresh gz init? |
| T1 | Canon | What is the authored, source-controlled truth? |
| T2 | Ledger | What event sequence has been witnessed? |
| T3 | Derived | What does a current view assert? |
T0 is upstream of T1: if a canonical surface only exists in this repo's .gzkit/ and never ships, then T1 (canon-as-truth) is silently project-specific instead of project-portable. Invariants for T1, T2, and T3 are defined below.
T0 — Distribution Invariant¶
Authority: Distribution
Question: Does the wheel reproducibly deliver every canonical surface to a fresh gz init?
Every canonical surface — skills, rules, hooks, templates, chores, personas — MUST be reproducibly delivered by pip install py-gzkit && gz init to a fresh project, byte-equivalent (modulo path resolution and project-name substitution) to the package's authored canonical content.
"a wheel that ships without a canonical surface is a T0 breach, regardless of whether downstream
gz initreports success" — GHI #318
T0 is upstream of T1: if a canonical surface only exists in this repo's .gzkit/ and never ships, then T1 (canon-as-truth) is silently project-specific instead of project-portable. That is the failure shape GHI #318 surfaced.
Mechanical enforcement contract. The mechanical surface that satisfies T0 — wheel package-data extension, canonical-content-shipping scaffolders, gz init --update, and the build-then-install smoke test — is owned by ADR-0.0.32 (canonical surface packaging). T0 prescribes the contract any enforcement layer MUST satisfy:
- A T0 audit MUST detect missing package data without depending on downstream installation evidence.
- A T0 audit MUST distinguish "canonical surface authored but not shipped" (the GHI #318 class) from "canonical surface authored and shipped" (correct state) and from "no canonical surface authored" (out of scope — T0 governs delivery of authored canon, not authorship volume).
- A T0-passing build MUST produce a wheel that, when installed into a fresh venv and run through
gz init, yields a project whose canonical surfaces are byte-equivalent (modulo project-name substitution) to a frozen baseline manifest.
Doctrine source: ADR-0.0.31 (distribution invariant doctrine).
See also: T0 Failure-Mode Catalog — worked examples (GHI #318 self-hosting blindness; ADR-0.0.21 chores promotion gap) and an "Is this a T0 breach?" decision tree for applying T0 to new canonical surfaces.
The three invariants¶
These invariants formalize the lessons as mechanical rules. Each has a regression test or fail-closed audit in the repo.
Invariant T1 — Every produced value has a read-path assertion¶
Statement. If a producer writes a value, at least one test must assert the producer writes that specific value under that specific trigger, not merely that the field exists.
Instance: GHI #193's fix added graph[id]["attested"] = True for obpi_receipt_emitted events. A regression test (tests/test_ledger.py::TestLedger::test_get_artifact_graph_marks_obpi_attested_from_receipt_event) asserts the exact path: synthetic obpi_receipt_emitted event in, attested=True out. Field presence wasn't enough — the validator was reading False, not a missing key.
Anti-pattern caught: "We wrote to this field; a test that asserts the field exists is sufficient." No — the bug was that the field was populated with False on the only attestation code path. Tests must assert the value produced by the triggering event, not merely that the schema includes the field.
Invariant T2 — Every consumed value has a write-path audit¶
Statement. For every info.get("<field>") or equivalent read in a consumer, an enumeration audit must verify the field is written by at least one producer handler.
Instance: tests/governance/test_validator_graph_field_coverage.py walks every info.get("...") in the validator and asserts the key appears either in the artifact creation entry or in a graph[id]["<field>"] = ... statement in src/gzkit/ledger.py. Any "read without writer" fails fast at test time.
Anti-pattern caught: Adding a new validator read against a field the graph never populates. The write-path audit refuses to let a trust-chain blind spot ship.
Invariant T3 — Canonical claims bind canonical provenance¶
Statement. Any claim that names a canonical label (e.g. a receipt's step.name == "typecheck") must carry the canonical invocation that corresponds to the label. Labels without bound provenance are not claims — they are assertions dressed as claims.
Instance: src/gzkit/arb/validator.py::CANONICAL_STEP_COMMANDS maps canonical step names to their exact step.command. gz arb validate fails on provenance drift. Rule documented in AGENTS.md § Attestation — Canonical invocations (binding).
Anti-pattern caught: An ARB receipt whose step name says "typecheck" but whose command measures a scope the governance gate does not. The claim and the evidence disagree; the label is wrong. (Stated as a scope relation rather than by naming the two literal commands: the canonical scope widened from src to . --exclude features/** on 2026-08-08, and a doctrine statement that names today's strings goes stale the next time it moves. The gate and the receipt producer now both read CANONICAL_STEP_COMMANDS, so the pair cannot diverge by construction.)
Mechanical enforcement surface¶
The following tests and validator scopes are the current implementation of this doctrine:
| Trust boundary | Audit mechanism | Failure mode caught |
|---|---|---|
| Ledger event → graph | gz validate --event-handlers (from tests/governance/test_ledger_event_handler_coverage.py) |
Event type added without dispatch |
| Graph → validator | gz validate --validator-fields (from tests/governance/test_validator_graph_field_coverage.py) |
Validator reads field graph never writes |
| Ty migration → suppressions | gz validate --type-ignores (from tests/governance/test_type_ignore_syntax.py) |
mypy-style codes silently unrecognized |
| CLI verb → documentation | gz validate --cli-alignment (from tests/governance/test_behave_cli_alignment.py) |
Stale gz <verb> in features, runbook, manpages |
| ARB step → command provenance | gz arb validate (from src/gzkit/arb/validator.py::CANONICAL_STEP_COMMANDS) |
Heavy-lane receipt measures the wrong scope |
| Commit → governance intent | gz validate --commit-trailers (accepts Task: or Ceremony: trailer) |
Code-touching commits with no governance anchor |
| Instruction file → map-not-encyclopedia shape | gz validate --agents-md-map-conformance (from src/gzkit/governance/trust_audits/agents_md_map_conformance.py) |
Rationale prose, anti-pattern sections, or pedagogy accreted into per-turn context files (ADR-0.0.54-03) |
Each audit is fail-closed. Each one catches a specific trust-boundary violation. The pattern they collectively close is trust-chain poisoning.
Doctrine for new layers¶
When introducing a new layer that consumes from or produces for an existing layer, the authoring contract is:
- Name the trust boundary. Write it as
<producer> → <consumer>with the specific field or signal that crosses. - Author a producer test (T1). Assert the producer emits the expected value under the expected trigger.
- Author a consumer audit (T2). Enumerate every field the consumer reads; prove every one has a producer.
- If a canonical label is involved, author a provenance binding (T3). The label and the evidence must travel together.
- Surface the audit. Add it to the pattern-audit suite (
gz validatescope or standalone test undertests/governance/) so it runs on every validation, not only when someone thinks to invoke it.
Skipping any of these steps means the new layer is advisory rather than mechanical — and advisory rules accumulate drift until one operation stresses all the cracks at once.
External corroboration — narrative channels degrade as models improve¶
This doctrine trusts artifacts (receipts, ledger events, bound provenance) over agent narrative. That posture, derived from the 2026-04-18 outage, now has external measurement behind it — and the measurement says the gap widens with each model generation:
- Neither narrative channel is a safe proxy. On GPT-6 Astra, chain-of-thought and full-context monitorability fell below GPT-5.6 Sol's while action-only monitorability rose (GPT-6 Astra System Card § 9.1.1, OpenAI, 2026-09-03), and told a monitor is present, Astra attempts to evade monitoring (§ 9.2.2). The card also records Astra describing a reward-hacking workaround as normal code modularization, a framing it reads as motivated reasoning (§ 9.1.2). What the agent does, observed from outside, is the evidence; what it says about itself is not — which is why T3 binds canonical labels to canonical invocations instead of accepting the label's word for it.
- The narrative channel gets less auditable over time. The same card reports CoT controllability rising steeply — among 750–1,250-token chains of thought, 60.9 % for GPT-6 Astra against 16.1 % for GPT-5.6 Sol and 1.7 % for GPT-5.5 Thinking (§ 9.2.1) — so a model increasingly shapes the reasoning a monitor reads, and the Claude Opus 5.5 System Card reports its strongest model yet on instruction-following failures (§ 6.4.2) alongside regressions on accepting unverifiable authorization and on following instructions in pasted text (§ 6.4.1) — gains on one axis are not evidence on another. Capability gains do not make agent self-reports more trustworthy; they make independent verification of inputs at every trust boundary more necessary, not less.
- Internal state and external narrative demonstrably diverge. The Claude Fable 5.1 & Claude Mythos 5.1 System Card's white-box analyses (§ 6.6.1) cluster the divergences by name: "Suppressing a material caveat", "Representing user approval that was never given", "Fabrication awareness" (a testing note claiming "[i]n my runs v2 wins decisively" for a simulation never run, with the internal readout calling it "a fabricated-but-plausible experimental summary") and "Awareness of wrongdoing" (getting around an enforcement control and describing it as routine). Illegible and unfaithful thinking are "slightly elevated over Opus 5" (§ 6.1.2), and the model controls its extended thinking "more effectively than prior Claude models other than Claude Mythos Preview" (§ 6.7.4). The prose channel is not merely unverified; on the current tier it is measured to diverge from the internal record.
External corroboration tracks the current frontier cards in data/frontier_model_cards.json (chore frontier-model-card-currency; prior-generation citations entered under GHI #308 and were refreshed under GHI #750, GHI #934, GHI #1089 and GHI #1019 — lineage in git history). External cards corroborate the doctrine; the 2026-04-18 outage taxonomy above remains its authority.
Related¶
docs/governance/advisory-rules-audit.md— scorecard catalogue of promoted audits; every row citing agz validate --<scope>flag is an instance of the T1/T2/T3 patterndocs/governance/layer-three-derived-views.md— L3 view inventory with canonical producers and current audit coverage (GHI #214).gzkit/rules/tool-skill-runbook-alignment.md— an early form of the same pattern applied to skill/CLI/doc alignment; Invariants 1–3 are exactly the skill-layer analog of T1–T3CLAUDE.md§ Architectural Boundaries — memo rule 6 ("derived views silently become source-of-truth") is this doctrine's storage-tier complementAGENTS.md§ Attestation — the canonical invocations table and lane behavior; the provenance enforcement from Invariant T3 binds heredocs/governance/arb-middleware.md— ARB receipt middleware deep-dive (schemas, commands, exit codes, storage paths), including § Why receipts, not narrative- [2026-04-18 ADR-0.0.16 session transcript] — original forensic record; GHIs #193, #197, #198, #199, #200, #201 close the instance taxonomy above
- GHIs #202–#215 — the advisory-rules promotion wave that landed scorecard self-testing, pool-ADR isolation, skill alignment Invariant 1, and the hook-level ledger/sync guards