Criteria register
Criteria Register — Robin Knowledge Architecture
← eval suite index
Criteria Register — Robin Knowledge Architecture
STATUS: PENDING OWNER ACCEPTANCE — no criterion below is binding until the owner has graded and accepted it. Generated 2026-08-31 from the anchor (v5, grade A) and companion (v3, grade A).
Sources:
- Anchor —
~/papers/robin-geometric-compositional-knowledge-architecture.md, §10 (Evaluation Plan). Frozen at grade A. - Companion —
~/papers/robin-perspectival-machinery.md (v3, grade A), §11 (Evaluation Plan), plus falsifiers stated in §1, §3.2, §6, §7, §8 — all of which resolve into §11 subsections; none states a criterion §11 does not carry.
Tiers:
- runnable — executable against the live system today, no new architecture (annotators and an NLI model may be required; those are inputs, not architecture).
- gated — requires machinery that does not exist yet (Claims, Dimensions, Initiatives, the artifact engine, the Router, the replay harness). The precondition is named per criterion.
Statistical protocol for every comparative criterion is the anchor's shared protocol (§10 preamble): two annotators, graded 0–3 judgments, Cohen's κ reported with re-adjudication below 0.4; one primary clause per bet at one-sided α = 0.05 under the two-look extension (nominal 0.03 per look); secondaries Holm-corrected as one family per subsection; every comparative test states a design effect and is sized for 80% power; cost accounting charges a request for every model call it triggers. The companion adopts this protocol and supplements it with three clauses the anchor leaves unstated (§11 preamble, marked as supplements): the family-wise level for Holm secondaries is 0.05; guard and sanity clauses are exempt from correction and tested in the harm direction; and non-inferiority means the one-sided 95% bound on the paired difference lies within the stated margin. Construction guarantees (counts the design says must be zero) are defects to fix, not statistics to test.
Anchor §10 criteria
CR-01 — Dimension-aware reranking gain
- Statement: The Dimension term improves ranking over the hybrid retriever plus standing penalty, and the improvement is attributable to geometry (survives the permutation control), within the fast-path cost bound.
- Metric: nDCG@10 on held-out split, arm (iii) minus arm (ii); secondary Recall@50, MRR. Permutation control: (iv) minus (ii). Cost: query-projection share of median fast-path cost.
- Dataset/inputs: ≥150 general queries per workspace across 3 pilot workspaces (~270 pooled held-out), graded relevance over pooled top-20, 40/60 validation/test split; ≥30 validation queries per fitted Domain.
- Threshold: Dimension gain positive and significant (paired bootstrap, one-sided), design effect 0.03 absolute nDCG@10; no workspace significantly degraded (secondary); permutation control retains at most half the gain; projection calls add ≤10% to median fast-path cost for ≤3-Domain scope (≤3K calls).
- Drop condition: Dimension term dropped if the gain fails the shared success rule, or the permutation control retains more than half of it, or the cost bound is exceeded and the gain vanishes with query projection disabled. Dropping also parks all Dimensions active only by the Useful screen.
- Source: Anchor §10.1.
- Tier: gated — requires Claims layer, active Dimensions, projection machinery, fitted β_{D,d} and μ.
CR-02 — Standing penalty gain
- Statement: The standing penalty μ alone improves ranking over the plain hybrid baseline.
- Metric: nDCG@10, arm (ii) minus arm (i) (secondary clause of §10.1's family).
- Dataset/inputs: Same query sets and arms as CR-01.
- Threshold: Standing gain positive and significant.
- Drop condition: Dropped otherwise, independently of the Dimension term's fate.
- Source: Anchor §10.1.
- Tier: gated — requires Claims layer with maintained standing.
CR-03 — Claims-in-index retrieval contribution
- Statement: Indexing Claims alongside Signals improves retrieval over a Signals-only index (arm (v), the retrieval half of the thesis alternative).
- Metric: nDCG@10, arm (i) minus arm (v).
- Dataset/inputs: Same query sets as CR-01; arm (v) is the hybrid baseline with Claims removed from the index.
- Threshold: None stated as pass/fail; decision rule: if (v) matches (i) within the noise of the test, Claims are not contributing to retrieval and the quality half of the thesis rests entirely on §10.2 (CR-04).
- Drop condition: Not a component drop; feeds CR-19's thesis verdict.
- Source: Anchor §10.1.
- Tier: gated — requires Claims layer.
CR-04 — Claims-layer artifact quality
- Statement: Artifacts rendered from persisted Claims are preferred over per-request reasoning with persistence disabled (baseline (c)), at comparable citation fidelity.
- Metric: Primary: blinded pairwise preference vs baseline (c), exact binomial test. Secondary: preference vs Signal-only generation (a) and GraphRAG-style (b); ALCE citation recall and precision (NLI-computed, 20% human check); precondition check; cost per artifact per arm (reported, judged in CR-14's ledger).
- Dataset/inputs: 55 artifact requests per pilot workspace, ≥10 Initiatives each, audience specified; ≥150 decided pairs per baseline comparison after tie exclusion; two raters per pair, third adjudicates.
- Threshold: Preference vs (c) significant, design effect 60% vs 50% null (~86+ wins of 150); ALCE citation precision and recall no lower than the best baseline; precondition check at 100% (renderer requirement — any miss is a defect).
- Drop condition: Result at or below 50%, or failing after extension: the Claim layer contributes nothing to artifacts that per-request reasoning does not; thesis then decided by the ledger alone. A tie with (a) is the Chen et al. outcome: drop the Claim layer in favor of Signal extraction investment. Preference win with citation-precision loss is a renderer redesign signal.
- Source: Anchor §10.2.
- Tier: gated — requires Claims layer, slow path, renderer.
CR-05 — Signal-branch contradiction detection
- Statement: The contradiction check detects injected contradicting Signals and lifts them into conflicts_with edges cheaply and quickly.
- Metric: Primary: Signal-branch recall (targets acquiring a Signal-branch edge to the lifted Claim of their contradicting Signal). Also: Signal-branch precision; candidate recall after floor τ_s; contradictions removed by the floor; model calls per injected Signal; wall-clock to contested transition. Scored on branch-tagged edges only.
- Dataset/inputs: Per workspace, 100 accepted Claims as targets; 1 contradicting + 1 distractor Signal each (200 Signals), injected in random order over a simulated week through the ordinary write path.
- Threshold: Recall ≥0.8; precision ≥0.7; candidate recall ≥0.9 with the floor removing none of the injected contradictions; model calls ≤ k+1 per injected object; median wall-clock ≤ Δ+1 min (Δ = 5 min).
- Drop condition: If recall cannot reach 0.6 at k = 50 with the floor removed, automatic standing changes are disabled and standing becomes a purely human-maintained field. Redesign: candidate recall <0.7 → raise k or lower τ_s; precision <0.5 → stronger classifier model or human queue.
- Source: Anchor §10.3.
- Tier: gated — requires Claims layer and the §3.4 contradiction check.
CR-06 — Claim-branch position filter
- Statement: The Claim branch of the contradiction check finds positioned opposing Claims, and the position filter raises candidate precision over similarity alone without hurting recall.
- Metric: Claim-branch precision and recall on branch-tagged edges; candidate recall; fraction of checks where the position filter applied; candidate precision with filter vs similarity-alone at the same k (over checks where the filter applied); model calls per injected Claim including projection.
- Dataset/inputs: Per workspace, 100 accepted target Claims each with ≥1 active Dimension at relevance ≥ τ_r; per target one contradicting Signal + opposing Claim justified by it, plus distractor pair, written as units through the ordinary path.
- Threshold: Precision ≥0.7 and recall ≥0.8 on Claim-branch edges; candidate recall ≥0.9; position filter raises candidate precision without lowering candidate recall below 0.9; ≤ k+1 check calls plus ≤ K projection calls.
- Drop condition: If the position filter lowers candidate recall below 0.9, the filter is removed and the Claim branch falls back to similarity alone (also removing the §5.3 signature from the check).
- Source: Anchor §10.3.
- Tier: gated — requires Claims layer, active Dimensions, projection.
CR-07 — Retraction minimal-change check
- Statement: Withdrawing Entries retracts exactly the Claims whose every justification depended on them, and restoring the Entries returns exactly those Claims to proposed.
- Metric: Set equality in both directions (no extra retractions; exact restoration to proposed, never auto-accepted).
- Dataset/inputs: Withdraw the Entries behind 20 accepted Claims per arm; restore them.
- Threshold: Exact in both directions (construction guarantee).
- Drop condition: None stated; a miss is a defect in the dependency network.
- Source: Anchor §10.3.
- Tier: gated — requires Claims layer and justification network.
CR-08 — Initiative attention gain
- Statement: Initiative attention a_I improves Initiative-scoped retrieval over fitted Domain coefficients alone (a_I = 0).
- Metric: nDCG@10 on Initiative-scoped held-out queries, attention arm vs a_I = 0 arm; paired bootstrap.
- Dataset/inputs: ≥200 Initiative-scoped queries per workspace (~350 pooled held-out at 60% split across 3 workspaces); attention rule run with η = 0.2, γ = 0.1 over each Initiative's slow-path history.
- Threshold: Positive and significant, design effect 0.02 absolute nDCG@10. If the pooled count falls short, the design effect is raised to what the count detects at 80% power and reported before scoring.
- Drop condition: No significant lift removes the attention rule, leaving Initiatives as Domain groupings with a_I = 0.
- Source: Anchor §10.4.
- Tier: gated — requires Initiatives machinery, Dimensions, fitted coefficients.
CR-09 — Domain scoping gain
- Statement: Initiative/Domain scoping beats global workspace retrieval.
- Metric: nDCG@10, scoped vs global (secondary clause of §10.4's family).
- Dataset/inputs: Same Initiative-scoped query sets as CR-08.
- Threshold: Positive and significant, powered for 0.05.
- Drop condition: No lift questions whether Domains earn their place in the fast path at all; read together with CR-18's classifier figures before anything else.
- Source: Anchor §10.4.
- Tier: gated — requires Initiatives machinery and the scoped fast path.
CR-10 — Cross-Domain Claim usefulness — REJECTED
- REJECTED (2026-08-31 acceptance pass): not measurable as stated — the anchor maps no cut on the 3-point scale to "rated-useful," and the single-rater metric sits outside the shared two-annotator/κ protocol with no agreement check, so two engineers scoring the same slow-path runs could reach opposite verdicts on the ≥1-per-3-runs threshold. The anchor is frozen; the fix (a stated scale cut and a rater count under the shared protocol) must come from the owner and live in this register before the criterion can bind.
- Statement: Initiatives produce cross-Domain Claims people find useful.
- Metric: Newly derived Claims per slow-path run whose home set spans ≥2 Domains, rated non-obvious and useful on a 3-point scale.
- Dataset/inputs: Slow-path runs under Initiatives during the §10.4 experiment; one rater per Claim.
- Threshold: ≥1 rated-useful cross-Domain Claim per 3 slow-path runs.
- Drop condition: None stated separately; part of §10.4's success bundle.
- Source: Anchor §10.4.
- Tier: gated — requires Initiatives machinery and the slow path.
- Note (measurability gap, flagged at the 2026-08-31 proxy screen; the anchor is frozen, so it persists): the anchor maps no cut on the 3-point scale to "rated-useful," and the single-rater metric sits outside the shared two-annotator/κ protocol with no agreement check. Owner must fix the scale cut and rater count before this criterion can bind.
CR-11 — Visibility rule decision
- Statement: Decide between the subset rule (§3.2) and the intersection rule with fitted penalty ν for Claim visibility under Initiatives.
- Metric: Held-out nDCG@10 on Initiative-scoped queries, each rule run with attention on.
- Dataset/inputs: Same as CR-08.
- Threshold: Decision rule, not pass/fail: intersection adopted iff significantly higher; otherwise the subset rule stays as the conservative default.
- Drop condition: n/a (either outcome keeps a rule).
- Source: Anchor §10.4.
- Tier: gated — requires Initiatives machinery.
CR-12 — Fast/slow compounding
- Statement: Accumulated slow-path work improves the fast path over time on a frozen corpus.
- Metric: Primary: fast-path nDCG@10 gain from m = 0 to m = 20 on the clean replay series, paired bootstrap over N queries. Secondary: no checkpoint significantly below its predecessor; HyDE escalation rate falls; p95 latency does not rise.
- Dataset/inputs: Clean replay series per workspace from a fresh Signals-only copy; m artifact requests and N ≥ 60 read queries per workspace (180 pooled) taken from one logged window in logged order (author-labeled → artifact runs, retrieve-labeled → reads); checkpoints m = 0, 5, 10, 20; coefficients fit once at final checkpoint; post-replay pooling/judging pass for new Claims.
- Threshold: Gain positive and significant, design effect 0.03.
- Drop condition: No significant gain by m = 20 means slow-path outputs are not reused — the compounding claim fails (feeds CR-19).
- Source: Anchor §10.5.
- Tier: gated — requires full stack (Claims, slow path, fast path) plus the replay harness.
CR-13 — Staleness under injected contradictions
- Statement: Standing propagation keeps stale Claims out of top results as structure accumulates, and beats a Graphiti-style temporal-graph arm on staleness at comparable ingest cost.
- Metric: Staleness rate: fraction of top-10 results that are accepted/proposed Claims with an injected contradicting Signal present longer than Δ (contested-and-surfaced does not count). Sanity: retracted/superseded Claims in top-10 must be zero (construction guarantee). Per-Signal ingest cost c_ing both arms.
- Dataset/inputs: Injected replay series: same artifact runs as CR-12 plus, per checkpoint, up to 100 annotator-written contradicting Signals against existing accepted Claims (<20 accepted Claims at a checkpoint → staleness reported as not measurable, not zero). Comparison arm: Graphiti-style incremental temporal fact graph over the same Signals.
- Threshold: Staleness below 2% at every measurable checkpoint.
- Drop condition: Staleness rising with m falsifies the thesis in the weaker sense (§1). If the Graphiti-style arm matches Robin's staleness at ingest cost no higher than Robin's, the justification network is not earning its complexity on this axis.
- Source: Anchor §10.5.
- Tier: gated — requires full stack plus the replay harness.
CR-14 — Cost ledger break-even
- Statement: Robin's total spend (slow-path writes + ingest checks + fast reads) falls below the alternative's at the read/write ratio pilot workspaces exhibit.
- Metric: Break-even read count N* = (W(m) + m·σ·c_ing − m·g_base) / (r_base − r_fast), with r_base the per-query cost of the cheapest baseline read arm reaching quality parity (parity judged at the reasoning arm's evidence depth |C_q|, nDCG/precision at that depth); projection and ingest lines reported separately.
- Dataset/inputs: Clean-series costs (W(m), r_fast at m = 20), injected-series c_ing, log-measured σ (Signals per artifact-generating request) and N̄ (reads per artifact-generating request); baseline arms: Signals-only hybrid and hybrid + reasoning-over-retrieved.
- Threshold: N* ≤ N̄·m at m = 20. (r_fast < r_base is a sanity check, not a success.)
- Drop condition: N* above N̄·m at observed ratios, or no break-even at all, falsifies the cost half of the thesis even if the quality half holds (feeds CR-19).
- Source: Anchor §10.5.
- Tier: gated — requires full stack plus replay harness and pilot logs.
CR-15 — Dimension meaningfulness
- Statement: The discovery pipeline promotes Dimensions people recognize as meaningful axes of comparison.
- Metric: Primary: fraction of promoted Dimensions rated meaningful by both raters ("would you use this axis to compare these Claims?" over a sample of projected Claims, κ reported). Also: promotion rate through validation; per-generator (model proposal vs Alshaikh-style disentanglement) promotion and meaningfulness rates.
- Dataset/inputs: Promoted Dimensions from the pilot; two raters; samples of projected Claims per Dimension.
- Threshold: ≥50% of promoted Dimensions rated meaningful by both raters.
- Drop condition: Below 50% replaces automatic discovery with human-authored Dimensions per Domain, keeping validation and projection unchanged.
- Source: Anchor §10.6.
- Tier: gated — requires the discovery pipeline and Claims to project.
CR-16 — Small-projector agreement
- Statement: Projection can be done by a small fine-tuned model in agreement with the frontier projector at a tenth of the cost.
- Metric: r_s or κ between small projector (fine-tuned on frontier bootstrap) and frontier projector, per Dimension, with confidence interval; per-projection cost ratio.
- Dataset/inputs: Held-out sample per Dimension of up to 200 Claims (≥30 per Domain).
- Threshold: Agreement meeting §5.4's Measurable thresholds on ≥3/4 of active Dimensions, at ≤1/10 the frontier projector's per-projection cost.
- Drop condition: Below threshold keeps projection on the larger model, making c_π the binding constraint on active-Dimension count (lower K, higher W(m) in CR-14's ledger).
- Source: Anchor §10.6.
- Tier: gated — requires Dimensions, projection machinery, and a trained small projector (note: one of the anchor's two fine-tuned components, in tension with a no-training posture).
CR-17 — Router accuracy
- Statement: The Router labels requests retrieve/reason/author accurately enough that downstream failures are attributable to the components under test, not misrouting.
- Metric: Primary: routing accuracy. Also: the two misroute rates separately (slow-misroute wastes compute; shallow-misroute returns a shallower answer).
- Dataset/inputs: 300 requests per workspace sampled from logs, labeled by two annotators (κ reported); a disjoint portion is the Router's fine-tuning set.
- Threshold: Accuracy ≥0.9; shallow-misroute rate ≤0.05.
- Drop condition: Below either bar, the Router's inputs are revisited before any downstream result is interpreted (redesign, not drop).
- Source: Anchor §10.7.
- Tier: gated — requires the Router component (fine-tuned; anchor §8 — the anchor's other trained component) and the fast/slow runtime it routes between.
CR-18 — Domain classifier quality
- Statement: Signal-to-Domain classification is accurate enough to serve as the fast path's hard scope filter.
- Metric: Per-Domain precision and recall of Signal-to-Domain assignment; macro-F1.
- Dataset/inputs: 500 Signals per workspace with Domain labels from two annotators. Signals, Domains, and the classifier exist in the deployed system.
- Threshold: Macro-F1 ≥0.85 and no Domain with recall below 0.75.
- Drop condition: Any Domain below the recall bar has its description and classification logic revised and Signals reclassified before §10.1/§10.4 run on that workspace (redesign, not drop).
- Source: Anchor §10.7.
- Tier: runnable — the Domain classifier is live (verified in code at the 2026-08-31 proxy screen: capture-time
domain-classify stage, domain_signals rows with source: 'classifier', per-domain classifierPrompt); only annotation is missing.
CR-19 — Anchor thesis falsification (composite)
- Statement: The thesis-level verdict over the component results.
- Metric/threshold: Quality half fails if CR-01/CR-03 (Signals-only arm matches full system on retrieval) and CR-04 (per-request reasoning matches on artifacts) both fail against the alternative. Cost half fails if CR-14 finds N* above observed N̄·m or no break-even. Weak failure if CR-13 shows staleness rising with accumulated structure. Any single component failing its own drop condition removes that component and leaves the thesis judged on the rest.
- Dataset/inputs: Aggregates CR-01 through CR-14.
- Drop condition: Any of the three failures returns Robin to the alternative (no persisted Claims; hybrid retrieval over Signals; per-request reasoning for artifacts).
- Source: Anchor §10.8.
- Tier: gated — composite over gated criteria.
Companion §11 criteria
CR-20 — Entry→Signal extraction completeness
- Statement: Extraction recovers the atomic content of an Entry uniformly across position (the eval the anchor's §10 misses; bounds the confounder under every anchor experiment).
- Metric: Primary: pooled proposition recall (proposition recovered iff ≥1 extracted Signal entails it, NLI-judged with 20% human check), cluster-bootstrap intervals over Entries. Guard (exempt from Holm per the protocol supplement, tested in the harm direction): extraction precision — fraction of extracted Signals entailed by their recorded source span (hallucination guard). Sole secondary (Holm is the identity): positional uniformity — one-sided 95% bound on first-minus-fourth-quartile recall difference, cluster-bootstrapped over Entries. Also: Signals per Entry per band, for the ledger.
- Dataset/inputs: 30 Entries per pilot workspace (90 pooled) sampled from real captures across confident/triage size bands, ≥half dense (≥1,000 words); two annotators decompose to atomic propositions at Chen et al. (2024) granularity, κ reported; an NLI model.
- Threshold: Pooled proposition recall ≥0.8; extraction precision ≥0.95 (guard); positional uniformity bound within 0.10.
- Drop condition: None — Entry is a primitive and extraction is not droppable; this eval bounds it. Redesign: recall <0.8 with uniform position → revisit prompt and refusal thresholds; uniformity bound beyond 0.10 → per-section multi-pass extraction, extra cost charged to the ledger.
- Source: Companion §11.1 (contract stated in §3.1; the historical positional-discard failure, ~45% of dense notes, motivates the uniformity clause).
- Tier: runnable — entries, signals, and capture reports (confident/triage bands) are live; needs annotators and an NLI model, no new architecture. Pre-registered to run first (§11.7). Caveat, re-verified in code 2026-08-31:
signals.provenance.span is declared (server/src/db/schema.ts:584) but the extraction persist path writes no span (packages/agent/src/stages/persist.ts), so the precision guard's input ("recorded source span") is unobtainable until span-writing lands — which §3.1's extraction contract already demands of the extractor — or entailment is respecified against the full Entry payload. The recall and uniformity clauses run today.
CR-21 — Stance-conflict exclusion: contamination
- Statement: The §6 exclusion prevents editorial disagreement from moving epistemic standing, and the failure mode it forecloses actually occurs (the rule is not dead machinery).
- Metric: Primary: off-arm contamination rate — fraction of opposed target Claims that become contested where every contradicting support traces only to wiki-origin Signals and no contradicting support's stance-bearing side carries an acceptance event through claim adjudication — exceeds the dead-machinery floor; cluster bootstrap over authored sets. The acceptance qualification is load-bearing: a contested transition driven by an accepted stance Claim is excluded from the count on purpose, reported on its own line, and read as the formalization path working. Construction guarantee (defect if violated): on-arm contamination zero absent claim adjudication — an on-arm contested transition with no acceptance event anywhere behind it is the defect; the fourth test set's accepted stance Claims contesting their targets is the mechanism working. Guard: Signal-branch recall/precision on the evidential injections matches across arms (paired per injection; a difference explained by lifted-Claim pool divergence between the separate workspace copies is recorded as such, any other is an implementation fault).
- Dataset/inputs: Per workspace: 20 sets of accepted Claims (~5 per set), a pair of ~300-word opposing wikis per set (40 wikis, captured through the ordinary path as Entries of source class wiki) — ~100 opposed targets per workspace, 300 pooled in 60 sets; plus the anchor's §10.3 Signal-arm injection set run in the same workspaces; two arms (exclusion off/on) on separate workspace copies.
- Threshold: Null hypothesis: off-arm contamination 2%; one-sided test at the protocol's per-look level, cluster bootstrap over the 60 authored sets (rejection region at these numbers: ≥6 contaminated of the ~100 effective targets, identical at 0.03 and 0.05); design effect: true rate 10% (power >0.9 at within-set correlation 0.5).
- Drop condition: Observed off-arm contamination below 2% → the check-path branching is dropped as dead machinery (origin mark kept: one field, feeds §4's formalization path) and the anchor's §3.4 stands as written. At/above 2% but not significant → branching kept on the asymmetry (guard costs field reads, the failure corrupts standing), bet reported as unconfirmed and re-tested at next pilot scale, not counted as a pass.
- Source: Companion §11.2 (mechanism in §6; falsifier pointer at §6's close).
- Tier: gated — requires the Claim layer, the §3.4 contradiction check, and the origin mark (plus workspace-copy apparatus for the two arms).
CR-22 — Revision succession: survival and removal retraction
- Statement: Accepted judgment formalized from a wiki survives edits that preserve the stance and dies with edits that remove it (§3.2's succession clauses).
- Metric: Survival — fraction of accepted stance Claims whose stance the revision preserved that are not retracted after the revision batch (a contested transition is not a survival failure; what succession prevents is retraction through the withdrawn revision Entry). Removal retraction — fraction whose stance the revision removed that are retracted within Δ plus the batch's processing time.
- Dataset/inputs: For 10 authored wikis per workspace (exclusion-on arm; succession is orthogonal to the exclusion), three drafted stance Claims per wiki accepted through claim adjudication; two revisions each through ordinary capture — one stance-preserving edit (wording, ordering, a fixed typo), then one edit removing one stated stance and keeping the rest.
- Threshold: Survival ≥0.95; removal retraction ≥0.95 (both Holm-corrected secondaries of §11.2's family). A survival miss is a succession false negative (the entailment judge failed to match a successor to prose still asserting the stance); a removal-retraction miss is a false positive (prose no longer asserting the stance was matched anyway, laundering the removal).
- Drop condition: Redesign: survival <0.95 → stronger succession matcher, or fall back to the alternative rule considered and not chosen (acceptance pins the justifying revision Entry, exempt from succession-withdrawal, drift stamp carrying the divergence); removal retraction <0.95 → tighten the entailment bar.
- Source: Companion §11.2 (mechanism in §3.2).
- Tier: gated — requires the Claim layer, claim adjudication, and wiki revision machinery.
CR-23 — Stance boundary audit: leak rate, shadow run, classifier fallback
- Statement: The provenance-drawn boundary (origin mark) is fine enough — wiki prose is not carrying genuine evidential conflict the exclusion suppresses, and the quoting-Entry leak stays within its priced residual.
- Metric: Leak rate — fraction of quoting Entries producing at least one conflicts_with edge. Shadow run of the classifier over stance-bearing Signals (on-arm, edge-writing disabled); two annotators label each contradicts verdict as editorial disagreement or genuine evidential conflict (κ reported).
- Dataset/inputs: 20 evidential Entries (meeting notes) per workspace, each quoting a wiki's stance with attribution, as the leak probe; the §11.2 on-arm shadow verdicts.
- Threshold: Leak rate below 0.10 → the leak is the priced residual of a boundary drawn in provenance, reported as such. At or above 0.10 → the mark is too coarse in the quoting direction and the per-Signal stance classifier is built for that direction too. Redesign trigger: ≥20% of shadow contradicts verdicts labeled genuine evidential conflict by both annotators → the mark must be replaced by a per-Signal stance classifier at extraction time (a new trained component, charged as such against the §11 posture note — the single place a third trained component could enter). The classifier's bar, fixed in advance: precision ≥0.7 on its positive label (the anchor's §10.3 bar for the contradiction classifier, adopted because both work the region where de Marneffe et al. found precision scarcest), with the anchor's fallback ladder (stronger model, then human queue) behind it.
- Drop condition: If the leak rate is at or above 0.10 and the classifier, when built, cannot hold the 0.7 precision bar, the perspectival half of the thesis fails (§1, §11.8): wiki prose leaves the extraction pipeline; wikis remain authored documents; the §4 formalization path dies.
- Source: Companion §11.2 (leak named in §6; thesis consequence in §1 and §11.8).
- Tier: gated — requires the Claim layer, contradiction check, origin mark.
CR-24 — Artifact engine: stance legibility with citation fidelity
- Statement: The engine renders stance-true artifacts from the same Claim set without loss of citation fidelity — stance conditions treatment, not truth.
- Metric: Primary: stance legibility — two blinded raters per render, each shown an artifact and both wikis of an opposing pair, asked which governed it; a render is legible when both raters name the governing wiki (the anchor's §10.6 both-raters convention), κ reported; scored with a cluster bootstrap over intents (renders arrive in pairs sharing an intent and a wiki pair and share raters within a workspace — a binomial over renders would inflate the nominal level under correlated guessing). Secondary: ALCE citation recall/precision of stance-conditioned renders non-inferior to the stanceless control, paired per intent. The anchor's precondition check on engine renders and on the display surfaces of 20 wikis per workspace (§2 extends renderer obligations to every family member).
- Dataset/inputs: 20 intents per workspace, each rendered under both wikis of a §11.2 opposing pair (40 artifacts per workspace, 120 pooled in 60 intent clusters) plus one stanceless render (empty governing set) per intent as control; frozen Claim set, standalone engine runs (no derivation stages), ≤1 governing wiki per render — the per-Claim conditioning assignment is never exercised (field arithmetic with a deterministic rejection rule; no experiment of its own).
- Threshold: Legibility above the 0.5 chance rate (conservative bound for the both-raters statistic: independent guessing sits at 0.25, perfectly correlated at 0.5 — valid at any rater correlation), one-sided at the protocol's per-look level, design effect 0.8 (what the test is powered for, not a second bar; power above the floor at 120 renders / 60 clusters even at full within-pair correlation); citation recall/precision non-inferior within 0.05 absolute; precondition check 100% (any miss a defect).
- Drop condition: If stance-conditioning costs citation fidelity beyond the margin and the redesign (move stance source from prose conditioning to the wiki's citation edges — the wiki selects and orders what renders, prose stops steering characterization) does not recover it, the governing-wiki input is dropped and the engine renders stanceless — the weakened second read path of §1. Redesign trigger: legibility point estimate <0.8 with citations intact.
- Source: Companion §11.3 (mechanism in §7; falsifier pointer at §7's close).
- Tier: gated — requires the artifact engine and the Claim layer.
CR-25 — ALCE citation-fidelity floor (current pipeline)
- Statement: ALCE citation recall and precision of current wiki bodies against their cited signals, pre-registered as the citation-fidelity floor the engine's renders must not come in below.
- Metric: ALCE citation recall and citation precision, NLI-computed with the 20% human check.
- Dataset/inputs: Wiki bodies and their cited signals as stored in the deployed system (
wikis.content; SIGNAL_CITED_BY_WIKI edges plus per-section citation declarations — verified in code at the 2026-08-31 proxy screen); an NLI model. - Threshold: None on the measurement itself — it establishes the floor; the engine's renders (CR-24) must clear it, and the render-time ALCE gate (§7 stage iv) uses it as the pre-registered floor.
- Drop condition: n/a (floor measurement).
- Source: Companion §11.3 (runnable-today floor form), §11.7.
- Tier: runnable — the deployed system stores wiki bodies and cited signals; no new architecture. Pre-registered to run first (§11.7).
CR-26 — Live-artifact staleness and stamp coverage
- Statement: The stamp-and-lazy-re-render machinery keeps live artifacts within one batch interval of the knowledge they project.
- Metric: Staleness rate — fraction of live-artifact reads serving a version that cites a Claim whose standing dropped, or a Signal whose support dropped, more than Δ plus one render time earlier without re-rendering. Stamp coverage — fraction of wikis and live artifacts citing a changed object that carry the stamp within Δ plus the batch's processing time, computed separately for the four sources: propagation (injected contradictions), user rejection, supersession, cited-Signal support withdrawal. Re-renders per artifact per checkpoint (the thrash measure the debounce bounds); mean re-render cost as a new ledger line (rendering stages only — a re-render runs no derivation, so the line is disjoint from the anchor's W(m) and adds no writes).
- Dataset/inputs: The anchor's §10.5 injected replay series, extended: 10 pinned live artifacts per workspace at m = 10, each citing ≥5 accepted Claims and ≥3 Signals that justify no Claim (the deployed norm the §3.2 Signal hook exists for); contradiction set constructed to include targets cited by pinned artifacts; reads of pinned artifacts replayed between checkpoints; an added checkpoint at m = 15 — for ten pinned-artifact citation targets per workspace, 5 user rejections and 5 explicit supersessions through ordinary write paths, plus withdrawal of the source Entries behind 10 cited Signals per workspace that justify no Claim (a support drop with no Claim standing moving). Artifact requests run the anchor's full slow path unchanged; the experiment adds rendering-side machinery only.
- Threshold: Staleness below 2% at every measurable checkpoint (mirroring the anchor's bar); stamp coverage 100% on each of the four sources (construction guarantee — the stamp rides standing and support writes; a miss names the write path that skipped the hook); re-renders ≤1 per artifact per interval (debounce guarantee).
- Drop condition: If staleness cannot be held under lazy or eager (stamp-time re-render for artifacts above a read-frequency threshold, cost difference charged to the ledger) policy, the live family is cut to terminal — artifacts regenerated on demand and never advertised as current.
- Source: Companion §11.4 (mechanism in §3.2 and §7; falsifier pointer at §7's close).
- Tier: gated — requires the artifact engine, the Claim layer, and the replay harness.
CR-27 — Label-free eviction: engagement as salience proxy
- Statement: Query engagement predicts the Useful screen's verdict well enough to order provisional Dimensions for eviction until β exists.
- Metric: Primary: Spearman rank agreement between trailing engagement at fit time and mean out-of-fold β_{D,d} from the first Useful screen, across Dimensions, bootstrap interval over Dimensions. Secondary (Holm-corrected): engagement's agreement exceeds the recency ordering's (age since activation, the naive baseline) — paired one-sided bootstrap of the coefficient difference, both computed on each resample (respects the correlation the orderings inherit from sharing the β ranking).
- Dataset/inputs: Every Dimension crossing the anchor's fitting minimum during the pilot with a full engagement window (90 days initially) before first fit; floor of 28 Dimensions pooled (below it: reported as underpowered, cap-hit human escalation remains the only eviction path — stated in advance so an underpowered result cannot read as a pass).
- Threshold: One-sided 97% bootstrap lower bound on the Spearman coefficient above zero (the protocol's nominal 0.03 per look under the two-look extension; the second look pools further evaluation cycles until they contribute as many Dimensions again as the first look, since the extension is by the same size and the per-look boundary assumes equally sized looks); design effect 0.5 (≈80% power at the 28-Dimension floor, Fisher-z approximation).
- Drop condition: Lower bound at or below zero → engagement is uninformative or anti-informative; eviction reverts to cap-hit human escalation only; the engagement log is kept only if §6's shadow-suggestion machinery wants it. Redesign: significant but weak (bound >0, point estimate <0.5) → rule demoted to tie-breaking beneath human escalation, reported as such.
- Source: Companion §11.5 (mechanism in §8, replacing the anchor's §5.5 β-ranked eviction, which breaks under autonomous creation; falsifier pointer in §8).
- Tier: gated — requires active Dimensions, the query-projection log, and ≥1 Domain at the fitting minimum.
CR-28 — The dial: auto-activation quality
- Statement: Dimensions activated autonomously (five gates, no Useful screen, no human review) are not materially worse, as judged by people, than Dimensions a human reviewed before activation — the measurable statement of auto-activation quality, and the pre-registered result that would flip the shipped default.
- Metric: Primary: autonomous arm's meaningfulness rate (anchor §10.6 protocol: two raters, both must say yes, κ reported). Secondary: non-inferiority of autonomous vs review arm. Both cluster-bootstrapped over Domains (the design randomizes Domains; Dimensions within a Domain share subject matter, Claim base, and rater context — an unclustered bound would declare non-inferiority more easily than the data license). Also reported: review-arm approval rate (approval ≥0.9 means review is latency, not a quality bar — itself evidence about the default); activation-projection spend per arm (gated count × c_π per activation at the arm's disposal rate); the §8 engagement-log line per arm.
- Dataset/inputs: Domains cluster-randomized between the autonomous and review dial positions for an evaluation cycle; sample floor 40 Domains per arm with ≥150 activated Dimensions per arm, accumulated over as many cycles as needed, stated in advance; power conditional on assumed within-Domain correlation 0.3 at ~4 Dimensions per Domain (design effect ≈1.9, effective sample ≈79 per arm, half-width ≈0.13 at worst-case rates near 0.5); observed correlation reported and power recomputed.
- Threshold: Autonomous meaningfulness ≥0.5 (the anchor's own §10.6 bar); non-inferiority within 0.20 absolute (one-sided 95% bound on the rate difference); ≈80% power for non-inferiority at equal true rates once the floor is met; below the floor the comparison is reported as underpowered and moves nothing.
- Drop condition: Primary failure flips the shipped default from autonomous to review. Non-inferiority failure (scored only at the sample floor) also flips it — but only when the review arm's approval rate is below 0.9; at ≥0.9 the result is reported without moving the default. Either flip is the autonomy half of the thesis failing, by design: the AI still authors (there is no manual end), but a human approves each Dimension.
- Source: Companion §11.6 (dial mechanism in §8).
- Tier: gated — requires the discovery pipeline and the five gates.
CR-29 — Companion thesis falsification (composite)
- Statement: The paper-level verdict over the companion's mechanisms.
- Metric/threshold: Perspectival half fails if CR-23's condition lands: the leak rate reaches its 0.10 threshold and the per-Signal classifier fallback, when built, cannot hold the 0.7 precision bar it inherits from the anchor's §10.3 (wiki prose leaves extraction; wikis remain authored documents; the §4 formalization path dies). Weaker engine-shaped failure if CR-24 shows stance-conditioning and citation fidelity cannot coexist (the wiki survives but not as a stance source; the engine renders stanceless — which the anchor effectively already promised). Autonomy half fails if CR-28 flips the default. Each mechanism also fails alone by its own drop condition — including CR-21's provision that the exclusion rule is removed as dead machinery if the contamination it forecloses does not occur — without taking the thesis with it.
- Dataset/inputs: Aggregates CR-20 through CR-28.
- Source: Companion §11.8 (restating §1's failure statement).
- Tier: gated — composite over gated criteria.
Observed failures with no criterion
Listed here rather than promoted to criteria, per the register's charter: these are failure modes the papers name without stating a pass threshold.
- Capture-report fidelity and refusal handling. The Entry extraction contract (companion §3.1) requires that a refused input be recorded on the Entry ("refusing is honest in a way silent truncation is not"), and the deployed extractor persists a capture report (extracted, dropped, band) per Entry. Neither paper states a metric or threshold for refusal correctness or the report's own accuracy — §11.1 measures completeness of what extraction attempted, and the companion itself notes "the report counts only what the breaker trimmed; nothing measures what extraction never proposed" (that gap is what CR-20 closes; the report's fidelity remains unmeasured).
- Evidential circularity probe. Companion §6 names a second hole the exclusion closes: a single-source Claim cited into a wiki can receive its own restatement back as an apparently independent second Entry and auto-accept on its echo (manufactured independence — the anchor's ancestry check cannot see it because the loop travels through prose). The supports-branch exemption forecloses this by construction, but §11.2 states no metric that counts supports-branch entries by stance-bearing Signals or auto-acceptances whose second "distinct Entry" is a wiki revision. If the owner wants the closure verified rather than trusted, a probe (cite a single-source Claim into a wiki, revise, check no auto-acceptance) needs a stated criterion.
Resolved since the v2 register: the quoting-Entry stance leak now carries its own threshold (0.10) and disposition, and the fallback classifier's precision bar is fixed (≥0.7), both in companion v3 §11.2 — the former item 1 is promoted into CR-23.
Historical failures that DO have criteria (not listed above): the positional discard of ~45% of dense notes is covered by CR-20's positional-uniformity clause; the broken β-ranked eviction rule under autonomous creation is covered by CR-27.
Acceptance review
Adversarial acceptance pass, 2026-08-31. Each criterion was screened for refutation on five tests: (1) measurable as stated, (2) threshold justified by the source paper, (3) inputs obtainable or construction stated, (4) traceable to a mechanism the papers specify with correct citation, (5) tier honest against the live codebase at /home/drew/.local/projects/master.withrobin.ai (HEAD 7f379974, 2026-08-28). This is a proxy screen only: the header's PENDING OWNER ACCEPTANCE status stands, and no verdict below binds the owner. Verdict count: 28 accepted, 1 rejected (CR-10, marked inline above).
- CR-01 — ACCEPT. Arms, nDCG@10, 0.03 design effect, permutation control, 10% cost bound, and drop condition all stated in anchor §10.1; gated tier honest (no Claims layer or projection machinery in code).
- CR-02 — ACCEPT. Exactly the secondary clause anchor §10.1 states, measurable as the (ii)−(i) significance test on the same arms as CR-01.
- CR-03 — ACCEPT. Honestly registered as a decision rule, not pass/fail; "matches within the noise" resolves to a defined significance test under the shared protocol and feeds CR-19 as §10.1/§10.8 specify.
- CR-04 — ACCEPT. Binomial design with 60% design effect (~86+/150 rejection region), ALCE secondaries, 100% precondition check, and the three-way drop/redesign reading are §10.2 verbatim.
- CR-05 — ACCEPT. Thresholds (0.8/0.7/0.9, ≤k+1 calls, Δ+1 min) and the recall-0.6-at-k=50 drop floor all in §10.3; annotator injection construction stated; branch-tagged scoring makes it attributable.
- CR-06 — ACCEPT. Filter-vs-similarity comparison at fixed k and the candidate-recall-below-0.9 removal condition are §10.3 verbatim, including the §5.3 signature consequence.
- CR-07 — ACCEPT. Construction guarantee (set equality both directions, restoration to proposed never auto-accepted), correctly registered as defect-not-statistic per the protocol preamble.
- CR-08 — ACCEPT. Design effect 0.02, ~350 pooled query sizing, η/γ settings, and the pre-registered raise-and-report fallback for a short count are all in §10.4.
- CR-09 — ACCEPT. Secondary of §10.4 powered for 0.05; the read-with-CR-18 drop interpretation is the anchor's own.
- CR-10 — REJECTED. Fails test (1); reason inline at the criterion. The measurability-gap note the register already carried is converted to a rejection: a criterion the register itself says cannot bind is not accepted at a proxy screen.
- CR-11 — ACCEPT. Decision rule with both outcomes defined and a conservative default; measurable as a significance comparison on CR-08's query sets.
- CR-12 — ACCEPT. Replay construction (checkpoints m = 0/5/10/20, one-time coefficient fit, post-replay pooling pass) fully stated in §10.5; gated tier honest — no replay harness exists.
- CR-13 — ACCEPT. The 2% bar, the not-measurable-below-20-Claims rule, the zero-retracted sanity guarantee, and the Graphiti-arm cost comparison are all §10.5.
- CR-14 — ACCEPT. The N* formula and the N* ≤ N̄·m condition are §10.5 verbatim; every input is measured on a named series or pilot logs, and the r_fast < r_base sanity role is preserved.
- CR-15 — ACCEPT. 50% both-raters bar and the human-authored-Dimensions fallback are §10.6 verbatim; two-rater metric sits inside the shared protocol.
- CR-16 — ACCEPT. "§5.4's Measurable thresholds" resolve to numbers in the anchor (r_s ≥ 0.7 scalar/ordinal, κ ≥ 0.6 categorical — verified at this pass), so the ≥3/4-of-active-Dimensions clause is checkable; the 1/10 cost ratio is §10.6's.
- CR-17 — ACCEPT. 0.9 accuracy and 0.05 shallow-misroute bars in §10.7; correctly registered as redesign-not-drop; the fine-tuned-component posture tension is flagged where it belongs.
- CR-18 — ACCEPT. Bars (macro-F1 ≥ 0.85, per-Domain recall ≥ 0.75) in §10.7; runnable tier re-verified at this pass:
packages/agent/src/stages/domain-classify.ts, domain_signals.source 'classifier' (server/src/db/schema.ts:1579), per-domain classifierPrompt (server/src/db/schema.ts:1484). Only annotation is missing, which the tier definition treats as input. - CR-19 — ACCEPT. Composite is a stated function of component verdicts (§10.8) and adds no thresholds of its own; CR-10's rejection does not touch it (CR-10 feeds §10.4's bundle, not the thesis conditions).
- CR-20 — ACCEPT, caveat standing. Recall and uniformity clauses are measurable, thresholded, and runnable per §11.1. The caveat was re-verified at this pass:
server/src/db/schema.ts:584 declares provenance.span; packages/agent/src/stages/persist.ts writes extractionJobId/model/sourceTitle/sourceEntities/sourceTerms and no span — so the companion's "source spans exist in the deployed system" overclaims, the register's correction is accurate, and the precision guard runs only after span-writing lands or the guard is respecified against the full Entry payload (which still guards hallucination, at the price of positional attribution). Tier honest because the register qualifies it precisely. - CR-21 — ACCEPT. Null 2%, design effect 10%, cluster bootstrap over 60 authored sets with the stated rejection region, the load-bearing acceptance qualification, and the full-range disposition (including the dead-machinery drop) are §11.2 verbatim; the authored-wiki construction is stated.
- CR-22 — ACCEPT. 0.95 survival/removal bars stated in §11.2 with their false-negative/false-positive readings; these are threshold clauses, not comparative tests, so the protocol's design-effect requirement is not dodged.
- CR-23 — ACCEPT. The 0.10 leak threshold, the 20% both-annotators shadow-run trigger, and the 0.7 classifier precision bar are all fixed in companion v3 §11.2; the thesis-level consequence matches §11.8; the third-trained-component charge is carried.
- CR-24 — ACCEPT. The 0.5 conservative null (with the rater-correlation argument that makes it valid), cluster bootstrap over 60 intents, 0.8 design effect, 0.05 non-inferiority margin, and 100% precondition check are §11.3 verbatim.
- CR-25 — ACCEPT. Floor measurement honestly registered without a threshold — its output is CR-24's bar and the §7 stage-iv gate's floor. Runnable tier re-verified at this pass:
wikis.content, SIGNAL_CITED_BY_WIKI edges (server/src/db/schema.ts:662), per-section citation declarations attached in server/src/lib/wikiSidecar.ts. - CR-26 — ACCEPT. 2% staleness, 100% stamp coverage per source, ≤1 re-render per interval, the four stamp sources, and the added m = 15 checkpoint construction are §11.4 verbatim; gated tier honest.
- CR-27 — ACCEPT. 97% one-sided bound at the per-look level, 0.5 design effect with Fisher-z power at the 28-Dimension floor, and pre-stated underpowered handling are §11.5; the 90-day engagement window the register cites is defined in companion §8 (verified).
- CR-28 — ACCEPT. 0.5 bar inherited from anchor §10.6, 0.20 non-inferiority margin, cluster floor (40 Domains / 150 Dimensions per arm) with explicit power arithmetic, and the approval-rate ≥ 0.9 qualifier on the flip are §11.6 verbatim.
- CR-29 — ACCEPT. Composite restates §11.8's failure statement exactly, including CR-21's dead-machinery provision; adds no thresholds of its own.
The two items under "Observed failures with no criterion" remain correctly unpromoted: neither paper states a threshold for them, and promoting them here would manufacture criteria the sources do not carry.