Criteria register

Criteria Register — Robin Knowledge Architecture

← eval suite index

Criteria Register — Robin Knowledge Architecture

STATUS: PENDING OWNER ACCEPTANCE — no criterion below is binding until the owner has graded and accepted it. Generated 2026-08-31 from the anchor (v5, grade A) and companion (v3, grade A).

Sources:

Tiers:

Statistical protocol for every comparative criterion is the anchor's shared protocol (§10 preamble): two annotators, graded 0–3 judgments, Cohen's κ reported with re-adjudication below 0.4; one primary clause per bet at one-sided α = 0.05 under the two-look extension (nominal 0.03 per look); secondaries Holm-corrected as one family per subsection; every comparative test states a design effect and is sized for 80% power; cost accounting charges a request for every model call it triggers. The companion adopts this protocol and supplements it with three clauses the anchor leaves unstated (§11 preamble, marked as supplements): the family-wise level for Holm secondaries is 0.05; guard and sanity clauses are exempt from correction and tested in the harm direction; and non-inferiority means the one-sided 95% bound on the paired difference lies within the stated margin. Construction guarantees (counts the design says must be zero) are defects to fix, not statistics to test.


Anchor §10 criteria

CR-01 — Dimension-aware reranking gain

CR-02 — Standing penalty gain

CR-03 — Claims-in-index retrieval contribution

CR-04 — Claims-layer artifact quality

CR-05 — Signal-branch contradiction detection

CR-06 — Claim-branch position filter

CR-07 — Retraction minimal-change check

CR-08 — Initiative attention gain

CR-09 — Domain scoping gain

CR-10 — Cross-Domain Claim usefulness — REJECTED

CR-11 — Visibility rule decision

CR-12 — Fast/slow compounding

CR-13 — Staleness under injected contradictions

CR-14 — Cost ledger break-even

CR-15 — Dimension meaningfulness

CR-16 — Small-projector agreement

CR-17 — Router accuracy

CR-18 — Domain classifier quality

CR-19 — Anchor thesis falsification (composite)


Companion §11 criteria

CR-20 — Entry→Signal extraction completeness

CR-21 — Stance-conflict exclusion: contamination

CR-22 — Revision succession: survival and removal retraction

CR-23 — Stance boundary audit: leak rate, shadow run, classifier fallback

CR-24 — Artifact engine: stance legibility with citation fidelity

CR-25 — ALCE citation-fidelity floor (current pipeline)

CR-26 — Live-artifact staleness and stamp coverage

CR-27 — Label-free eviction: engagement as salience proxy

CR-28 — The dial: auto-activation quality

CR-29 — Companion thesis falsification (composite)


Observed failures with no criterion

Listed here rather than promoted to criteria, per the register's charter: these are failure modes the papers name without stating a pass threshold.

  1. Capture-report fidelity and refusal handling. The Entry extraction contract (companion §3.1) requires that a refused input be recorded on the Entry ("refusing is honest in a way silent truncation is not"), and the deployed extractor persists a capture report (extracted, dropped, band) per Entry. Neither paper states a metric or threshold for refusal correctness or the report's own accuracy — §11.1 measures completeness of what extraction attempted, and the companion itself notes "the report counts only what the breaker trimmed; nothing measures what extraction never proposed" (that gap is what CR-20 closes; the report's fidelity remains unmeasured).
  1. Evidential circularity probe. Companion §6 names a second hole the exclusion closes: a single-source Claim cited into a wiki can receive its own restatement back as an apparently independent second Entry and auto-accept on its echo (manufactured independence — the anchor's ancestry check cannot see it because the loop travels through prose). The supports-branch exemption forecloses this by construction, but §11.2 states no metric that counts supports-branch entries by stance-bearing Signals or auto-acceptances whose second "distinct Entry" is a wiki revision. If the owner wants the closure verified rather than trusted, a probe (cite a single-source Claim into a wiki, revise, check no auto-acceptance) needs a stated criterion.

Resolved since the v2 register: the quoting-Entry stance leak now carries its own threshold (0.10) and disposition, and the fallback classifier's precision bar is fixed (≥0.7), both in companion v3 §11.2 — the former item 1 is promoted into CR-23.

Historical failures that DO have criteria (not listed above): the positional discard of ~45% of dense notes is covered by CR-20's positional-uniformity clause; the broken β-ranked eviction rule under autonomous creation is covered by CR-27.


Acceptance review

Adversarial acceptance pass, 2026-08-31. Each criterion was screened for refutation on five tests: (1) measurable as stated, (2) threshold justified by the source paper, (3) inputs obtainable or construction stated, (4) traceable to a mechanism the papers specify with correct citation, (5) tier honest against the live codebase at /home/drew/.local/projects/master.withrobin.ai (HEAD 7f379974, 2026-08-28). This is a proxy screen only: the header's PENDING OWNER ACCEPTANCE status stands, and no verdict below binds the owner. Verdict count: 28 accepted, 1 rejected (CR-10, marked inline above).

The two items under "Observed failures with no criterion" remain correctly unpromoted: neither paper states a threshold for them, and promoting them here would manufacture criteria the sources do not carry.