We measure structure (graph density, backlinks, freshness, fidelity labels, retrieval) but content accuracy has only ever been spot-checked manually (T1 found a real factual error; T3 found a wiki self-contradiction that a downstream model inherited verbatim). Now that raw/ is verbatim (PR #22), claim-level verification is finally mechanically possible. Build grimoire audit: (1) sample N claims per article, check support in the cited verbatim raw source (LLM-judged, but against ground truth text — auditable); (2) contradiction scan across articles; (3) citation-granularity score (T2's one real content loss vs deep-research: 104 inline URLs vs per-article sourcing); (4) auto-generated held-out retrieval eval per wiki. Output: a quality report artifact in wiki/.compile/ + badges in present + a section in grimoire_coverage_gaps. This becomes guarantee G6: 'the KB measures its own accuracy and shows you the number.'
We measure structure (graph density, backlinks, freshness, fidelity labels, retrieval) but content accuracy has only ever been spot-checked manually (T1 found a real factual error; T3 found a wiki self-contradiction that a downstream model inherited verbatim). Now that raw/ is verbatim (PR #22), claim-level verification is finally mechanically possible. Build
grimoire audit: (1) sample N claims per article, check support in the cited verbatim raw source (LLM-judged, but against ground truth text — auditable); (2) contradiction scan across articles; (3) citation-granularity score (T2's one real content loss vs deep-research: 104 inline URLs vs per-article sourcing); (4) auto-generated held-out retrieval eval per wiki. Output: a quality report artifact in wiki/.compile/ + badges in present + a section in grimoire_coverage_gaps. This becomes guarantee G6: 'the KB measures its own accuracy and shows you the number.'