Skip to content

Commit aeb75d7

Browse files
djbclarkclaude
andcommitted
bin: lint cross-references in the prose corpus
The corpus navigates by `§N` and by relative link across 71 documents and 544 sections, and nothing has ever checked that those pointers land. Two bad section refs were written by hand in a single session, both in top-authority documents, which is what prompted this. Stated plainly in the docstring, because it limits the tool's worth: this catches DANGLING references, not MIS-AIMED ones. A `§14` that should have been `§7` resolves fine and passes silently. That is the class that actually bit us; only reading the target catches it. Sections are identified the four ways the corpus writes them — ATX headings, already-dotted headings, a bare `### A.` composed with its parent (`16.A`), the paper's bold pseudo-headings, and the map's `- **§14.2**` list items. Missing that last form was a bug in the first draft that produced 75 false positives, nearly all of them the same `§14.2`; one word of the corpus disagreeing with a tool is a finding, the whole corpus disagreeing is a bug in the tool. Qualifier matching is deliberately tight for the same reason: a loose window read "E1 R4/§9.8" as an E1 reference when §9.8 belongs to the map. Findings are suppressed when an unqualified ref resolves in some sibling document, since bare `§5.1` meaning "E1 §5.1" is established house prose, not a dangling pointer. 47 findings on first clean run, all left unfixed — the fix pass is a separate, operator-gated job. Most sit in `reviews/` and `deprecated/`, which README.md marks as an evidence trail not to be rewritten. Five are in live design documents, and two of those are real: the reconciliation cites "Grok §9.12" twice, and the Grok opinion has sections 0-13 with no 9.12 anywhere. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent d0c306f commit aeb75d7

1 file changed

Lines changed: 160 additions & 0 deletions

File tree

‎bin/xref_lint.py‎

Lines changed: 160 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,160 @@
1+
#!/usr/bin/env -S uv run --script
2+
# /// script
3+
# requires-python = ">=3.11"
4+
# dependencies = []
5+
# ///
6+
"""Lint the cross-references in tendcf's prose corpus.
7+
8+
This corpus is pointer-dense — the guide, the implementer map, the paper,
9+
the adjudications and GLOSSARY.md all navigate by `§N` and by relative
10+
link — and nothing has ever checked that those pointers land.
11+
12+
Two layers, cheapest first:
13+
14+
1. every relative markdown link resolves to a file that exists;
15+
2. every `§N` reference resolves to a section that exists — in the
16+
document that makes it, or, when the reference is qualified
17+
("guide §4", "E1 §5.6", "map §13"), in the document it names.
18+
19+
WHAT THIS CANNOT DO, stated up front so nobody trusts it further than it
20+
goes: it catches *dangling* references, not *mis-aimed* ones. A `§14`
21+
that should have been `§7` resolves fine and passes silently — both
22+
sections exist. That class is only caught by reading the target, and it
23+
is the class that has actually bitten this repo. This lint is a floor,
24+
not a substitute for checking what you cite.
25+
26+
Sections are identified the way the corpus writes them, which is three
27+
ways:
28+
29+
- `## 7. Title` -> 7
30+
- `### 9.2 Title` -> 9.2 (already dotted, taken as-is)
31+
- `### A. Title` under 16 -> 16.A (composed with its parent)
32+
- `**8.8 Title**` -> 8.8 (the paper's open questions are
33+
bold pseudo-headings, not ATX)
34+
35+
Exit 1 on any finding. No dependencies: this must run in a bare checkout.
36+
"""
37+
38+
from __future__ import annotations
39+
40+
import re
41+
import sys
42+
from pathlib import Path
43+
44+
ROOT = Path(__file__).resolve().parent.parent
45+
DOCS = ROOT / "docs"
46+
47+
# Qualifiers the corpus uses when pointing at a sibling document.
48+
QUALIFIERS = {
49+
"guide": "paper/tendcf-architecture-guide.md",
50+
"paper": "paper/tendcf-architecture-paper.md",
51+
"map": "architecture/architecture-DEFINITIVE-v3.md",
52+
"e1": "architecture/e1-adjudication-xhigh-2026-08-15.md",
53+
"fable": "architecture/goal-file-schema-opinion-fable.md",
54+
"grok": "architecture/goal-file-schema-opinion-grok.md",
55+
"gemini": "architecture/goal-file-schema-opinion-gemini.md",
56+
}
57+
58+
# `§5.x` is a literal placeholder ("references of the form E1 §5.x"), not a
59+
# pointer at a section called "x".
60+
PLACEHOLDER = re.compile(r"\.x$", re.IGNORECASE)
61+
62+
ATX = re.compile(r"^(#{1,6})\s+([0-9]+(?:\.[0-9A-Za-z]+)*|[A-Z])\.?\s+\S")
63+
BOLD = re.compile(r"^\*\*([0-9]+\.[0-9A-Za-z]+)\s+\S")
64+
# `- **§14.2** Title` — the map defines its residue subsections as list items.
65+
ITEM = re.compile(r"^\s*[-*]\s+\*\*§?([0-9]+\.[0-9A-Za-z]+)\*\*")
66+
# A §ref, qualified only when a document name IMMEDIATELY precedes it.
67+
# The window must stay tight: "E1 R4/§9.8" cites the map's §9.8, not E1's.
68+
REF = re.compile(
69+
r"(?:\b(guide|paper|map|E1)\s+)?§\s?([0-9]+(?:\.[0-9A-Za-z]+)*)",
70+
re.IGNORECASE,
71+
)
72+
LINK = re.compile(r"\[[^\]]*\]\(([^)#\s]+)(?:#[^)]*)?\)")
73+
FENCE = re.compile(r"^\s*```")
74+
75+
76+
def sections(path: Path) -> set[str]:
77+
"""Every section id the document defines."""
78+
found: set[str] = set()
79+
parent = None
80+
for line in path.read_text(encoding="utf-8").splitlines():
81+
m = ATX.match(line)
82+
if m:
83+
label = m.group(2)
84+
if label.isdigit():
85+
parent = label
86+
found.add(label)
87+
elif "." in label:
88+
found.add(label)
89+
parent = label.split(".")[0]
90+
elif parent: # a bare letter: `### A.` under `## 16.`
91+
found.add(f"{parent}.{label}")
92+
continue
93+
b = BOLD.match(line)
94+
if b:
95+
found.add(b.group(1))
96+
continue
97+
i = ITEM.match(line)
98+
if i:
99+
found.add(i.group(1))
100+
return found
101+
102+
103+
def prose_lines(path: Path):
104+
"""Yield (lineno, text), skipping fenced code so examples aren't linted."""
105+
in_fence = False
106+
for n, line in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
107+
if FENCE.match(line):
108+
in_fence = not in_fence
109+
continue
110+
if not in_fence:
111+
yield n, line
112+
113+
114+
def main() -> int:
115+
docs = sorted(DOCS.rglob("*.md"))
116+
index = {p: sections(p) for p in docs}
117+
qualified = {
118+
k: (DOCS / v) for k, v in QUALIFIERS.items() if (DOCS / v).exists()
119+
}
120+
findings: list[str] = []
121+
122+
for path in docs:
123+
rel = path.relative_to(ROOT)
124+
own = index[path]
125+
for n, line in prose_lines(path):
126+
for target in LINK.findall(line):
127+
if target.startswith(("http://", "https://", "mailto:")):
128+
continue
129+
if not (path.parent / target).exists():
130+
findings.append(f"{rel}:{n}: dead link -> {target}")
131+
132+
for qual, ref in REF.findall(line):
133+
if PLACEHOLDER.search(ref):
134+
continue
135+
q = qual.lower()
136+
if q and q in qualified:
137+
where, name = index[qualified[q]], f"{qual} §{ref}"
138+
else:
139+
where, name = own, f"§{ref}"
140+
if ref in where:
141+
continue
142+
# An unqualified ref that lands in some sibling document is
143+
# ambiguous prose, not a dangling pointer. Only report refs
144+
# that resolve nowhere at all.
145+
if not q and any(ref in s for s in index.values()):
146+
continue
147+
findings.append(f"{rel}:{n}: {name} resolves nowhere")
148+
149+
for f in findings:
150+
print(f)
151+
print(
152+
f"\n{len(docs)} documents, "
153+
f"{sum(len(s) for s in index.values())} sections, "
154+
f"{len(findings)} finding(s)."
155+
)
156+
return 1 if findings else 0
157+
158+
159+
if __name__ == "__main__":
160+
sys.exit(main())

0 commit comments

Comments
 (0)