PR #21's 10/10 was measured on the same 10 questions the ranking was tuned against (codex read T3-brains.md including the answer key). Reproducible ≠ generalizable. Before the T3 rerun becomes public benchmark material: (a) generate a NEW question set from raw sources by a model that has not seen the ranking code or the old set; (b) pre-register expected slugs; (c) grade mechanically (slug match), not by model judgment; (d) run on a second wiki on an unrelated topic (non-docs-site sources) to test generalization.
PR #21's 10/10 was measured on the same 10 questions the ranking was tuned against (codex read T3-brains.md including the answer key). Reproducible ≠ generalizable. Before the T3 rerun becomes public benchmark material: (a) generate a NEW question set from raw sources by a model that has not seen the ranking code or the old set; (b) pre-register expected slugs; (c) grade mechanically (slug match), not by model judgment; (d) run on a second wiki on an unrelated topic (non-docs-site sources) to test generalization.