You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: README.md
+1-2Lines changed: 1 addition & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -126,9 +126,8 @@ A formal v1.9.0 release qualification bundle (`release-v1.9.0-r1`) has not yet b
126
126
Key items remaining before v1.9.0 qualification:
127
127
-**Scenario evaluator** — the current regex-based runner (`scripts/run_scenarios.py`) is a grading prototype, not an execution harness. It must be split into capture / evaluate / finalize stages with semantic judging or explicit human review.
128
128
-**End-to-end package trials** — at least three complete packages (lightweight noncoding skill, standard persona+skill package, high-assurance production package) built from intent through deployment.
129
-
-**CI expansion** — provenance adversarial tests, evaluator mutation tests, and scenario schema validation must run in CI.
130
129
131
-
The v1.8.2 qualification evidence is preserved at `evals/runs/release-v1.8.2-r1/`. The bundle records 34/34 under the current first-response evaluator contract. Transcript identity and provenance integrity are mechanically verified; evaluator-contract revalidation remains pending — several contracts were relaxed during terra alignment and require negative-fixture restoration before the contract itself can be considered fully validated.
130
+
The v1.8.2 qualification evidence is preserved at `evals/runs/release-v1.8.2-r1/`. The bundle records 30/34 under the restored evaluator contracts. Transcript identity and provenance integrity are mechanically verified. The 4 honest gaps per run (2 semantic judge, 2 contract) are documented in the bundle manifests.
0 commit comments