The canonical results for cap-evolve. Every number here is derived from a committed
run artifact (examples/*/run_full/*.json) or, where noted, from a held-out run whose
artifact is committed separately. The README's Results section is a short snapshot of
this page.
SkillsBench results live on their own branch. The full per-task sweeps, the evolved
skill packages, and the analysis write-ups are on
skillsbench-history
— that branch is the source of truth for SkillsBench optimization output. See
SkillsBench sweeps below for what is there and
which numbers to quote.
Each result is labeled by split discipline:
- fit metric —
train == val == test(no holdout); the test number is not held out and the engine logs asplits_warning. Useful to show the loop works; not a generalization claim. - held-out — test ids the optimizer never saw, scored exactly once at
finalize.
Reward is mean task reward in [0, 1]. Where we quote externally reported results that use 0–100% units, we label them explicitly as percentages.
Gains are given as absolute and relative %.
Read the caveats before quoting this one. It is the only entry on this page whose run artifact is not committed to this repo, so it is the only one that cannot be re-derived from
examples/*/run_full/*.json. It is recorded here because the README quotes it and points at this page; it is flagged rather than dropped so the gap is visible instead of implied.
- Capabilities:
[skill-package, system-prompt]. Harness: Harbor (each task runs a fullclaude-codeagent inside an isolated Docker container). - Optimizer:
claude-code@claude-opus-4-6. - Agent under test:
claude-sonnet-4-6via Harbor. - Tasks: 119 val tasks from SWE-bench Verified.
- Split:
train == val == test == 119— fit metric, no holdout. The optimizer saw every task it was scored on, so this is not a generalization claim.
| reward (119 tasks) | Δ vs baseline | |
|---|---|---|
| Baseline (seed prompt + skill) | 0.580 (58.0%) | — |
Best candidate (cand_0002), val |
0.765 (76.5%) | +0.185 / +31.9% relative |
2 of 7 iterations accepted: iter 1 +0.118 (0.580 → 0.698), iter 2 +0.067 (→ 0.765).
Per-task: 24 improved, 2 regressed, 93 unchanged.
- Fit metric, no holdout.
train == val == test, so the +0.185 measures how well the optimizer fit 119 known tasks, not how it generalizes. Every held-out number on this page is a stronger claim than this one. - No committed artifact. There is no
run_full/for this benchmark. The numbers above are transcribed fromsite/results.html; there is no run dir, commit hash, trial count, or cost figure in the repo to check them against.ci/benchmarks/swebench/holds task lists and split ids, not results. - The published figure disagrees with these numbers, and I could not resolve which is
authoritative.
site/assets/rh_swe_bench.png— the image README andsite/results.htmlboth embed next to this result — is a cross-model / cross-harness bar chart. Its bars readSonnet 4.6 Claude Code optimized with cap-evolve 73.1,Opus 4.6 Claude Code 63.3,Sonnet 4.6 Claude Code 55.7, and three RedHatAI/NVIDIA-Nemotron rows at 30.8 / 22.4 / 21.6. It contains neither 58.0 nor 76.5, and73.1/55.7appear nowhere else in the repo. The chart is clearly the same body of work — its optimized bar is labelled "Sonnet 4.6 Claude Code optimized with cap-evolve" and the agent under test above isclaude-sonnet-4-6— so this is two scorings of one optimization, not two unrelated experiments. What is unresolved is which scoring is authoritative: either the chart is a later re-measurement and 58.0 → 76.5 is stale, or the two use different subsets/scoring and are not comparable. Until someone with the run says which, treat both pairs as unconfirmed. Note the chart is the weaker of the two (+17.4 pp vs +18.5 pp), so this is not a case of a figure flattering the text.
| val | test | notes | |
|---|---|---|---|
| Seed prompt | 0.0 | — | no [CALC] marker |
Optimized (mock adds [CALC]) |
— | 1.0 | gate-accepted, test sealed |
Deterministic, no model call. Asserted by core/tests/test_e2e_slice.py and reproduced
by bash examples/toy_calc/run.sh.
Artifact: examples/tau2_airline/run_full/
(final.json, static dashboard under ui/). Reproduce: REPRODUCE_tau2.md.
- Capability: airline policy + tools optimized jointly (
[system-prompt, tools]). - Optimizer:
claude-code@claude-opus-4-6. - Runner + user simulator:
openai/gpt-oss-120bvia IBM RITS. - Tasks / trials: all 50 airline tasks · 10 trials each.
- Split:
train == val == test == 50— fit metric (no holdout). - Algorithm / gate:
hill-climb --focus all, 10 iterations, paired significance gatek_se 0.2. - tau2-bench commit:
8ebb7499622fc2be9b9d510d6f7a7653461f4f29.
| reward (50 tasks · 10 trials) | Δ vs baseline | |
|---|---|---|
| Baseline (seed policy + tools) | 0.536 | — |
Best candidate (cand_0007) — val |
0.712 | +0.176 / +32.8% relative |
cand_0007 — sealed test (fit metric) |
0.694 pass@1 (pass² 0.584) | — |
Accepted iterations (the rest were rejected by the gate as within-noise):
iter 1 +0.046 (0.536→0.582), iter 3 +0.052 (→0.634), iter 5 +0.036 (→0.670),
iter 6 +0.014 (→0.684), iter 7 +0.028 (→0.712). 5 of 10 iterations accepted.
What changed: deep in-code tool edits (tools.py 593 → 832 lines; policy 166 → 233
lines), not just prompt tweaks — five trajectory-verified before→after edits in
OPTIMIZATION_EXAMPLES.md; curated walkthrough in
examples/tau2_airline/DEMO.md.
Same benchmark and capability, run with a real holdout split (split_ids.json,
train=val=30, test=20) so the test number is a genuine generalization result.
| split | baseline | optimized | Δ |
|---|---|---|---|
| val (30 tasks) | 56.7 | 70.0 | +13.3 pp / +23.5% relative |
| sealed test (20 tasks, scored once) | 30.0 | 47.5 | +17.5 pp / +58.3% relative |
The held-out
run_fullartifact for this run is committed separately. Until it lands, treat these figures as the reported held-out result; the reproducible artifact-backed run above is the no-holdout fit metric.
See COMPARISON.md for how this +58.3% within-run relative held-out
gain sits next to external tool-optimization work (EvoTool, Evolutionary Context Search)
— with the important caveat that those use different benchmark versions, models, splits,
and budgets and are not an apples-to-apples comparison.
The first run driven in agent orchestration mode (orchestration_mode: agent,
algorithm_skill: agent-optimize): the conversational agent understood the benchmark, ran the
baseline, then proposed the policy edits itself, gated every candidate on the full val split, and
sealed the test once — see AGENT_ORCHESTRATION.md. Reproduce:
REPRODUCE_tau2.md.
- Capability: airline policy (
system-prompt) optimized by the agent itself (no per-iteration optimizer subprocess). - Runner + user simulator:
aws/gpt-oss-120bvia the IBM ete litellm proxy. - Split: 30 train == 30 val (fit) · 20 sealed test (held out, disjoint) — user-pinned ids.
- Candidates: 5 proposed, gate-decided each round; winner
cand_5(payment-construction discipline via thecalculatetool was the load-bearing edit; plus a basic-economy scope fix, a don't-transfer-in-scope rule, and a worked booking example). - Gate: significance / paired,
k_se 1.0.
Headline (single-trial num_trials: 1 — the pipeline default and the basis for the deterministic head-to-head):
| split | baseline (seed) | best (cand_5) |
Δ |
|---|---|---|---|
| val (30, fit) | 0.500 | 0.633 | +0.133 / +26.7% relative — gate-significant (Δ > k·SE) |
| sealed test (20, held-out, scored once) | 0.400 | 0.550 | +0.150 / +37.5% relative |
Stable re-evaluation (num_trials: 3, for honesty about variance):
| split | baseline | best | Δ |
|---|---|---|---|
| val (30, fit) | 0.544 | 0.644 | +0.100 / +18.4% — paired-significant (Δ/SE = 1.80) |
| test (20, held-out) | 0.467 | 0.400 | −0.067 (SE 0.105 — not significant) |
Head-to-head vs deterministic orchestration (same split, same gpt-oss-120b, same proxy): the
bounded deterministic hill-climb run (claude-code optimizer) proposed candidates reaching val
0.567 but none cleared the gate, so its best stayed the seed and its sealed test was 0.35 (Δ 0).
Agent mode produced the only gate-accepted improvement and the only positive held-out test.
Honest reading. τ²-Bench airline is high-variance at num_trials: 1: the single-trial numbers
above are real observations but noisy (the same policy drew val 0.63–0.73 across runs). The stable
n=3 numbers are the sober view — a paired-significant val fit gain that, because train == val,
is fitting rather than generalization, and a held-out test that is flat within noise. The
requested 0.78 / >75% val target was not reached: gpt-oss-120b's stable val ceiling on this set
is ~0.64 (the opus-driven, 10-trial no-holdout run above peaked at 0.712). What agent mode did show
honestly: it drove the whole loop itself, gated on val, sealed the test once, beat the deterministic
run head-to-head, and improved the single-trial held-out test +37.5%. A genuine, stable held-out gain
here needs a stronger runner model or tool-level edits (which drove most of the lift in the
no-holdout run).
τ²-Bench airline — agent-optimize on gpt-oss-120b, OFFICIAL 30(=val)/20 split, 5 rounds — first accepted edit after four null results, and the four measurement defects that caused them
The fifth agent-mode attempt on airline, and the first to accept anything. It is documented in full because the four preceding null results were not caused by a weak optimizer: they were caused by four defects in the measurement and acceptance machinery. All four are now fixed and pinned by tests.
- Capability: airline policy + tools (
[system-prompt, tools]),orchestration_mode: agent. - Runner + user simulator:
aws/gpt-oss-120b(both roles) via the IBM ete litellm proxy (OpenAI-compatible). Provider-default sampling — noreasoning_effortoverride — so the numbers stay comparable to the four prior runs and torun_full. - Split: tau2's own
data/tau2/domains/airline/split_tasks.json— train 30 (= val 30) / test 20, disjoint,test_used: falseuntil the single sealed measurement. - Gate: paired,
k_se 1.0. The no-regression veto is off (defect 1). - Every round evaluates
ctl_null_i<n>, a byte-identical copy of the current best, so each round reports its own noise floor and the control cannot be skipped or overwritten. - Cost: 1620 gating rollouts + the sealed measurement, ~101M runner tokens, ~6.6 h of runner time. The proxy reports $0/call (unmetered, not free), so the run is bounded by rollouts.
The identical, unchanged best candidate was re-measured once per round on the same 30 tasks × 3 trials:
| round | tag (byte-identical copy of the best) | val |
|---|---|---|
| 1 | cand_toolguard (the artifact itself) |
0.6778 |
| 2 | ctl_null |
0.6667 |
| 3 | ctl_null_i1 |
0.6778 |
| 4 | ctl_null_i3 |
0.6444 |
| 5 | ctl_null_i4 |
0.7000 |
Mean 0.669, SD 0.021, spread 5.6 points — from re-measurement alone. In round 5 the control scored the highest val number of the entire run. Any single-round delta below ~5 pp on this benchmark is indistinguishable from noise, which is precisely why four earlier runs alternately accepted nothing and chased phantom regressions. The accepted edit below is +8.9 pp — 4.3 SD — which is why it is believable.
1. The no-regression veto rejected noise. It vetoed any mean gain that dropped a val task the
parent had passed. run_agentoptv4's dashboard records "paired Δ > k·SE but VETOED by
regression" for both candidates that passed the significance test, and a byte-identical seed
copy "vetoed at an EXACTLY equal mean". This run measured the mechanism directly: ctl_null,
a byte-identical copy of the seed, reported 4 regressions (tasks 10, 20, 4, 5). Regressions are
now REPORTED as diagnosis and never veto (--veto-regressions restores the old behaviour). The
round-1 winner's regressions [10, 20] were a strict SUBSET of the control's — noise. The old
rule would have vetoed the only accepted edit this benchmark has produced in agent mode.
2. A starved evaluation reported a confident zero. At TAU2_MAX_CONCURRENCY=300 the proxy
queues rather than serves: per-call latency went ~20 s → ~200 s and 292 of 300 rollouts ended
TerminationReason.TIMEOUT with a median trace of six messages, so val "measured" 0.0067
and nothing objected. TIMEOUT now routes into the same path as INFRASTRUCTURE_ERROR in the run
adapter and templates/adapters/tau2_bench/adapter.py, so the gate returns "INDECISIVE: only
2% of val tasks produced a valid score (< 60% required). The evaluation measured the
infrastructure, not the edit." Pinned by check_timeout_honesty.py. It fired for real later in
the run when a candidate's environment failed to build.
3. The learning signal named the tool, not the defect. Feedback read Failed action(s): update_reservation_flights while argument-value errors are the majority failure mode. The
argument-level localizer that existed was calling an undefined helper and silently falling
back through an except, and it read the trace from metadata while this adapter stores it on
Rollout.trace. Fixed, and feedback now reads "get_reservation_details: agent used
reservation_id='MSJ4OA'; cancel_reservation: agent used reservation_id='LU15PA'" — the agent
inspected one reservation and cancelled a different one. Re-deriving costs zero rollouts
(scoring is deterministic on persisted rollouts); resignal.py rewrote 50 of 90 baseline
feedback strings and is verified to touch score.feedback only.
4. The subset-screen ladder could not pay for itself. breakeven_kill_rate is 0.5 at
val_n 12; across four runs the screen killed 0 of 8 promoted candidates and produced one
documented false positive. At val_n <= 30 it is now skipped and full val is paid directly.
Ten gated candidates over five rounds, each with a per-round control. The baseline's per-task
pass RATE (k/n) put 6 tasks in a 0.0–0.3 defect band, 13 unstable, 11 solid; 34 of 37 failing
rollouts were database-state mismatches, splitting into "required write never called" (40
failed checks) and "called with wrong argument values" (30).
| candidate | surface / form | val | paired Δ | verdict |
|---|---|---|---|---|
cand_toolguard |
tools — in-code preconditions refusing illegal writes, naming the legal next call | 0.6778 | +0.0889 | ACCEPT |
cand_lean |
tools — return payloads cut 33%, eligibility cut 64% | 0.6888 | +0.0111 | reject (best absolute) |
cand_argsdoc |
tools — positive-recipe docstrings + derived_facts returns |
0.6333 | +0.0444 | reject |
cand_autofix |
tools — derive a computable argument instead of refusing | 0.6666 | −0.0111 | reject (= control) |
cand_composite |
tools — +3 composite tools, equivalence to primitives proven | 0.6556 | −0.0222 | reject |
cand_minimal |
tools — ALL return payloads removed, refusal text kept | 0.6444 | −0.0333 | reject (= control) |
cand_correct |
tools — refusals lead with a ready-to-copy corrected call | 0.6444 | −0.0333 | reject |
cand_merge |
tools — guards + docs merged (additivity test) | 0.6444 | −0.0333 | reject |
cand_enable |
tools — eligibility gains enabled_by / next_legal_actions |
0.5777 | −0.1000 | reject |
cand_lean2 |
tools — payloads and refusal text trimmed | 0.5777 | −0.1000 | reject |
cand_writeflow |
policy — 37-line structural Write Runbook | 0.5777 | −0.0111 | reject (= control) |
| split | seed | best (cand_toolguard) |
Δ |
|---|---|---|---|
| val (30 tasks × 3 trials, gating currency) | 0.5889 ± 0.0814 | 0.6778 ± 0.0805 | +8.9 pp / +15.1% rel |
| sealed TEST (20 tasks × 10 trials, scored once, both arms) | 0.460 ± 0.0881 | 0.545 ± 0.0960 | +8.5 pp / +18.5% rel |
pass^1 = 0.545, pass^2 = 0.454. Solid val tasks 11 → 16. The val gain (+8.9) and the held-out
gain (+8.5) agree to within a point, which is the signature of a real edit rather than an overfit
one — and it is the first positive sealed-test delta agent-mode has produced on this benchmark
(prior attempts: 0.417, 0.750 with no accepted change, 0.500, 0.500).
Three conclusions, each from a controlled comparison rather than an argument:
- Constraining behaviour works; informing it does not. The only accepted edit was in-code
preconditions.
cand_writeflow, a careful structural policy rewrite, scored identically to the byte-identical control (0.5777 both) — a third independent confirmation on this benchmark that the airline policy is not the binding constraint for a mid-tier runner. - Return payloads are dead weight; the corrective text is load-bearing. Removing every added
return payload was neutral (
cand_minimal= its control exactly), while adding to them cost up to −0.10 (cand_enable). But trimming the refusals'WHAT TO DO INSTEADclauses cost the same −0.10 (cand_lean2— its only difference from the neutral ablation). So the guard's message pays for itself and its payload does not. - Four separate mechanisms aimed at the six remaining 0.0 tasks all failed: composite
payloads, derived arguments, enabling-path hints, and stronger corrective wording. Trace
evidence says those six are decision failures at escape hatches — e.g. on one task the
agent read a basic-economy reservation, called
transfer_to_human_agents, and never issued the write, even though upgrading the cabin first would have made the cancellation legal. The tool layer cannot reach that, and the policy layer measurably does not move this model.
The run was asked to reach >0.90. It did not, and the reason is quantified rather than asserted: 0.6778 → 0.90 requires all six remaining 0.0 tasks to convert plus the unstable band to stabilise, and every mechanism tried against those six measured at or below its own control. This is consistent with the two independent prior findings on this benchmark — this repo's own ceiling of 0.712 (50 tasks × 10 trials) and a predecessor framework's conclusion that "there is no honest path to >0.90 for gpt-oss-120b on tau2 airline" (its best stable numbers being 0.727 at 3 trials and 0.518 at 10). Reaching 0.90 here would require a stronger runner, a stronger user simulator (a swap measured elsewhere as worth up to +16 pp and deliberately NOT made here, to keep comparability), or a different measurement basis — not a better optimizer.
τ²-Bench airline — agent-optimize on gpt-oss-120b, disjoint 26/12/12, num_trials: 5 — fourth null result; trial-averaging fixed the significance test and the no-regression veto became the blocker
Fourth attempt on the same benchmark, split and models. v3 had established that at
num_trials: 1 a byte-identical copy of the seed measured paired Δ̄ = −0.0833 and was
rejected, so no edit of any quality could be detected. v4 changed exactly one
substantive spec field — num_trials: 1 → 5 — leaving gate_mode: paired and
gate_k_se: 0.2 untouched, because the bar was never the problem. Nothing cleared the
gate. What v4 adds is the measurement that says why, and this time the cause is a
different one from v3's.
- Spec:
examples/tau2_airline/capevolve.agentopt.v4.gptoss.yaml— a copy ofcapevolve.agentopt.gptoss.yamlwithnum_trials: 5,max_metric_calls: 600and a rollout-denominatedstop_condition(the proxy does not meter dollars). - Agent + user simulator:
aws/gpt-oss-120bvia the IBM litellm proxy. Optimizer:aws/claude-opus-5(the conversational agent itself). - Split: disjoint 26 train / 12 val / 12 test (
agentopt_split.json),paired,k_se 0.2,num_trials: 5— sopass^kis defined here, unlike v3. - Cost: 550 rollouts (360 in the loop + 130 train + 60 sealing test), 5 gated
candidates, 158 min wallclock / 172 min runner time. Dollars remain UNMETERED
(
usd: 0.0is missing data;spend.pyreportsrunner_spend_metered: false), so the budget unit is rollout counts. - Artifacts, including
events.jsonl, the five per-candidate gate verdicts, and the two analysis scripts, inrun_agentopt_v4/.
| split | n | trials | seed | best (= seed) |
paired Δ̄ | pass^1 | pass^2 |
|---|---|---|---|---|---|---|---|
| train (never gated; diagnosis surface) | 26 | 5 | 0.5308 ± 0.0818 | 0.5308 ± 0.0818 | 0.0 | 0.5308 | 0.3962 |
| val (the gate) | 12 | 5 | 0.5667 ± 0.1180 | 0.5667 ± 0.1180 | 0.0 | 0.5667 | 0.4167 |
| test (sealed once) | 12 | 5 | 0.5000 ± 0.1254 | 0.5000 ± 0.1254 | 0.0 | 0.5000 | 0.3750 |
Every 0.0 is by construction — best_id == seed, so measure.py scored one
capability on both sides and emits that warning itself. This is a null result with a
diagnosed cause, not a 0.000 improvement. The requested goal of train ≥ 0.90 was neither
reached nor approached: train is 0.5308, and 0.90 on n=26 needs 24/26, i.e. fixing ~10 of
the 12 failing train tasks while breaking none. That was never achievable in this run and
is not claimed.
v3 ran one null-edit control. v4 ran three, each a byte-identical copy of the seed
(diff -r clean — and candidate_diffs.txt contains no entry for any of them, which is the
artifact-level proof), each evaluated on full val at 5 trials and put through the same
gate_check.py:
| measurement of the SAME capability | val | paired Δ̄ vs the frozen seed | no-regression veto |
|---|---|---|---|
seed (the gate's baseline) |
0.5667 | — | — |
c0_null5 |
0.4667 | −0.1000 | fired: 8, 12, 20, 40 |
c0_null5b |
0.5167 | −0.0500 | fired: 8, 12, 32, 40 |
c0_null5c |
0.5667 | −0.0000 | fired: 8, 40 |
Four measurements of one unchanged capability: mean 0.5292, SD 0.0479. v3's four equivalent measurements at 1 trial were mean 0.6042, SD 0.1423. So trial-averaging worked: the per-eval spread fell 3.0× (better than the √5 = 2.24 predicted, on 4 points each, so treat the ratio as approximate). Pooling all four evals gives the run's best estimate of the seed's val score, 0.5292 ± 0.1100 over 20 trials/task.
The significance half of the gate is now sound and its SE is predictable:
noise_power.py derives, from
the baseline's own measured per-task rates, that a null Δ̄ should have SD 0.0632 at 5
trials — and the five gates measured SE 0.0628 / 0.0702 / 0.0696 / 0.0672 / 0.0796. The same
formula gives 0.1414 at 1 trial, against v3's observed 0.1423. The model of the noise is
right.
But c0_null5c is the finding. It measured an exactly equal val mean to the seed —
Δ̄ = −0.0000, the significance test cannot fault it — and the gate still rejected it,
because two tasks' fractional rewards had dipped. That is the whole v4 result in one row.
Correction, and it changes the conclusion. The first version of this section said
gate_check.pyandharness"both" define a regression as any strictly lower per-task reward. They did not, and the difference is the whole bug:
harness._candidate_task_impactvetoes onlypar[t] >= 1.0 - eps and cand[t] < par[t] - eps— the parent must have measured-and-passed, which is whatSKILL.md:221specifies.gate_check.regressionsvetoed on any strict drop from any parent level, while its own docstring claimed to "mirror the harness's no-regression rule exactly".So the trap was specific to agent-optimize's gate, which was silently stricter than every other algorithm's, and it did not apply to
hill-climb/gepa/skillopt. Simulated against this run's measured per-task val rates:
num_trialsany-drop (old gate_check)parent-passed ( harness, and nowgate_check)1 0.889 0.889 5 0.983 0.428 10 0.990 0.129 The old rule got worse as trials rose, so no trial count could fix it. The harness rule converges, which is the behaviour a variance-aware gate must have.
gate_checkhas been fixed to match, pinned bycore/tests/test_regression_gate.py, which also asserts the harness predicate's source text so the two cannot drift apart again.
At num_trials: 1 rewards are 0/1, so only real flips register. At 5 trials every reward is
a fifth, so noise alone produces small drops — and the parent's reference vector is frozen
from a single 5-trial draw at baseline time. Whichever tasks the baseline happened to
over-measure become veto triggers for every later candidate.
Val task 8 is the worked example: the baseline drew 4/5 = 0.80; over the 15 later trials of the same capability (the three null controls) its rate is 5/15 = 0.33, and its best estimate over all four unchanged evals is 9/20 = 0.45. Task 8 therefore regressed in all five gates of this run — including all three byte-identical seed copies. (Honest limit: a Fisher exact test of the baseline window against the later windows gives p = 0.127 for task 8, so this is regression to the mean off a lucky draw, not a demonstrated drift over time. No val task showed a significant baseline-vs-later shift.)
Simulating the veto against the measured per-task rates shows the trap closes as trials rise — the opposite direction from the significance test:
num_trials |
SD of a null Δ̄ | P(no-regression veto fires on a NULL edit) |
|---|---|---|
| 1 | 0.1414 | 0.80 |
| 5 | 0.0632 | 0.95 |
| 10 | 0.0447 | 0.97 |
| 20 | 0.0316 | 0.98 |
Under the old any-drop rule no trial count could make this gate accept anything: raising
trials sharpened the significance test and simultaneously made the veto near-certain. With
gate_check corrected to the harness rule, that specific trap is gone.
Fixing the veto does not make this run's candidates acceptable — it reveals that the significance bar was never doing any work. At 5 trials:
| quantity | value |
|---|---|
| null-edit Δ̄ spread (SD of 4 measurements of the same capability) | 0.0479 |
the gate's bar at gate_k_se: 0.2 |
0.0134 |
cA_partial / cB_becabin measured Δ̄ |
+0.0167 |
The bar sits ~3.6× below the noise floor, so both real candidates "cleared significance" on a Δ̄ that is one sixth of the span the null controls span by chance. An over-strict veto was the only thing preventing false accepts, and it was rejecting byte-identical seeds to do it. Neither +0.0167 is evidence of improvement, and this run must not be read as one.
What the numbers actually require, for anyone running a stochastic benchmark:
num_trials: 10— enough for a 2-task val gain (0.1667) to sit at >3 SD. A 1-task gain needsnum_trials ≥ 26(312 rollouts per candidate); a 3-task gain needs only 3.gate_k_sewell above 0.2.gate_check's own default is 1.0; the tau2 agent-optimize spec set 0.2, which on this benchmark is far inside the noise. It was deliberately not changed here — loosening or tightening the bar to reach a desired verdict is precisely the move this project exists to prevent, and the right value follows from a measured noise floor, not from a target. The measurement now exists: k·SE should exceed ~0.05 at 5 trials, i.e. k ≳ 1 on the paired SE this run observed.- a no-regression tolerance is still worth having for the harness rule at high trial counts, though it is no longer the binding constraint (0.129 at 10 trials).
Both were narrow policy.md edits (diffs in candidate_diffs.txt), aimed at the two val
tasks the pooled 20-trial estimate showed to be genuinely near-zero and therefore worth
winning: task 24 (0.05) and task 32 (0.20).
| candidate | val | paired Δ̄ | bar 0.2·SE |
veto | verdict |
|---|---|---|---|---|---|
cA_partial — a barred part of a multi-part request is not grounds for transfer: refuse that part, complete the rest |
0.5833 | +0.0167 | 0.0134 | fired: 8 | reject |
cB_becabin — a basic-economy flight change may go via a confirmed cabin change first, barred if any segment has flown |
0.5833 | +0.0167 | 0.0159 | fired: 8, 32, 40 | reject |
Both cleared the significance bar and neither is evidence of an improvement: +0.0167 is
one sixth of the spread the three null controls span (−0.1000 … −0.0000). Clearing a
0.2·SE bar of 0.013 is not a meaningful test at this noise level. Per-task, the edits did
essentially nothing to what they aimed at:
cA_partialexecuted but missed.check_transfers.pyshows transfers on task 24 fell 2/5 → 1/5 andbook_reservationwas called on both sides, so the edit changed behaviour — but task 24 stayed 0/5. The v3 single-trial transcript this edit was designed from ("agent transferred instead of booking") was not representative: at 5 trials task 24's modal failure is a wrongbook_reservationargument set, not abandonment. Diagnosing a stochastic benchmark from one rollout produced a correct-looking edit aimed at the wrong defect.cB_becabinmoved its target the wrong way, 0.20 → 0.00 on task 32, within the null range (the controls gave 0.40 / 0.00 / 0.20 there). It did not break the paired case it was scoped to protect: task 36, where refusing is correct because a segment has already flown, held at 1.00 across all five trials.
Both re-checked from rollouts rather than taken on trust: val task 40 and train task
7 are user-simulator defects, not capability defects — the simulator answers the agent's
confirmation request and appends ###STOP### in the same turn, ending the episode before
the agent can act (train task 7 also leaks a <reasoning> block into the user turn). Task
40 is not, however, hopeless: at 5 trials it scores 0.20–0.40, so the defect is
intermittent. Val tasks 24 (0.05) and 44 (0.00) are the genuinely stuck ones, and 44 is a
five-reservation itinerary-arithmetic task that no one-paragraph rule will fix.
No subset screen was run. At val_n 12 the tier-1 floor of 6 tasks gives a breakeven kill
rate of 0.5, and with two candidates a screen could not pay for itself; full val was paid
directly for each.
τ²-Bench airline — agent-optimize on gpt-oss-120b, disjoint 26/12/12 — third null result; the gate rejects a byte-identical copy of the seed
Third attempt on the same spec, split and models as the section below, built to convert that run's diagnosis into a measured gain. Nothing cleared the gate again, and this time the run establishes why with a control the two earlier runs never ran: a null-edit candidate — a byte-identical copy of the seed capability — was evaluated on full val and put through the honest gate, and the gate REJECTED it.
- Spec:
examples/tau2_airline/capevolve.agentopt.gptoss.yaml,orchestration_mode: agent,algorithm_skill: agent-optimize, capabilities[system-prompt, tools]. Same file as the v2 run exceptstall: 3 → 5and the prosestop_condition's "3 rejects in a row" → "5" (the prose ceiling, not the spec field, was what actually stopped v2 at 3 of 8 iterations).gate_mode,gate_k_seandnum_trialswere NOT touched. - Agent + user simulator:
aws/gpt-oss-120bvia the IBM litellm proxy. Optimizer:aws/claude-opus-5(the conversational agent itself). - Split: disjoint 26 train / 12 val / 12 test (
agentopt_split.json),num_trials: 1,paired,k_se 0.2. - Cost: 134 rollouts (122 in the loop + 12 sealing test;
metric_calls: 134), 5 iterations, ~2 h of runner wall-clock. Runner dollars remain UNMETERED on this path (usd: 0.0is missing data, not a free run —spend.pyreportsrunner_spend_metered: false), so the budget unit here is rollout counts. - Artifacts, including
events.jsonl, inrun_agentopt_v3/.
| split | n | seed | best (= seed) |
paired Δ̄ |
|---|---|---|---|---|
| train (never gated; diagnosis surface) | 26 | 0.5385 ± 0.0997 | 0.5385 ± 0.0997 | 0.0 |
| val (the gate) | 12 | 0.6667 ± 0.1421 | 0.6667 ± 0.1421 | 0.0 |
| test (sealed once) | 12 | 0.5000 ± 0.1508 | 0.5000 ± 0.1508 | 0.0 |
Every 0.0 is by construction — best_id == seed, so finalize scored one
capability on both sides. measure.py emits that warning itself.
The requested target — train mean ≥ 0.90 — was not reached and was not approached. Train stayed at its baseline 0.5385 because no candidate was ever accepted, and train may never gate. Nothing here is progress toward 0.90. For the record, 0.90 on n=26 requires 24/26, i.e. fixing 10 of 12 failing train tasks while breaking none; at least one of those 12 (task 7) is unfixable by any capability edit, because the user simulator ends the episode on its second turn.
Five candidate measurements (four distinct edits, one of them measured twice) all came back
at exactly 0.4167 on full val while the seed measured 0.6667. Unrelated edits landing on
the same mean was implausible enough to test the
pipeline itself, so an unmodified copy of the seed (c0_nulledit, verified byte-identical
with diff -r) was run through the same evaluate → gate_check.py path:
NULL-EDIT (byte-identical to seed) gate verdict: reject
paired Δ̄=-0.0833 <= 0.2·SE=0.0167 (SE=0.0833, n=12)
regressions ['8'] paired_n 12
A known-zero-effect candidate produced a negative Δ̄ and tripped the no-regression
veto. The bar is not the problem — 0.0167 is negligible. The problem is that at
num_trials: 1 a single per-task reward vector is noisy enough that zero change registers
as minus one task. Any real +1-task gain has to clear noise of the same magnitude, in
the same direction, on the same measurement. That is a measurement-power limit, and it
explains all three null results on this benchmark better than any of the behavioural
diagnoses offered for v1 and v2.
c5_guards_only changes only tool bodies — no policy text, no tool docstring, no schema.
Auditing its 12 val rollouts for the guard error strings shows not one guard fired on any
val task:
| val task | guarded tool called | guard fired |
|---|---|---|
| 12, 16 | update_reservation_baggages, update_reservation_flights |
no |
| all others | none | no |
So on this split c5_guards_only is functionally identical to the seed — same prompt,
same tool schemas, and no added code path ever executed. Its true Δ is 0 by construction.
The gate measured Δ̄ = −0.2500 with regressions on tasks 8, 16 and 20.
That gives four measurements of a capability that is functionally identical to the seed on val:
| measurement | val |
|---|---|
| v2 seed | 0.7500 |
| v3 seed (the gate's baseline) | 0.6667 |
c0_nulledit (byte-identical copy) |
0.5833 |
c5_guards_only (no guard ever fired) |
0.4167 |
Mean 0.6042, SD 0.139, range 4 of 12 tasks. The consequence is blunt: 0.4167 is inside the observed range of the seed itself, so none of the five candidate results below can be called evidence of harm, and the −0.25 deltas the gate computed are not measurements of the edits. The gate behaved correctly given its input; the input cannot support the question.
| replication | result |
|---|---|
| val, 4 functionally-identical-to-seed measurements | 0.7500 / 0.6667 / 0.5833 / 0.4167 — mean 0.6042, SD 0.139 |
| val, tasks that flip across the 3 byte-identical draws | 2 of 12 (tasks 8 and 32) |
| train, 2 draws of the identical seed capability (v2, v3) | mean identical at 0.5385 both times, but 6 of 26 per-task rewards flipped (23%) — tasks 10, 11, 15, 30, 38, 49 |
val, one candidate (c2_toolguard) measured twice, once concurrently and once serially |
0.4167 both times, 2 of 12 per-task flips — so process concurrency is not a confound |
Three val tasks (24, 40, 44) fail in every one of 11 measurements and no candidate ever fixed any of them. So the gate's usable signal is 3 immovable failures against 2–6 coin-flip tasks.
Every reject is a genuine full-val paired gate decision (paired_n: 12,
coverage: 1.0), recorded with the new machine-readable reject_basis: gate. This closes
the v2 run's weakest point, where gate_check.py never fired on real data.
| candidate | lever | val | Δ̄ | bar (k·SE) |
regressed | fixed |
|---|---|---|---|---|---|---|
c1_verify |
prompt: verification obligations only | 0.4167 | −0.2500 | 0.0261 | 8 12 20 |
— |
c2_toolguard |
tools: DB-checked policy guards + a search docstring | 0.4167 | −0.2500 | 0.0261 | 8 12 20 |
— |
c3_finish |
prompt: act-on-confirmation + finish-every-part | 0.4167 | −0.2500 | 0.0261 | 8 12 20 |
— |
c5_guards_only |
tools: the guards alone, policy byte-identical | 0.4167 | −0.2500 | 0.0261 | 8 16 20 |
— |
c6_onerule |
prompt: one sentence, verification only | 0.5000 | −0.1667 | 0.0298 | 12 20 |
— |
c0_nulledit (control, not a proposal) |
none — byte-identical to seed | 0.5833 | −0.0833 | 0.0167 | 8 |
— |
Not one candidate fixed a single val task — that part is a real, repeated observation. The "regressed" column is not: every candidate, and both controls, "broke" task 8, a knife-edge pass the seed itself wins in only 2 of 8 measurements. Read the Δ̄ column as what the harness reported, not as an effect of the edit.
c6_onerule was a deliberate dose-response probe: one sentence of verification-only
policy text instead of a block. It scored 0.5000 — higher than every multi-rule candidate
and the only candidate that kept task 8 — but still below the seed draw it was gated
against, and still a reject. Suggestive of a text-volume effect on this weak reader, and
nothing more: one measurement, inside the noise band established above.
The two hypotheses this run was built to test both failed, and failed distinguishably:
- Verification-only prompt rules (
c1_verify, andc3_finishwith an explicit clause that a confirmation never authorises a forbidden action) did not raise the val mean, and the measurement cannot say whether they lowered it. What they do falsify is v2's explanation for the task-12 regression: v2 blamed "compliance obligations", yet task 12 regressed here underc1_verify, a rule set carrying none. Across all 11 measurements in this run and v2, task 12 passes only 4 of 11 times; it is a fragile task, not a casualty of a compliance rule. - The tools lever was exercised for the first time in isolation.
c5_guards_onlymoves four policy invariants that tau2's API deliberately does not enforce into the tool bodies — cancellation eligibility (24h / airline-cancelled / business / insured, checked against the DB rather than the user's claim), no reduction of checked bags, basic-economy flights immutable, route and trip type immutable — each raising an error that names the rule and the actual values checked. The guards were unit-checked to fire on real DB cases the seed silently accepted (e.g. it cancelled reservation3RK2T9, created 2024-05-02, basic economy, uninsured, on the user's unverified claim that it was "booked ten hours ago") and to leave a legitimately cancellable reservation alone. They are correct, they would have blocked exactly the wrong writes in three failing tasks (train 48, train 49, val 44) — and none of them fired anywhere on val, which is what makesc5_guards_onlythe strongest control in the run rather than a result about tools. The tools lever is therefore still untested for effect: it was exercised, audited and shown correct, but this split never gave it an opportunity to act.
The guard error strings appear in zero val rollouts across all three tools candidates
(c2_toolguard, c2_toolguard_solo, c5_guards_only). The cases the guards exist for do
occur in the data — the seed illegally cancels reservation 3RK2T9 in train tasks 48 and
49 and cancels an already-flown reservation in val task 44 — but on val the guarded paths
are either never reached or reached only on legitimate calls. So half the capability
surface under optimization (capabilities: [system-prompt, tools]) is invisible to the
split that decides acceptance. Diagnosing on train and gating on val is the right
discipline, and here it means a correct tools edit can never be accepted, however good it
is. A tools-lever run on this benchmark needs a val split chosen to contain
policy-invariant cases, or capabilities narrowed to what val can actually score.
The v2 section reports task 40 as "the agent replies the name has been updated without
ever calling the tool — a hallucinated success". Re-reading the persisted rollout,
that text is in the user message, not the assistant's: the user simulator (also
gpt-oss-120b) emitted Yes.<reasoning>Agent will likely confirm and process.</reasoning> Your request has been processed. The passenger name on reservation 3RK2T9 has been updated … ###STOP###. It answered, then role-played the agent's reply, then terminated the
episode. Task 40 is a user-simulator defect, not an agent defect, and no edit to the
policy or tools can fix it. Same for train task 7, which the simulator ends on its second
turn. That is 1 of 12 val tasks and 1 of 26 train tasks that are unwinnable by
construction — and one of the three val tasks that fail in all 11 measurements.
- The null-edit control is now part of the protocol. Evaluating a byte-identical copy of the parent through the same evaluate → gate path costs one val eval and bounds the harness's own false-reject rate. Both earlier null results on this benchmark are reinterpretable in its light; neither ran it.
commit.py --reject-basisrecords the driver's disposition (gate|screen_kill|ceiling|budget|infra) machine-readably next to the screen's own verdict.screen.py'sdecisionis authoritative only as the screen's statistical verdict — by invariant 1 it can only saykillorpromote, and "promote" never meant "was then evaluated on full val". In v2 that ambiguity made two artifacts read as a contradiction (screen: promotebeside a commit note saying "not promoted to full val"); both were true. Every reject in this run carriesreject_basis: gate, so the log asserts that a full-val paired gate actually ran.- The gate was exercised on real data six times (
paired_n: 12,coverage: 1.0each), closing v2's weakest provenance point, wheregate_check.pynever fired on a real full-val eval and two rejects rested on an arithmetic ceiling. - Skipping the screen ladder was the right call and is now measured, not argued. At
val_n 12the tier-1 floor of 6 gives a breakeven kill rate of 0.5, and with 2 of 12 val tasks flipping on an identical capability a 6-task screen carries no signal. All six candidates went straight to full val;screen_ledgeris empty by design, not by neglect.
The optimizer did not find an improvement, and this run shows the harness could not have
recognised one — nor could it correctly reject a non-change. Two controls, one
byte-identical and one functionally identical on the scored split, were both rejected by
the honest gate. Concretely: the gate rejects a copy of the seed, half of
val is coin-flips, and 3 of 12 val tasks are immovable (one of them because the user
simulator breaks). The correct next step is not a looser gate — k_se was left at
0.2 and the bar was never the obstacle. It is num_trials > 1 so each per-task reward is
a mean rather than a single Bernoulli draw, and a wider val, in that order. Until then any
accept on this configuration would be a coin flip dressed as a result.
τ²-Bench airline — agent-optimize on gpt-oss-120b, disjoint 26/12/12 — second null result, sharper diagnosis
Kept as history, superseded by the v3 section above. Its central mechanism claim about task 40 is wrong — the "hallucinated success" text is in the user simulator's message, not the agent's; see "v2's central mechanism claim was wrong" above. Its conclusion (no accept) stands and was reproduced.
Re-run of the section below with the agent+user-simulator switched from claude-haiku-4-5
to aws/gpt-oss-120b to restore headroom, and with the algorithm's own defects fixed
first. Nothing cleared the gate again, but this time the mechanism was read directly
out of the transcripts rather than inferred from means, and the run cost 68 rollouts
instead of 110.
- Spec:
.capevolve/project/capevolve.agentopt.yaml(derived fromexamples/tau2_airline/capevolve.agentopt.yaml),orchestration_mode: agent,algorithm_skill: agent-optimize, capabilities[system-prompt, tools]. - Agent + user simulator:
aws/gpt-oss-120bvia the IBM litellm proxy. Optimizer:aws/claude-opus-5(the conversational agent itself — no per-iteration subprocess). - Split: disjoint 26 train / 12 val / 12 test (
agentopt_split.json),num_trials: 1,paired,k_se 0.2. - Architecture change: diagnosis on train (26), acceptance gated on val (12),
test sealed once. Previously
diagnosewas hardcoded to val, so train was unreachable and the run was fitting the split it was judged on. - Cost: 68 rollouts, 48 min of runner wall-clock, ~75 min end to end. Runner
dollars are UNMETERED on this serving path — the proxy returns no cost, litellm logs
model isn't mapped yet, and the ledger recordsusd: 0.0. That 0.0 is missing data, not a free run; rollout counts are the honest unit here. Pre-runcap-evolve estimatesaid $50.39 expected (runner $10.52 + optimizer $39.86) from prior calibration — not comparable, and reported only for the record. - Stopped on its own
stallrule (3 consecutive rejects), at iteration 3 of 8.
| split | n | seed | best (= seed) |
paired Δ̄ |
|---|---|---|---|---|
| train (never gated; diagnosis surface) | 26 | 0.5385 ± 0.0997 | 0.5385 ± 0.0997 | 0.0 |
| val (the gate) | 12 | 0.7500 ± 0.1306 | 0.7500 ± 0.1306 | 0.0 |
| test (sealed once) | 12 | 0.7500 ± 0.1306 | 0.7500 ± 0.1306 | 0.0 |
Every 0.0 is by construction — best_id == seed, so finalize scored one
capability on both sides. measure.py emits that warning itself. Artifacts, including
events.jsonl, in run_agentopt_v2/.
The requested target — train mean ≥ 0.90 — was not reached, and was never approached. Train stayed at its baseline 0.5385 because no edit was ever accepted. Nothing here should be read as progress toward it.
Switching to gpt-oss-120b did what it was supposed to: baseline val fell from 0.8333
(haiku, 10/12) to 0.7500 (9/12), and train sits at 0.5385 (14/26) — 12 failing train
tasks to work with instead of 2 val tasks. And the gate bar was never the obstacle:
For exactly one val task flipping +1 out of
npaired deltas, Δ̄ = 1/n and SE = 1/n exactly, so Δ̄/SE = 1.0000 for every n. Atk_se 0.2, n=12, a single clean flip gives Δ̄ = 0.0833 against a bar of 0.0167 — it clears comfortably. Widening val would not have helped; for two flips the ratio actually falls slightly with n (1.483 at n=12 vs 1.439 at n=30). Both null results are failures to move the mean, not failures of resolution.
Val's 3 failures (24, 40, 44) and 11 of train's 12 share one cluster signature
(database state does not match), so train-driven edits could bank — the overlap was
checked for free before any spending. Within it, two sub-modes: the required write was
never called (7/12 train, 3/3 val), or it was called with wrong arguments (6/12
train, update_reservation_flights four times). Task 40's transcript is the purest case:
the agent asked for confirmation, the user said "Yes", and the agent replied "the
passenger name has been updated" without ever calling the tool.
Three candidates, each rejected, each for a reason visible in the rollouts:
| candidate | edit | screen decision | fixed | regressed |
|---|---|---|---|---|
cand_r1_disc |
+70-line "Execution discipline" section (policy) | promote (inconclusive) | 0 of 3 | 12 |
cand_r2_short |
+18-line policy section + once-and-only-once / last-resort contracts in 6 tool docstrings | kill (Δ̄ −0.333, SE 0.211) | 0 of 3 | 12, 16 |
cand_r3_lookup |
+13 lines, policy only, one scoped rule (look it up; don't transfer) + an explicit "this does not widen what is permitted" clause | promote (inconclusive) | 0 of 3 | 12 |
All three regressed task 12, and none fixed anything. Task 12's transcripts say why,
and it is not noise. Under the seed, the user asks to upgrade one passenger to business
and add bags; the agent correctly refuses the partial upgrade (cabin must be uniform),
the user falls back to bags only, and the agent makes one correct write → 1.0. Under
cand_r2_short, the same "act on confirmation / answer every part" pressure made the
agent comply with the request it should have refused — it upgraded both segments to
business and added bags → wrong DB state → 0.0. cand_r3_lookup reproduced it even with
an explicit precedence clause telling it not to.
So on this model the edits traded policy compliance for eagerness. Pushing
gpt-oss-120b to act more decisively makes it act wrongly on the tasks whose correct
answer is a refusal — a genuine capability trade-off, correctly refused three times. A
secondary effect is visible too: cand_r2_short drove task 44 from 20 to 54 turns of
runaway search_direct_flight calls, and its screen took 10 minutes against the seed's 2.
- A screen can prove a reject arithmetically. When the tier-1 subset already covers
every val task the parent fails, the unscreened remainder is all tasks the parent
passes, so it can only stay level or regress — and the candidate's best conceivable
full-val mean is computable. For
cand_r1_disc: at most 8/12 = 0.667 against the parent's 0.750, best-case Δ̄ = −0.0833. The gate needs Δ̄ > k·SE ≥ 0, so no full-val eval could have accepted it. This is nowsubsample.full_val_ceiling(), reported in every screen artifact, and it escalates a promote to a provable kill when the ceiling is strictly below the parent. It cannot ever conclude "accept", so honesty invariant 1 is intact. Consequence for this run: no candidate was ever taken to a full-val eval, sogate_check.pynever fired on real data. All three rejects rest on screen evidence, one statistical and two arithmetic — weaker provenance than a full-val gate rejection, and stated as such. - A narrow val makes the ladder uneconomic, and the artifact now says so.
savings.breakeven_kill_rate=fired / full_val_rollouts= 0.5 at val 12. With 1 kill in 3 screens the recorded ledger isnet_rollouts −6— screening was still a net cost under its own accounting, which assumes every promote gets paid for. - The subset floor moved from 3 to 6. The 3-task tier-1 screen in the run below
reported
fixed: ["44"]for a candidate full val showed never fixed 44. At 6 the holdout caught the task-12 regression on all three candidates. - val predicted test here. val 0.7500 and sealed test 0.7500 for the same seed, on the same index-stride split that previously gave val 0.8333 / test 0.4167. Train (0.5385) is materially harder than either — the stride split balances draw order, not difficulty.
diagnosewas hardcoded torollouts/val— train was unreachable by every one of the five algorithms, so "diagnose on train, gate on val" was undocumentable and unrunnable. Now--split train|val, default unchanged.spend.py/measure.pyreadproject/capevolve.yamlby filename, butcap-evolve run --specsupports any filename. Every agent-mode run of a variant spec silently reportedpredicates: []— the whole re-read-your-constraints discipline no-opping without a word. Nowspecfile.spec_for_run()reads the pathclialready logs into the run dir'srun_configevent.- Unmetered runner spend read as $0/rollout, so
usd_neededwas 0.0, amax_usdceiling could never block anything, and any fan-out came backaffordable: true. Now reported asrunner_spend_metered: falsewith the ratenull. - A train/test-qualified score goal was parsed into
target_val_score, which is only ever checked against the full-val mean — so"reach train mean >= 0.9"would have enforced a val bar while reporting it as the train one. Now reported inambiguous. commit.pyaccepted a duplicate candidate id. The tag collision documented below was possible because nothing checked. It now refuses a tag that already carries an accept/reject event, readingevents.jsonlso the guard holds across processes.measure.py --train onre-ran the seed's whole train split even with complete rollouts already on disk. Candidate dirs are immutable snapshots, so those rollouts are measurements of exactly that capability; reuse saved 26 rollouts here.
The gate reads *__<tag>__t*.json, and screens write <tag>__screenN. Probed
adversarially — a screen rollout claiming 1.0 and a full-val rollout claiming 0.0 for the
same candidate — the full-val read returns 0.0; and on the live run dir, a full-val
read of cand_r1_disc (which has only screen rollouts) returns n_scored: 0 rather than
6. Both probes are permanent checks in
skills/algorithms/agent-optimize/scripts/check.py.
Two independent runs, two models, seven candidates, zero accepts. The consistent finding
is not that the optimizer is weak but that the airline policy is not the binding
constraint for these models on these tasks — the seed policy already states the rules
the failures violate, and adding restatements of them makes a mid-tier model less
compliant, not more. The next thing worth testing is not more prose but the lever that
carried the no-holdout run: tool-level change (composite or guard-railed tools that
make the correct write the easy one), and a runner strong enough that the compliance
trade-off does not bite. Reporting a 0.90 train number here would have required either
gating on train or loosening k_se; neither was done.
Kept as history, superseded by the
gpt-oss-120bsection above. This is the earlierclaude-haiku-4-5run. It is retained in full because it is real evidence about the method — and because the second run reproduced its central finding (no accept) with a different model and a different diagnosis, which makes the pair more informative than either alone. The defects it exposed (tag collision, screen width, the--mode pairedCLI, thework/dir) are fixed; see the section above for what changed.
The run that exercised subset screening and prose-constraint parsing for the first time. Nothing cleared the gate. Recorded because a null result with a diagnosed cause is evidence, and because two integrity findings came out of it.
- Spec:
examples/tau2_airline/capevolve.agentopt.yaml,orchestration_mode: agent,algorithm_skill: agent-optimize, capabilities[system-prompt, tools]. - Agent + user simulator:
claude-haiku-4-5via the IBM Anthropic-compatible gateway. - Split: disjoint 26 train / 12 val / 12 test,
num_trials: 1,paired,k_se 0.2. - Cost: $12.98 ($10.58 runner including screens + $2.40 optimizer), 5 iterations,
110 metric calls. Stopped on its own
stallrule (3 consecutive rejects), not on a cap.
| split | n | seed | best (= seed) |
paired Δ̄ |
|---|---|---|---|---|
| train | 26 | 0.6154 ± 0.0973 | 0.6154 ± 0.0973 | 0.0 |
| val | 12 | 0.8333 ± 0.1124 | 0.8333 ± 0.1124 | 0.0 |
| test (sealed once) | 12 | 0.4167 ± 0.1486 | 0.4167 ± 0.1486 | 0.0 |
Every 0.0 is by construction, not by measurement — best_id == seed, so finalize
scored the same capability on both sides. measure.py emits that warning itself.
| candidate | val | Δ̄ | bar (k·SE) |
SE | n | regressed | fixed |
|---|---|---|---|---|---|---|---|
cand_tools |
0.8333 | +0.0000 | 0.0246 | 0.1231 | 12 | 24 |
8 |
cand_r1 |
0.5833 | −0.2500 | 0.0359 | 0.1794 | 12 | 16 20 24 36 |
8 |
cand_r2 |
0.7500 | −0.0833 | 0.0167 | 0.0833 | 12 | 32 |
— |
cand_r3 |
0.4167 | −0.4167 | 0.0297 | 0.1486 | 12 | 12 16 20 24 36 |
— |
Re-derived from the persisted rollouts in
run_agentopt/val_per_task.json;
the seed fails exactly two tasks, 8 and 44.
cand_tools is churn: mean identical to the seed, different tasks passing (fixed 8, broke
24). That is the failure mode the no-regression veto exists for, now observed a third time
across runs. No candidate ever fixed task 44 — only cand_tools and cand_r1 moved
anything, both unlocking task 8 and both breaking task 24 to do it. The edits were audited as
general, not task-keyed. This is a real capability trade-off, correctly refused.
Every candidate edited both the prompt and the tool code — policy/policy.md +2 to +11
lines and tools/tools.py +71 to +87 across 3 hunks (new _to_minutes() /
_scheduled_duration_hours() helpers, get_flight_status returning duration_hours with a
+1 next-day-arrival rule). So the qualitative claim — it edits prompts and tool code
jointly — is demonstrated. The statistical claim is not.
- The effect size is below the noise floor. Baseline is 10/12, so only 2 tasks are
winnable and one flipped task is Δ̄ ≈ 0.083 against a paired SE of 0.08–0.18. This split
cannot resolve a one-task gain. A measurement-power limit, not an optimizer failure.
k_se 0.2was already permissive and was not loosened to manufacture an accept. num_trials: 1collapses the paired gate on identical trials (SE = 0→ STRICT fallback, emitted as agate_warning).- Subset screening was pure overhead here — a measured loss. 4 screens, 0 kills, 4
promotes, 12 rollouts fired, 0 avoided,
net_rollouts −12,screen_usd $1.367. All four returnedinconclusive: true. Worse,cand_r3's screen produced a false positive: its subset was[16, 44, 8]withregressed: ["16"]andmean_delta 0.0, which requires a+1on 44 or 8 to average out — yet full val showscand_r3fixed neither. So a 3-task screen read an improvement that does not exist. That vindicates the rule that a screen may never accept, and undercuts 3 as a sufficient triage width. The ladder behaved as designed; this run does not demonstrate the cost win it exists for. Correction (2026-08-16): an earlier revision of this section claimedscreens/*.jsonleavesfixed: null. That was wrong. The artifacts do record it, underpaired.fixed—cand_r3__screen1.jsonsayspaired.fixed: ["44"]andpaired.regressed: ["16"]. The top-levelregressedkey (fromscreen_decision) has nofixedsibling, which is what the earlier audit read. The false positive is therefore directly auditable from the artifact, not inferred. - val does not predict test on this split. val 0.8333 vs sealed test 0.4167 for the same seed capability, despite index-stride stratification. Second observation across runs.
Two concurrent optimizer drivers both tagged a candidate cand_r2. The event log holds two
reject cand_r2 events with different notes describing different edits, but only one
evaluate cand_r2, one screen cand_r2, and one set of 12 rollouts — so one edit was rejected
on the other's evidence, and the second snapshot overwrote the first in candidates/cand_r2.
Direction of harm matters: it produced a spurious reject, not a spurious accept, and
best_id is seed, so no reported number rests on it. This is agent-optimize's own
unique-tag invariant (rollouts are <task>__<tag>__t<k>.json, so a shared tag silently
collapses two candidates into one) being violated by real concurrent agents. By contrast the
<tag>__screenN isolation held perfectly under the same pressure — 12 files per candidate
tag, 3 per screen tag, no bleed.
The run only works with TAU2_AGENT_MODEL / TAU2_USER_MODEL set to an anthropic-looking
model string; otherwise rits.llm_args_for() falls through and demands RITS_API_KEY. Neither
appears in the spec. This broke the first finalize attempt (the seal survived — the failure
preceded mark_test_used()). Worth recording in the spec or the adapter docstring.
Same benchmark, capability, split, and algorithm as the held-out run above, with a self-hosted open model (Qwen 2.5 14B-Instruct via vLLM on OpenShift) replacing the Claude runner.
- Capability: airline policy + tools (
[system-prompt, tools]). - Optimizer:
claude-code@claude-sonnet-4-6(Vertex AI). - Runner + user simulator:
Qwen/Qwen2.5-14B-Instructvia vLLM on OpenShift. - Tasks / trials: all 50 airline tasks · 5 trials each.
- Split: train=val=30, test=20 — held-out.
- Algorithm / gate:
hill-climb --focus all, 10 iterations, paired significance gatek_se 0.3.
| split | baseline | optimized | Δ |
|---|---|---|---|
| val (30 tasks) | 0.200 (20.0%) | 0.387 (38.7%) | +0.187 / +93.5% relative |
| sealed test (20 tasks, scored once) | 0.170 (17.0%) | 0.240 (24.0%) | +0.070 / +41.2% relative |
3 of 10 iterations accepted. The optimizer added code-level enforcement of cancellation rules, input validation guards, and pre-flight status checks — hardening tool implementations rather than rewriting policy prose.
Same model and split as above, with all three capability types optimized jointly.
- Capability:
[skill-package, system-prompt, tools]. - Optimizer:
claude-code@claude-sonnet-4-6(Vertex AI). - Runner + user simulator:
Qwen/Qwen2.5-14B-Instructvia vLLM on OpenShift. - Tasks / trials: all 50 airline tasks · 5 trials each.
- Split: train=val=30, test=20 — held-out.
- Algorithm / gate:
hill-climb --focus all, 10 iterations, paired significance gatek_se 0.3.
| split | baseline | optimized | Δ |
|---|---|---|---|
| val (30 tasks) | 0.273 (27.3%) | 0.520 (52.0%) | +0.247 / +90.5% relative |
| sealed test (20 tasks, scored once) | 0.120 (12.0%) | 0.270 (27.0%) | +0.150 / +125.0% relative |
3 of 10 iterations accepted, 7 rejected. The optimizer edited tool code, policy rules, and skill prose jointly — combining code-level guards with policy clarifications and structured methodology in SKILL.md.
Best held-out test gain across all Qwen 14B runs (+125%). On a self-hosted 14B model,
[skill-package, system-prompt, tools] with hill-climb outperformed the tools-only run
(+41.2%) when paired with hill-climb's conservative gating.
The sections below are the SkillsBench runs whose artifacts are committed on main. They
are not the whole picture: the large per-task sweeps, their evolved skill packages, and
the analysis write-ups live on the
skillsbench-history
orphan branch, which is the source of truth for SkillsBench optimization output.
Why a separate branch: those artifacts are SkillsBench output, where cap-evolve is merely
the tool — not contributions to the library, which is what this tree holds. They don't
belong on benchmark-history either; that branch is machine-written (regenerated by
ci/benchmarks/lib/record.py aggregate, pruned by benchmark-history-prune.yml), so
hand-authored content placed there would eventually be deleted. Dashboard metrics for
SkillsBench still go through the normal benchmark-history pipeline — see
ci/benchmarks/PUBLISHING.md.
What's on it:
| path | what |
|---|---|
results/ |
per-sweep result tables; results/results.json is the canonical per-task ledger |
recipes/ |
the capevolve.*.yaml + split_ids.*.json needed to rerun a sweep |
artifacts/ |
the seed/best skill packages those recipes produced |
insights/ |
analysis write-ups, including the pass^k derivation |
evidence/ |
proof bundles for specific claims (e.g. the optimizer writing a 465-line solver from scratch) |
If you are quoting a SkillsBench headline number, read
results/task-by-task-87/summary.md on that branch first. The 87-task sweep's
pass rate is 64/87 = 73.6% under a metric that takes the better of val and test, but 45 of
those 64 passes are val-only and only 34 of 87 tasks have a test number at all. The
test-where-available cut is 62/87 = 71.3%, and raw-only pass^5 is 67.1%. The
per-task ledger, the caveats, and the cuts that disagree are all documented there.
A 3-iteration hill-climb over the same four shared office-document skills as the
baseline section above, starting from the run_baseline_opus frozen baseline
(reused via --reuse-baseline). Same fit-metric split (train == val == test =
all 87), num_trials: 1, paired gate k_se: 0.2.
Artifacts under .capevolve/run_opus_optimize3/ (gitignored, per-run): SUMMARY.md,
summary.json, baseline.json, final.json, events.jsonl, PATCH_NOTE.md,
rollouts/val/*.json (348: 87 seed + 87 × 3 candidates), rollouts/test/*.json
(174: 87 optimized + 87 baseline-seed).
| baseline (seed) | best (cand_0001) |
Δ | |
|---|---|---|---|
| val_reward (mean) | 0.281 ± 0.048 | 0.357 ± 0.050 | +0.076 (+27.2% relative) |
| pass_at_1 (fully-passing tasks / 87) | 23 (26.4%) | 28 (32.2%) | +5 tasks (+22% relative) |
| test_reward (fit metric = val by construction) | 0.281 * | 0.357 | +0.076 |
* The finalize step's baseline-test re-evaluation broke overnight (VPN drop; every
task returned 0 across a 5-hour eval); test_baseline_reward here is manually
spliced from the properly-scored seed val in run_baseline_opus. Under the fit-
metric split (train == val == test = same 87 tasks) the seed's test reward is
identical to its val reward by construction, so the splice is exact.
Documented in a PATCH_NOTE.md file inside the (gitignored) run dir on the
recording host; the raw broken artifacts remain untouched for audit.
| iter | candidate | parent | val | Δ vs parent | accepted? |
|---|---|---|---|---|---|
| 1 | cand_0001 |
seed |
0.357 | +0.077 | ✓ (paired gate: Δ > 0.2·SE) |
| 2 | cand_0002 |
cand_0001 |
0.325 | −0.033 | ✗ regressed |
| 3 | cand_0003 |
cand_0001 |
0.170 | −0.187 | ✗ large regression |
The optimizer converged to cand_0001; two follow-up attempts both regressed
and were correctly rejected. Optimizer spend: ~$32 total (well below the $400 cap).
8 tasks newly passing under cand_0001 (5 non-office + 3 office):
bike-rebalance, energy-ac-optimal-power-flow, energy-market-pricing,
exceltable-in-ppt (xlsx+pptx), grid-dispatch-operator,
paper-anonymizer, paratransit-routing, weighted-gdp-calc (xlsx).
3 tasks regressed from baseline: citation-check,
crystallographic-wyckoff-position-analysis, pptx-reference-formatting (pptx).
Net: +5 tasks (23 → 28) on pass_at_1. Two of the newly-passing tasks and
one of the regressed tasks are shared-office-skill tasks — the optimizer's edits
to xlsx/pptx had real cross-task effect.
- Fit metric, not held-out.
test == valby construction; thetest_deltaabove just re-affirms the val_delta at higher confidence. Not a generalization claim. - Single trial.
num_trials: 1— one draw per task per iteration. The paired gate still ran cleanly on the accepted iteration (Δ = +0.077 > 0.007 = 0.2·SE). - Not directly comparable to EvoSkills' 71.1%. Different paradigm (shared 4-skill package vs per-task skills), different setup (BenchFlow strips each task's own bundled skills and mounts ours). This is the same shared-skill approach as the baselines above.
Artifact: examples/skillsbench/run_full/
(report.md, final.json). Reproduce: REPRODUCE_skillsbench.md.
- Capability: the four shared office-document skill packages (
docx/pptx/xlsx/pdf). - Agent under test:
claude-sonnet-4-6in a Docker sandbox. - Optimizer:
claude-code@claude-opus-4-8. - Tasks / trials: 7 val tasks (
train == val) · 3 trials; 3 sealed test tasks. - Iterations: 7 (best
cand_0004, 4 accepted).
| reward | Δ | |
|---|---|---|
| Baseline — val | 0.333 | — |
Optimized (cand_0004) — val |
0.714 | +0.381 / +114% relative |
| Baseline seed skills — sealed test | 0.556 | — |
| Optimized skills — sealed test (held-out) | 0.667 | +0.111 / +20.0% relative |
Test was scored once on the sealed split for both baseline and optimized skills, so
the improvement is on tasks the optimizer never saw. The optimizer edited all four
SKILL.md bodies and added executable scripts, then stopped on a real ceiling
(diagnosing two unsolved tasks as broken oracles rather than overfitting them).
Five earlier rounds of agent-optimize on tau2-bench airline (aws/gpt-oss-120b, official
30/30/20 split) spent ~1500 rollouts to produce 10 learning steps and 1 accepted edit. Each
round bought exactly one bit per candidate — accept or reject — for 300 rollouts. This round
changed the shape of the search rather than the edits, and the numbers below are what that
bought.
A full-val gate costs val_n x n_trials rollouts per candidate. One task at n_trials
costs n_trials — a 30x cheaper feedback loop, aimed at the unit the defect actually lives in.
So: fan out one optimiser per defect task (and one per unstable cluster), let each iterate
against its own task plus a canary of tasks measured 1.0, merge the results with a git 3-way
merge, and gate the merge once against a null control. Ten optimisers ran; the first wave of
nine produced ~54 measured learning steps against the previous shape's 10.
New scripts, all contract-tested (scripts/check.py, 41 assertions):
taskeval.py (per-task eval + traces + partial credit), merge_taskopt.py (3-way merge with
per-branch bases), mechanisms.py (shared finding ledger), plus --gate-against control on
round.py.
Per-task, n=10, canaries 0 and 3 at 1.0 throughout:
| task | baseline | tuned alone | in the 3-way merge |
|---|---|---|---|
| 1 | 0.90 | 1.00 | 1.00 |
| 38 | 0.60 | 1.00 | 1.00 |
| 40 | 0.10 | 1.00 | 1.00 |
| 20 | 0.60 | 0.90 | 0.60 |
| 21 | 0.00 | 0.70 | 0.50 |
| 23 | 0.50 | 0.40 | 0.30 |
| 42 | 0.00 | 0.20 | — |
| 17 | 0.60 | 0.50 | excluded (conflict + net negative) |
| 33 | 0.44 | 0.44 | no edit survived |
| 7 | 0.00 | 0.00 | no edit survived |
A "74% merge retention" figure was computed from this table, reported, and then withdrawn.
An optimiser dispatched to repair the apparent interference re-measured the merged bytes twice
more and found the giveback was re-measurement noise: task 20 reads 0.6 / 0.9 / 0.5 across
three independent n=10 runs of byte-identical files (pooled 20/30 = 0.667), and task 21 reads
0.5 / 0.4 / 0.4 / 0.4 against the single solo reading of 0.70. It also verified by diff that
every owned mechanism survives the merge structurally intact, and killed two interference
hypotheses by direct computation rather than eval. At n=10 the standard error on a task near 0.5
is ~0.16, so any per-task difference below ~0.3 is indistinguishable from noise — which means the
tuned alone and in the merge columns above are hypotheses, not attributions, and only the
largest movements (task 40's 0.10 -> 1.00, task 38's 0.60 -> 1.00, task 21's pooled 0.00 -> 0.43)
survive that scrutiny.
The between-phases check is still worth its ~70 rollouts, but as a smoke test: it catches a merge that dropped an edit or broke a canary outright. To claim a per-task delta, pool every run of those bytes rather than quoting the last one.
- A feedback helper failed silently. A localizer called a method that does not exist; the
AttributeErrorwas swallowed by a bareexcept, so every failed numeric communicate check degraded to the generic "1 required piece(s) of information were not clearly communicated". An optimiser read that as "the checker is unsatisfiable" and spent seven rounds instructing the agent to state a value it was already stating. The repaired signal distinguishes never stated a figure from stated one and it was wrong; re-deriving it across 125 persisted rollouts cost zero rollouts. A check'sinfofield often is the expected value (airline stores a bare"1628"), so the message reports the agent's own figures and never the expected one. - A screening band is not a baseline. Ten tasks whose 3-trial rates summed to 2.33 measured
4.04 at
n=10— the screen understated the artifact by 1.71 task-equivalents (0.057 of val), and one task went the other way (0.33 -> 0.10). Every "0.0 DEFECT" label was suspect: three of them measured 0.30, 0.444 and 0.60. So the parent's true val atn=10was never 0.6778; that was a 3-trial number, and deltas were being computed against it. - Canaries chosen from the same small sample. Two canaries read 1.0 at 3 trials; one
measured 0.67 at 10, the other 0.667/0.333/0.0/0.333/0.333 across five independent 10-trial
runs.
canary_mean == 1.0was therefore unreachable for reasons no candidate caused, and several optimisers burned iterations chasing it. - Tool-return enrichment corrupted the signal without touching the score. A candidate nested summary objects under the key the adapter's id-extractor reads, so feedback claimed valid reservation ids were "not among the user's reservations". Reward was never affected — it comes from the harness's own checks, which never read a tool return — but one optimiser spent a round renaming a key and another concluded the name was capping its score.
- In-code guards beat prose, again. Every accepted edit this round was tool-side. Prose additions measured at or below control, and three were actively harmful: a "make the remaining calls NOW" hint suppressed intermediate reads; "ask which of the listed trips they mean" made the user simulator invent a non-existent trip; asserting a date-direction convention scored 0.200 against a 0.444 baseline.
- A guard must fire on a decision, not on a tool. One that refused on every first cancellation dropped a canary 1.0 -> 0.333 and pushed wall time 299s -> 1493s. Re-keying the same guard from per-user to per-contested-decision was worth 0.0 -> 0.333 by itself, and softening its wording flipped over-writing into under-writing.
- Concurrency does not compose. Nine optimisers at concurrency 8 put ~72 requests against a proxy sustaining 24-90; a per-task eval went from ~4 minutes to ~50. Nothing errored — latency grew, so it reads as "the model got slower".
- Findings belong in a ledger, not the coordinator's head. Four of nine optimisers
independently rediscovered one cause and two wrote colliding fixes for it.
mechanisms.pyrecords mechanism / evidence / owner / status, soverifiedmeans rebase onto that copy,proposedmeans someone owns it,rejectedmeans a retry must be structurally different. The round's ledger holds 41 findings — 16 verified, 16 rejected. - Memorisation must be enforced by a script. Diff each candidate against the base, grep the added lines only, and skip literals the base already contains — both filters matter, since a whole-file grep flagged three clean candidates on the harness's own example payment ids. It caught one real case: a docstring enumerating a task's own cities as if it were a general rule.
- The coordinator may audit the spec; an optimiser may not. Reading
evaluation_criteriato answer "is this task winnable, and is the optimiser chasing the right criterion?" is measurement integrity. Relaying a criterion the agent can evaluate from the conversation is legitimate; relaying a value, an id, or an expected write is not. That distinction revived two tasks their optimisers had written off. If the fix cannot be stated without naming the answer, the task is not winnable and that is the finding.
cand_fan — the union merge of five optimisers' verified work, 596 added lines, audited clean —
was evaluated on all 30 val tasks at 10 trials against a byte-identical ctl_null_i5 control
measured in the same round.
| arm | val (30 x 10) | paired delta | threshold (k_se 1.0) | verdict |
|---|---|---|---|---|
parent cand_toolguard |
0.7130 ± 0.0677 | — | — | — |
ctl_null_i5 (byte-identical copy) |
0.7130 | 0.0000 | — | control |
cand_fan |
0.7256 | +0.0126 | 0.0507 | reject |
The per-task fan-out produced no measurable val gain. +0.0126 is a quarter of the gate's own
threshold, and smaller than the +0.0352 phantom the precision mismatch alone produces (below).
Committed as a rejection; best_id stays cand_toolguard.
Three things the control earned in this one round:
- The parent's true val is 0.7130, not 0.6778. The 0.6778 everything had been compared against was a 3-trial number. Every delta computed against it was inflated by ~0.035.
- A 10-trial candidate gated against a 3-trial parent manufactures +0.0352. Demonstrated, not
argued: the byte-identical control scored
acceptagainst the stored parent —gate_delta+0.0352 over a 0.0337 threshold — while its true delta is exactly 0.0000. This is what--gate-against controlexists to prevent, and the round that introduced it caught it live. - A regression list at
n=10is noise. The control reported four regressed tasks and the candidate reported four — identical counts, disjoint sets, one artifact provably unchanged. With the no-regression veto still on, a copy of the parent would have been vetoed for the same reason as the candidate. That is the strongest justification the veto removal has received.
A defect in the new tooling also surfaced here and is fixed: under --gate-against control the
reported noise_floor_from_control computed to 0.0 by construction, because the control is the
reference. Claiming zero re-measurement noise is the most dangerous number the script can print; it
now reports the control's delta against the stored parent instead.
Where an optimiser reported a task unwinnable, the coordinator audited the spec (measurement integrity, never relayed as a value) and both audits changed the diagnosis without moving the score:
- Task 7 — the communicate check was called unsatisfiable after an optimiser stated twelve distinct candidate totals with the database passing and got 0.0 every time. The audit confirmed the check is satisfiable and the required figure is none of the twelve, so the agent's scope is wrong. But the task is bounded near 0.6 regardless: 3-4 of 10 rollouts die when the user simulator terminates right after a mandatory question, and the reward needs both components. Closed as a finding after 13 rounds across two optimisers.
- Task 42 — an optimiser ran a careful elimination over 30 scored trials and concluded the task needed hardcoding. The audit found its route classifier keyed off the reservation rather than the segment flying that date, which inverts the answer for a round trip; the correction was relayed as a per-segment, metro-level criterion. A third optimiser confirmed the fix by hand, and the score still did not move — the residual failure is downstream: in 30 traced trials the agent cancels one conflicting booking, then presents a table and asks instead of issuing the second cancel.
The second is the round's best methodological lesson: elimination evidence is only as good as the classifier feeding it. A one-line classification bug, visible in seconds by printing the helper's output next to its input, cost three eval rounds and produced a false verdict about the benchmark.
At 0.7130 on the official 30-task val at 10 trials, this artifact sits at the best honest airline
number this repository has ever produced — the previous best, 0.712, came from a 50-task fit
run with no holdout. Thirteen per-task optimisers, ~60 measured learning steps, and 77 recorded
findings moved it by an amount indistinguishable from re-measurement. Combined with the two tasks
shown to be structurally bounded (one at ~0.6 because the user simulator terminates before a
mandatory question can be answered and resolved), the evidence points to ~0.71-0.73 being the
ceiling for gpt-oss-120b on tau2-bench airline via policy and tool edits. Reaching 0.90 was not
possible within the stated constraints — no model change, no reasoning_effort change, no grader
access — and nothing in this round suggests a policy-or-tools edit that would close a 4 SE gap.
The sealed-test column is unchanged from the previous run, because no candidate was accepted:
best_id is still the artifact already measured there. Two
things bound it. The 20 test tasks were scored once in the previous run, so a test figure here is
their second scoring — no candidate was ever selected on them, but the disclosure belongs
with the number. And the per-task rates above are training numbers by construction: each
optimiser tuned against the task it is scored on. Only the full-val gate against its own control,
and the sealed test, are evidence.
The previous section closed by reporting that ~0.71-0.73 looked like the ceiling and that 0.90 was not reachable under the stated constraints. That conclusion was drawn from aggregate gate deltas, which is the wrong instrument for the question. Recomputing per task from the stored rollouts — free, since the 300-rollout arms were already on disk — gives a different and more useful picture.
Parent (cand_toolguard), exact rates from the byte-identical control arm at n=10:
| rate | tasks | count |
|---|---|---|
| 0.00 | 7, 42 | 2 |
| 0.10 | 23, 40 | 2 |
| 0.20 | 21, 39 | 2 |
| 0.40 | 20 | 1 |
| 0.50 | 14, 17, 33 | 3 |
| 0.80 | 9, 10, 12, 15, 38 | 5 |
| 0.90 | 5, 11 | 2 |
| 1.00 | (thirteen others) | 13 |
val = 0.7100, headroom = 8.70 task-equivalents over 30 tasks. Reaching 0.90 requires 5.70 of them — 65% of everything remaining — and 5.4 sits in six tasks. So the target is demanding, and the honest statement is that it is not costed as impossible; the earlier flat "unreachable" overstated what the aggregate numbers could support.
cand_fan scored +0.0126 against a 0.0507 bar and was rejected. Per task it contained both:
| direction | movement | total |
|---|---|---|
| gains | 40 0.10→1.00, 21 0.20→0.80, 39 0.20→0.50, 33 0.50→0.80, 38 0.80→1.00, 5 0.90→1.00 |
+2.1 |
| losses | 10 0.80→0.10, 9 0.80→0.40, 14 0.50→0.30, 17 0.50→0.30 |
−1.6 |
The gate saw neither, because a gate answers one bit about a sum. Keeping only the gaining half measures ~0.78. A merge must therefore be selected on a panel of the below-1.0 tasks before it is gated; tasks already at 1.0 cannot contribute a gain, so they are pure cost during selection (they remain essential in the gate, which is what protects them).
Attribution of the losses was also free: diffing the stored failure feedback per task showed the regressed tasks had identical feedback strings in both arms at different frequencies. The edit shifted a tendency rather than introducing a bug — a different thing to fix, and invisible from the means. The same pass priced infrastructure noise at 5 of 2040 val rollouts (0.25%), concentrated on one task: small, but measured small.
cand_fan merged 5 branches; ten branches carried verified mechanisms. The fixes for the two
tasks scoring 0.00 were not in it — they finished after the merge was cut. Re-merging exposed
a deeper fault:
| merge strategy | branches retained | result |
|---|---|---|
| whole-file 3-way | 4 / 10 | 6 verified branches dropped as "conflicts" |
whole-file + --union-on-conflict |
10 / 10 | did not parse; 5 duplicated defs |
per-function (funcmerge.py) |
7 / 7 leaves | parses, 14 tools register, audit clean |
The "conflicts" were not disagreements. Every optimiser had added one state field to the same
__init__ and one independent guard call to the same tool method right after the same existing
check — adjacent lines of a shared insertion point, which diff3 cannot distinguish from rival
rewrites. Merging per function makes independent additions stop interacting. Three details cost
real work to learn:
- Pick the trunk of a contested function by which branch changed that function most, not by
whose task holds the most headroom. The branch owning a full task-equivalent (task 7) had added
exactly one line to the contested
cancel_reservation; its real fix was elsewhere. Ranking by headroom discarded the branch that had actually rewritten the return value and kept nothing. - Dropping a losing branch's rewrite must not drop its insertions. Doing so would have discarded task 42's guard call to settle a disagreement about a money string.
- A helper can survive a merge with no call site — dead code that costs context and buys
nothing. Verify the call (
grep -c '_check_foo(reservation)'), never the definition. Relatedly, a ledgertouchesfield named a function (_remaining_upcoming) that no branch ever defined, so the merged artifact must be checked against the code, not against the ledger's description of itself.
reasoning_effort — the last untouched runner lever — does nothing here. It had never been
set in any prior run, and the reason turned out to be missing wiring rather than choice: litellm
validates parameters against its own model registry and rejects reasoning_effort client-side for
openai/aws/gpt-oss-120b, so the request never leaves the process. Forced through with
allowed_openai_params, it is real — reasoning tokens 3 / 91 / 327 for low / provider-default
/ high on an identical prompt (the gpt-oss "harmony" Reasoning: high system message does not
work through this gateway: 91 → 107, i.e. noise). Measured on the parent across all 17 headroom
tasks:
| arm | mean over the 17 headroom tasks | paired Δ |
|---|---|---|
| provider default (n=10) | 0.488 | — |
reasoning_effort=high (n=5) |
0.497 | +0.0088 |
Per-task deltas scatter from +0.60 to −0.50, on exactly the tasks independently measured as high-variance. This is a null result, and it retires the lever.
Turns are not the constraint, and failures are not truncations. Across 600 val rollouts,
596 end USER_STOP and exactly one hit MAX_STEPS, median trace 18 of 100 allowed messages.
Passing rollouts have median trace 18 / 9 agent turns; failing ones 24 / 12, with identical
narration-only counts. The agent is not cut off — it gets more chances and still commits to a
wrong argument. So edits aimed at brevity or at reminding the agent to keep going are aimed at
the wrong mechanism; the lever is decision quality at the moment of the call.
One thing checked and not established: single model calls are perfectly deterministic at temperature 0 (six identical completions by sha1), yet per-task rates sit at 0.8 and byte-identical artifacts re-measure 0.4 vs 0.222 on task 23. The variance therefore enters through the multi-turn conversation rather than the sampler, but its cause is not identified here and should not be attributed to sampling.
The most expensive defect of this round was not a policy mistake. funcmerge carried
_check_bags_before_cabin_change and its call site, and left behind the class attribute both
needed. The live tool return was:
Error: 'AirlineTools' object has no attribute 'CABIN_LADDER'
tau2's tool layer turns the AttributeError into a string, the agent reads it, abandons the bag
change, and the reward records a missing write — indistinguishable from the agent deciding not
to act. All seven merge-derived candidates and all four per-task working copies carried it, so
these readings are floors, not values:
| measurement | reported | status |
|---|---|---|
cand_all (17-task panel) |
0.494 | measured with a live crash |
cand_lift (17-task panel) |
0.553 | measured with a live crash |
abl_noshift / abl_nochrono |
0.533 / 0.567 | contaminated; trade-off conclusion retracted |
The retraction matters as much as the bug. Those two ablations returned near-identical per-task results, which was read as evidence that the two guards are coupled on one code path. A shared crash explains the same pattern equally well, so the "guard is a net −0.20 trade-off" conclusion does not survive and was withdrawn.
Three process facts came out of it, in descending order of how much they should change practice.
The detector worked and the reader failed. The dropped_additions audit printed
CABIN_LADDER = ("basic_economy", "economy", "business") as the third line under that branch, and
the coordinator read past it. An advisory list that names a crash is not enough; the crash class
now hard-fails — funcmerge refuses to write a result in which a constant-shaped attribute
read off self is undefined.
Making that check safe took two corrections, both of which are the general lesson. Instance
fields are routinely declared with annotations (self.x: set[str] = set()), which is
ast.AnnAssign, not ast.Assign; collecting only the latter reported six valid fields as
undefined and refused a good merge. And hard-failing on every unresolved self.NAME breaks on
inheritance — the class under merge has a base class, and an inherited method is not resolvable
from one file. Only UPPER_CASE names hard-fail; the rest are advisory. A hard check with false
positives is worse than no check.
It was found by a per-task optimiser reading a live trace, not by any aggregate. Four full-panel measurements had already passed through it. Nothing about the means was anomalous, because the crash's signature is the same signature as the defect everyone was already hunting. That is the strongest argument in this whole run for keeping per-task traces in the loop: an aggregate cannot distinguish "the agent chose wrong" from "the tool raised".
A related contamination the same episode created, worth recording because it is a coordination
hazard rather than a code one: the fix was applied in place to four live working copies while
their evaluations were running, and taskeval builds the toolkit per rollout. Any eval spanning
the write mixes crashed and clean rollouts. One optimiser established by grepping its own traces
that contamination was confined to a single trial and kept its run; another found its run
straddled the write by two minutes and discarded it. Both responses were right, and neither was
available without per-rollout traces on disk.
| arm | measurement | note |
|---|---|---|
parent (cand_toolguard) |
0.7100 val, n=10 | byte-identical control arm, 300 rollouts |
| task 7 | 0.00 → 0.30 | n=10, reproduced 3x incl. crash-fixed; ceilinged |
| task 17 | 0.50 → 0.70 | n=10, canaries 1.0; ceilinged |
cand_best |
0.7333 val, n=5 | paired +0.0867 vs same-batch control, accept |
| arm | val (30 tasks x 5 trials) | paired Δ | SE | bar (k=1.0) | verdict |
|---|---|---|---|---|---|
ctl_final (byte-identical parent, same batch) |
0.6467 | — | — | — | control |
cand_best |
0.7333 | +0.0867 | 0.0548 | 0.0548 | accept |
Both arms were launched in the same batch at equal concurrency, so they share endpoint conditions — the one comparison this benchmark has repeatedly shown you cannot fake, since a 10-trial candidate gated against a 3-trial parent manufactured a phantom +0.0352 with byte-identical code.
cand_best composes: the seven-branch per-function merge, the lifted docstring guidance, u33's
policy paragraph, u33's next_step block restored byte-exact, r2t7's task-7 mechanism, an
own-booking cancel guard, and the CABIN_LADDER crash fix. Thirteen tasks improved — 40
0.00→0.80, 33 0.20→0.80, 17 0.00→0.60, 39 and 15 both 0.60→1.00.
Three caveats belong next to that number, not below it.
The absolute level is ambiguous and the delta is not. The same byte-identical control read
0.6467 here at n=5, while the parent measured 0.7100 at n=10 over 300 rollouts. Anchored
to the 300-rollout parent, cand_best is ≈0.797; anchored to today's control it is 0.7333. The
paired delta is the statistic that survives either anchor, so the claim is +0.0867, not a
level.
It clears the bar by a hair — 0.0867 against 0.0548, about 1.58 SE, one-sided p ≈ 0.06. An accept, not a rout.
And then the null retracted it. Re-running the byte-identical control a second time, on the same seeds, produced this:
| arm | reading | paired Δ vs control run 1 | verdict at k_se=1.0 |
|---|---|---|---|
| control, run 1 | 0.6467 | — | — |
| control, run 2 — byte-identical, same seeds | 0.7267 | +0.0800 | "accept" |
cand_best |
0.7333 | +0.0867 | accept |
cand_best vs control run 2 |
— | +0.0067 | reject |
A byte-identical null passed the same gate at nearly the same magnitude, and the candidate's verdict flips depending on which control reading it is compared against. So no gain is demonstrated. The accept above is withdrawn.
What this is not: seeds were identical across all three runs, temperature is 0, and single model calls are perfectly deterministic (six identical completions by hash). This is run-to-run nondeterminism entering through the multi-turn conversation, whose cause is still unidentified. It follows that a determinism check cannot bound it, and neither can more trials inside the same seed block — the whole arm has to be run again.
The retraction above is a symptom; this is the disease. Diffing the two byte-identical control runs per task — same bytes, same seeds, temperature 0, n=5:
| identical code, run 1 vs run 2 | value |
|---|---|
| mean per-task |difference| | 0.160 |
| median | 0.200 |
| tasks that moved at all | 19 / 30 |
| tasks that moved >= 0.40 | 3 |
| worst single-task swing | 0.60 (task 23, 0.20 -> 0.80) |
Task 23 moved 0.20 -> 0.80, task 14 0.40 -> 0.80, task 11 0.60 -> 1.00, all with no change
to the code. The per-task k/n signal that the whole fan-out design rests on is, at this trial
count, mostly noise. At n=10 the floor is roughly 0.11 — still large against a claimed 0.20 step.
This downgrades a large fraction of what this round reported: every n=5 panel comparison (both
merge panels, both guard ablations, the cand_best accept and its improved/regressed lists) sits
at or under the floor, and the per-task lists are close to uninformative.
One nuance partially rescues the method. The variance is concentrated in particular tasks rather than spread uniformly. Across five byte-identical readings of one 12-task subset (means 0.5333 / 0.6500 / 0.5667 / 0.6500 / 0.6500), tasks 0 and 46 read 1.00 in every run while tasks 23, 14 and 11 carried nearly all the movement. So a canary set drawn from demonstrably stable tasks is trustworthy even though per-task rates in general are not — which is why every result in this round could still report "canaries 1.0" meaningfully, and why canaries must be chosen from repeated measurements at the real trial count rather than from a 3-trial screen (an earlier defect this run had already had to correct).
What survived is exactly what was established without a rate. Three findings rest on structure rather than on a delta, and none of them moved:
- the
CABIN_LADDERcrash — found in live tool returns, confirmed by its error string going 3 -> 0; - the
Returns:docstring section never reaching the model — 5469 of 12929 characters, counted off the rendered schema; reward_basismislabeling — the tool named in the loudest feedback line was invoked in 0 of 300 rollouts while the task still scored 0.8.
Every finding that rested on a rate difference was later retracted or downgraded. The practical conclusion for the algorithm is not to abandon per-task work but to stop treating the rate as the evidence: the per-task loop's real output is a diagnosis verifiable without the metric — a wrong argument visible in a trace, a tool that raised, text that never arrived — and the rate is only a hint about where to look next.
The practical floor: run-to-run paired noise at n=5 over 30 tasks is ≈0.08, which is larger than every effect measured in this session. Every n=5 panel comparison reported above — the guard ablations, the composition result, the merge panels — sits at or below it and must be read as unresolved rather than as evidence. The one comparison that survives is the seed-matched composition test, and only because its two arms differ by a known edit rather than by a re-run.
The discipline this buys, and the reason it is in the skill now: evaluate the control twice before believing any candidate, and set the bar from the null's own spread rather than from a formula. A gate whose bar is smaller than the null's re-run delta is not a gate.
Five previously-solid tasks dropped, three of them canaries (0 1.00→0.60, 3 and 27
1.00→0.80). A regression list at n=5 is noise-dominated on this benchmark — a byte-identical
control has itself reported four — but three canaries moving together is reported rather than
waved away.
Two tasks were audited to a ceiling and closed. Task 7 is pinned at 0.30 because its DB and
COMMUNICATE halves are each winnable at 0.3–0.5 but land in different rollouts, and the
residual — a turn spent on the confirmation the policy mandates — now measures negative from both
the policy-prose surface and the tool-return surface. Task 17 is pinned near 0.70 on three
failures that share no cause, two of which are user-simulator artifacts (the simulator appends its
stop token to the same message that grants confirmation; and it volunteers a wrong reservation id
while describing the right route).
The retracted accept was replaced with the estimator that should have been used from the start: repeat the whole paired comparison on distinct seed blocks, both arms in the same batch, and take the error across runs.
| seed block | cand_best |
control | paired Δ |
|---|---|---|---|
| 0–4 | 0.7333 | 0.6467 | +0.0867 |
| 0–4 (control re-run only) | — | 0.7267 | null +0.0800 |
| 100–104 | 0.6867 | 0.6667 | +0.0200 |
| 200–204 | 0.7133 | 0.7067 | +0.0067 |
| 300–304 | 0.7067 | 0.7067 | 0.0000 |
| combined | +0.0283, SE 0.0199, t = 1.42 — NOT DEMONSTRATED |
The deltas decay monotonically to exactly zero: +0.0867, +0.0200, +0.0067, 0.0000. The first run — the one that produced the accept — was the outlier, and each subsequent run moved toward no effect. That is the signature of a null, not of a small real gain.
The deltas shrink run over run; the first was the outlier, and it was the one that produced the accept. Pooling every reading of each arm:
| arm | readings | mean | spread |
|---|---|---|---|
| control (byte-identical parent) | 0.6470 / 0.7270 / 0.6670 / 0.7070 | 0.6870 | 0.0800 |
cand_best |
0.7330 / 0.6870 / 0.7130 | 0.7110 | 0.0460 |
| parent, independent 300-rollout n=10 | — | 0.7100 | — |
cand_best's pooled mean (0.7110) is essentially the parent's own best-measured value (0.7100), and
the +0.024 gap to the n=5 control pool comes mostly from the control reading low rather than the
candidate reading high.
So the honest conclusion is that the composed artifact does not measurably beat the artifact it
was built from — after a seven-branch per-function merge, lifted docstring guidance, a policy
paragraph, a byte-exact restored next_step block, a task-7 mechanism worth +0.30 on its own task,
an own-booking cancel guard, and a runtime-crash fix.
Two things are worth separating out, because they are the transferable part.
The per-task work produced real, verifiable mechanisms. Task 7 went 0.00 → 0.30 and task 17 0.50 → 0.70 and task 14 0.50 → 0.70, each reproduced on its own base with canaries intact. What failed was composition: adding one optimiser's verified task-14 edits to an artifact already carrying three others' work measured -0.0617 on seed-matched arms, and task 14 itself fell. A per-task gain is verified against one base and is not transitive to another.
The ceiling that mattered was the instrument, not the model. The binding constraint on this run was never a missing idea; it was that a paired full-val comparison at an affordable trial count cannot resolve an effect below roughly 0.08, and per-task rates move 0.16 on identical code. Under that floor, "reach 0.90 in a few iterations" is not a hard target — it is an unmeasurable one. A credible verdict on a +0.05 effect here needs several full paired runs; certifying +0.19 would need the effect to exist first, and nothing in 150 recorded findings suggests a policy-or-tools edit of that size.
Superseded — read "The noise was never mysterious — it was the trial count" below. This section's direction survives (conc 25 genuinely carried excess noise above the binomial floor) but its blanket attribution to load does not: at conc 8 the residual spread is 1.27x the binomial SE, i.e. nothing left to explain. Kept as written because the correction is the interesting part.
The last experiment of the round was the most useful one. If the re-measurement noise were inherent to the model or the benchmark, nothing about this setup could be trusted again. It is not: it is substantially a function of endpoint load.
Same bytes, same seeds, same 12 tasks, control measured twice at each concurrency:
| identical bytes and seeds | conc 25 | conc 8 |
|---|---|---|
| arm-level |delta| between the two runs | 0.1167 | 0.0333 |
| mean |per-task| movement | 0.250 | 0.100 |
| tasks that moved at all | 10 / 12 | 5 / 12 |
Tasks 9, 17, 40, 39 and 15 each moved 0.20–0.40 at conc 25 and were perfectly repeatable at conc 8. Arm-level noise fell 3.5x and per-task noise 2.5x.
Two caveats, both registered in the ledger before the result landed so the reading could not be chosen after the fact: the two low-concurrency runs ran sequentially, so load and elapsed-time drift are confounded; and 12 tasks × 2 runs makes a variance comparison thin. The direction was consistent across all three metrics, which is why it is worth acting on, but it is not settled.
The operational conclusion is "search fast, gate slow." Per-task exploration can run at high concurrency because its product is a mechanism verified from a trace rather than a rate. The accept decision must run at a concurrency where the null actually reproduces — roughly 3x the wall clock, for the one evaluation whose answer is load-bearing. Every gate in this round, including the retracted accept, ran at conc 25 or above.
This also qualifies the earlier conclusion of this section. The instrument was the ceiling at the concurrency used, not inherently; a repeat of this round with low-concurrency gating would be able to resolve effects that this one could not. That is the single most actionable thing the round produced, and it cost 120 rollouts to find.
Last reviewed: 2026-08-19.
Two 10-trial readings of near-identical artifacts give 600 val rollouts. Pooled per-task at n=20, val is 0.7167, which is 5.50 task-equivalents short of 0.90. Twelve tasks hold 8.05 of available headroom, so the target is arithmetically reachable; nothing about the distribution says it is easy.
The two arms illustrate the noise problem rather than a gain: cand_fan 0.7233 vs
ctl_null_i5 0.7100, paired Δ +0.0133, while individual tasks swung 0.90 (task 40: 0.10 → 1.00)
and 0.70 (task 10: 0.80 → 0.10) between them. Tasks 7 and 42 read 0.00 in all 40 rollouts —
the only headroom in the set that is certain rather than inferred.
Classifying all 600 rollouts:
| signature | count | share of the 170 failures |
|---|---|---|
| database state mismatch | 162 | 95% |
| wrong argument value | 110 | 65% |
| required write never made | 71 | 42% |
| communication miss | 41 | 24% |
Reward is strictly binary — 430 passes, 170 failures, zero partial — so a task's rate is a pass probability. Failing episodes run longer than passing ones (median 24 messages vs 16) and only 1 of 170 failures hit the step cap: the agent is not running out of steps, the simulated user stops. Turns are a scarce currency, which is the argument for repairing a recoverable argument slip inside the tool instead of bouncing it back.
User messages containing leaked <reasoning> skew 32 failures : 7 passes. In 21–22 episodes
the simulator emitted **Per-component ceilings, which would otherwise look like a failed mechanism.** Reward requires every component in reward_basis`, so a task whose COMMUNICATE rate is below 1.0
cannot be lifted past it by any database fix. At n=40, COMMUNICATE is exactly 1.0 on seven of nine
tasks — but task 23 sits at 0.667 and task 14 at 0.730. So task 23's realistic headroom is
0.436, not 0.769, and task 14's is 0.216, not 0.486. A perfect DB fix on task 23 still
leaves it near 0.67. Its communication misses are arithmetic/scope errors rather than silence: the
agent stated 13–19 distinct figures on task 23, and up to 56 on task 14, while still missing
the required one.
###STOP###` in the same message as reasoning that explicitly planned to continue ("we must wait for agent's third message. Continue."). Task 7 is worst hit: 15 of its 27 observed failures carry a leak.
I expected a harness bug — is_stop() is a bare substring test, so a stop token discussed inside
leaked reasoning would end the episode spuriously. Measured: only 1 of 26 cases has the token
inside the reasoning block. The other 25 sit after </reasoning>, so the simulator model is
genuinely stopping and this must not be patched away. It caps the achievable score at ≈0.958.
The check cost ten minutes and prevented shipping a "fix" for a bug that was not there.
Rollouts containing a guard refusal pass at 0.767 against 0.702 for those without, so the
guard programme was worth it. Two guards run the other way: segment_departs_after_previous_arrival
(1 pass / 5 fail) and payment_adds_up_to_total_price (0 / 4). Both fire on 4–6 rollouts and
fire on the hardest itineraries, so the rates are confounded with difficulty — motivation, not
proof.
The mechanism behind it is visible in a trace. On task 10 the chronology guard rejected an itinerary and the agent's recovery was to silently move the date the user had asked for. A refusal that names no valid alternative invites an unauthorised workaround. The reusable form is "refuse and name a valid option", or a deterministic auto-repair where one exists.
- Feedback pointed at evidence that did not exist. A DB-only divergence produced
"…See the per-action detail below for the specific wrong argument." followed by nothing,
because no gold action had mismatched. Task 10 failed 11/11 with exactly that text. It now
states that this is not a wrong value but an extra/duplicated write or a write side effect —
payment_historyis appended by every successful update and no retry removes it — and lists the agent's own writes in order. Unit-tested on both branches and ported to the template. - Per-process concurrency was mistaken for load.
TAU2_MAX_CONCURRENCYis per process, so a gate at--conc 8running beside four exploring optimisers at--conc 12is a ~56-in-flight measurement wearing a low-load flag. One such reference run was launched this round and killed rather than recorded. This qualifies the earlier conc-25-vs-conc-8 result, whose low-conc arms happened to be sequential and therefore genuinely quiet.
round.pynow defaults--concurrencyto 8 and reportsmeasurement_concurrencyplus aconcurrency_warningin the result when run above 12, so an unresolvable verdict cannot be read as a clean one.- New
integrate.py: folds verified branches in one at a time, measuring after each, with canaries inside the objective, sub-floor steps recorded askept_provisionallyrather than as gains, and the subset objective self-labelled as upward-biased by selection. This exists because a one-shot merge of ten verified branches measured −0.0617, with the merged task's own fix regressing 0.40 → 0.20. - SKILL.md gained the fan-out evidence contract: a parallel optimiser's deliverable is a mechanism with trace proof, not a rate. A fan-out is by construction the high-load regime, so briefing K subagents to "measure whether your edit helps" asks for the one thing they cannot provide. Load-independent evidence (the guard fired; the next action changed; the bad payload now succeeds; the delivered docstring contains the keys) replaces it.
- Contract check: 60 → 67 checks, all green.
Six optimizers ran in parallel against real traces. Five produced edits; all five pass the toolkit's own runtime self-check and a scan for hardcoded task/user/reservation literals.
| candidate | surface | mechanism |
|---|---|---|
c_readfirst |
tools | a write is refused until get_reservation_details has been called for that reservation; nudges reading every booking when the request is a cleanup |
c_nextstep |
tools | book_reservation names any still-active reservation on the same route/date, excluding ones booked in this same conversation so a split party does not flag its own siblings |
c_argrepair |
tools + descriptions | amount aliases and "$255"/"255.00" parsed; a single missing amount inferred from the real bill; a mismatched remainder moved to the one credit card; nested object keys named exactly in the docstrings |
c_t910 |
tools + policy | adds a read-only preview_reservation_change dry-run reusing the real pricing and preconditions, plus policy requiring preview before any write |
c_batch |
policy | treat each message as the only one you get: batch every missing field and every confirmation into one message |
preview_reservation_change adds a tool, i.e. changes the action space. It is read-only, leaks
no gold, and reuses the existing pricing path — disclosed here rather than buried.
Gating is running serialized at --conc 8, with two byte-identical controls on identical
seeds ahead of the candidates, because the gap between those two controls is the only honest bar.
The round's central methodological result, and it corrects a claim this file made one section earlier. Two byte-identical controls, same seeds, serialized at concurrency 8, over the 10 stable target tasks:
| value | |
|---|---|
| ctlA target mean | 0.3066 |
| ctlB target mean | 0.3844 |
| observed |difference| | 0.0778 |
| predicted binomial SE at n=10 over 10 tasks | 0.0615 |
| ratio | 1.27 |
Mean per-task movement was 0.0978 against a binomial prediction of 0.1445 — less movement than chance requires. There is nothing left to explain, and the earlier blanket "load-induced noise" row is superseded in the ledger.
The concurrency result survives in a narrower form: at conc 25 per-task movement was 0.250,
genuinely above the 0.1445 floor, so dropping to conc 8 removed a real excess and exposed the
floor beneath it. Lowering concurrency fixed what it could; the rest is n.
One precision, because it matters for honesty: at temperature 0 with identical seeds a fully deterministic system would return identical arms, so this is not sampling error in the textbook sense. The defensible claim is that the variation is indistinguishable in magnitude from independent per-rollout coin flips. Its physical cause remains unidentified — MoE batching, seed races and set-ordering were each tested and disproved in earlier rounds. The value of the binomial framing is negative: no further mechanism needs positing, and more trials is the remedy either way.
"Noise is high, so narrow the task set" is backwards for judging an artifact. A task sitting at 1.00 contributes signal to the mean with almost no variance, so removing it discards a free denominator:
| arm | rollouts/arm | SE of paired difference |
|---|---|---|
| 12 hard tasks, n=10 | 120 | 0.0496 |
| full val 30 tasks, n=10 | 300 | 0.0262 |
| full val 30 tasks, n=20 | 600 | 0.0185 |
Full val at n=10 is nearly twice as precise as the hard subset. The four prior gate rounds ran at full val n=5 (SE 0.0371), so their task set was never the problem — they ran at a concurrency carrying excess noise on top of that, and chased effects smaller than the sum.
The two questions therefore need opposite designs, and conflating them was the real error:
- does this mechanism work → only the tasks where it fires, high
n, per-task two-proportion test. A mechanism firing on two tasks is diluted to nothing in a 30-task mean. - what is the artifact worth → full val, both arms in one batch, because the 1.00 tasks are free precision.
One control at n=40 on the 9-task union, then each mechanism on its own tasks at the same seeds, serialized at conc 8. Per-task two-proportion z. Canaries clean at 1.00 on every arm.
| mechanism | task | control | candidate | Δ | z | |
|---|---|---|---|---|---|---|
destguard refuse round-trip destination == origin |
23 | 0.231 | 0.359 | +0.128 | +1.24 | |
| 14 | 0.514 | 0.553 | +0.039 | +0.34 | ||
searchdate echo the searched date into results |
33 | 0.487 | 0.600 | +0.113 | +1.01 | pre-registered |
| 21 | 0.700 | 0.650 | −0.050 | −0.48 | exploratory | |
| 10 | 0.425 | 0.375 | −0.050 | −0.46 | exploratory | |
baggage disclose free-bag allowance on cabin change |
17 | 0.375 | 0.450 | +0.075 | +0.68 | |
| 21 | 0.700 | 0.718 | +0.018 | +0.18 | ||
| 33 | 0.487 | 0.575 | +0.088 | +0.78 | ||
t910 preview tool + decline-if-over-limit |
10 | 0.425 | 0.941 | +0.516 | +4.68 | resolvable |
| 9 | 0.417 | 0.129 | −0.288 | −2.61 | resolvable regression | |
nextstep name the still-active same-route booking |
23 | 0.231 | 0.333 | +0.102 | +1.00 |
The headline is the collective direction, not any single row. 8 of 9 pre-registered task-level
deltas are positive; a one-sided sign test gives p = 0.0195. No individual z reaches 2 except
t910's pair. Including the two exploratory tasks that were not pre-registered, both negative, it
is 8/11 and p = 0.113 — so the pre-registration carries the result, and it is legitimate only
because every target was written into the ledger before its arm ran.
Method worth keeping: when individual effects sit below the noise floor, pre-registered directional predictions plus a sign test demonstrate them at a fraction of the cost. An aggregate mean cannot — see the certification arithmetic below.
De-duplicated (two mechanisms each claim tasks 23 and 33, so they cannot both be additive):
| accounting | per-task sum | val |
|---|---|---|
| upper bound, fully additive | +0.791 | +0.0264 |
| conservative, best single mechanism per task | +0.601 | +0.0200 |
To certify +0.02 at 2 SE on full val requires n ≈ 178 per task — about 5,300 rollouts per arm, ~26 hours per arm at the measured 17.3 s/rollout. At practical budgets the effect is 0.47 SE (n=10), 0.95 SE (n=40), 1.50 SE (n=100). A full-val gate cannot return a significant result at any budget available here, and quoting its point estimate as a gain is precisely the error the four earlier rounds made.
So the honest resolution: the mechanisms are individually sound (deterministic proof they engage, correct incidence with pass/fail asymmetry) and collectively positive (p = 0.0195), while their aggregate magnitude sits below the measurement floor of the benchmark used to judge them.
c_readfirst — refuse a write until get_reservation_details has been called for that reservation
— did not score; it made tasks fail outright, 30 times on task 42 and 28 on task 23.
tau2's Environment.replay re-runs the agent's own trace against a fresh environment as a
determinism check, and deliberately skips every non-mutating tool ("Non-mutating tools (reads,
thinks, etc.) don't change state — skip them"). The guard's precondition is set by a read, so
in replay it always fires, the write returns PRECONDITION_FAILED instead of the recorded success,
and strict=True raises.
General rule: any guard whose precondition is established by a non-mutating tool is incompatible with this evaluator. Write-established state is fine, because writes are replayed. The available workaround — detect replay and relax the guard — was refused: behaving differently under evaluation is cheating however it is framed.
c_t910 produced the round's only large per-task effect (task 10 +0.516, z +4.68) alongside a
resolvable regression (task 9 −0.288, z −2.61). It bundled five changes, so neither was
attributable. Both ablations changed exactly one thing, with the rest byte-identical.
| variant | task 10 | task 9 | net | canaries |
|---|---|---|---|---|
c_t910 full bundle |
+0.516 (z 4.68) | −0.288 (z −2.61) | +0.228 | clean |
c_preview − "finish everything" paragraph |
+0.375 (z 3.44) | −0.235 (z −2.12) | +0.140 | clean |
c_nonoop − no-op-change refusal |
+0.498 (z 4.71) | −0.147 (z −1.32) | +0.351 | task 1 → 0.80 |
Ablation 1 refuted its own hypothesis. I predicted the "finish everything that was asked"
paragraph caused the task-9 regression, because task 9's failures under c_t910 write on a
reservation the control never touches. Removing it left task 9 regressed and cost task 10 0.141 —
so that paragraph was mildly helpful and irrelevant. Without that null I would have shipped the
wrong fix and kept the real cause. The two variants are also not separable from each other (task 10:
0.941 vs 0.800 is z 1.80).
Ablation 2 found the cause. Removing only the change_must_change_something refusal halved the
regression while task 10 kept its gain. The general lesson, now in SKILL.md: a guard that forbids
the harmless option can force the harmful one. Task 9's correct behaviour is to make no change;
refusing the no-op left the agent with only real changes to choose from, and it chose one. The
failure is not the guard's logic but its closure — before adding a refusal, name the action set it
leaves behind and check that "do nothing" is still reachable.
And the canary vetoed it. Re-measuring task 1 at n=20 on both arms: parent 1.00 (20/20),
c_nonoop 0.85 (17/20). The single z is only −1.80, but task 1 reads exactly 1.00 for the parent
and for all seven other candidates this round, and c_nonoop alone breaks it in two independent
readings. So the guard is a genuine trade-off — −0.141 on task 9, +0.150 on task 1, net a wash —
and it is retained, because a task at 1.00 is worth more than a marginal gain on one at 0.417.
Consequence for the shipped artifact: c_win keeps the guard, and therefore carries c_t910's
task-9 regression of −0.288. That is a known, measured cost of the round's largest gain, reported
rather than omitted.
c_win = destguard + searchdate + baggage + nextstep + argrepair + the read-only
preview_reservation_change tool, with c_t910's policy. Six branches merged per-function with
zero conflicts and no undefined attributes; 15 tools; toolkit self-check passes; every mechanism
verified live on the merged artifact (date echoed, destguard firing, argrepair repairing an
aliased-id-no-amount call, baggage note present, preview registered). c_readfirst is excluded —
it is unmeasurable under tau2's replay.
One conflict during assembly was my own doing: the literal-cleanup regex had edited a
payment_id docstring line in two branches, producing a false conflict between collateral edits.
funcmerge refused to write rather than union it — the behaviour that was missing when an earlier
round shipped an unparsable file — and reverting c_searchdate to its minimal one-line change
cleared it.
audit_no_memorization.py was refusing clean candidates for three reasons that were not
memorisation: hand-written placeholders (credit_card_0000000), the seed's own documented example
(HAT001, which lives in reference/data_model.py and was missing from the base context), and
noqa: BLE001, a linter suppression matching the confirmation-code shape. All three fixed in the
audit rather than by editing measured artifacts — changing bytes after measuring them is the worse
error. All 8 candidates now clean, and the authoritative test — intersecting every identifier
against the live tau2 database — reports zero real identifiers anywhere.
Two independently-seeded blocks, 30 val tasks at n=10 each, serialized at conc 8, arms alternated within each block so drift could not land on one arm.
| block | control | c_win |
paired Δ | SE | z |
|---|---|---|---|---|---|
| seed 0 | 0.7326 | 0.7367 | +0.0041 | 0.0248 | +0.16 |
| seed 100 | 0.7533 | 0.7200 | −0.0333 | 0.0265 | −1.26 |
| pooled | −0.0146 | 0.0181 | −0.81 |
The blocks disagree in sign, so by the rule fixed before the gate ran, the artifact-level effect is not detectable at this budget — no accept, and the pooled point estimate is negative anyway. This is not a drift artifact: the two byte-identical control blocks, three hours apart, differ by 0.0207, below the 0.0262 binomial SE. The instrument was working; the artifact was not.
The informative cut is which per-task deltas replicate across both blocks:
| direction | tasks | deltas |
|---|---|---|
| gains that replicate | 10, 33 | +0.80/+0.40, +0.20/+0.10 |
| losses that replicate | 9, 11, 15, 20, 39, 47, 49 | −0.30/−0.10, −0.10/−0.30, −0.10/−0.20, −0.14/−0.40, −0.30/−0.10, −0.30/−0.10, −0.10/−0.20 |
Two replicated gains against seven replicated losses. Tasks 11, 15, 47 and 49 were high-scoring tasks the artifact damaged, and none of them was in the nine-task canary set — the canary set was too small, and picking canaries only from tasks the mechanisms were expected to touch left the rest of the suite unguarded.
Worse for the mechanism case: task 20 replicates a loss (−0.14/−0.40) despite argrepair
measuring +0.220 (z +2.35) alone at n=40. A resolvable per-task gain did not survive
integration. That is the non-composition result again, this time with a properly powered
per-mechanism measurement on one side of it.
From block 0 alone I proposed that destguard and nextstep interfere on task 23: both measured
individual gains at n=40 (+0.128, +0.102) while the merged artifact read −0.50 there. Block 100
reads task 23 at +0.20 — the sign flips. It was n=10 sampling noise, not interference.
Six of the moving tasks flip sign between blocks (23, 43, 17, 41, 27, 12). That is the correct null expectation at n=10, and a direct warning against reading a mechanism story out of a single block's per-task table — which is exactly what I did before the second block landed. The planned n=40 interference test was cancelled as no longer informative.
Exactly one mechanism: the read-only preview_reservation_change tool, whose effect on task 10 has
now replicated in three independent measurements — +0.516 at n=40, +0.80 in block 0, +0.40
in block 100 — alongside its measured cost on task 9 (−0.288 at n=40, −0.30/−0.10 at full val).
Tasks 11, 15, 39, 47 and 49 are targeted by no mechanism at all, so their damage in c_win is
unattributed. c_t910 alone was never measured on them, which is the one remaining question worth
600 rollouts: does the minimal artifact carrying the real gain also carry the collateral?