44evalbench.eval() exits 0 whenever a run completes, including when the agent
55made no tool calls at all, so the exit code alone cannot gate CI.
66
7- Tier 1 (threshold): trajectory_matcher must hit THRESHOLDS -- the agent really
8- reached the expected tools.
9-
10- Tier 2 (non-zero): telemetry scorers must report more than 0. They swallow
7+ Tier 1 (non-zero): telemetry scorers must report more than 0. They swallow
118parse failures and return 0.0 with an explanation rather than raising, so a
129plain liveness check passes even when a CLI renames a token field -- the exact
1310drift this build exists to catch. Every scenario makes at least one MCP call,
1411so 0 tokens or 0 latency can only mean the scorer failed to read the output.
1512
16- Tier 3 (liveness): every remaining scorer must emit a row per scenario, with no
13+ Tier 2 (liveness): every remaining scorer must emit a row per scenario, with no
1714comparison_error and a numeric score. This covers the LLM judges without ever
1815gating on their verdict, which would make the build flaky.
16+
17+ trajectory_matcher is deliberately liveness-only, not thresholded: which tools
18+ an agent reaches for varies run to run, and a harness may shell out instead of
19+ calling the MCP tool. tool_call_latency > 0 is what proves tools were used.
1920"""
2021import csv
2122import json
2728HARNESSES = ["agy_cli" , "claude_code" , "codex_cli" , "gemini_cli" ]
2829RUN_CONFIG_DIR = ".ci/run_configs"
2930EVALSET = ".ci/harness_smoke.evalset.json"
30- THRESHOLDS = {"trajectory_matcher" : 100.0 }
3131POSITIVE = {
3232 "turn_count" ,
3333 "agent_steps" ,
@@ -102,11 +102,6 @@ def check(harness, scenario_ids):
102102 problems .append (f"{ scorer } : errored on { sid } -- { error [:120 ]} " )
103103 elif score is None :
104104 problems .append (f"{ scorer } : non-numeric score for { sid } " )
105- elif scorer in THRESHOLDS and score < THRESHOLDS [scorer ]:
106- problems .append (
107- f"{ scorer } : { sid } scored { score :.1f} , "
108- f"need >= { THRESHOLDS [scorer ]:.0f} "
109- )
110105 elif scorer in POSITIVE and score <= 0 :
111106 problems .append (
112107 f"{ scorer } : { sid } reported 0 -- scorer could not read "
@@ -117,9 +112,7 @@ def check(harness, scenario_ids):
117112
118113def main ():
119114 scenario_ids = expected_scenario_ids ()
120- gated = ", " .join (f"{ k } >= { v :.0f} " for k , v in THRESHOLDS .items ())
121115 print (f"Scenarios: { len (scenario_ids )} | Harnesses: { len (HARNESSES )} " )
122- print (f"Thresholded: { gated } " )
123116 print (f"Must be > 0: { ', ' .join (sorted (POSITIVE ))} " )
124117 print ("All other scorers are liveness-checked (ran, no error, "
125118 "numeric score)\n " )
0 commit comments