Skip to content

Commit 1f56b46

Browse files
committed
Add support for per-prompt execution timeout (prompt_timeout_seconds)
Supports setting prompt_timeout_seconds across Agent CLI generators (gemini_cli, claude_code, codex_cli, agy_cli) to terminate prompt turns that spin indefinitely. Configurable per scenario JSON, per model YAML, or per run_config YAML. - Extracts process_timeout context manager and _kill_process_group helper to encapsulate process group termination (start_new_session=True + os.killpg) across all CLI generators. - Fixes Codex CLI generator to support timeout in CLICommand, create_command, and _run_codex_cli. - Cleans up unused imports across CLI generators. - Preserves partial stdout and stream output captured before timeout to retain execution trajectories. TAG=agy CONV=4229555b-deeb-44bb-a481-c751dab59476
1 parent 1f23314 commit 1f56b46

19 files changed

Lines changed: 399 additions & 59 deletions

‎AGENTS.md‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -166,6 +166,7 @@ simulated_user_model_config: datasets/model_configs/gemini_2.5_pro_model.yaml
166166

167167
runners:
168168
agent_runners: 10
169+
prompt_timeout_seconds: 300
169170

170171
scorers:
171172
trajectory_matcher: {}
@@ -183,11 +184,12 @@ reporting:
183184
```
184185
185186
### 2. Model Configuration (e.g., `gemini_cli_model.yaml`)
186-
Specifies tested version, model ID, and environment variables.
187+
Specifies tested version, model ID, environment variables, and prompt timeout.
187188

188189
```yaml
189190
gemini_cli_version: "@google/gemini-cli@0.36.0"
190191
generator: gemini_cli
192+
prompt_timeout_seconds: 180
191193
env:
192194
GOOGLE_CLOUD_PROJECT: "my-evaluation-project"
193195
GOOGLE_CLOUD_LOCATION: "us-central1"
@@ -211,7 +213,8 @@ Contains the test cases.
211213
"conversation_plan": "Ensure the agent accurately calls list_instances. Verify the output is returned correctly.",
212214
"expected_trajectory": ["cloud-sql__list_instances"],
213215
"env": { "GOOGLE_CLOUD_PROJECT": "my-evaluation-project" },
214-
"max_turns": 4
216+
"max_turns": 4,
217+
"prompt_timeout_seconds": 120
215218
}
216219
]
217220
}

‎docs/agy_cli_agent_testing.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -167,6 +167,7 @@ Specifies the generator, model label, execution timeouts, and environment:
167167
| `generator` | Yes | Must be `agy_cli` |
168168
| `model` | Optional | Model label (e.g. `"Gemini 3.1 Pro (Low)"` or `"Gemini 3.5 Flash (Medium)"`). Omit to use agy's default. |
169169
| `timeout` | Optional | CLI turn timeout string (e.g. `"20m"`, passed to `--print-timeout`). Defaults to 5m. |
170+
| `prompt_timeout_seconds` | Optional | Timeout limit in seconds for each CLI prompt turn (e.g., `180`). |
170171
| `env` | Optional | Environment block. Set `GOOGLE_CLOUD_PROJECT` (see below); `GOOGLE_CLOUD_LOCATION` defaults to `global`. |
171172
| `setup` | Optional | Tool setup block for `mcp_servers`, `skills`, or `fake_mcp_servers`. |
172173

‎docs/claude_code_agent_testing.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -206,6 +206,7 @@ The model config defines the Claude Code CLI version, model, auth, environment,
206206
| `vertex_project_id` | If `use_vertex` | GCP project for Vertex AI |
207207
| `vertex_region` | If `use_vertex` | Vertex region (e.g., `us-east5`) |
208208
| `env` | Optional | Environment variables passed to the CLI process |
209+
| `prompt_timeout_seconds` | Optional | Timeout limit in seconds for each CLI prompt turn (e.g., `180`) |
209210
| `setup.mcp_servers` | Optional | MCP server configurations (see [MCP Servers](#mcp-servers)) |
210211
| `allowed_tools` | Optional | List of tool names to allow (e.g., `["Bash", "mcp__cloud-sql"]`) |
211212

‎docs/codex_cli_agent_testing.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -213,6 +213,7 @@ The model config defines the Codex CLI version, model, auth, sandbox/approval po
213213
| `profile` | Optional | Codex profile name (forwarded as `--profile <name>`) |
214214
| `json_flag` | Optional | `"--json"` (default, newer Codex versions) or `"--experimental-json"` (older versions). Codex requires NDJSON for the eval pipeline to extract tool calls and tokens. |
215215
| `pricing` | Optional | Per-model rates used to compute `cost_usd` per turn. See [Pricing & Cost Tracking](#pricing--cost-tracking). |
216+
| `prompt_timeout_seconds` | Optional | Timeout limit in seconds for each CLI prompt turn (e.g., `180`) |
216217
| `env` | Optional | Environment variables passed to the CLI process (e.g., `GOOGLE_CLOUD_PROJECT` for Cloud SQL MCP) |
217218
| `setup.mcp_servers` | Optional | MCP server configurations (see [MCP Servers](#mcp-servers)) |
218219
| `setup.config` | Optional | Free-form key/value pairs written to the top of `~/.codex/config.toml`. Merged on top of the default `forced_login_method = "api"`. |

‎docs/configs/model-config.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -15,6 +15,7 @@ These settings are passed to all generators, regardless of the specific engine u
1515
| `max_tokens` | Optional | N/A | Specifies the maximum number of tokens the model can generate in a single output. |
1616
| `execs_per_minute` | Optional | `60` | Sets the maximum number of executions allowed per minute. If not provided, it defaults to `60`. This helps throttle the rate of query generation. |
1717
| `max_attempts` | Optional | `3` | Specifies the maximum number of attempts for query generation in case of failures. Defaults to `3` if not provided. |
18+
| `prompt_timeout_seconds` | Optional | N/A | Timeout limit in seconds for each CLI prompt execution turn (e.g. `180`). Overrides run config `runners.prompt_timeout_seconds`. |
1819

1920
## GCP Specific Configuration
2021

‎docs/configs/run-config.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -19,6 +19,7 @@ This section defines the primary resources used during evaluation, including the
1919
| `num_trials` | Optional | Number of trials to run for each prompt. |
2020
| `scenarios` | Optional | A list of specific scenario IDs to run (only applies to scenario-based agentic datasets like `gemini-cli-format` or `cortado-format`). Defaults to empty (runs all scenarios). |
2121
| `scenario_pattern` | Optional | A glob pattern of scenario IDs to run (only applies to scenario-based agentic datasets). Defaults to None (runs all scenarios). |
22+
| `runners` | Optional | Dictionary configuring concurrency (`agent_runners`, default: 10) and prompt timeouts (`prompt_timeout_seconds`, e.g. `300` seconds per CLI execution turn). |
2223
---
2324

2425
## 2. Prompt and Generation Modules

‎docs/gemini_cli_agent_testing.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -255,6 +255,7 @@ The evalset JSON file defines the test scenarios. Each scenario represents an ag
255255
| `max_turns` | Yes | Maximum number of conversation turns before the evaluation stops |
256256
| `env` | Optional | Per-scenario environment variables (merged with model config env) |
257257
| `kind` | Optional | Category label (e.g., `"tools"`) |
258+
| `prompt_timeout_seconds` | Optional | Timeout limit in seconds for each CLI prompt turn in this scenario. Overrides model config and run config timeout settings. |
258259

259260
#### Tool name format
260261

‎evalbench/evaluator/agentevaluator.py‎

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -141,6 +141,12 @@ def process_scenario(
141141
"Declared env file not found in session: %s", src_path
142142
)
143143

144+
prompt_timeout = (
145+
scenario.get("prompt_timeout_seconds")
146+
or getattr(self.generator, "prompt_timeout_seconds", None)
147+
or self.config.get("runners", {}).get("prompt_timeout_seconds")
148+
)
149+
144150
session_id = None
145151
for turn in range(max_turns):
146152
logging.info(
@@ -153,6 +159,7 @@ def process_scenario(
153159
resume=(turn > 0),
154160
session_id=session_id,
155161
cwd=resolved_work_dir,
162+
timeout=prompt_timeout,
156163
)
157164
try:
158165
result = self.generator.safe_generate(cli_cmd)

‎evalbench/generators/models/agent_cli.py‎

Lines changed: 75 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -1,13 +1,17 @@
11
from abc import abstractmethod
2+
from contextlib import contextmanager
3+
import logging
4+
import os
5+
import shutil
6+
import signal
7+
import subprocess
8+
import threading
29

310
from mcp import types as mcp_types
411

512
from . import mcp_client
613
from .generator import QueryGenerator
714
from .tool_naming import canonical_tool_name
8-
import logging
9-
import os
10-
import shutil
1115

1216

1317
class AgentCliGenerator(QueryGenerator):
@@ -91,7 +95,7 @@ def version(self) -> str:
9195
@abstractmethod
9296
def create_command(
9397
self, cli: str, prompt: str, env: dict = None, resume: bool = False,
94-
session_id: str = None, cwd: str = None,
98+
session_id: str = None, cwd: str = None, timeout: float | int = None,
9599
):
96100
raise NotImplementedError("Subclasses must implement this method")
97101

@@ -110,3 +114,70 @@ def extract_tools(self, stdout: str) -> list:
110114
@abstractmethod
111115
def extract_skills(self, stdout: str) -> list:
112116
raise NotImplementedError("Subclasses must implement this method")
117+
118+
119+
def parse_timeout_seconds(timeout: float | int | str | None) -> float | None:
120+
if timeout is None:
121+
return None
122+
if isinstance(timeout, (int, float)):
123+
return float(timeout)
124+
if isinstance(timeout, str):
125+
s = timeout.strip()
126+
unit = 1.0
127+
if s.lower().endswith("s"):
128+
s = s[:-1]
129+
elif s.lower().endswith("m"):
130+
s = s[:-1]
131+
unit = 60.0
132+
elif s.lower().endswith("h"):
133+
s = s[:-1]
134+
unit = 3600.0
135+
136+
try:
137+
return float(s) * unit
138+
except ValueError as e:
139+
logging.warning("Failed to parse timeout string %r: %s", timeout, e)
140+
return None
141+
142+
143+
def _kill_process_group(proc: subprocess.Popen):
144+
"""Terminates proc's process group with SIGKILL, falling back to proc.kill()."""
145+
try:
146+
os.killpg(os.getpgid(proc.pid), signal.SIGKILL)
147+
return
148+
except (ProcessLookupError, OSError) as e:
149+
logging.warning("os.killpg failed for pid %s: %s; trying proc.kill()", proc.pid, e)
150+
except Exception as e:
151+
logging.warning("Unexpected error in os.killpg for pid %s: %s; trying proc.kill()", proc.pid, e)
152+
153+
try:
154+
proc.kill()
155+
except (ProcessLookupError, OSError) as e:
156+
logging.warning("proc.kill() failed for pid %s: %s", proc.pid, e)
157+
except Exception as e:
158+
logging.warning("Unexpected error in proc.kill() for pid %s: %s", proc.pid, e)
159+
160+
161+
@contextmanager
162+
def process_timeout(proc: subprocess.Popen, timeout: float | int | str | None):
163+
"""Context manager that sets a timer to kill ``proc``'s process group on timeout.
164+
165+
Yields a callable ``is_timed_out() -> bool``.
166+
"""
167+
timed_out = False
168+
timer = None
169+
timeout_sec = parse_timeout_seconds(timeout)
170+
if timeout_sec:
171+
def _on_timeout():
172+
nonlocal timed_out
173+
timed_out = True
174+
_kill_process_group(proc)
175+
176+
timer = threading.Timer(timeout_sec, _on_timeout)
177+
timer.start()
178+
179+
try:
180+
yield lambda: timed_out
181+
finally:
182+
if timer:
183+
timer.cancel()

‎evalbench/generators/models/agy_cli.py‎

Lines changed: 21 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
1-
from .agent_cli import AgentCliGenerator
1+
from .agent_cli import AgentCliGenerator, process_timeout
22
from .tool_naming import canonicalize_agy_tool_name, parse_agy_mcp_tool_call
33
import subprocess
44
import os
@@ -35,12 +35,13 @@ def _shred_credential(path: str) -> None:
3535

3636

3737
class CLICommand:
38-
def __init__(self, cli, prompt, env=None, resume=False, cwd=None):
38+
def __init__(self, cli, prompt, env=None, resume=False, cwd=None, timeout=None):
3939
self.cli = cli
4040
self.prompt = prompt
4141
self.env = env if env else {}
4242
self.resume = resume
4343
self.cwd = cwd
44+
self.timeout = timeout
4445

4546

4647
class AgyCliGenerator(AgentCliGenerator):
@@ -872,17 +873,27 @@ def generate_internal(self, cli_cmd):
872873
return self._run_agy_cli(cli_cmd)
873874

874875
def _execute_cli_command(
875-
self, command, env=None, cwd=None
876+
self, command, env=None, cwd=None, timeout=None
876877
) -> subprocess.CompletedProcess:
877878
try:
878-
return subprocess.run(
879+
proc = subprocess.Popen(
879880
command,
880-
stdin=subprocess.DEVNULL, capture_output=True,
881+
stdin=subprocess.DEVNULL,
882+
stdout=subprocess.PIPE,
883+
stderr=subprocess.PIPE,
881884
text=True,
882-
check=False,
883885
env=env,
884886
cwd=cwd if cwd else self.fake_home,
887+
start_new_session=True,
885888
)
889+
with process_timeout(proc, timeout) as is_timed_out:
890+
stdout, stderr = proc.communicate()
891+
892+
if is_timed_out():
893+
stderr = (stderr + "\n" if stderr else "") + f"Error: Command timed out after {timeout} seconds."
894+
return subprocess.CompletedProcess(command, 124, stdout or "", stderr)
895+
896+
return subprocess.CompletedProcess(command, proc.returncode, stdout or "", stderr)
886897
except FileNotFoundError:
887898
return subprocess.CompletedProcess(
888899
command, 127, "", f"Error: Command not found: {command[0]}"
@@ -901,10 +912,10 @@ def _run_agy_cli(self, cli_cmd: CLICommand):
901912
command = self._base_agy_command(
902913
self.agy_bin, cli_cmd.prompt, cli_cmd.resume, self.model,
903914
output_format="stream-json", log_file=self.cli_log_path,
904-
timeout=self.timeout,
915+
timeout=cli_cmd.timeout or self.timeout,
905916
)
906917
cwd = cli_cmd.cwd if cli_cmd.cwd else self.fake_home
907-
result = self._execute_cli_command(command, env=env, cwd=cwd)
918+
result = self._execute_cli_command(command, env=env, cwd=cwd, timeout=cli_cmd.timeout or self.timeout)
908919

909920
# Parse whenever agy emitted a stream, even on a non-zero exit: a
910921
# timed-out/errored run still ends in a ``result`` event carrying real
@@ -1212,7 +1223,7 @@ def safe_generate(
12121223

12131224
def create_command(
12141225
self, cli: str, prompt: str, env: dict = None, resume: bool = False,
1215-
session_id: str = None, cwd: str = None,
1226+
session_id: str = None, cwd: str = None, timeout: float | int = None,
12161227
) -> CLICommand:
12171228
# The executable is always this session's sandbox binary
12181229
# (self.agy_bin); the ``cli`` argument -- the agent_version label "agy"
@@ -1222,4 +1233,4 @@ def create_command(
12221233
# environment are layered in once at invocation time by
12231234
# ``_run_agy_cli`` via ``_merged_env``.
12241235
return CLICommand(cli=self.agy_bin, prompt=prompt, env=env or {},
1225-
resume=resume, cwd=cwd)
1236+
resume=resume, cwd=cwd, timeout=timeout)

0 commit comments

Comments
 (0)