Skip to content

fix(agents): refuse oversized prompts and surface empty model replies instead of silent failures - #8084

Open
LUOSENGWA wants to merge 1 commit into
agentscope-ai:mainfrom
LUOSENGWA:fix/window-guard-empty-reply
Open

LUOSENGWA wants to merge 1 commit into
agentscope-ai:mainfrom
LUOSENGWA:fix/window-guard-empty-reply

Conversation

@LUOSENGWA

Copy link
Copy Markdown
Contributor

What Problem This Solves

When a prompt exceeds the model's context window (observed: 196,602 tokens against a 196,608 window with the 5%/4096 output reserve), some providers respond 200 with completion_tokens=0: no overflow error, no visible output, no recovery — the room just looks like a hung agent. Two independent gaps made this invisible:

  1. the compression gate in scroll/manager.py relies on the local count_tokens estimate, which can undercount (observed ~0.60x of the provider-side count), so an over-window prompt that stays below the trigger ratio passes through silently — and the output reserve (5%/4096) was never applied as a hard pre-send budget;
  2. _reasoning has no fallback for a final message with zero meaningful blocks: an empty reply is yielded as a normal completion.

Fix

Two layers:

  • Pre-send window guard (scroll/manager.py compress()): the trigger gate now also checks a hard budget of window minus max(configured output reserve, declared request output cap). Over-budget input takes the existing forced-compaction path (the same is_forced_compaction mechanism the manual/forced path already uses), and a final recount after the pipeline refuses to send what still does not fit, raising the existing ContextWindowUnfitError — which already propagates to the channel as a visible message (the error string carries CONTEXT_UNFIT plus both token numbers). With no declared output cap the budget equals the existing effective hard limit, so behavior is byte-for-byte unchanged.
  • Empty-reply fallback (react_agent.py _reasoning()): a final message with zero meaningful blocks (text / tool_call / data / reasoning — reusing the provider-level _block_has_meaningful_content) is retried once through recover_from_context_overflow with a continuation message (the same pattern as the existing INTERRUPT_AND_CONTINUE stop action); a second empty reply surfaces a visible bilingual warning carrying the prompt_tokens / completion_tokens when available. Tool-call-only and thinking-only turns are untouched, and at most one retry happens per episode.

Evidence

Local verification (Python 3.11, PYTHONPATH=src):

  • tests/unit/agents/test_window_guard_and_empty_reply.py (new, 9 tests, fully mocked — no network, no real LLM): pre-check rejects oversized input without a provider call; under-budget passes through; the declared max_tokens tightens the budget; below-trigger + no cap passes through (pins the pre-existing behavior case B will address later); below-trigger + over hard budget forces compaction and then refuses; output-reserve config defaults unchanged; the unfit error message carries both numbers for the visible channel; the empty-reply fallback retries exactly once and then surfaces the warning; tool-call-only and thinking-only replies pass through.
  • Full tests/unit/agents/ regression: 3854 passed / 4 skipped / 0 failed — zero collateral breakage.
  • One pre-existing test alignment: test_reasoning_fallback_bridge.py fed its fake a zero-content final message while the fake model's real output was "ok"; the fake now mirrors the real output, since a zero-content final message is by design intercepted by the new fallback (the test's actual subject — fallback metadata surviving the model→event bridge — is unchanged).

Scope (v1)

The warning rides the normal outgoing-message path, so it works on every channel without per-channel code. The declared output cap is read from the request-side max_tokens / max_completion_tokens (the field providers populate on model.parameters), not from a new config surface.

Out of scope (follow-ups)

  • count_tokens calibration against provider-side counts (the ~0.60x undercount case) — a separate PR with its own measurement cases;
  • extending the hard budget to a user-configurable reserve;
  • gate file-ization / per-room notice dedup.

问题

prompt 超窗时(实测:196,602 tokens 对 196,608 窗 + 5%/4096 输出预留),部分 provider 返回 200 + completion_tokens=0:无错误、无输出、无恢复——房间里只剩一个看似卡死的 agent。两个独立缺口使其不可见:

  1. scroll/manager.py 的压缩门依赖本地 count_tokens 估算,而它会低估(实测约为 provider 侧计数的 0.60x)——低于 trigger 比例的超窗 prompt 静默放行;且输出预留(5%/4096)从未被用作发送前硬预算;
  2. _reasoning 对零有意义块(text/tool_call/data/reasoning 全无)的最终消息无兜底:空回复被当作正常完成 yield。

修复

两层:

  • 发送前窗校验(scroll/manager.py compress()):trigger 门同时检查硬预算 = 窗 − max(配置输出预留, 请求声明的输出上限)。超预算输入走既有强制压缩路径(is_forced_compaction 同机制),流水线走完后复计一次,仍超则抛既有 ContextWindowUnfitError——它本来就会以可见消息到达房间(错误串含 CONTEXT_UNFIT 与两个 token 数)。未声明输出上限时预算=既有 effective hard limit,行为逐字节不变(零回归)。
  • 空回复兜底(react_agent.py _reasoning()):零有意义块的最终消息(复用 provider 级 _block_has_meaningful_content)经 recover_from_context_overflow + continuation 消息重试一次(照既有 INTERRUPT_AND_CONTINUE stop action 模式);第二次空回复发出带 prompt_tokens/completion_tokens(可取时)的中英双语可见警告。仅 ToolCall / 仅 Thinking 回合不受影响;每个 episode 至多重试一次。

证据

本地验证(Python 3.11,PYTHONPATH=src):

  • tests/unit/agents/test_window_guard_and_empty_reply.py(新,9 例,全 mock 无网络无真实 LLM):预检拒发超窗输入且不打 provider;预算内放行;声明 max_tokens 收紧预算;trigger 下+无上限=放行(钉住案 B 后续处理的残差现状);trigger 下+超硬预算=强制压缩后拒发;输出预留配置默认值不变;unfit 错误串携带两个数字直达可见通道;空回复兜底恰好重试一次后出可见警告;仅 tool_call / 仅 thinking 放行。
  • tests/unit/agents/ 全量回归:3854 passed / 4 skipped / 0 failed——零连带破坏。
  • 一处既有测试对齐:test_reasoning_fallback_bridge.py 的 fake 给最终消息填了空 content,而 fake 模型真实输出是 "ok";fake 现与真实输出一致——因为零 content 最终消息按新设计会被兜底拦截(该测试真正的主题——fallback 元数据在 model→event 桥上的存活——不变)。

范围(v1)

警告走正常 outgoing-message 路径,所有通道零特例代码。声明输出上限从请求侧 max_tokens / max_completion_tokens(provider 填入 model.parameters 的字段)读取,不新增配置面。

不在本 PR 范围(后续)

  • count_tokens 对 provider 侧计数的校准(~0.60x 低估案)——独立 PR,自带测量案;
  • 硬预算/输出预留用户可配;
  • gate 文件化 / 同房间通知去重。

… instead of silent failures

When a prompt exceeds the model's context window (e.g. 196,602 tokens
against a 196,608 window with the 5%/4096 output reserve), some
providers respond 200 with completion_tokens=0: no overflow error, no
visible output, no recovery — the room just looks like a hung agent.
The compression gate relied on the local count_tokens estimate, which
can undercount (observed ~0.60x of the provider count), so an
over-window prompt below the trigger ratio passed through silently.

Two-layer fix:

- Pre-send window guard (scroll/manager.py compress()): the trigger
  gate now also checks a hard budget of window minus max(configured
  output reserve, declared request output cap); over-budget input
  takes the forced-compaction path, and a final recount after the
  pipeline refuses to send what still does not fit, raising the
  existing ContextWindowUnfitError (which already reaches the channel
  as a visible message). With no declared cap the budget equals the
  existing effective hard limit and behavior is byte-for-byte
  unchanged.
- Empty-reply fallback (react_agent.py _reasoning()): a final message
  with zero meaningful blocks (text/tool_call/data/reasoning) is
  retried once through context recovery with a continuation message;
  a second empty reply surfaces a visible bilingual warning carrying
  the usage numbers. Tool-call-only and thinking-only turns are
  untouched; at most one retry per episode.

9 new unit tests (mocked model, no network); the pre-existing
reasoning-fallback-bridge tests now feed their fake a final message
matching the model's real "ok" output, since a zero-content final
message is by design intercepted by the new fallback.
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Welcome to QwenPaw! 🐾

Hi @LUOSENGWA, this is your 7th Pull Request.

📋 About PR Template

To help maintainers review your PR faster, please make sure to include:

  • ✅ Description - What this PR does and why
  • ✅ Type of Change - Bug fix / Feature / Breaking change / Documentation / Refactoring
  • ✅ Component(s) Affected - Core / Console / Channels / Skills / CLI / Documentation / Tests / CI/CD / Scripts
  • ✅ Checklist:
    • Run and pass pre-commit run --all-files
    • Run and pass relevant tests (pytest or as applicable)
    • Update documentation if needed
  • ✅ Testing - How to test these changes
  • ✅ Local Verification Evidence:
    pre-commit run --all-files
    # paste summary result
    
    pytest
    # paste summary result

Complete PR information helps speed up the review process. You can edit the PR description to add these details.

🙌 Join Developer Community

Thanks so much for your contribution! We'd love to invite you to join the official QwenPaw developer group! You can find the Discord and DingTalk group links under the "Developer Community" section on our docs page:
https://qwenpaw.agentscope.io/docs/community

We truly appreciate your enthusiasm—and look forward to your future contributions! 😊

We'll review your PR soon.

This branch is waiting to be deployed

2 waiting deployments
maintainer-approved — f0f1a598 Waiting Oct 2, 2026 by LUOSENGWA via Maintainer Approval #10072
ai-review-approved — f0f1a598 Waiting Oct 2, 2026 by LUOSENGWA via AI Review Approval #4512
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/L 500-999 changed lines (additions + deletions)

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

1 participant