Skip to content

fix(providers): surface finish_reason length truncation in chat response metadata - #8096

Open
wxhking wants to merge 1 commit into
agentscope-ai:mainfrom
wxhking:fix/surface-length-finish-reason
Open

wxhking wants to merge 1 commit into
agentscope-ai:mainfrom
wxhking:fix/surface-length-finish-reason

Conversation

@wxhking

@wxhking wxhking commented Oct 3, 2026

Copy link
Copy Markdown

Description

When a generation is cut off by the output cap, the provider reports
finish_reason="length" on the terminal chunk, but the base stream parser
never reads it — the stop reason is silently dropped, and a truncated answer
is indistinguishable from a complete one (issue #8085).

This change surfaces the provider finish_reason as truncation metadata on
the final ChatResponse for both paths of the OpenAI-compatible provider
layer:

  • _SanitizedStream records the last non-empty finish_reason seen on the
    raw chunks. The terminal choice chunk usually produces no yielded delta,
    so the value is kept on the stream wrapper and read once the final
    accumulated response arrives — the same delta-only sidecar pattern the
    relay already uses for tool-call extras.
  • _relay_stream_tool_call_extras writes metadata["finish_reason"] and
    emits one warning line when the captured value is length.
  • _parse_completion_response does the same for the non-streaming path.

Behavior is unchanged for stop, tool_calls, and every other stop reason.

One boundary, consistent with the staged proposal in #8085: the pinned
AgentScope release drops ChatResponse.metadata when converting model
output into agent events (the reason FallbackChatModel publishes routing
transparency through an out-of-band sink). In this first stage the metadata
is therefore observable by direct consumers of the provider layer, and the
warning log line is the signal visible everywhere. Propagating the flag
through the agent loop (e.g. reusing the fallback-notice sink pattern) and
rendering a channel notice are left as the follow-up stage for maintainers
to decide.

Related Issue: Fixes #8085

Security Considerations: None — the change only adds a diagnostic
metadata entry and one warning log line; no auth, config, or environment
handling is touched.

Type of Change

  • Bug fix

Component(s) Affected

  • Core / Backend (app, agents, config, providers, utils, local_models)
  • Tests

Checklist

  • I ran pre-commit run --all-files locally and it passes
    (all hooks pass on the two touched files; the --all-files run reports
    3 pre-existing mypy errors in src/qwenpaw/cli/shutdown_cmd.py and
    pre-existing pylint W0613/E0611 findings in tests/unit/loop/,
    tests/unit/plugins/computer_use/, tests/unit/sandbox/, and
    tests/unit/services/ — none touch this change)
  • If pre-commit auto-fixed files, I committed those changes and reran checks
  • I ran tests locally (pytest or as relevant) and they pass
  • Documentation updated (if needed)
  • Ready for review

Testing

uv run python -m pytest tests/unit/providers -q

New tests in tests/unit/providers/test_openai_stream_toolcall_compat.py:

  • streaming finish_reason="length" → final response metadata
    {"finish_reason": "length"} + warning log line
  • streaming finish_reason="stop" / "tool_calls" → metadata untouched
    (controls)
  • non-streaming finish_reason="length" / "stop" via the public model
    call path

Evidence

$ uv run python -m pytest tests/unit/providers -q
893 passed, 1 skipped, 2 warnings in 47.40s

$ uvx pre-commit run --files src/qwenpaw/providers/openai_chat_model_compat.py tests/unit/providers/test_openai_stream_toolcall_compat.py
check python ast / docstring / encoding pragma / private key / trailing whitespace / trailing commas: Passed
mypy / black / flake8 / pylint: Passed

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

Welcome to QwenPaw! 🐾

Hi @wxhking, this is your 5th Pull Request.

🙌 Join Developer Community

Thanks so much for your contribution! We'd love to invite you to join the official QwenPaw developer group! You can find the Discord and DingTalk group links under the "Developer Community" section on our docs page:
https://qwenpaw.agentscope.io/docs/community

We truly appreciate your enthusiasm—and look forward to your future contributions! 😊

We'll review your PR soon.


Tip

⭐ If you find QwenPaw useful, please give us a Star!

Star QwenPaw

Staying ahead

Star QwenPaw on GitHub and be instantly notified of new releases.

Your star helps more developers discover this project! 🐾

@veveyluo

veveyluo commented Oct 4, 2026

Copy link
Copy Markdown

+1 for surfacing truncation info in response metadata.

This is the same observability gap I described in #8103 (daemon silent fallback). Users have no way to know:

  1. When a response was truncated (finish_reason: length)
  2. When the daemon silently fell back to a different model

Both cases leave users confused about what actually happened. Adding metadata like:

{
  "metadata": {
    "finish_reason": "length",
    "truncated": true,
    "model_used": "actual-model-id",
    "model_requested": "requested-model-id"
  }
}

...would make debugging much easier and build user trust.

Related: #8103 (fallback notification), #7738 (kwargs filtering — different issue but same "silent failure" pattern).

This branch is waiting to be deployed

1 waiting deployment
ai-review-approved — 7caafb93 Waiting Oct 3, 2026 by wxhking via AI Review Approval #4518
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/S 50-199 changed lines (additions + deletions)

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

truncation: surface finish_reason="length" when the output is cut off (currently the stop reason is dropped silently)

3 participants