Skip to content

fix(console): recover boot from failed entry loads with watchdog error surface - #8102

Open
wxhking wants to merge 1 commit into
agentscope-ai:mainfrom
wxhking:fix/console-boot-watchdog
Open

wxhking wants to merge 1 commit into
agentscope-ai:mainfrom
wxhking:fix/console-boot-watchdog

Conversation

@wxhking

@wxhking wxhking commented Oct 4, 2026

Copy link
Copy Markdown

Description

Boot watchdog for the console: when the entry chunk fails to load (stale cache after an upgrade 404s old hashed assets, network stall, CDN hiccup), the static boot splash now surfaces an error state with a Reload button instead of hanging forever, with one automatic reload attempt to self-heal the stale-cache case.

Related Issue: Fixes #8094

Root cause (main today): the .qwenpaw-boot splash in console/index.html is static — nothing ever replaces it if the module entry fails. main.tsx's INITIAL_RENDER_TIMEOUT_MS only guards i18n readiness; there is no window.onerror / unhandledrejection handling for the boot phase. src/utils/lazyWithRetry.ts already retries page-level lazy chunks — the gap is the entry module itself plus any error surface.

What this PR adds (issue suggestions 1+2 only):

  • console/public/bootWatchdog.js (new, no-build vanilla IIFE): capture-phase error listener (SCRIPT/LINK resource failures only — IMG failures such as the boot logo are ignored), plus error / unhandledrejection listeners and a 15s boot timeout (covers stalls that never emit an error event). Any pre-boot failure shares one budgeted path: update the splash, auto-location.reload() once (sessionStorage qwenpaw:boot-retries), then show a persistent error surface with the failing URL and a manual Reload Console button. Exposes window.__qwenpawBootWatchdog.disarm(); idempotent.
  • console/src/main.tsx: one line after render() calls disarm() — sets the booted flag, clears the timer/listeners, and refunds the retry budget so the next cold start keeps its one auto-retry.
  • console/index.html: loads the watchdog as a classic script before the module entry, plus light/dark CSS for the error surface.

Why public/ instead of an inline script: the packaged Tauri app enforces script-src: 'self' (src-tauri/tauri.conf.json), which blocks inline scripts. A public/ asset is same-origin, unhashed (so it loads fresh alongside a fresh index.html even in the stale-cache scenario), and Vite copies it verbatim.

Security Considerations: No security impact. The watchdog only reads its own DOM nodes and sessionStorage, renders static strings (textContent — the failing resource URL is not HTML-interpreted), and reloads the current page. No new network calls, no data leaves the device.

Out of scope (follow-ups, per the issue's items 3–5): backend health polling with periodic auto-reload, Cache-Control response headers for hashed assets, and a tray-menu Reload entry.

Evidence

All gates run locally on Windows (console/):

  • npm run test:run src/utils/bootWatchdog.test.ts — 14 passed (new tests execute the real shipped public/bootWatchdog.js in jsdom via new Function(...) injection; covers: script-tag-before-entry assertion, SCRIPT/LINK vs IMG filtering, error/rejection surfacing, post-boot silence, boot flag on disarm, auto-reload-once + budget exhaustion on second timeout, manual button reload after budget spent, retry-budget refund on successful boot, idempotency, missing-splash resilience).
  • src/utils regression: 33 files / 257 passed.
  • npx tsc -b --noEmit clean; npx prettier --check clean; npx eslint clean on all touched files.
  • npm run build passes (verify:initial-bundle OK); dist/index.html contains <script src="/bootWatchdog.js"></script> before the deferred module entry.

Real-browser E2E (built dist/, module entry src pointed at a 404 asset to reproduce the issue's stale-cache scenario):

  1. Boot splash appears → watchdog auto-reloads once → second failure shows the error surface: title "Console failed to load", the failing asset URL, and a working Reload Console button (verified getComputedStyle: display:flex; flex-direction:column, dark-mode palette applied).
  2. Unmodified dist/ boot: window.__qwenpawBooted === true, no error surface, app renders normally — success path unchanged.

Relationship to existing PRs: #5569 (stalled since June) targets a native-layer instant splash to hide startup latency; this PR is the web-layer safety net for failed entry loads — orthogonal file intents, different problem. lazyWithRetry.ts is untouched (it keeps handling page-level chunks).

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Documentation
  • Refactoring

Component(s) Affected

  • Core / Backend (app, agents, config, providers, utils, local_models)
  • Console (frontend web UI)
  • Channels (DingTalk, Lark, QQ, Discord, iMessage, etc.)
  • Skills
  • CLI
  • Documentation (website)
  • Tests
  • CI/CD
  • Scripts / Deploy

Checklist

  • I ran pre-commit run --all-files locally and it passes — left unticked honestly: this PR touches only console/ files; the repo's prettier pre-commit hook excludes ^console/, and the Python hooks do not apply to JS/HTML files. The console-side gates were run instead (see Testing).
  • If pre-commit auto-fixed files, I committed those changes and reran checks — N/A, no pre-commit auto-fixes
  • I ran tests locally (pytest or as relevant) and they pass
  • Documentation updated (if needed)
  • Ready for review

Testing

All commands run locally on Windows in console/:

npm run test:run src/utils/bootWatchdog.test.ts
# Test Files  1 passed (1)
#       Tests  14 passed (14)

npm run test:run src/utils
# Test Files  33 passed (33)
#       Tests  257 passed (257)

npx tsc -b --noEmit      # clean
npx prettier --check .   # clean (console-scoped; repo prettier gate)
npm run lint             # eslint clean on all touched files
npm run build            # full build passes; verify:initial-bundle OK
# dist/index.html contains <script src="/bootWatchdog.js"></script> before the deferred module entry

New tests execute the real shipped public/bootWatchdog.js in jsdom (injected via new Function(...), mirroring the source-reading pattern of src/styles/uiFontSizeCoverage.test.ts): SCRIPT/LINK vs IMG error filtering, error/rejection surfacing, post-boot silence, boot flag on disarm, auto-reload-once + budget exhaustion on the second timeout, manual button reload after the budget is spent, retry-budget refund on successful boot, idempotency, and missing-splash resilience.

Evidence

Real-browser E2E against the built dist/ (module entry src pointed at a 404 asset — reproduces the issue's stale-cache scenario):

  1. Boot splash appears → watchdog auto-reloads once → second failure renders the error surface: title "Console failed to load", the failing asset URL, and a working Reload Console button. Verified in-page via getComputedStyle(document.getElementById('qwenpaw-boot-error')): display:flex; flex-direction:column, dark-mode palette applied.
  2. Unmodified dist/ boot (happy path): window.__qwenpawBooted === true, #qwenpaw-boot-error absent, app renders normally — success path unchanged; the watchdog stands down.
# happy-path probe in the built app:
{"booted":true,"surface":false,"appRendered":true}
# broken-entry probe:
{"className":"qwenpaw-boot__error","display":"flex","flexDirection":"column","present":true}

Additional Notes

  • Relationship to feat(desktop): eliminate startup white screen with instant splash loading #5569 (stalled since June): that PR targets a native-layer instant splash to hide startup latency; this PR is the web-layer safety net for failed entry loads. Orthogonal intent, different problem — file overlap is index.html/main.tsx only in passing.
  • src/utils/lazyWithRetry.ts is untouched: it keeps handling page-level lazy chunks; the gap covered here is the entry module itself plus the missing error surface.
  • Why public/ instead of an inline script: the packaged Tauri app enforces script-src: 'self' (src-tauri/tauri.conf.json), which blocks inline scripts. A public/ asset is same-origin, unhashed (loads fresh alongside a fresh index.html even in the stale-cache scenario), and Vite copies it verbatim.
  • Follow-ups (issue suggestions 3–5, intentionally out of scope): backend health polling with periodic auto-reload, Cache-Control response headers for hashed assets, tray-menu Reload entry.

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

Welcome to QwenPaw! 🐾

Hi @wxhking, this is your 6th Pull Request.

🙌 Join Developer Community

Thanks so much for your contribution! We'd love to invite you to join the official QwenPaw developer group! You can find the Discord and DingTalk group links under the "Developer Community" section on our docs page:
https://qwenpaw.agentscope.io/docs/community

We truly appreciate your enthusiasm—and look forward to your future contributions! 😊

We'll review your PR soon.


Tip

⭐ If you find QwenPaw useful, please give us a Star!

Star QwenPaw

Staying ahead

Star QwenPaw on GitHub and be instantly notified of new releases.

Your star helps more developers discover this project! 🐾

@github-actions github-actions Bot added the size/M 200-499 changed lines (additions + deletions) label Oct 4, 2026
@LUOSENGWA

Copy link
Copy Markdown
Contributor

Independent reproduction — and our case exercises the 15 s timeout branch, not the stale-chunk branch.

We reproduced a related permanent-splash failure on 2.2.2-beta.4 (headless Linux Docker): the backend event loop stalled (a plugin synchronous-call freeze — that side is #7840/#7842), so the boot's API-dependent steps never completed. The console stayed pinned on "LOADING CONSOLE" with nothing actionable in F12. The entry chunk itself loaded fine; the backend simply stopped answering. So this is the "stall that never emits an error event" path the 15 s boot timeout targets.

With this PR, that scenario would end with a visible error state and a working Reload Console button instead of a dead splash — in our case the single auto-reload at 15 s can still hit the ongoing stall (our blocking calls run up to 20 s), but the manual button is the recovery path once the backend unsticks.

Boundary note with our own PR #8108: #8108 covers page-level lazy chunks (routing them through lazyWithRetry), this PR covers the entry module plus the error surface. Complementary, no file overlap.


2.2.2-beta.4 无头 Docker 独立复现相关故障:后端事件循环被插件同步调用冻结(#7840/#7842 侧),启动依赖 API 的步骤永远完不成,控制台停在 LOADING CONSOLE、F12 无信息。入口 chunk 本身加载正常、只是后端不响应——走的正是 15s 超时针对的"不发出任何错误事件的卡死"分支。有本 PR,该场景会落到可见错误面 + 可用 Reload 按钮(15s 自动重载一次时我们的阻塞可能还没结束,但手动按钮在后端解冻后就是恢复路径)。与我们的 #8108 边界:#8108 管页面级 lazy chunk,本 PR 管入口模块 + 错误面,互补、零文件重叠。

This branch is waiting to be deployed

1 waiting deployment
ai-review-approved — a5b54274 Waiting Oct 4, 2026 by wxhking via AI Review Approval #4524
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/M 200-499 changed lines (additions + deletions)

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

[Bug] Console boot splash (LOADING CONSOLE) has no retry and no error surface; stale WebView2 cache after update can permanently block boot

2 participants