You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
- add nudge breakdown line test asserting rendered reasoning category
- fix 2 pre-existing vacuous acp_status tests (partial mocks dropped by
filterMessages → passed with zero visible messages)
- tighten protected-reasoning assertion to exact value
- system prompt breakdown example percentages now sum to 100%
1086/1086 tests passing
Copy file name to clipboardExpand all lines: devlog/2026-09-08_reasoning-in-context-estimate/WORKLOG.md
+19-3Lines changed: 19 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -26,16 +26,32 @@
26
26
27
27
### Tests
28
28
-`tests/inject-utils-pure.test.ts`: +4 — reasoning counted in `reasoningTokens`/`total`; total formula includes reasoning; mixed message (msgTotal vs messageTokens + largestRanges footprint); no-reasoning regression guard.
29
-
-`tests/protection-aware-stats.test.ts`: +1 — reasoning on a protected message counted in `protectedTokens`.
29
+
-`tests/protection-aware-stats.test.ts`: +1 — reasoning on a protected message counted in `protectedTokens` (exact `reasoningTokens === 200`).
-**Code reviewer (nit)**: system prompt example percentages now sum to 100% (were 121% pre-existing, 131% after first edit).
42
+
43
+
Not addressed (documented, non-blocking):
44
+
- Composition-vs-range divergence (code reviewer minor): `buildCompressibleRanges` range tokens still exclude reasoning (intentional — ranges = compressible amounts; the pipeline's min-size check `countMessageCharacters` also excludes reasoning, so adding it there risks phantom "Range too small" rejections per #37). Consequence: "Effective compressible: ~X" (nudge) and overview totals now include reasoning while per-range lines don't. Documented in PR description; candidate follow-up issue (source-tagged).
45
+
- Reasoning-only messages render with `text` label in the drilldown (`toolName || "text"`); `classifyMessageType` would say `reasoning` — label-semantics change, out of scope.
46
+
- Per-message `dcp-message-id` token annotation (`countMessageCharacters`, token-utils.ts) still excludes reasoning — pre-existing, out of scope, candidate follow-up.
31
47
32
48
## Verification
33
49
34
50
-`npm run typecheck` — clean.
35
-
-`npm run test` — **1085/1085 pass** (was 1077 on master; +8 new).
51
+
-`npm run test` — **1086/1086 pass** (was 1077 on master; +9 new).
36
52
-`npm run build` — clean.
37
53
-`npm run format:check` — repo-wide pre-existing Prettier drift (423 files fail on clean master, incl. all 7 touched files); CI does not run format checks; no reformat to keep the diff minimal.
38
-
- Test-input fidelity note: `filterMessages` (`lib/messages/shape.ts:14-24`) drops messages lacking `info.sessionID`/`info.time.created` — the new acp_status tests use complete mocks. (Pre-existing tests at `tests/acp-status.test.ts:286/:307` use partial mocks and only assert section headers, so they pass with zero visible messages — flagged for the review, not fixed here.)
54
+
- Test-input fidelity note: `filterMessages` (`lib/messages/shape.ts:14-24`) drops messages lacking `info.sessionID`/`info.time.created` — all acp_status tests (new + 2 pre-existing fixed per review) use complete mocks.
0 commit comments