Skip to content

Reduce token activity heatmap update costs - #4328

Open
Yuxin-Qiao wants to merge 5 commits into
steipete:mainfrom
Yuxin-Qiao:perf/spend-activity-updates
Open

Yuxin-Qiao wants to merge 5 commits into
steipete:mainfrom
Yuxin-Qiao:perf/spend-activity-updates

Conversation

@Yuxin-Qiao

@Yuxin-Qiao Yuxin-Qiao commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Annual token-activity updates repeatedly created medium-date formatters and recalculated the same calendar dates and coverage. Cache normalized dates, visible indices, and coverage in each immutable series; reuse at most 16 immutable formatter contexts keyed by calendar, locale, and time zone; build weekly aggregates only in weekly/cumulative modes.

Maintainer cleanup shares the existing saturating-add operation, removes redundant navigation wrappers, and keeps direct fixture construction in the tests. Production diff versus reviewed main c3c6ce06120: +63/−70 lines. Contributor history is preserved. Thanks @Yuxin-Qiao!

Measured helper elapsed time

Apple M3 Ultra, Swift 6.4 debug/native backend, 365 synthetic dates, Gregorian Asia/Shanghai, zh_Hans_CN; median of 21 samples after three warmups. These measurements describe helpers, not whole-app frame rates.

Operation Before After Reduction
365 medium-date labels 41.120750 ms 0.471375 ms 98.85%
Dates, coverage and weekly aggregation 1.564250 ms 0.295834 ms 81.09%
Create series and read coverage 1.035000 ms 0.816917 ms 21.07%

Verification

  • make check: passed with zero violations on retry. The first attempt hit an unchanged subprocess-cleanup fixture's exit-observation race; no check was weakened.
  • Scrubbed swift test --build-system native --jobs 4 -Xswiftc -gnone --filter 'SpendActivityBenchmarkTests|SpendActivityPerformanceTests|SpendActivityHeatmapTests|SpendActivityAppearanceTests|SpendActivityReadabilityRenderTests|SpendDashboardLongRangeRenderTests': 41 Swift Testing tests passed. The optional long-range screenshot test was skipped without its separate proof flag.
  • The initial native navigation test missed its first click. An isolated SpendActivityReadabilityRenderTests retry passed all four tests. All 120 before/after captures have identical dimensions and decoded RGBA bytes across language, theme, width, mode, partial data, and recorded zero.
  • Isolated CODEXBAR_ACTIVITY_BENCHMARK=1 swift test --skip-build --build-system native --jobs 4 -Xswiftc -gnone --filter SpendActivityBenchmarkTests: passed and produced the table above. No wall-clock assertions were added.
  • Independent review: no actionable P0–P2 findings.

Final combined source proof at 8f564053f16d0b1a027e35828dca44c49df34d2c:

source Scripts/test_environment.sh
CODEXBAR_TEST_SUITE_TIMEOUT=600 ./Scripts/test.sh \
  --swift-command "$PWD/../reports/spend-prs-23-evidence/swift-native" --direct-workers 4

The wrapper forwards test/build to swift with --build-system native --jobs 4 -Xswiftc -gnone. 1,592 selections, 144/144 groups successful on the first attempt, zero retries and zero timeouts. The earlier 180-second invocation stopped on an unrelated publication-suite timeout; source and assertions were unchanged for the successful rerun. The pushed integration branch is triage/20260921-spend-prs-23; its later commit adds only the contributor's runtime-proof documents.

Synthetic before/after

Before:

Annual token activity before caching

After, pixel-identical:

Annual token activity after caching

These are production-component renders with synthetic fixtures. Reproduction and scope: docs/proof/spend-activity-performance.md.

Refs #4297.

The contributor subsequently added proof-only commit 1e2b5b25492c535612795e580bb05089f95257a9; Sources, Tests, and WidgetExtension are unchanged from tested maintainer head f7a9968e3c4ebf8e958ebb86627f4b2ff10287fe. Its additional full-settings runtime capture is in docs/proof/spend-activity-performance/. The read-only verifier passed: 11 checkpoints, 731 date/context checks, zero mismatches, 9,375 cache hits and two formatter creations. The diagnostic app was not independently launched here.

CI for proof-only head 1e2b5b254 was queued at that status check. The previous head's run was cancelled when this proof commit arrived.

Main synchronization

Commit 197422ae094239be93242dd7dda28364071ab878 merges main at f8b75cf2a and resolves the sole CHANGELOG.md conflict by retaining both sets of release notes. The activity heatmap and app entry source remain byte-identical to the pinned diagnostic build, and the committed runtime verifier still passes. New upstream provider and spend-trend changes are outside that recorded runtime capture.

After synchronization, make check passed (2,853 Swift files, zero lint violations); git diff --check passed. The full repository runner ./Scripts/test.sh --direct-workers 4, using the default 180-second group timeout and scrubbed test environment, exited successfully: 1,600 selections, 145/145 groups successful on the first attempt, zero failures, retries, or timeouts. Total elapsed time was 313.5 seconds. The preceding serial make test and one-worker direct invocations were deliberately interrupted to switch execution modes; those interrupted invocations are not counted as passed.

The automated review of synchronized head 197422ae0 accepts the logs and reports no actionable findings. GitHub lint, Linux x64/ARM64/musl, security checks, and the macOS compatibility build passed on that head; both macOS test shards remained queued before the next synchronization.

Subsequent main synchronization

Commit a13ab86eab173b22fdc0b1b374651475af7db55b merges main at 08eb56931 after further upstream changes again overlapped in CHANGELOG.md. Both release-note sets are retained. The activity heatmap and app entry source remain byte-identical to the pinned diagnostic build, and the runtime verifier passes. The original capture does not claim runtime validation of newly imported provider, menu, or account changes.

make check passed after this synchronization (2,870 Swift files, zero violations), and git diff --check passed. The complete repository runner ./Scripts/test.sh --direct-workers 4 also exited successfully on this head with the unchanged default 180-second group timeout: 1,607 selections, 145/145 groups successful on the first attempt, zero failures, retries, or timeouts. Total elapsed time was 945.6 seconds, including 513.7 seconds of build/discovery and 425.0 seconds of execution. The test environment was scrubbed and Keychain access suppressed.

Automatic review revision 7 covers a13ab86ea, accepts the runtime logs, and reports no actionable findings or contributor blockers. GitHub lint, Linux x64/ARM64/musl, and security checks have passed; both macOS test shards plus the macOS compatibility build are queued. The PR is mergeable and awaiting maintainer review, and has not merged.

@clawsweeper

clawsweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P3 Low-risk cleanup, docs, polish, ergonomics, or speculative feature. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Oct 7, 2026
@clawsweeper

clawsweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed October 8, 2026, 5:22 AM ET / 09:22 UTC (Revision 8).

ClawSweeper review

What this changes

Cache annual token-activity dates, coverage counts, and localized date formatters, and calculate weekly totals only when needed.

Review scores

Measure Result What it means
Overall readiness 🦞 diamond lobster (5/6) A focused, owner-supported optimization with full-settings runtime proof, compatibility coverage, and no actionable defects.
Proof confidence 🦞 diamond lobster (5/6) Sufficient (logs): The freshly built macOS full-settings bundle exercised the production heatmap through native UI actions using synthetic history. Logs show formatter and date-cache reuse, aggregation only in the selected modes, successful selection and refresh, and zero recorded date mismatches across time-zone changes. The changed source matches the pinned capture; helper timings support reduced work without claiming whole-app frame rates.
Patch quality 🦞 diamond lobster (5/6) No actionable review findings were identified.

Product

Kind: Performance · Worth it: Yes
User problem: Annual token activity updates and date inspection can take longer than necessary.
Reason: Comparative helper measurements show substantial savings with preserved output and net-negative production growth. The area owner’s cleanup commit confirms the direction.

Merge readiness

✅ Ready for maintainer review

This PR is useful, sufficiently proven, and has no actionable review findings or additional contributor blockers. Current main still performs the repeated work it addresses.

Likely related people: steipete and Yuxin-Qiao are high-confidence routing candidates from merged heatmap history.

Priority: P3
Reviewed head: a13ab86eab173b22fdc0b1b374651475af7db55b

Before merge

None.

Findings

None.

Agent review details

How this fits together

CodexBar’s Usage & Spend settings page turns retained token history into daily, weekly, and cumulative activity charts. This change reduces repeated chart calculations while preserving dates, totals, coverage, and selection.

flowchart LR
A[Retained token history] --> B[Annual activity snapshot]
C[Reporting calendar] --> B
B --> D{Selected chart mode}
D --> E[Daily chart]
D --> F[Weekly or cumulative totals]
G[Cached localized date labels] --> E
G --> F
Loading

Technical review

Best possible solution:

Retain the bounded immutable caches and lazy aggregation while preserving existing chart output and reporting-calendar semantics.

Do we have a high-confidence way to reproduce the issue?

Not applicable as a functional bug reproduction: source confirms repeated work on main, and comparative measurements plus runtime counters demonstrate the optimization.

Is this the best way to solve the issue?

Yes: immutable snapshot data, bounded formatter reuse, and mode-specific aggregation address the measured costs without adding settings or another data source.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 08eb56931c8e.

Provenance checked

  • Sources/CodexBar/SpendActivityHeatmap.swift: annual dates and coverage keeps the original intent (3efd823: Normalize calendar arithmetic to local day starts so midnight DST does not lose activity totals or coverage.)
  • Sources/CodexBar/SpendActivityHeatmap.swift: localized date formatting keeps the original intent (6538ac3: Use the configured reporting calendar and invalidate the heatmap when that calendar changes.)
  • Sources/CodexBar/SpendActivityHeatmap.swift: chart modes, totals, and navigation keeps the original intent (521af81: Provide daily, weekly, and cumulative activity over the shared scan cache while preserving coverage, localization, keyboard navigation, and accessibility.)

Testing

Proof path: shipped entry point. Added test files: 3.

Security

None.

Evidence

What I checked:

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • Yuxin-Qiao: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Review metrics

Metric Value Why it matters
Production versus test changes Production +63/-70, net -7 lines; tests +169 lines across 3 files The measured optimization reduces shipping code overall while adding compatibility coverage and an opt-in benchmark.

Labels

Label changes:

No label changes.

Label justifications:

  • P3: This is a bounded performance improvement with preserved output and no demonstrated urgent user-facing regression.
  • rating: 🦞 diamond lobster: Overall readiness is 🦞 diamond lobster; proof is 🦞 diamond lobster and patch quality is 🦞 diamond lobster.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR.
  • proof: sufficient: Contributor real behavior proof is sufficient.

Rating scale

6/6 🦀 challenger crab · 5/6 🦞 diamond lobster · 4/6 🐚 platinum hermit · 3/6 🦐 gold shrimp · 2/6 🦪 silver shellfish · 1/6 🧂 unranked krab. Overall follows the weaker of proof and patch quality; ✨ marks media proof (a screenshot, video, or linked artifact) that directly shows the changed behavior.

Workflow

ClawSweeper edits this one comment on every review. Comment @clawsweeper re-review for a fresh review only; repair and merge need explicit maintainer commands such as @clawsweeper autofix or @clawsweeper automerge.

History

Review history (7 earlier review cycles)
  • reviewed 2026-10-07T15:30:39.872Z sha d95267f :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-08T00:48:55.022Z sha f7a9968 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-08T01:34:24.249Z sha 1e2b5b2 :: needs maintainer review before merge. :: none
  • reviewed 2026-10-08T01:55:10.567Z sha 1e2b5b2 :: needs maintainer review before merge. :: none
  • reviewed 2026-10-08T06:04:34.391Z sha 197422a :: needs maintainer review before merge. :: none
  • reviewed 2026-10-08T06:31:11.825Z sha 197422a :: needs maintainer review before merge. :: none
  • reviewed 2026-10-08T09:07:16.227Z sha a13ab86 :: needs maintainer review before merge. :: none

steipete and others added 2 commits October 7, 2026 17:43
Keep the contributor date and formatter caches, reuse saturating totals and navigation data, and move fixture-only construction into tests. Add reproducible helper measurements and document pixel-identical native render coverage. Sync the reviewed main baseline without rewriting contributor history.

Co-authored-by: Yuxin Qiao <104957188+Yuxin-Qiao@users.noreply.github.com>
@clawsweeper clawsweeper Bot added proof: sufficient Contributor real behavior proof is sufficient. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. and removed status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P3 Low-risk cleanup, docs, polish, ergonomics, or speculative feature. proof: sufficient Contributor real behavior proof is sufficient. rating: 🦞 diamond lobster Very strong PR readiness with only minor maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants