Skip to content

feat(spend): show native Codex turn performance - #4304

Merged
steipete merged 15 commits into
steipete:mainfrom
Yuxin-Qiao:feat/spend-token-speed
Oct 8, 2026
Merged

steipete merged 15 commits into
steipete:mainfrom
Yuxin-Qiao:feat/spend-token-speed

Conversation

@Yuxin-Qiao

@Yuxin-Qiao Yuxin-Qiao commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Native Codex sessions in Usage & Spend now show weighted whole-turn output, median model-first-token latency, and median completed-turn duration beneath the existing cost-ranked header. Performance details start collapsed and expose sample coverage, percentiles, cache fraction, and model/effort groups. Sessions without valid samples show no performance line.

Timing is joined to owned, deduplicated requests in the existing ledger and filtered by completion day. Failed/incomplete turns, malformed supplied start times, invalid counters, and mismatched turn totals are excluded. Missing first-token observations stay unavailable. Billing keeps its request dates, saved prices, and existing range totals; whole-turn throughput includes reasoning, tools, and waits and is not streaming generation speed.

Parser revision 12 (0d8f9504f8e63d0f) adopts compatible predecessor stores and uses bounded backfill. The regression coverage and synthetic renderer remain; historical proof diaries and private-input/benchmark harnesses have been replaced by a shared fixture and a concise metric contract.

Thanks @Yuxin-Qiao! The original contributor history is preserved.

Verification

The Swift wrapper sources Scripts/test_environment.sh, disables real Keychain access, and forwards test/build to swift with --build-system native --jobs 4 -Xswiftc -gnone.

  • Red on the original contributor production code combined with main: /tmp/pr-4304-swift test --filter 'CostUsage(Turn|RequestLedger|Store|CoverageCompatibility)|SpendSessionPerformance' reported 239 tests in 19 suites and five failures, all from the new malformed-start parameter cases.
  • Green: /tmp/pr-4304-swift test --filter 'CostUsageTurnPerformanceTests|CostUsageTurnPerformanceDetailsTests|CostUsageTurnCompletionDayTests|CostUsageRequestLedgerMigrationTests|CostUsageStoreTests|CostUsageCoverageCompatibilityTests|SpendSessionPerformanceTests|SpendDashboardSessionRowTests|ProviderArchitectureGatekeeperTests|LocalizationLanguageCatalogTests' passed 205 tests in 10 suites. All five malformed-start cases pass while retaining their 220 billed tokens; missing/null start times remain supported.
  • make check: passed, zero violations in 2,861 Swift files. The initial check caught a system/pinned formatter mismatch; the final code uses the repository-pinned formatter.
  • env -u CODEXBAR_ALLOW_TEST_KEYCHAIN_ACCESS CODEXBAR_TEST_SUITE_TIMEOUT=900 ./Scripts/test.sh --swift-command /tmp/pr-4304-swift --direct-workers 4: passed all 1,602 selections in 145 groups in direct mode: 144 groups passed first attempt, one recovered on a fresh group retry, zero timeouts (2,062.1 s). The initial run with the default 180-second harness limit exited 124; larger cost suites exceeded that limit, and an Antigravity fixture assertion recovered on isolation. The successful rerun changed only the harness timeout, not selections or assertions.
  • GitHub CI for fbec7506af51130390afd9b19896a8ee9d4d393c was cancelled; a fresh required-check run is still needed before merge.
  • Independent Codex autoreview of the final diff: no actionable P0–P2 findings.
  • Synthetic headless renders at 320/420/820 points, English/Chinese, light/dark, including privacy masking and a session without timing. The running app and real accounts were not used; no installed-window responsiveness claim is made.

Synthetic UI

Before the rank restoration (original proposed row):

Synthetic session rows before retaining cost ranks

After retaining cost ranks; the third example has no timing samples:

Synthetic session rows with cost ranks and optional timing

Narrow dark layout:

Synthetic session rows at 320 points in dark mode

@clawsweeper

clawsweeper Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P3 Low-risk cleanup, docs, polish, ergonomics, or speculative feature. proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Oct 6, 2026
@clawsweeper

clawsweeper Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Codex review: blocked before merge. Reviewed October 8, 2026, 2:58 AM ET / 06:58 UTC (Revision 9).

ClawSweeper review

What this changes

Adds completed-turn performance measurements and expandable statistics to native Codex sessions in Usage & Spend.

Example: Select one day in Usage & Spend

  • Before: Session rows show spending and token totals without turn-performance measurements.
  • After: The demonstrated session shows 12 timed turns, 1.0 s first-token latency, 17.3 tok/s whole-turn output, and 18.8 s duration.

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) Useful native behavior is demonstrated, accounting compatibility is covered, and no actionable code finding remains.
Proof confidence 🦞 diamond lobster (5/6) ✨ media proof bonus Sufficient (screenshot): Historical installed Settings-window evidence exercises production scanning, metrics, disclosure, and completion-day selection using isolated synthetic logs, with 24 range turns becoming 12 per day and matching reference values. Current synthetic renders supplement the preserved runtime proof; predecessor adoption and bounded-backfill coverage support stored-ledger compatibility.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Product

Kind: Feature · Worth it: Yes
User problem: Users can inspect session spending but cannot see how long completed Codex turns took or compare observed performance.
Reason: Existing local records can provide useful observations at bounded cost. The area's recorded owner integration accepts the direction while preserving ranked spending headers.

Merge readiness

⛔ Blocked before merge - 2 items remain

This PR remains a useful, distinct feature with accepted owner direction and sufficient historical native-app proof. No concrete patch defect was identified.

Priority: P3
Reviewed head: fbec7506af51130390afd9b19896a8ee9d4d393c

Before merge

  • Resolve merge risk (P1) - GitHub reports merge conflicts, and integration with fetched current main has not been validated.
  • Complete next step (P2) - Resolve the reported merge conflicts against current main and refresh review of the resulting integration.

Findings

None.

Agent review details

How this fits together

CodexBar reads local Codex session logs into a cached usage ledger and displays estimated spending in Settings. Completed-turn timing joins those records to provide performance observations alongside billing totals.

flowchart TD
  A[Local Codex session logs] --> B[Usage and completion parser]
  B --> C[Cached usage ledger]
  C --> D[Validate completed turns]
  D --> E[Filter by completion day]
  E --> F[Performance statistics]
  C --> G[Session billing totals]
  F --> H[Usage and Spend session rows]
  G --> H
Loading

Technical review

Best possible solution:

Expose qualified native turn observations through the existing ledger while preserving cost ranks, billing dates, saved prices, and unavailable-data semantics.

Do we have a high-confidence way to reproduce the issue?

Not applicable as a bug reproduction; historical installed-window checks demonstrate the requested measurements, disclosure, and day selection.

Is this the best way to solve the issue?

Yes: extending the existing local ledger preserves authoritative accounting and avoids a competing parser or data source.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 4614415bbed7.

Merge-risk options

Maintainer options:

  1. Decide the mitigation before merge
    Expose qualified native turn observations through the existing ledger while preserving cost ranks, billing dates, saved prices, and unavailable-data semantics.
  2. Pause or close
    Do not merge this PR until maintainers decide whether the risk is worth taking.

Provenance checked

Testing

Proof path: shipped entry point. Added test files: 16.

Security

None.

Evidence

What I checked:

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • Yuxin-Qiao: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • urda: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Review metrics

Metric Value Why it matters
Swift line changes Production +1152/-193; tests +1630/-4 The introduced delta includes already-merged integrations; remaining growth supports ledger timing, statistics, and native presentation.

Labels

Label changes:

No label changes.

Label justifications:

  • P3: This adds analytical information to an existing dashboard without repairing a blocked core workflow.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🦞 diamond lobster and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR.
  • proof: sufficient: Contributor real behavior proof is sufficient.
  • proof: 📸 screenshot: Contributor real behavior proof includes screenshot evidence.

Rating scale

6/6 🦀 challenger crab · 5/6 🦞 diamond lobster · 4/6 🐚 platinum hermit · 3/6 🦐 gold shrimp · 2/6 🦪 silver shellfish · 1/6 🧂 unranked krab. Overall follows the weaker of proof and patch quality; ✨ marks media proof (a screenshot, video, or linked artifact) that directly shows the changed behavior.

Workflow

ClawSweeper edits this one comment on every review. Comment @clawsweeper re-review for a fresh review only; repair and merge need explicit maintainer commands such as @clawsweeper autofix or @clawsweeper automerge.

History

Review history (8 earlier review cycles)
  • reviewed 2026-10-06T09:06:55.339Z sha 616d2c0 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-06T13:38:46.741Z sha a6f79c3 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-06T15:27:08.563Z sha fcf836d :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-07T13:51:54.098Z sha 45b1a62 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-07T14:13:50.120Z sha bcba22d :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-07T15:24:28.929Z sha b28134d :: needs maintainer review before merge. :: none
  • reviewed 2026-10-08T01:25:13.731Z sha 5d6cdd8 :: blocked before merge. :: none
  • reviewed 2026-10-08T06:07:10.247Z sha fbec750 :: needs maintainer review before merge. :: none

@Yuxin-Qiao
Yuxin-Qiao force-pushed the feat/spend-token-speed branch from 616d2c0 to ce92708 Compare October 6, 2026 13:16
@clawsweeper clawsweeper Bot removed the proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. label Oct 6, 2026
@Yuxin-Qiao
Yuxin-Qiao marked this pull request as ready for review October 6, 2026 13:24
@clawsweeper clawsweeper Bot added proof: sufficient Contributor real behavior proof is sufficient. proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. and removed status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Oct 7, 2026
@clawsweeper clawsweeper Bot added rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. labels Oct 8, 2026
Integrate the existing contributor branch with current main without rewriting
its history. Keep native turn timing separate from billing and saved pricing,
retain the cost-ranked header, and exclude malformed supplied start times.
Parser revision 12 preserves predecessor stores for bounded timing backfill.

Replace historical proof diaries and opt-in private/benchmark harnesses with
a shared synthetic turn fixture and concise metric documentation. Keep
behavioral coverage and headless production-row rendering, and add the
Unreleased changelog credit.

Co-authored-by: Yuxin Qiao <104957188+Yuxin-Qiao@users.noreply.github.com>
@clawsweeper clawsweeper Bot added rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Oct 8, 2026
# Conflicts:
#	CHANGELOG.md
#	Sources/CodexBar/Resources/ar.lproj/Localizable.strings
#	Sources/CodexBar/Resources/ca.lproj/Localizable.strings
#	Sources/CodexBar/Resources/de.lproj/Localizable.strings
#	Sources/CodexBar/Resources/en.lproj/Localizable.strings
#	Sources/CodexBar/Resources/es.lproj/Localizable.strings
#	Sources/CodexBar/Resources/fa.lproj/Localizable.strings
#	Sources/CodexBar/Resources/fr.lproj/Localizable.strings
#	Sources/CodexBar/Resources/gl.lproj/Localizable.strings
#	Sources/CodexBar/Resources/id.lproj/Localizable.strings
#	Sources/CodexBar/Resources/it.lproj/Localizable.strings
#	Sources/CodexBar/Resources/ja.lproj/Localizable.strings
#	Sources/CodexBar/Resources/ko.lproj/Localizable.strings
#	Sources/CodexBar/Resources/nl.lproj/Localizable.strings
#	Sources/CodexBar/Resources/pl.lproj/Localizable.strings
#	Sources/CodexBar/Resources/pt-BR.lproj/Localizable.strings
#	Sources/CodexBar/Resources/ru.lproj/Localizable.strings
#	Sources/CodexBar/Resources/sv.lproj/Localizable.strings
#	Sources/CodexBar/Resources/th.lproj/Localizable.strings
#	Sources/CodexBar/Resources/tr.lproj/Localizable.strings
#	Sources/CodexBar/Resources/uk.lproj/Localizable.strings
#	Sources/CodexBar/Resources/vi.lproj/Localizable.strings
#	Sources/CodexBar/Resources/zh-Hans.lproj/Localizable.strings
#	Sources/CodexBar/Resources/zh-Hant.lproj/Localizable.strings
@steipete
steipete merged commit a88f081 into steipete:main Oct 8, 2026
4 checks passed
@steipete

steipete commented Oct 8, 2026

Copy link
Copy Markdown
Owner

Thanks @Yuxin-Qiao! Merged in a88f081 with your commits preserved: native Codex session rows in Usage & Spend can now show whole-turn output throughput, median first-token latency and the number of valid timed turns, with billing semantics and cost ranks unchanged (verified against the 0.73.0 ledger work after rebasing and regenerating the parser fingerprint) and a reproduced timing-validation gap closed in review. Ships in the next release.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P3 Low-risk cleanup, docs, polish, ergonomics, or speculative feature. proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. proof: sufficient Contributor real behavior proof is sufficient. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants