You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 4d1d3ec
Browse filesBrowse the repository at this point in the historyBrowse files
# ADR-63664: Surface AWF steering counters in compact usage data
2
+
3
+
**Date**: 2026-09-26
4
+
**Status**: Draft
5
+
**Deciders**: gh-aw maintainers
6
+
7
+
---
8
+
9
+
### Context
10
+
11
+
The PR adds support for reading AWF API proxy steering events from multiple log file layouts and schema variants, then backfills those counts into compact usage artifacts and audit output. Today, detailed steering behavior requires raw firewall logs, while compact usage summaries and usage-only audits do not expose per-event steering totals. The changed files span JavaScript artifact generation, Go audit backfill logic, schema updates, tests, and documentation, which indicates a cross-cutting product decision rather than a localized bug fix. The decision is how gh-aw should represent steering behavior when only compact usage artifacts are available.
12
+
13
+
### Decision
14
+
15
+
We will aggregate AWF steering events by normalized event name inside the compact usage activity summary and propagate those counters into audit token-usage output when detailed firewall analysis is unavailable. We will support the log filename and schema variants visible in the PR evidence, including `events.jsonl`, `event-logs.jsonl`, and top-level or nested event-name fields. We chose this because it preserves steering visibility for `gh aw audit --artifacts usage` and related usage-only flows without requiring raw firewall log downloads.
16
+
17
+
### Alternatives Considered
18
+
19
+
#### Alternative 1: Keep steering analysis only in raw firewall log processing
20
+
21
+
This was a realistic option because gh-aw already derives detailed steering information from firewall artifacts when those logs are present. It was not chosen because the PR evidence explicitly adds steering data to `usage/activity/summary.json` and backfills audit results from compact usage artifacts, showing that raw-log-only visibility is insufficient for the intended audit workflows.
22
+
23
+
#### Alternative 2: Publish only a single total steering-event count
24
+
25
+
This was considered because earlier code paths already tracked `total_steering_events`, and a single aggregate is simpler to compute and document. It was not chosen because the PR updates both JS and Go paths to preserve per-event counters such as `token_steering` and `timeout_steering`, which provides more actionable inspection of AWF behavior than a single total.
26
+
27
+
### Consequences
28
+
29
+
#### Positive
30
+
- Usage-only audit flows can report steering behavior even when raw firewall logs are unavailable.
31
+
- Operators gain per-event steering counters, which makes it easier to distinguish token, timeout, model, or other steering causes.
32
+
- The implementation becomes more robust across AWF versions by recognizing multiple log filenames and event-name field variants.
33
+
34
+
#### Negative
35
+
- Steering parsing logic now exists in both JavaScript artifact-generation code and Go audit-analysis code, which increases maintenance overhead.
36
+
- Supporting multiple historical file layouts and schema variants adds complexity and ongoing compatibility expectations.
37
+
- Compact usage artifacts and schemas become broader, which increases documentation and regression-test surface area.
38
+
39
+
#### Neutral
40
+
- Audit consumers now need to understand both `total_steering_events` and `steering_event_counts` fields in token-usage output.
41
+
- The change does not replace detailed firewall analysis; it supplements it when only usage artifacts are available.
42
+
- Additional tests are required to keep the JS and Go aggregation paths aligned as AWF logging evolves.
43
+
44
+
---
45
+
46
+
*ADR created by [adr-writer agent]. Review and finalize before changing status from Draft to Accepted.*
Copy file name to clipboardExpand all lines: docs/src/content/docs/reference/artifacts.md
+9Lines changed: 9 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -259,6 +259,13 @@ Its `activity/summary.json` file uses the `usage-activity-summary/v1` schema. Th
259
259
"filtered_tool_counts": { "issue_read": 2 },
260
260
"filtered_reason_counts": { "integrity": 2 }
261
261
},
262
+
"steering": {
263
+
"total_events": 3,
264
+
"event_counts": {
265
+
"token_steering": 2,
266
+
"timeout_steering": 1
267
+
}
268
+
},
262
269
"working_set": {
263
270
"measurement_state": "measured",
264
271
"rebuild_factor": 3.9017857142857144,
@@ -325,6 +332,8 @@ label name, and GitHub database or node ID when returned by the API.
325
332
326
333
The conclusion job derives `gateway` and `integrity` from MCP gateway logs, falling back to `rpc-messages.jsonl` when `gateway.jsonl` is unavailable. These compact aggregates let `gh aw logs --artifacts usage` report MCP call, payload-size, duration, failure, and integrity-filter metrics without downloading raw logs. Cross-run reports include `runs_with_filtered_events`; the existing logs report summary remains the source for the total number of runs.
327
334
335
+
The `steering` section aggregates AWF API proxy events by normalized event name. `gh aw audit --artifacts usage` exposes these counters in `firewall_token_usage.steering_event_counts`, so steering behavior can be inspected without downloading raw firewall logs.
336
+
328
337
`rebuild_factor` is `cumulative_input_tokens / peak_input_tokens`, where each invocation contributes the canonical `input_tokens` value from the agent `token_usage.jsonl` record. Cache-read and cache-write fields are not added because provider normalization has already produced that logical input count. The factor is omitted when `measurement_state` is `unavailable`; `partial` means usable records were measured but malformed or unsupported records were ignored.
329
338
330
339
Working-Set Rebuild Factor measures cumulative context reconstruction relative to peak invocation context. It is an efficiency/trajectory metric, not a measurement of semantic coherence debt and not a predictor of task success. It cannot identify missing task facts or classify outcome quality. The metric is conceptually inspired by [“The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks”](https://arxiv.org/abs/2608.16630), while deliberately limiting the implementation to observable token traffic.
0 commit comments