Skip to content
Merged
Show file tree
Hide file tree
Changes from 46 commits
Commits
Show all changes
51 commits
Select commit Hold shift + click to select a range
27f449f
test(fuzz): add the differential fuzz harness and the pure-helper fuzzer
claude Aug 9, 2026
879a03f
test(fuzz): add the protobuf codec differential fuzzer
claude Aug 10, 2026
78f1660
test(fuzz): add decoder robustness fuzzing under byte mutation
claude Aug 10, 2026
71af124
test(fuzz): generate the wire-fidelity cases from the schema
claude Aug 10, 2026
4623ea0
test(fuzz): fuzz the bridge anti-corruption layer and the event buffer
claude Aug 10, 2026
4edf11a
test(fuzz): fuzz the closed-domain argument boundary
claude Aug 10, 2026
6853d65
build(fuzz): wire the fuzz suite into CI, npm scripts and the docs
claude Aug 10, 2026
81998b7
test(fuzz): tighten the oracle after review
claude Aug 10, 2026
919a4b8
test(fuzz): credit the audit that already tracked the proto gaps
claude Aug 10, 2026
6996078
test(fuzz): fix a crash, an unreachable entry, and four blind spots
claude Aug 10, 2026
5b5347f
test(fuzz): make the trap guard cover every decode, and pin the packi…
claude Aug 10, 2026
208947a
test(fuzz): close six oracle holes found in review
claude Aug 10, 2026
86a9e02
test(fuzz): scope packability per message, not across the schema
claude Aug 10, 2026
a0c91bb
test(fuzz): make seven targets actually reach the code they name
claude Aug 10, 2026
2cedd89
test(fuzz): pin the not-encoded entry to the outcome it describes
claude Aug 10, 2026
be6c893
test(fuzz): stop the oracle guessing where the schema can tell it
claude Aug 10, 2026
efe09cd
test(fuzz): report the direction that was invisible, and stop the flakes
claude Aug 10, 2026
cedac01
test(fuzz): weight crypto lengths toward the ones the ciphers accept
claude Aug 10, 2026
c8f09c5
test(fuzz): make three targets run the code they claim to cover
claude Aug 10, 2026
f776a00
test(fuzz): stop four checks claiming more than they verified
claude Aug 10, 2026
627378e
test(fuzz): run the successful paths, and fail on lost coverage
claude Aug 10, 2026
a204a57
test(fuzz): make nine allowlist entries prove what they claim
claude Aug 10, 2026
b77946a
test(fuzz): repair a generator regression, and stop three loose matches
claude Aug 10, 2026
8956734
test(fuzz): stop the harness losing findings, and parse balanced groups
claude Aug 10, 2026
f4daa5a
test(fuzz): keep the signed zero, the array order, and the deep chains
claude Aug 10, 2026
229e634
test(fuzz): four more predicates that check the outcome, and a dead c…
claude Aug 10, 2026
9ef1534
test(fuzz): record why the oneof target does not compare ordering
claude Aug 10, 2026
5d40709
test(fuzz): stop three canonicalisations agreeing when they should not
claude Aug 10, 2026
d25c199
test(fuzz): anchor the hex fold, and make the oneof premise self-chec…
claude Aug 10, 2026
918f5a6
test(fuzz): compare repeated fields in order, and make three checks r…
claude Aug 10, 2026
20b5181
test(fuzz): order-preserving subset, schema-aware round trips, and a …
claude Aug 10, 2026
89bc565
test(fuzz): bound shrinking by the clock, and reach five branches not…
claude Aug 10, 2026
855cf1a
docs(fuzz): list what the suite does not cover, where planners will s…
claude Aug 10, 2026
6e83224
test(fuzz): narrow three allowlist entries, and reach four more branches
claude Aug 10, 2026
0d75d71
test(fuzz): compare oneof winners, and stop three oracles trusting th…
claude Aug 10, 2026
339c7c7
test(fuzz): compare the whole singleton encoding, and stop folding "0…
claude Aug 10, 2026
8b00212
test(fuzz): correlate buffered event identities, and find two merge-p…
claude Aug 10, 2026
03a58f9
test(fuzz): serialise reports without invoking the traps in what they…
claude Aug 10, 2026
cc26669
test(fuzz): let the round trip decide what the copy strategy explains
claude Aug 10, 2026
da42e1c
test(fuzz): drop two corpus files a measurement script left behind
claude Aug 10, 2026
d8ea5e1
test(fuzz): pin the omission and renumbering entries to what they mea…
claude Aug 10, 2026
fd6ab9e
test(fuzz): revert formatter churn in three unrelated audit files
claude Aug 10, 2026
0f92629
test(fuzz): make an unresolved schema path a finding, not a skip
claude Aug 10, 2026
7e9f5ab
test(fuzz): close three classifiers that excused what they should report
claude Aug 10, 2026
f07b816
test(fuzz): stop three checks from trusting a key they should not
claude Aug 10, 2026
bebf7a5
test(fuzz): stop three checks from normalising away their own subject
claude Aug 10, 2026
4418fd1
fix(fuzz): repair the renumbering classifier that was excusing nothing
claude Aug 10, 2026
5e60d67
fix(event-buffer): keep an id-less chats.upsert out of a history set
claude Aug 10, 2026
48b701c
test(fuzz): record findings, write failures and issue ownership honestly
claude Aug 10, 2026
94cc655
test(fuzz): bound crash details, and count them separately
claude Aug 10, 2026
62db42c
fix(fuzz): stop a clean seed closing an unresolved finding
claude Aug 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
261 changes: 261 additions & 0 deletions .github/workflows/fuzz.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,261 @@
name: Nightly Fuzz

# Two modes, on purpose.
#
# Every pull request already runs the fuzz suite through `npm test`, with a fixed
# seed and small per-target budgets. That run is deterministic: it cannot fail
# because of an unlucky draw, which is the only way a fuzz suite survives contact
# with a CI system people have to trust.
#
# This job is where the searching happens. It varies the seed per run, raises the
# budgets, and enforces the parts of the known-divergence registry that would be
# hostile on a pull request — entries past their review date, and entries that no
# longer excuse anything.

on:
schedule:
# 03:17 UTC, off the hour so it does not queue behind everything else.
- cron: '17 3 * * *'
workflow_dispatch:
inputs:
seed:
description: 'Fuzz seed (defaults to the run id)'
required: false
type: string
mode:
description: 'smoke or deep'
required: false
default: 'deep'
# A choice, not free text: the runner rejects anything else outright, and
# a rejected dispatch is better than one that silently runs a smoke pass.
type: choice
options:
- deep
- smoke
Comment thread
coderabbitai[bot] marked this conversation as resolved.

# One fuzz run at a time in the *repository*, not per ref. The tracking issue is
# repository-wide, so two runs on different refs — a nightly and a manual
# dispatch on a branch — would each look it up, each find nothing, and each
# create one. Every later run then comments on and closes only the lowest
# number, leaving the duplicate open forever holding a stale report. Queued
# rather than cancelled: the in-flight run's result is the one worth keeping.
concurrency:
group: fuzz
cancel-in-progress: false

permissions:
contents: read
issues: write

jobs:
fuzz:
runs-on: ubuntu-latest
timeout-minutes: 45

steps:
- name: Checkout code
uses: actions/checkout@v4
Comment thread
jlucaso1 marked this conversation as resolved.

- name: Setup Node.js
uses: actions/setup-node@v4
with:
node-version: '24'
cache: 'npm'

- name: Install dependencies
run: npm ci

- name: Run the fuzz suite
id: fuzz
env:
# A fresh seed per run is the point: the fixed-seed smoke run on every
# PR has already searched its own corner exhaustively.
FUZZ_SEED: ${{ inputs.seed || github.run_id }}
FUZZ_MODE: ${{ inputs.mode || 'deep' }}
FUZZ_TIME_BUDGET_MS: '180000'
FUZZ_REPORT_DIR: fuzz-reports
FUZZ_STRICT_ALLOWLIST: '1'
# --expose-gc turns on the WASM handle-leak probe, which skips without it.
# --test-timeout bounds a decoder that stops returning: the runner's own
# slowMs is measured after the check completes, so it cannot see an input
# that never completes. The parent process owns this timer, so it fires
# even when the child's event loop is blocked by a synchronous WASM loop.
run: node --expose-gc --test --test-timeout=1800000 "./src/__fuzz__/**/*.test.ts"
continue-on-error: true
Comment thread
jlucaso1 marked this conversation as resolved.
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.

- name: Summarise
id: report
if: always()
run: |
{
node scripts/fuzz/report.ts fuzz-reports --markdown --fail-on-stale \
${{ steps.fuzz.outcome == 'success' && ' ' || '--run-failed' }}
Comment on lines +91 to +92

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Skip stale-entry enforcement for smoke dispatches

When a manual dispatch selects mode: smoke, the runner executes roughly 1/25 of the randomized cases used by deep mode, but this step still passes --fail-on-stale. report.ts consequently treats any known divergence that the smaller sample merely did not reach as a stale registry entry and exits nonzero, causing an otherwise healthy smoke run to update the tracking issue and fail the final gate. Restrict stale-entry enforcement to deep runs, or explicitly mark smoke reports as incomplete for this check.

Useful? React with 👍 / 👎.

} > fuzz-summary.md 2>&1 && echo "clean=true" >> "$GITHUB_OUTPUT" || echo "clean=false" >> "$GITHUB_OUTPUT"
cat fuzz-summary.md >> "$GITHUB_STEP_SUMMARY"
Comment thread
coderabbitai[bot] marked this conversation as resolved.

- name: Upload reports
if: always()
uses: actions/upload-artifact@v4
with:
name: fuzz-reports-${{ github.run_id }}
path: |
fuzz-reports/
fuzz-summary.md
retention-days: 30

# The fuzz outcome is part of the condition, not just the report: a run can
# fail on something the report has no findings for — a harness crash, a
# suite-level assertion — and an issue is exactly what those need too.
#
# Runs on every outcome, not only failures. A clean run has something to
# say too: without it the tracking issue stayed open forever holding a
# report the nightly had since disproved, which is the fastest way to
# teach people to ignore it.
- name: Update the tracking issue
if: always()
uses: actions/github-script@v7
with:
script: |
const fs = require('node:fs')
const summary = fs.readFileSync('fuzz-summary.md', 'utf8')
// The seed is a free-form workflow input, so the reproduction command
// has to quote it or it is not the command that ran: `nightly run`
// would make the shell treat `run` as the program, and a `;` would
// append a second command to whatever the reader pastes. Same rule as
// the runner's own replay hint.
const shellQuote = value =>
/^[\w.:@/+=-]+$/.test(value) ? value : `'${String(value).replaceAll("'", String.raw`'\''`)}'`
// The same condition the step used to be gated on, now a value: the
// report found nothing *and* the run itself did not fail.
const clean = ${{ steps.report.outputs.clean == 'true' && steps.fuzz.outcome == 'success' }}
const seed = process.env.FUZZ_SEED
const mode = process.env.FUZZ_MODE
const budget = process.env.FUZZ_TIME_BUDGET_MS
// Stable across runs, so the issue this job owns is identifiable by
// something other than a label anyone can apply. The seed moved into
// the body: it changes nightly, and a title that changes cannot be a key.
const title = 'Nightly fuzz: open findings'
const marker = '<!-- nightly-fuzz-tracking-issue -->'

// One open issue per topic, updated rather than duplicated: a fuzzer
// that opens a fresh issue every night trains people to close them
// unread.
//
// Matched on the marker alone. `listForRepo` returns pull requests as
// well as issues, and the generic `fuzz` label is one anybody can put on
// anything — so the label by itself would let the nightly report land in
// an unrelated thread, or on a PR, and the API's default ordering makes
// which one unstable as more labelled items are opened.
//
// And the label cannot be part of the *query* either, only of the
// answer. A label is editable by anyone with write access: strip `fuzz`
// from the open tracking issue and a label-scoped search stops returning
// it before the marker is ever consulted, so the next failing night files
// a duplicate and no clean night can ever find and close the original.
// The marker lives in the body, which is the one part of the issue this
// job writes and nothing routine edits. It still *applies* the label on
// creation, for people who browse that way — it just never trusts it.
//
// The cost is listing open issues rather than a label slice, which is
// why it paginates. The lowest number wins, so the choice does not depend
// on page order either.
const open = await github.paginate(github.rest.issues.listForRepo, {
owner: context.repo.owner,
repo: context.repo.repo,
state: 'open',
per_page: 100
})
const owned = open
.filter(item => !item.pull_request && (item.body ?? '').includes(marker))
.sort((left, right) => left.number - right.number)
Comment on lines +184 to +193

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Require bot ownership for marker-matched issues

Any repository user can open an issue whose body contains this publicly visible marker, but the lookup treats the marker alone as proof that the workflow owns the issue. If that issue has the lowest number, a failing nightly posts its report there instead of creating or updating the real tracker; a clean nightly will comment on and close every such user-authored issue. Restrict matches to issues created by the workflow's bot identity (and preferably the expected title) before updating or closing them.

Useful? React with 👍 / 👎.


const body = [
marker,
`Seed: \`${seed}\` · mode \`${mode}\``,
'',
summary,
'',
`Run: ${context.serverUrl}/${context.repo.owner}/${context.repo.repo}/actions/runs/${context.runId}`,
// The run's own flags, or the reproduction can pass where the run
// failed: without --expose-gc the leak probe skips, without
// FUZZ_STRICT_ALLOWLIST an expired entry is a warning rather than a
// failure, and the budget decides how many inputs each target gets
// through before it truncates. The mode is read back rather than
// hard-coded to `deep`, because a manual dispatch can run `smoke`
// and `npm run fuzz:deep` would then reproduce a different run.
`Reproduce locally: \`FUZZ_SEED=${shellQuote(seed)} FUZZ_MODE=${mode} FUZZ_TIME_BUDGET_MS=${budget} FUZZ_STRICT_ALLOWLIST=1 node --expose-gc --test --test-timeout=1800000 "./src/__fuzz__/**/*.test.ts"\``,
'',
'---',
'_Generated by [Claude Code](https://claude.ai/code)_'
].join('\n')

// A clean run closes the issue rather than opening one. Commented
// first, so the thread records *why* it closed and against which
// seed — a bare close leaves whoever reopens it guessing.
// Every owned issue, not just the lowest. The concurrency group is
// repository-wide now so two runs cannot race into duplicates, but one
// may already exist from before that, or from a hand-filed issue
// carrying the marker — and closing only `owned[0]` would leave it open
// forever holding a report the nightly has disproved.
if (clean) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep unresolved findings open across clean seeds

When a nightly opens this issue for a seed-specific finding, a later clean run—especially a manual smoke dispatch—closes it even though that run samples different inputs and the workflow never enables FUZZ_RECORD, so the failing input is not guaranteed to be replayed. Thus an unresolved regression can disappear from the tracker merely because the next seed did not hit it; only close after replaying the prior reproducer successfully or after an explicit resolution.

Useful? React with 👍 / 👎.

for (const issue of owned) {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue.number,
body: [
marker,
`The nightly is clean on seed \`${seed}\` (mode \`${mode}\`), so this is closed.`,
'',
`Run: ${context.serverUrl}/${context.repo.owner}/${context.repo.repo}/actions/runs/${context.runId}`,
'',
'It reopens by itself: the next run with findings files against this same marker, or creates a fresh issue if this one is closed.',
'',
'---',
'_Generated by [Claude Code](https://claude.ai/code)_'
].join('\n')
})
await github.rest.issues.update({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue.number,
state: 'closed',
state_reason: 'completed'
})
}
return
}

if (owned.length > 0) {
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: owned[0].number,
body
})
} else {
await github.rest.issues.create({
owner: context.repo.owner,
repo: context.repo.repo,
title,
body,
labels: ['fuzz']
})
}
env:
FUZZ_SEED: ${{ inputs.seed || github.run_id }}
# Mirrored from the run step, so the reproduction command names the mode
# and budget the findings were actually produced under.
FUZZ_MODE: ${{ inputs.mode || 'deep' }}
FUZZ_TIME_BUDGET_MS: '180000'

# The fuzz and summarise steps deliberately do not fail on the spot, so the
# artifacts get uploaded and the issue gets filed first. Without a final gate
# the job would then finish green holding findings, which is the one outcome
# that would make the whole nightly pointless.
- name: Fail the job when the run was not clean
if: always() && (steps.fuzz.outcome != 'success' || steps.report.outputs.clean != 'true')
run: |
echo "fuzz outcome: ${{ steps.fuzz.outcome }}"
echo "report clean: ${{ steps.report.outputs.clean }}"
exit 1
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,14 @@ so existing integrations can migrate with minimal changes. See
| Key management | JS auth state | Rust `PersistenceManager` |
| Auto-reconnect | Manual `startSock()` loop | Transient drops retried in Rust (fibonacci backoff); terminal ones still yours |

Compatibility is checked rather than assumed: a declaration audit against
upstream's `.d.ts`, a wire-fidelity audit of the send path, ~50 behavioural
compatibility suites, and a
[differential fuzz suite](src/__fuzz__/README.md) that generates its own inputs
from the proto schema and compares the two libraries directly. Differences the
fuzzers find are recorded with a reason and a review date, and known open ones
are listed in `src/__fuzz__/harness/divergence.ts`.

## Documentation

The full API reference and guides live in the
Expand Down
5 changes: 5 additions & 0 deletions package.json
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@
"lib/**/*",
"!lib/**/*.map",
"!lib/**/__tests__/**",
"!lib/__fuzz__/**",
"!lib/**/*.test.*",
"!lib/**/*.test-e2e.*"
],
Expand All @@ -70,6 +71,10 @@
"prepack": "npm run build && node scripts/check-pack.ts",
"prepare": "npm run build",
"test": "node --test",
"fuzz": "node --test --test-timeout=600000 ./src/__fuzz__/**/*.test.ts",
"fuzz:deep": "FUZZ_MODE=deep node --expose-gc --test --test-timeout=1800000 ./src/__fuzz__/**/*.test.ts",
"fuzz:record": "FUZZ_RECORD=1 node --test --test-timeout=600000 ./src/__fuzz__/**/*.test.ts",
"fuzz:report": "node scripts/fuzz/report.ts",
"test:compat-auditor": "node --test scripts/compatibility/__tests__/audit.test.ts",
"typecheck:compat-auditor": "npm run build --silent && npm run compat:check-waproto --silent && npm run compat:layers --silent && tsc -p scripts/compatibility/tsconfig.json",
"test:e2e": "NODE_TLS_REJECT_UNAUTHORIZED=0 ADV_SECRET_KEY=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA= node --expose-gc --test --test-concurrency=1 ./src/__tests__/e2e/*.test-e2e.ts"
Expand Down
Loading
Loading