Skip to content

test(bellwether): read the tree bellwether unpack writes, checked against sets.toml - #2834

Merged
slin1237 merged 2 commits into
mainfrom
test/bellwether-consumers-at-scale
Oct 7, 2026
Merged

slin1237 merged 2 commits into
mainfrom
test/bellwether-consumers-at-scale

Conversation

@hello-alexmcc

Copy link
Copy Markdown
Collaborator

Description

Problem

bellwether #31 stores each benchmark set as <set>.jsonl.zst in Git LFS, and makes the plain tree bellwether unpack writes the root that consumers read. smg's two bellwether tests still point at fixtures/. Their .jsonl filter skips every compressed set without a word. With Qwen3-8B's BFCL sets recorded:

  • the tokenizer test read 63 of 6,124 cases and passed;
  • the gateway's render parity test read 47 of 3,688 render cases and passed.

Solution

Both tests read the unpacked tree, which #2822's workflow writes with bellwether unpack, and hold what they read to each model's sets.toml:

  • A compressed set or a Git LFS pointer under the root fails the run, and the message gives the unpack command.
  • The cases read from each set must be the cases sets.toml lists. Each of these fails the run:
    • a listed set the tree lacks;
    • a set file the table does not list;
    • sets with no sets.toml beside them.
  • Each case is compared as its line is read. The run prints each set with its count and a tally per model.
  • A model whose tokenizer does not load is listed at the end instead of stopping the run.
  • A tokenizer download is written beside its final name and renamed into place, so an interrupted write no longer leaves a short file that the next run trusts.

Changes

  • crates/tokenizer/tests/bellwether_fixtures.rs: the unpacked tree, the sets.toml checks, per-line comparison, and the tally.
  • model_gateway/tests/bellwether_render_parity.rs: the same for the render sets it replays, plus the atomic tokenizer download.

Test Plan

On eb42d3a..92c1e1c, on main 200ed03, with a target directory of its own and exit codes captured unpiped:

  • cargo +nightly fmt --all -- --check: 0.
  • cargo +1.98.0 clippy -p llm-tokenizer -p smg --tests -- -D warnings: 0.
  • BELLWETHER_FIXTURES=<bellwether main, unpacked> cargo +1.98.0 test -p llm-tokenizer -p smg --test 'bellwether_*' -- --nocapture: 0, with 5 passed in each suite:
    • sets_without_a_sets_toml_stop_the_run
    • compressed_sets_and_lfs_pointers_are_found_under_the_root
    • the_cases_read_must_be_the_cases_sets_toml_lists
    • a_model_whose_tokenizer_does_not_load_fails_the_run_after_the_others
    • encode_and_incremental_decode_match_the_reference / render_fixtures_match_the_reference_byte_for_byte
  • With Qwen3-8B's BFCL sets: the tokenizer test reads all 6,124 cases and the render parity test all 3,688, from the unpacked tree.
  • A run over the 80 checkpoint groups of a local scale recording is in progress. Its results will follow on this PR.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

…nst sets.toml

smg-project/bellwether#31 stores a benchmark set as <set>.jsonl.zst in
Git LFS, and makes the plain tree `bellwether unpack` writes the root
consumers read. Pointed at fixtures/, as this test's doc said, the jsonl
filter skipped every compressed set without a word: with Qwen3-8B's BFCL
sets recorded, the run read 63 of 6,124 cases and passed.

The doc and the skip notice now point at the unpacked tree, and:
- a compressed set or a Git LFS pointer under the root fails the run
  with the unpack command;
- the cases read from each render and parse set must be the cases the
  model's sets.toml lists; a listed set the tree lacks, a set file the
  table does not list, or sets with no sets.toml beside them fail the
  run;
- each case is compared as its line is read, not after a model's whole
  directory is read, and the run prints each set with its count, the
  differences and a tally per model, not a line per case;
- a model whose tokenizer does not load no longer ends the run: it is
  listed when the run fails at the end, after the other models are
  compared.

Unit tests on small fixture trees cover the refusal, the comparison with
sets.toml, a tree without sets.toml and a model that does not load.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
…t sets.toml

The render parity test of the chat request path gets the tokenizer
fixture test's changes, for the render sets it replays. It points at the
tree `bellwether unpack` writes and refuses compressed sets and Git LFS
pointers. It holds the cases read from each render set to the model's
sets.toml, and compares each case as its line is read. It prints the
differences with a tally per model, and lists a model whose tokenizer
does not load at the end instead of stopping there. With Qwen3-8B's BFCL
sets recorded, it read 47 of 3,688 render cases from fixtures/ and
passed; it now fails there with the unpack command, and reads all 3,688
from the unpacked tree.

The tokenizer download is now written beside its name and renamed into
place, as the tokenizer fixture test does, so an interrupted write no
longer leaves a short file that the next run trusts.

Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
@github-actions github-actions Bot added tokenizer Tokenizer related changes tests Test changes model-gateway Model gateway crate changes labels Oct 7, 2026
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Team
  • Run ID: 136d01a2-c4f6-4651-92a2-e364408fb8ee
📥 Commits

Reviewing files that changed from the base of the PR and between 200ed03 and 92c1e1c.

📒 Files selected for processing (2)
  • crates/tokenizer/tests/bellwether_fixtures.rs
  • model_gateway/tests/bellwether_render_parity.rs

Included review availability: This review used your included allowance. 3 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 5 reviews per hour.


📝 Summary

Summary by CodeRabbit

  • Tests
    • Expanded tokenizer and rendering parity checks to validate unpacked JSONL fixtures, set listings, and case counts, and to detect missing, extra, or compressed files.
    • Comparisons now continue for other models when one tokenizer fails to load, while reporting the failure.
    • Added coverage for fixture mismatches, missing set tables, and tokenizer-load failures.

Walkthrough

Both Bellwether test suites now validate unpacked JSONL fixture trees, stream cases during comparison, and report fixture or tokenizer-loading failures after processing available models. The render parity test also stages downloaded tokenizer files before moving them to their destination.

Changes

Bellwether fixture validation and comparison

Layer / File(s) Summary
Unpacked fixture tree validation
crates/tokenizer/tests/bellwether_fixtures.rs, model_gateway/tests/bellwether_render_parity.rs
Both suites detect compressed sets and Git LFS pointers, compare plain JSONL files and case counts with sets.toml, and read cases line by line. Tests cover missing tables, missing or unlisted sets, and count mismatches.
Model comparison and failure reporting
crates/tokenizer/tests/bellwether_fixtures.rs, model_gateway/tests/bellwether_render_parity.rs
Both suites record tokenizer-load failures and continue comparing other models. Reports collect set and case mismatches. The render parity test writes downloads to a .part path before renaming them.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Other

Merge Risk: ⚪ Minimal · up to 92c1e

This change only affects the opt-in Bellwether fixture tests. They now read the unpacked fixture tree, reject packed files and Git LFS pointer files, and report tokenizer-load failures after the remaining models have been compared. No production code paths change, and no blocking issues were found, so the change is ready to merge.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly describes the main change: reading the tree produced by bellwether unpack and checking it against sets.toml.
Description check ✅ Passed The description explains the problem, the changes to both Bellwether tests, and the reported validation. It is directly related to the changeset.
Docstring Coverage ✅ Passed Docstring coverage is 88.89% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 54 functions across 2 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

tally: &mut Tally,
known: &BTreeMap<&str, &str>,
) {
self.seen.insert(fixture.id.clone());

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This file and crates/tokenizer/tests/bellwether_fixtures.rs share most of their code, but this record drops the insert result. The tokenizer version asserts it ("{id}: recorded more than once"). Now that each model reads several set files and each set's line count is checked against sets.toml, an id that appears in two sets (e.g. a BFCL set and common) passes here without a word. When that happens:

  • tally.cases counts both lines, but seen keeps one, so the per-model tally and the closing N cases match disagree.
  • differences.insert overwrites the first occurrence's reason with the second's.

Asserting it as the tokenizer test does would keep the two in step:

Suggested change
self.seen.insert(fixture.id.clone());
assert!(
self.seen.insert(fixture.id.clone()),
"{}: recorded more than once",
fixture.id
);

provenance: Value,
}

/// One `[<kind>.<set>]` table of `<slug>/sets.toml`, which `bellwether record`

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: This new doc correctly says <slug>/sets.toml, but the unchanged struct docs above it still name bellwether's own tree: Manifest (line 73, fixtures/<slug>/manifest.toml), RenderCase (line 82) and ParseCase (line 100). The same goes for Manifest (line 131) and Fixture (line 139) in model_gateway/tests/bellwether_render_parity.rs. The module docs and assertion messages already dropped the fixtures/ prefix. The root is now the unpacked tree (fixtures-plain/ by default), and bellwether's fixtures/ is exactly the directory these tests refuse, so <slug>/manifest.toml and <slug>/render/<set>.jsonl would keep those docs from pointing readers at the wrong tree.

@slin1237
slin1237 merged commit eda1134 into main Oct 7, 2026
29 of 46 checks passed
@slin1237
slin1237 deleted the test/bellwether-consumers-at-scale branch October 7, 2026 02:13
@hello-alexmcc

Copy link
Copy Markdown
Collaborator Author

These tests at scale: 80 checkpoint groups

I ran this branch's two tests over a local bellwether recording: 80 checkpoint groups (#54's group primaries) and the corpus of every importer branch, sampled where sources are large. The fixtures were unpacked with bellwether unpack (5,785 sets, 43.6 GB plain).

Tokenizer test: 5,173,149 cases, 34,863 differ, all known

  • Encode: 2,246,887 of 2,281,750 match.
    • tinyllama differs in all 34,109 of its cases. smg puts a ▁ after </s> (bellwether#39, finding 6).
    • qwen-drive-1.0-4b differs in 754 Thai and Devanagari cases. transformers' Qwen2Tokenizer pattern lacks \p{M} (bellwether#39, finding 7).
  • Incremental decode: all 2,891,399 match.
  • sets.toml: every set's cases matched its table.

Render-parity test: 2,281,750 cases, 2,065,153 match, 216,597 differ

cause issue cases groups
chat_template.json preferred over the tokenizer's template (qwen-drive; step3, which #2820 does not name, renders the system turn twice) #2820 38,411 2
DeepSeek-V4.1 reasoning effort, 50 vs 75 bellwether#39 f5 34,106 1
same text, other ids (tinyllama) bellwether#39 f6 34,105 1
tool-definition keys smg does not model are dropped #2836 (new) 33,187 55
Current date: line missing #2819 22,571 1
no tools renders undefined, not None (olmo-3: every request without tools) #2821 17,004 1
tool-call arguments as an object where the reference keeps the string (smg does what vLLM does) bellwether#27 14,457 4
strftime_now undefined (apertus) #2819 11,559 1
+ on string and map (DeepSeek R1, V3, V3.1) #2783 10,848 3
empty content rejected with 400 bellwether#13 141 71
add_generation_prompt ignored #2780 65 65
parameters: null rendered as {} #2837 (new) 50 50
object arguments rejected bellwether#12 48 48
continue_final_message #2779 45 45

Two new defects came out of this, each with a minimal Qwen3-8B fixture that reproduces it. #2836 is the larger: it touches 55 of 71 groups through BFCL multi-turn's response and glaive's function-level required.

Found while running it

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-gateway Model gateway crate changes tests Test changes tokenizer Tokenizer related changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants