Skip to content

Commit b2fda84

Browse files
authored
docs: correct the server-dense threshold count in the final measurement (#2231)
## Summary One-line correction to `docs/benchmark_results/unified-engine-final-gb10-2026-10-08.md` (PR #2226). The server engine at B=1 with dense storage measured -1.44, -1.14, -1.24 and -2.05 percent against the pre-epic CLI. That misses the -1.0 percent threshold in all four cells, not in three as the summary line said. The table was already correct. The follow-up is #2227. ## Test plan - [x] Text-only change; the table values are unchanged Part of #2166
1 parent ad84435 commit b2fda84

1 file changed

Lines changed: 1 addition & 1 deletion

File tree

‎docs/benchmark_results/unified-engine-final-gb10-2026-10-08.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -23,7 +23,7 @@ Paired per-round decode tok/s delta against `base-cli` (median, range over 5 rou
2323
| Llama-3.2-1B 4-bit | 8192 | +0.19 % (+0.03..+0.82) | -2.05 % (-2.25..-1.50) | +0.06 % (-0.36..+0.42) |
2424

2525
- **`mlxcel generate` passes in every cell.** The engine client decodes at least as fast as `CxxGenerator`; three of the four cells clear the null range upward.
26-
- **The server engine at B=1 with dense storage misses the threshold in three of four cells.** That is the path `mlxcel run` and the chat REPL take since #2173 (one slot, dense). Phase 0 measured the same gap before the epic (server dense within 1.0 to 1.5 percent of the CLI), so it is the scheduler's per-tick overhead rather than something the epic added. But `run` used `CxxGenerator` before #2173, so for `run` users this is a 1 to 2 percent decode regression. Follow-up needed.
26+
- **The server engine at B=1 with dense storage misses the threshold in all four cells.** That is the path `mlxcel run` and the chat REPL take since #2173 (one slot, dense). Phase 0 measured the same gap before the epic (server dense within 1.0 to 1.5 percent of the CLI), so it is the scheduler's per-tick overhead rather than something the epic added. But `run` used `CxxGenerator` before #2173, so for `run` users this is a 1 to 2 percent decode regression. Follow-up needed.
2727
- **TTFT at 256 tokens on `generate`**: +4.4 % on Qwen3 (null range -3.9..+10.9, unresolved) and +9.0 % on Llama (+4.2..+12.7 against a null of -4.4..+8.3), about 1 ms. At 8192 tokens TTFT is unchanged (-0.4 %).
2828

2929
Medians (tok/s, TTFT ms):

0 commit comments

Comments
 (0)