qwisp runs Qwen3.6-35B-A3B (MoE) locally on Apple Silicon, streaming expert weights from flash so the model fits far beyond your RAM — with a bit-exact lossless mode as the reference. It works today, but almost all published numbers come from one dev machine (M1 Max / 64GB) plus SSD-throttle approximations. We need real hardware.
What we need most
| hardware |
why |
| 8 GB / 16 GB Macs (any M-series) |
the streaming tiers (C=64 / C=128) — only throttle-approximated today |
| MacBook Air / base MacBooks with 256 GB SSD |
slow-NAND floor (~1.5 GB/s reads) — the tier we can least reproduce |
| M2 / M3 / M4 of any config |
GPU-generation scaling |
| anything else |
every row helps — resident 32GB+ included |
Unstable output is wanted data too. If a row says **LOOPY** (0.31), please post it — run-to-run stability variance on streaming tiers is exactly what we're trying to map. There is no wrong result.
How (~15 min, mostly download)
brew install penta2himajin/qwisp/qwisp
qwisp pull # downloads the model (~20 GB) + writes config
qwisp benchtest # ~1-2 min; prints a markdown report
At the end, benchtest prints a one-click URL — cmd+click it in the terminal and the issue form opens pre-filled. Just press Submit. (Or paste the report into a new benchtest report manually.)
Please use v0.3.3 or later (qwisp version) — it reports the numeric stability ratio, so all rows share one format.
Notes:
- Plug into AC power if you're on a laptop (battery throttles the GPU; benchtest records which).
- The benchmark is greedy/deterministic and needs no account or telemetry — you see exactly what you post.
Results so far
| chip |
RAM |
disk read |
tier / mode |
code-256 |
nl-256 |
long-600 |
provenance |
report |
| M1 Max (32c GPU) |
64 GB |
3.5 GB/s |
resident · strict |
87.1 tok/s |
91.5 |
85.2 |
dev machine |
#36 |
| your Mac here |
|
|
|
|
|
|
community |
|
Community rows are aggregated automatically into bench/RESULTS.md (median decode/TTFT + LOOPY rate per hardware group) — your submitted issue is ingested by a bot PR. Dev-measured and throttle-approximated figures live in the README Performance table.
Questions / problems
If qwisp pull or benchtest fails outright, open a regular issue with the error output — install failures are as valuable as benchmark rows right now.
Thank you! 🙏
Discussion thread on r/LocalLLM: https://www.reddit.com/r/LocalLLM/comments/1uwa85m/i_built_a_singlemodel_engine_that_streams_moe/
qwisp runs Qwen3.6-35B-A3B (MoE) locally on Apple Silicon, streaming expert weights from flash so the model fits far beyond your RAM — with a bit-exact lossless mode as the reference. It works today, but almost all published numbers come from one dev machine (M1 Max / 64GB) plus SSD-throttle approximations. We need real hardware.
What we need most
Unstable output is wanted data too. If a row says
**LOOPY** (0.31), please post it — run-to-run stability variance on streaming tiers is exactly what we're trying to map. There is no wrong result.How (~15 min, mostly download)
At the end, benchtest prints a one-click URL — cmd+click it in the terminal and the issue form opens pre-filled. Just press Submit. (Or paste the report into a new benchtest report manually.)
Please use v0.3.3 or later (
qwisp version) — it reports the numeric stability ratio, so all rows share one format.Notes:
Results so far
Community rows are aggregated automatically into
bench/RESULTS.md(median decode/TTFT + LOOPY rate per hardware group) — your submitted issue is ingested by a bot PR. Dev-measured and throttle-approximated figures live in the README Performance table.Questions / problems
If
qwisp pullorbenchtestfails outright, open a regular issue with the error output — install failures are as valuable as benchmark rows right now.Thank you! 🙏
Discussion thread on r/LocalLLM: https://www.reddit.com/r/LocalLLM/comments/1uwa85m/i_built_a_singlemodel_engine_that_streams_moe/