Skip to content

Call for testers: benchmark qwisp on your Apple Silicon Mac (8/16GB + slow-NAND MacBooks especially wanted) #38

Description

@penta2himajin

qwisp runs Qwen3.6-35B-A3B (MoE) locally on Apple Silicon, streaming expert weights from flash so the model fits far beyond your RAM — with a bit-exact lossless mode as the reference. It works today, but almost all published numbers come from one dev machine (M1 Max / 64GB) plus SSD-throttle approximations. We need real hardware.

What we need most

hardware why
8 GB / 16 GB Macs (any M-series) the streaming tiers (C=64 / C=128) — only throttle-approximated today
MacBook Air / base MacBooks with 256 GB SSD slow-NAND floor (~1.5 GB/s reads) — the tier we can least reproduce
M2 / M3 / M4 of any config GPU-generation scaling
anything else every row helps — resident 32GB+ included

Unstable output is wanted data too. If a row says **LOOPY** (0.31), please post it — run-to-run stability variance on streaming tiers is exactly what we're trying to map. There is no wrong result.

How (~15 min, mostly download)

brew install penta2himajin/qwisp/qwisp
qwisp pull        # downloads the model (~20 GB) + writes config
qwisp benchtest   # ~1-2 min; prints a markdown report

At the end, benchtest prints a one-click URL — cmd+click it in the terminal and the issue form opens pre-filled. Just press Submit. (Or paste the report into a new benchtest report manually.)

Please use v0.3.3 or later (qwisp version) — it reports the numeric stability ratio, so all rows share one format.

Notes:

  • Plug into AC power if you're on a laptop (battery throttles the GPU; benchtest records which).
  • The benchmark is greedy/deterministic and needs no account or telemetry — you see exactly what you post.

Results so far

chip RAM disk read tier / mode code-256 nl-256 long-600 provenance report
M1 Max (32c GPU) 64 GB 3.5 GB/s resident · strict 87.1 tok/s 91.5 85.2 dev machine #36
your Mac here community

Community rows are aggregated automatically into bench/RESULTS.md (median decode/TTFT + LOOPY rate per hardware group) — your submitted issue is ingested by a bot PR. Dev-measured and throttle-approximated figures live in the README Performance table.

Questions / problems

If qwisp pull or benchtest fails outright, open a regular issue with the error output — install failures are as valuable as benchmark rows right now.

Thank you! 🙏


Discussion thread on r/LocalLLM: https://www.reddit.com/r/LocalLLM/comments/1uwa85m/i_built_a_singlemodel_engine_that_streams_moe/

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    call-for-testersRequests for community hardware testingparkedWorkstream paused/on hold — infra shipped, not actively worked

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions