Speculative decoding with the model’s own MTP head.
Three draft tokens per round, verified in one pass of the full model.
Meld Turbo runs Qwen3.8-Flash-Next — a Mixture-of-Experts model with 125B parameters, 6B of them active per token — on a single Mac, behind an OpenAI-compatible API. No cloud, no API key, nothing leaves your Mac.
The fastest published result for this model on M2-generation Macs, to our knowledge, as of October 2026.
Numbers are end-to-end through the OpenAI-compatible API with speculative decoding on. Results from other Macs are very welcome.
How they were measured| Generation | ~40–47 tok/s on code and English, ~41 tok/s on Chinese |
|---|---|
| Prompt processing | ~300 tok/s |
| Context | 128K tokens |
| Memory | ~45–48 GB, locked in RAM |
| On disk | ~68 GB (39 GB weights + 29 GB n-gram table) |
Models this size usually live on GPU clusters. A MoE model only reads a small part of its weights per token, and Apple silicon puts fast unified memory right next to the GPU.
Three draft tokens per round, verified in one pass of the full model.
The draft head scores 106K common tokens instead of all 248K, with no change in acceptance — including for Chinese.
Low-bit K-quants are integer-ALU bound when verifying several tokens on Apple GPUs. Q8_0 decodes cheaply, at a KL divergence of only 0.0026.
Kernels for 2-bit experts and for the small-batch matrix–vector products that verification is made of, as patches on llama.cpp.
Setup fetches llama.cpp, applies the patches, and builds with Metal. Model downloads and the one-time Q8_0 re-encode are in the README.
64 GB Macs may work with a smaller context (CTX=16384) and the original Q2_0 file. Untested and slower — reports welcome.
Full instructions on GitHub$ git clone https://github.com/MeldlabsAI/meld-turbo
$ cd meld-turbo
$ ./setup.sh
# Download the models and re-encode the
# dense layers (see README), then:
$ ./serve.sh
# API http://127.0.0.1:8080/v1
# Chat http://127.0.0.1:8080
Point any OpenAI-compatible client at http://127.0.0.1:8080/v1. Any API key and any model name will do. Thinking is on by default; turn it off per request when you want faster replies.
$ curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"messages": [{"role": "user",
"content": "Write a haiku about unified memory."}],
"chat_template_kwargs": {"enable_thinking": false}
}'
We would rather tell you up front. Short benchmarks are optimistic; these are the trade-offs today.
The GPU holds its top clock for a few seconds, then settles about 20% lower. Long sessions run slower than short benchmarks.
The experts are 2-bit. Quality against the original BF16 model has not been evaluated.
At ~300 tok/s of prompt processing, a long first prompt takes a while. Later turns reuse the prompt cache.
One request at a time by default.
Benchmarks from other Macs, bug reports, and patches are all welcome.
Star on GitHub