MELD TURBO · OPEN SOURCE · v0.1 · MIT

A 125B model.
On a MacBook.

Meld Turbo runs Qwen3.8-Flash-Next — a Mixture-of-Experts model with 125B parameters, 6B of them active per token — on a single Mac, behind an OpenAI-compatible API. No cloud, no API key, nothing leaves your Mac.

The fastest published result for this model on M2-generation Macs, to our knowledge, as of October 2026.

Real time, not sped up — Qwen3.8-Flash-Next (125B) writing code on a MacBook Pro M2 Max, 96 GB, fully offline.
125BParameters6B active per token
~45Tokens / second~40–47 on code and English
128KContextTokens, on one Mac
0Bytes sent outOffline after download
RESULTS

Measured end to end,
through the API.

Numbers are end-to-end through the OpenAI-compatible API with speculative decoding on. Results from other Macs are very welcome.

How they were measured
MacBook Pro 14″ · M2 Max (38-core GPU) · 96 GB · ~400 GB/s
Generation~40–47 tok/s on code and English, ~41 tok/s on Chinese
Prompt processing~300 tok/s
Context128K tokens
Memory~45–48 GB, locked in RAM
On disk~68 GB (39 GB weights + 29 GB n-gram table)
HOW IT GETS THERE

Four changes.
One usable local model.

Models this size usually live on GPU clusters. A MoE model only reads a small part of its weights per token, and Apple silicon puts fast unified memory right next to the GPU.

3 drafts / round

Speculative decoding with the model’s own MTP head.

Three draft tokens per round, verified in one pass of the full model.

3.5 → 2.4 ms

A smaller draft vocabulary.

The draft head scores 106K common tokens instead of all 248K, with no change in acceptance — including for Chinese.

+19% generation

Dense layers re-encoded to Q8_0.

Low-bit K-quants are integer-ALU bound when verifying several tokens on Apple GPUs. Q8_0 decodes cheaply, at a KL divergence of only 0.0026.

3 kernel patches

Metal kernels for verification.

Kernels for 2-bit experts and for the small-batch matrix–vector products that verification is made of, as patches on llama.cpp.

QUICK START

From clone
to local API.

Setup fetches llama.cpp, applies the patches, and builds with Metal. Model downloads and the one-time Q8_0 re-encode are in the README.

Apple silicon, Max-class GPU or better
96 GB of unified memory or more
macOS 14+, Xcode command-line tools, CMake
Python 3 with numpy, ~110 GB free disk

64 GB Macs may work with a smaller context (CTX=16384) and the original Q2_0 file. Untested and slower — reports welcome.

Full instructions on GitHub
meld-turboZSH
$ git clone https://github.com/MeldlabsAI/meld-turbo
$ cd meld-turbo
$ ./setup.sh

# Download the models and re-encode the
# dense layers (see README), then:
$ ./serve.sh

# API   http://127.0.0.1:8080/v1
# Chat  http://127.0.0.1:8080
USE IT

Bring your own client.

Point any OpenAI-compatible client at http://127.0.0.1:8080/v1. Any API key and any model name will do. Thinking is on by default; turn it off per request when you want faster replies.

Chat UI in the browser at 127.0.0.1:8080
Coding agents such as pi, via an OpenAI-compatible provider
Scripts and SDKs that speak the OpenAI API
LISTEN=0.0.0.0 to serve your local network
Agent setup on GitHub
request.shZSH
$ curl http://127.0.0.1:8080/v1/chat/completions \
    -H 'Content-Type: application/json' \
    -d '{
      "messages": [{"role": "user",
        "content": "Write a haiku about unified memory."}],
      "chat_template_kwargs": {"enable_thinking": false}
    }'
KNOWN LIMITS

What to expect.

We would rather tell you up front. Short benchmarks are optimistic; these are the trade-offs today.

Sustained load

The GPU holds its top clock for a few seconds, then settles about 20% lower. Long sessions run slower than short benchmarks.

Quality

The experts are 2-bit. Quality against the original BF16 model has not been evaluated.

Long first turns

At ~300 tok/s of prompt processing, a long first prompt takes a while. Later turns reuse the prompt cache.

Concurrency

One request at a time by default.

Credits. Built on llama.cpp and Unsloth’s MTP branch of it. The model is by the Qwen team, the GSQ-RCO quantizations are by ISTA-DASLab, the MTP head GGUF is by Unsloth, and the draft vocabulary comes from Strata. Model weights are not included and are under their own licenses. Meld Turbo is MIT licensed.
BUILD IN THE OPEN

Run it. Break it. Tell us.

Benchmarks from other Macs, bug reports, and patches are all welcome.

Star on GitHub