NEW · MELD TURBO 0.1 — A 125B MODEL ON YOUR MAC

Frontier AI.
Yours to run.

Meld Labs is an AI research lab building infrastructure for inference, RL post-training, agent sandboxes, and training. Run it on hardware you control — from a single MacBook to GPU servers in your company’s private cloud. Your models and data stay yours. Our managed cloud is next.

01 / OUR FIRST RELEASE

A 125B model.
On a MacBook.

Meld Turbo runs Qwen3.8-Flash-Next — a Mixture-of-Experts model with 125B parameters, 6B of them active per token — on a single Mac, behind an OpenAI-compatible API. No cloud, no API key, nothing leaves your Mac.

Parameters
125B
Tokens / second
~45
Context
128K
Real time, not sped up — Qwen3.8-Flash-Next (125B) writing code on a MacBook Pro M2 Max, 96 GB, fully offline.
03 / NOTES FROM THE FRONTIER

Research worth sharing.

All research
EFFECTIVE BANDWIDTH · 4-TOKEN MAT-VEC · Q3_K → Q8_0
MELD TURBO · PERFORMANCE NOTES

Where the time goes
on an M2 Max.

Why dense K-quants stall at small batches, how a smaller draft vocabulary pays off, and everything we tried that didn’t help.

Read the report
LET’S BUILD WHAT COMES NEXT

Bring your big idea.

Investors, collaborators, and builders — we’d love to hear from you.

Get in touch