RESEARCH

Not just what we build.
How we think.

Reproducible measurements and engineering notes from our work on inference, RL post-training, agent sandboxes, and training — published with the code, including what didn’t work.

  1. INFERENCE · MELD TURBO

    Where the time goes on an M2 Max

    Why dense K-quants stall at small batches, how a smaller draft vocabulary pays off, what sustained load does to GPU clocks, and everything we tried that didn’t help.

Have a question about our numbers? Open an issue on the project, or write to us. Get in touch