Maple-Preview
This intrigues me quite a bit, since I’ve been looking into ternary models myself. It isn’t hard (at all) to beat Gemma 4 on quality, or to push tokens quickly through consumer hardware if you accept the quality trade-offs, but there may be a sweet spot at the intersection of quantization, model size and–more importantly–memory bandwidth where local models become genuinely useful without what is now massively expensive prosumer hardware.
As usual, I have quibbles with the benchmarks–Qwen’s MoE models, especially with MTP, are hard to beat at this scale, and Maple’s own table has Qwen3.5 35B-A3B ahead overall. I’m taking the 218 tokens/s and SOTA framing with a large pinch of salt, but this is another useful data point on the number of optimization techniques people are trying.
Update, two days later: I took a stab at getting it to run on AVX2 and it’s… OK. Nothing special, and clearly I’m hamstrung by the hardware, but it works. If they release bigger weights I’ll be ready.