Is ternary (1.58-bit) LLMs making a come back?

I'm just thinking, ever since microsoft announced bitnet, this sub (and myself) has been hoping for massive ternary models. In the last month alone, prismML dropped 27B ternary (though I've read community experience suggested it sometimes didn't hold up to it's benchmarks), Deepgrove dropped their ternary maple-20b-a1b which from my experience works really well and clocks like 100 tok/s on an iphone, and Doses AI dropped pestle-27b-ternary medical specialised which beats medgemma-27b nearly across the board. The common problem across all of them is long-horizon agentic coding/work, but i really think that's because all of these are new small labs that haven't prioritised RL-maxxing yet - they have indicated this is their next step though. I'm hopeful, and it seems like we could be very close to a massive ternary model MoE that's actually competitive with qwen3.8 at coding and agentic work. Or have most folks lost faith in ternary architecture?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论