1-bit / 2-bit / Ternary / Bitnet Models - Updates & Tracking

Bonsai / Ternary Bonsai During April Bonsai came with bunch of models .... 1-bit & 1.58-bit(Ternary) versions. And last month(July) they released 27B models in same versions. Last month itself, Bonsai-27B was able to run on all backends mainline. But Ternary-Bonsai-27B was not ready on all backends. This month, PRs got merged for CUDA & Vulkan on mainline. Also an Optimization PR for CUDA got merged so +15-40% tg, +8% pp. github.com/PrismML-Eng/Bonsai-demo - Demo fork github.com/PrismML-Eng/llama.cpp - Custom fork BitCPM-CANN huggingface.co/collections/openbmb/bitcpm-cann Tencent - Hy-MT1.5 - Mixed Meta Translation Model Version 1.5 huggingface.co/collections/tencent/hy-mt15 Maple-Preview DeepGrove/maple-preview - 20B-A1B - 200+ t/s on Mac Mini M4 & 120+ t/s on iPhone. llama.cpp PR #27000 - CPU backend github.com/deepgrove-ai/llama.cpp - Custom llama.cpp fork github.com/deepgrove-ai/mlx-lm-deepgrove - Custom MLX fork Mach-1-Additive-35B Mach-1-Additive-35B - A3B - Up to 120 t/s on Consumer Laptop. Mach-1-Additive-35B-Multimodal github.com/SyzygyResearch/llama.cpp-mach1 - Custom llama.cpp fork From their recent tweet : Currently they're cooking new ones based on Laguna-S-2.1 & Qwen3.8-27B . Neutrino-8B huggingface.co/FermionResearch/Neutrino-8B github.com/fermionresearch/llama.cpp - Custom llama.cpp fork Pestle-27B-Ternary huggingface.co/Doses-AI/Pestle-27B-Ternary-GGUF - Medical research model github.com/DosesAI/mortar.cpp - Custom inference - CPU, CUDA, Metal Image Models: huggingface.co/collections/prism-ml/bonsai-image huggingface.co/Green-Sky/bonsai-image-ternary-4B-GGUF huggingface.co/Green-Sky/bonsai-image-binary-4B-GGUF huggingface.co/clark-labs/clark-air-sana-1.6b-1.58bit Abliterated Models: huggingface.co/Hikari07jp/Ternary-Bonsai-27B-Abliterated-LowDeg-GGUF huggingface.co/Hikari07jp/Maple-Preview-TQ2-Abliterated huggingface.co/dealignai/Bonsai-27b-1bit-CRACK-GGUF Other Misc items: huggingface.co/GoAutomateAI/terna-e2b-GGUF huggingface.co/Danny-Dasilva/Bonsai-27B-antidoom-1bit-DSpark huggingface.co/Danny-Dasilva/Ternary-Bonsai-27B-antidoom-DSpark Some Open/Ongoing llama.cpp (related) PRs: ggml-cpu : add STQ1_0 ternary quantization with ARM NEON vec_dot kernel- #22836 ggml/cpu: skip zero-scale blocks in TQ1_0 and TQ2_0 vec_dot kernels- #23439 ggml-cpu: add x86 VNNI Q2_0 dot product -- 3x speed improvement for VNNI-compatible CPUs- #26348 Notes: Didn't include old models(Before 2026). Let me know if I missed any models, I'll update thread. Included custom forks to check their progress. I'll be updating this thread after seeing any similar type models. Disclaimer : This thread is mainly for Poor GPU Club .

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论