Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based "quality tested" claims — every number is reproducible from the files in the repo. 1. Hy3 — Tencent's 295B MoE (21B active) LordNeel/Hy3-GGUF Base: tencent/Hy3 — 295B MoE, 21B active, 262K context, Apache 2.0. Converted from BF16 with a Hy3-enabled llama.cpp branch, imatrix from a cust
评论
?
参与讨论