Every Model That Can Be Run On 10-16GB VRAM Ranked

It's been 4 years since c.ai first hallucination model, and yet, we're nowhere good enough at LLMs in terms of spontaneity/interesting hallucination features. Fortunately, last year brought new architecture that allow those within 10-16GB VRAM range to run something usable up to 64k Context Windows via Kobold or LLAMA. Requirements: 10-16GB VRAM - 26-32GB RAM. How I Evaluated: Speed; If tokens per second are at least 10t/s. Coherency: If dialogue/monologue/POV makes any internal sense asides from basic logic (robotic feeling) or behaviorism. JP/PT-BR/US-EN: If it can handle/interpret typos, nuance or harsh verbiage. Slop Patterns: From "casting long shadows" to "smells like ozone and burnt sugar" frequency. MOE vs FULL: If rather pick a MOE instead of a FULL model. Censorship; All these models have either none or too low refusal rates even at NSFL topics, worry not. Instruction/Engineering: I've tested WITH and WITHOUT Pandora Instructions, if a model can handle 23 intertwined rules it can handle anything/not need a prompt at all. Side Note: I've ignored all CoT models, waste of tokens, slower, not needed, not at all. Do prompt nicely and it's the same. Just pick any quant higher than Q_4_K, size doesn't matter ( ͡° ͜ʖ ͡°). The Catch: I've done a small knowledge test involving animes/manga/movies/music on 70B, 120B and 300B models and found out only models above 200B actually have this sort of data, even a 120B will rarely have the information about the media you're interest at. But who can even run it? The List (Best to Worse/Why): Gemma 4 26B-A4B: Qwen 35B-A3B but with better dialogue/monologue, directive, slightly smarter/precise and funnier. Qwen 3.6 35B-A3B (Abliterated): Fast, accurate, and decently fun, yet may overfit previous inputs/output format/patterns without proper directive instructions, or speak gibberish/ask too many questions because of directive. Gemma 4 Instruct 19B-A1B Heretic (Abliterated): 26B-A4B but a bit more rigid, plain, yet keeps accuracy and logic. The Omega Directive 14B: Follows instructions, has low slop pattern rate and writes nicely, yet might become predicable after some context fill. GLM-4.7-Flash 31B-A3B (Abliterated): Performs similar to Qwen 35B-A3B but more straightforward, less interpretative of nuance/verbiage. Qwen3 30B A1.5B 64K High Speed NEO MAX: Basically 35-A3B with increased reasoning. MN CaptainErisNebula Chimera Heretic Uncensored 12B (Abliterated): Its weird, a bit rushy trying to fill everything at once, full of patterns, but its the most proactive of them all, unlike to get stuck in a loop, has good dialogue/monologue yet ain't suitable for Q&A roles. Mistral Nemo Instruct 2407 HERETIC HI Claude Opus 12B: It's decent, slop seems to have been rephrased properly, yet may overfit impersonation POV, or try to rizz you up mostly, not suitable for Q&A roles. Gemma 3 12b GLM-4.7 Flash INSTRUCT Hybrid Heretic (Uncensored): Interestingly impersonative, accurate, yet may overfit previous inputs' formatting/style with frequent slop patterns or lack humour/too prim. Ministral Instruct 2512 Absolute Heresy 14B: Mistral sucks at following complex rules (regex, status panel, commands, lorebook, etc), or turn {{char}} into a rizzler, despite being the best one regarding dialogue/monologue/spontaneity, ironically. GLM 4.6V Flash Hybrid ?B: UD but with decreased slop frequency, even less refusals. GLM 4.6V Flash UD 9B: Can feel robotic/repetitive depending on the personality of your {{char}} as making too many questions, despite high accuracy at following all PI 23 rules + regex commands. Gemma 4 E4B UD: For the desperate, it's tiny, fast, accurate and smart, but will prolly repeat itself/have slop patterns. LFM2 12B A1B SpeedDemon The Deckard II HERETIC (Uncensored): Lighting fast, highly accurate and smart, but lacks sauce. Overall Review: GLM tends to be pessimist and too logical, Qwen tends to be too inquisitive and fancy, Gemma tends to be funny and creative, Mistral tends to be flirty and oddly specific, Llama tends to be sloppy and predicable, this happens across all of their finetunes regardless if Omega, Heretic, Deckard, Dark, Horror, RP, Style, etc. Settings: Kobold can't load these, use LLAMA for QWEN 35B (20,6GB) and GLM 31B (17,7GB). Other models all running at fastest speed possible with 40-64k - FlashAttention/FastFowarding/SWA/SmartCache + MMPROJ + TTS if any in Kobold. Models I'm aware but discarded after trial cuz either too slow (massive), poor or have better alternatives: Mistral 24B, GPT 20B OSS, Kimi K2, Phi-4, Skyfall, DeepSeek, Cydonia, Magistry, Magnus Cydons, etc. Not Tested: Goetia 26B-A4B v1.7, Gemma 4 26B-A4B StyleTune V2, Promethean Dawn 26B-A4B and Fenrir-X 26B-A4B. What if I have 24GB VRAM (e,g, RTX 3090)? If going lower try Omega Darker Gaslight The Final Forgotten Fever Dream 24B, Qwen3 24B-A4B Freedom HQ Thinking Heretic NeoMAX-D_AU, Cydonia 24B, Magistry 24B, Magnus Cydons 24B, Erebus 12B Instruct 2608, Mistral RP 24B, DeepSeek 33B, Velvet Eclipse 2x12B or Velvet Eclipse 4x12B. END

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论