What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?
Hi, I'm wondering what settings you are using in order to run Qwen3.8-Flash-Next on your devices? I'm especially interested in setups with 96GB VRAM. I'm not quite sure if llama.cpp does offload the embeddings to RAM or disk with my settings. I would like to offload them to RAM in order to avoid too much performance penalty. These are the settings I use and which work the best at the moment: [qwen3.8-flash-next] model = /mnt/kyouma/1TB/ML/models/unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf mmproj = /mnt/kyouma/1TB/ML/models/unsloth/Qwen3.8-Flash-Next-GGUF/mmproj-F16.gguf split-mode = layer flash-attn = on load-mode = none lazy-mode = auto fit = on fit-ctx = 262144 cache-ram = 94208 parallel = 1 temp = 1.0 top-p = 0.95 top-k = 20 min-p = 0.00 repeat-penalty = 1.0 presence-penalty = 0.0 I get about 15t/s on tg and 100-200t/s on pp at a context size of 130000. My setup: 3x AMD MI50 (32GB) 512GB DDR4 RAM 2x Intel E5-2683 v4 I'm using ROCm. But Vulkan has similiar speed, maybe a little bit slower.