Final revision of zzpanic vllm-radiance for 1 x R9700 GPU, qwen3.8-27b
Summary of project status metrics Since StillDeadcode has released his new radiance inference engine , the writing is on the wall for the old vllm-radiance builds, people have moved on to newer and better. I can see a few more single digit percent speedups available but probably won't chase them. I thought I'd spend the last 5% of effort and finish off my personal build. This borrows from quite a few people who have done work in the Launch80 Discord and is by no means my sole work. I've personally worked through a few things: - startup caching for vllm to prevent recompiling various things on each restart - a working kv-offload feature for vram-sysram-disk, with memory thinning to save storage. This makes multi-agent work possible on a single card with long context. - the occasional bugfix zzpanic/qwen3.6-vllm-gfx1201-launchers