Production Qwen 3.6-27B VLLM config?

Hi everyone, I've spent about four days now trying to find the eight configuration for running Qwen 27b in production using VLLM but have been getting significant decreases in performance with the FP8 safetensors in comparison to the llama.cpp variants, especially under load/high concurrency. Our quality benchmarks drop by almost 20% in absolute terms. Would some mind sharing their VLLM config they use in production that execute correctly? I'm at a loss at this point :) Thanks in advance! (Update) Current c

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论