Qwen3.8-27B FP8 dual GPUs

Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB) = 48 GB VRAM 32 GB DDR5 6200 RAM Arch Linux, KDE on the 5090 (takes ~1-1.5 GB VRAM) daily driver OrcaRouter Qwen3.8-27B Uncensored, block-FP8 (28.75 GiB), on vLLM 0.30.0 with pipeline parallel across both cards 4070 = rank 0 with layers 0-19 + vision encoder, 5090 = rank 1 with layers 20-63 + lm_head + MTP. Full 262,144 context, MTP K=3, fp8 KV, 2 slots. Benchmarks (vLLM 0.30.0, vllm bench serve) 1K in / 512 out, c1, random 81.3 tok/s, ITL 36.7 ms (~100 tok/s on code) MTP acceptance ~87% 32K prefill 3,160 tok/s 145K fresh prefill 68.5-68.8 s 258K prefill 169 s, 5090 peak 31,472 MiB The 4070Ti Super was collecting dust for 2 months because i thought Pcie x1 will slow it but apparently not

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论