what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn?
we have an HPC cluster with 70 nodes, each with an A30 (roughly a 3090, same 24GB VRAM) and 1TB DDR4 RDIMM system RAM. I feel like I saw a "twin 3090 and 128GB RAM" recipe around here recently that could do tensor parallel and I can't find it anymore.
评论
?
参与讨论