Measuring PCIe transfer under dual GPU with pipeline & tensor llama.cpp

Hello all, sharing some data points: In this setup there are 2 cards connected direct to motherboard via PCIE 3 16x slots. Running llama.cpp. Ubuntu 24.04, single xeon motherboard The cards are 1x RTX 3090 24GB power limit 250W and 1x Titan RTX 24GB power limit 225W nvidia-smi with Qwen3.6-27B-UD-Q4_K_XL.gguf loaded at 180k context: preview.redd.it/t28n6nxiooch1.png My test today was comparing this setup using tensor parallel

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论