7900 XTX 24GB + RX 6800 16GB for local LLMs? Worth it with PCIe x2?
Hey guys, Just bought a PC mainly for local experimenting with AI/coding + plus the occasional gaming sesh and it hasn't even arrived yet 😅 9950X, 7900 XTX 24GB, B650 Tomahawk, 32GB DDR5 (likely going 64GB+), planning to run Linux/llama.cpp. I've noticed used RX 6800 16GB (cannot afford more )cards are still accessible, which got me curious about adding one eventually for LLMs. That would give me 40GB VRAM across the two GPUs, but the second PCIe slot on my board is only PCIe 4.0 x2. I know I can run models larger than 24GB by offloading into system RAM, so I'm less interested in whether this would simply let me "run a 70B model." What I'm really wondering is what useful model/quantization tier does going from 24GB to ~40GB GPU-resident actually unlock? Has anyone run a 7900 XTX + RX 6800 (gfx1100 + gfx1030) with llama.cpp under ROCm or Vulkan? I'd be especially interested in real-world performance with larger models/MoEs, how much the x2 link matters once the weights are loaded, and how 40GB distributed VRAM compares with just putting more RAM in the 9950X system and accepting some CPU offload. Basically: does an extra cheap 16GB AMD card materially expand what this machine can do, and if so, which models are actually worth running with it? Actual experience/benchmarks with mixed AMD GPUs would be great.