How are you all managing multiple GPUs and models across your homelab? My setup is becoming a mess

I have been running local LLMs in my homelab for a while now and I am hitting a wall with how fragmented everything is. Right now I have LM Studio on a Ryzen 9 9950X with an RX 7900 XTX, another LM Studio on an i5-10400 with an RTX 3050, and Ollama on my MacBook Pro M3 - and honestly I cannot tell you off the top of my head which model version is loaded where, or what is eating VRAM/unified memory without checking each machine individually. There is no single view across all three boxes (AMD + NVIDIA + Apple), and monitoring means SSHing in per machine. Curious how everyone else handles this: What is your stack for serving + managing models across multiple machines/GPUs? How do you keep track of which model is where and what it is doing right now? Do you bother with energy/cost tracking, or is that overkill for a homelab? Happy to share my current setup and trade notes if it is useful.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论