current homelab setup for local AI experimentation:
- hermes box hosted on a Framework Desktop Mainboard AI Max+ 395
- 5090 eGPU running Qwen 3.8 27B for fast tok/s LLM use - sometimes 150+ tok/s
- 2x DGX Spark: running Deepseek v4 Flash 0731 - better but slower model
- pi 5 for monitoring
- Mac mini as a dev box - use Herdr and ohmypi/codex/claude depending on the use case
- housed in a 10" DeskPi mini rack (mostly)
Hermes is defaulted to local AI but with a homegrown routing plugin hitting a small low TTFT model (Arch-Router) to decide whether to go local or upgrade to cloud/frontier. Trying to get to 100% local over time, but right now probably more like 60-70%
The Sparks are for batch processing background runs (all the cron jobs, longer dev builds, etc)
do you need all of this? Absolutely not lol. I started with the mac mini and couldn't help myself but to add over time!
Also I regret getting the eGPU so I wouldn't recommend that to anyone
评论
?
参与讨论