current homelab setup for local AI experimentation:

  • hermes box hosted on a Framework Desktop Mainboard AI Max+ 395
  • 5090 eGPU running Qwen 3.8 27B for fast tok/s LLM use - sometimes 150+ tok/s
  • 2x DGX Spark: running Deepseek v4 Flash 0731 - better but slower model
  • pi 5 for monitoring
  • Mac mini as a dev box - use Herdr and ohmypi/codex/claude depending on the use case
  • housed in a 10" DeskPi mini rack (mostly)

Hermes is defaulted to local AI but with a homegrown routing plugin hitting a small low TTFT model (Arch-Router) to decide whether to go local or upgrade to cloud/frontier. Trying to get to 100% local over time, but right now probably more like 60-70%

The Sparks are for batch processing background runs (all the cron jobs, longer dev builds, etc)

do you need all of this? Absolutely not lol. I started with the mac mini and couldn't help myself but to add over time!

Also I regret getting the eGPU so I wouldn't recommend that to anyone

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论