Self-hosting LLMs on budget hardware: general principles, hardware, benchmarks and frontends

Hello, I've been self-hosting LLMs on various budget hardware for a while (6x RTX 3060 12 GB, Intel Arc Pro B60 24 GB, RX 9070 XT, etc). Over the last few months, I wrote about it in 4 articles: General principles Hardware and inference optimization CPU+RAM offloading, MoE, prefill speed and benchmarks (or "Why some influencers sell unrealistic use cases") Frontends and example of complete configuration I hope it may be useful to some people :-)

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论