i would like to share my experience. working with huge LLMs and poor Machine
hello people i wanted to share my experience with big and huge models (usually 100B+ models and 200B+ models and more) my laptop specs is very poor I7-8750H 20G Ram GTX 1050 Mobile 4G Vram but what nearly saved me is my NVMe from samsung i have 512G NVMe from samsung and yes as you expected. i run these huge models while throwing most of the parameters in my NVMe but i strictly use MoE models, Dense models will kill my machine always used mmap. and throwing experts in my CPU with Quantized KV Cache (Q4_0) a
评论
?
参与讨论