I’m experimenting with locally running AI
Right now, I’m experimenting with locally running AI (i.e., on my computer or graphics card). I have an Nvidia P1000 card with only 4 GB of memory, so it’s a relatively weak and outdated GPU. Even so, low-quantization models like Qwen 3.5 4Bit run locally on it. They run, but very slowly (4 tokens per second). It’s also interesting that Qwen 3.5 from Alibaba “thinks” in Chinese. That’s interesting to me, though for the Qwen developers, of course, it’s normal. I tested the llama.cpp and Docker Model Runner e
评论
?
参与讨论