M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context

I would like to stay at q4, and serve the model to various coding harnesses. None of the models or tools I have tried will run without running out of memory. I am hoping for some advice. oMLX? MTPLX? llama.cpp? Has anyone gotten a good Qwen 3.8 27B running ok on an M1 Max 32GB MacBook Pro?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论