Some testing on RTX Pro 4500 (With Oculink) on PrismaQuant, INT4 Autoround and NVFP4 W4A4 quantized model

This little beast has been around for a while after asking about whether it's possible to set up in this subreddit . Beelink SER 8 8745 HS, AooStar eg01, and RTX Pro 4500 32GB The Sakamakismile model I've been running was throwing tool call errors and getting stuck in thinking loops in both Opencode and Cline across vLLM 0.22, 0.23, and 0.24 for some reasons, and even swapping to the Froggeric Chat Template didn't seem to improve things (it was relatively OK for roughly 2 days of intense using, then the pro

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论