Qwen3.8-Flash-Next on Strata

Hey! 👋 I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next. Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights. Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only. github.com/Niko1221/Strata Will be happy for any feedback and pull requests you could give! 👀

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论