Are you running Qwen 3.8 27b or Qwen Flash Next?
Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx? Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: blog.jetbrains.com/junie/2026/08/qwen-for-junie Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.
评论
?
参与讨论