Qwen3.8 27B vs Qwen3.6 27B vs Qwen3.5 27B, a slight improvement in oneshotting ability across generations.

Ran Qwen3.8 27B through all 35 oneshot prompts on oneshotlm and compared them to Qwen3.6 27B and Qwen3.5 27B. Here are the results. 3.6 was not able to create a working 2048 game but 3.8 did. 3.8's pelican is accurate 3.8 was able to get Wolfenstein mostly correct whereas 3.6 did not draw anything. I used Sonnet 5 to evaluate the outputs of all these models and looks like the average rating increased across the 27B generations 2.46 -> 2.74 -> 3.00.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论