peculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B

I have a long list of my own domain specific eval questions that I run to validate which models I can rely on: coding, coding (numpy/pandas), data analytics decision making, local RAG, and voice assistant. It's made up of the types of things I'm likely to deal with on the daily. The test questions vary in dififculty and composition: easy, medium, hard. System: M1 Max 32c 32GB with context 128K for Dirk and 110K for Swift. Swift doesn't have XL. So I had to test with the L quant to stay as close as possible. I ran the eval (using my tuieval tool) on peculiar-ragdoll's Dirk-Qwen3.8-27B-UD-Q4_K_XL and Swift-1.5-Qwen3.8-27B-Q4_K_L loaded with a modified version of Splash. The "amalgam" is a local I made out of incoai/Splash 1.1 and paperniuk's apple7-m1-kernels. It is modififed a little but not in ways that would alter model performance. I only merged and tweaked for some memory features I like from llama.cpp such as fit context check at the start of a load and personal QoL updates re auto-context manipulations that I don't want to think about, etc. To say this result surprised me is quite an understatement. It's blown my mind. When I did the first test a couple of days ago with only 44 questions, I thought it must be a prompt caching issue I missed that Dirk was benefiting from. I validated it is not and ran it against a lot more questions to certify it. It's a legit test outcome. Dirk-Qwen is much sharper at getting to decisions and responses. The "be brief" instruction that gets passed each time in the chat templste is doing more magic than I had anticipated. It also gets more answers correctly with way less time consumed. What trips up Swift-1.5 are mostly hard questions. It tries and tries until the 16,384 max token limit per question is reached and it fails with truncation. Even when you ignore the 16,384 truncation failures and compare the other questions, Dirk token usage comes out on top. Snipped view... this basically goes on pattern for another 191 unique questions. preview.redd.it/otg3d1rkbqsh1.png More importantly, this behavior is not just in question answering. You can see it in actual code refactor tasks. On an unrelated note: tne model that has been able to pass a 100% of my eval packs is Opus 5.5. Deepseek Flash 4.1 fp32 got them all right except three.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论