Engineer / developer observations of Gemma4-31B, Qwen3.8-27B, and 6.1-Sol for software engineering work

For what it's worth. I'm a mechanical engineer by education and software engineer in practice over the past ~20 years. Playing around with local models has been a recent hobby. Today I decided to do a few tests on things that I'd consider analogous to "real" work that I do, trying out some different combinations of models and harnesses. Not to make this overly formal, but going into it the idea was: Models: Gemma4-31B vs Qwen3.8-27B at the same quantization (an Unsloth flavor of Q4). Harnesses: Codex CLI vs OpenCode CLI Types of Work: "Tell me about this code" and "Let's make something new" GPT 6.1-Sol will enter the picture later. Harness differences were largely inconclusive at this time so I won't waste words on that. Repository Inspection & Analysis Using a reference repository I'm very familiar with, the prompt was effectively, "Read-only pass, tell me about such-and-such library, its core abstractions, from a consumer standpoint what I have to be cognizant of in situations X, Y, and Z, and what are the overall strengths and weaknesses of the approach. Keep it under 900 words." Qwen and Gemma both came up with equally good analyses and answers, with slightly different takes. Of note, Gemma4 did so far more efficiently than Qwen. More focused on the task as explicitly stated, with far fewer server calls (~10 compared to ~20-30 depending on harness). Qwen seemed more curious and had initiative to dig deeper than the specific request and put together a slightly more complete-picture assessment. But again, at the end, both responses were equally good. I might lean Gemma4 as a preference here just on an efficiency basis. Authoring a New Project I thought of an application that's concisely scoped, doesn't rely on a legacy codebase or significant external dependencies, and would probably take me ~1-2 days to hammer out by hand. ~8 paragraphs of prompt covering general concept, user experience in a couple different roles, design constraints and future-proofing needs. C# language with a reusable back end and WPF front end. This showed more tangible differences. Gemma4 once again seemed very efficient. Very fast planning and executed quickly. I'd give the end result a B-. It was clearly going to take a few iterations of feedback to converge on a usable solution, but that's not bad! After two passes of revisions I figured that was enough of an evaluation and put it aside. Qwen3.8 doesn't seem nearly as "smart" as Gemma4 but far more "persistent." Many more small errors as it went; missing using statements, assorted other small mistakes, like frequently tripping over its own two feet as it went, but aiming for a higher target and sticking with it. Definitely took substantially longer. First-cut product was a step better than Gemma's and it only took 1 revision to get it to where I had something I could viably use. I'd give a B+. Solid result and an easy pick over Gemma. While Qwen3.8 was doing its thing making the lights in my room flicker and dim, for grins I started up the ChatGPT desktop app, grabbed 6.1-Sol, and gave it the same prompt. Result? A+. Just an entirely different tier of functionality and polish. Really understanding the user context for interaction. One-shot solution that feels like, "Welp, that's a wrap - ship it." Definitely interesting seeing the differences between the two open-weight models I was evaluating. And for consumer hardware, I thought the results were very workable. Just clearly not in the same conversation as the frontier lab stuff. I think an interesting follow-up would be to put these to the task of real developer hell - working through old legacy code bases. There should be a benchmark for that! Maybe picking up old Doom or Command & Conquer source code will be a future test.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论