Tested 15 local models for agent/tool use.. Bonsai 27B was last!
Tested 15 local models for agent/tool use — Bonsai 27B was last I wanted to see which local models are actually usable for agent/tool-use tasks, so I ran 15 of them through the same benchmark instead of guessing from model hype. The benchmark was Toolery . It doesn't just check whether a model can call a tool.. the scenarios require the model to actually complete tasks under different constraints. 143 scenarios 3 trials per scenario 429 trials per model Easy, Medium, Hard, and Very Hard tiers Everything served locally through LM Studio No API costs 30k token context window for every model Temperature 0.8 for every model Otherwise I kept the default settings that each model came with in LM Studio The exception was qwen3.8-flash-next , where I was using the Strata , but I still set the temperature to 0.8 Concurrency 4 Timeout scale 4.0 Results: # Model Score Easy Medium Hard Very Hard 1 qwen/qwen3.8-27b 71.8% 96.7 93.3 71.6 41.7 2 mellum2-12b-a2.5b-thinking 71.4% 88.3 91.1 71.6 45.8 3 google/gemma-4-26b-a4b 70.4% 92.5 93.3 60.8 50.0 4 granite-4.2-8b 67.8% 85.8 88.9 62.7 45.8 5 qwen3.8-flash-next-iq2_xs 64.5% 89.2 91.9 62.7 30.6 6 mistralai/devstral-small-2-2512 63.9% 75.8 83.7 58.8 45.8 7 ornith-1.5-9b 62.3% 84.2 83.7 60.8 34.7 8 lfm2.5-8b-a1b 61.6% 72.5 74.1 55.9 51.4 9 ornith-1.5-35b-a3b 61.0% 85.8 84.4 57.8 31.9 10 meta/muse-glimmer 60.0% 87.5 92.6 52.0 26.4 11 qwen3.8-27b-gsq-rco 59.0% 84.2 85.9 51.0 31.9 12 gpt-oss-20b 58.6% 81.7 86.7 53.9 27.8 13 google/gemma-4-12b-qat 56.7% 84.2 83.0 46.1 31.9 14 qwen3-coder-30b-a3b-instruct 54.3% 60.0 73.3 51.0 37.5 15 prism-ml/bonsai-27b 50.5% 65.8 64.4 58.8 22.2 preview.redd.it/hc91j4zu6ith1.png So yeah, Bonsai was last overall . It was also last on the Very Hard scenarios, with only 22.2% . But the more interesting part is why it failed. The failure mode Bonsai had 148 failed trials. 103 of those were budget_violated . That means the model often selected the correct tool and got the correct result, but then made one more tool call than the scenario allowed. Only 30 failures were classified as wrong-tool-choice failures. So the problem was not always: I have no idea which tool to use. It was more like: I know what I am doing, let me call this one more time. And then it violated the tool budget. That's actually the part I found most interesting. The model can sometimes execute the task correctly, but it doesn't know when to stop. For comparison: qwen3.8-27b ranked first with 71.8% qwen3.8-27b had 84 failed trials total granite-4.2-8b had the fewest failures overall, with 69 Speed was also not great Bonsai was the slowest model in this run. Total wall-clock time was 6,208 seconds . The next slowest model took around 5,750 seconds , while scoring about 17 points higher. The qwen flash model was around 9 seconds per scenario . Important caveat This is the original Bonsai 27B , not the newer Bonsai 2 27B . The compression itself is still interesting. Getting a 27B-class model into roughly 4 GB with the 1-bit version is pretty impressive, and the model does retain a lot of the original capability. But this benchmark is only testing a specific thing: How does this particular model behave when it has to complete tool-use/agent tasks under constraints? It is not a general intelligence benchmark, and it doesn't prove that Bonsai is a bad model. The result is more specific: The compression is impressive, but the original 1-bit Bonsai 27B had surprisingly weak agent/tool-use behavior in this test. Especially when it came to knowing when to stop. I would also like to run Bonsai 2 through the exact same benchmark. Since it is based on the newer Qwen3.8 27B family, that comparison should be much more interesting than comparing the original Bonsai against everything without separating the versions. Limitations This was: one benchmark one harness one run per model 30k context window temperature 0.8 mostly the default LM Studio settings for each model one adapter configuration local serving with quantization and context settings that were not perfectly normalized So treat this as a snapshot.