How can i limit reasoning effort on the qwen3.5 and gemma4 models?

So, umhh, I am working on an agentic coding platform, and I need to make qwen3.5 and gemma4 models out of controlled reasoning chains. For example, at low, the model should prioritize finding the quickest solution and prioritize speed, and at high and xhigh, the model should absolutely hit that limit it has and figure out the question no matter what. I mean... I can use DeepSeek v4 flash and do something with it, because it is more controllable from a system prompt. I find it easy to control the 12B gemma4'

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论