Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s

Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10 KB goes to a file max 4 subagents, appendOnlyContext on More info in the post, happy to answer any questions.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论