Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s
Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10 KB goes to a file max 4 subagents, appendOnlyContext on More info in the post, happy to answer any questions.
评论
?
参与讨论