DeepSeek-V4-Flash on Ollama Cloud Pro enters infinite thinking loops / repetitive CoT with high reasoning effort in Agent frameworks (Pi / DeepSeek Harness)

Environment & Setup Provider: Ollama Cloud Pro Model: DeepSeek-V4-Flash (0731 checkpoint) Client / Frameworks: Pi Coding Agent, DeepSeek Harness Settings: reasoning_effort: high The Issue When setting reasoning_effort to high , the model frequently enters an infinite loop of overthinking and repetitive generation. The internal thinking stream ( blocks) continuously repeats the same phrases and verification loops. The final response repeats identical text chunks indefinitely until it hits the max token cutoff. This occurs consistently across both Pi Coding Agent and DeepSeek Harness during coding and multi-step reasoning tasks. Does anyone have any solutions?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论