Better Models: Worse Tools

A rare case where I want to put a post of mine here. Reason being that I was hunting down tool calling behavior regressions with the latest generation of Anthropic models and I found the resulting behavior both puzzling and quite problematic. Those models appear to be strongly RL'ed on their own Claude Code harness which is closed source, and when you come close in tool declarations but slightly off, you can now expect to get broken tool call behavior when older models did not yet have that defect.

Thought this might be interesting for folks here.

Comments

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论