I wired 4 models together in Claude Code. It backfired 4 ways on Terminal-Bench

I wired 4 models together in Claude Code. It backfired 4 ways on Terminal-Bench 图片 1

Common wisdom says to put a strong model like Fable in charge and let cheaper models do the work. I tested it and the result was not what I expected: seventh place at twice the cost of the top single-model run.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论