I wired 4 models together in Claude Code. It backfired 4 ways on Terminal-Bench

Common wisdom says to put a strong model like Fable in charge and let cheaper models do the work. I tested it and the result was not what I expected: seventh place at twice the cost of the top single-model run.
评论
?
参与讨论