Aftermarket Harnesses

The harness now moves the coding benchmark more than the model does. Endor Labs' Agent Security League found GPT-5.5 scored 61.5% functional correctness in Codex & 87.2% in Cursor, & Claude Opus 4.7 scored 87.2% in Claude Code & 91.1% in Cursor. Input tokens are 86-98% of OpenRouter volume, so the harness controls most of the bill through cache discipline. First-party co-design buys real cache hit rates, but a third-party harness can match them.
评论
?
参与讨论