My eval said a perfect MCP server was broken. It was the eval that was lying.

Three rounds of calibration, four real servers, $0.60 in API costs — how I made an LLM-powered tool-selection benchmark fair enough to publish, and what it revealed about static lint rules predicting live model behavior.
评论
?
参与讨论