Anthropic made Sonnet 5 intro pricing permanent ($2/$10) — but is the real cost-per-task still higher than alternatives? Looking for real usage data
Anthropic announced yesterday that Claude Sonnet 5’s introductory pricing is now permanent: $2 per million input tokens and $10 per million output tokens. I had been operating under the assumption that Sonnet (especially via the Claude subscription plans) was one of the more cost-effective options for heavy daily use. After seeing the reaction on X, I’m less sure about that when you measure by actual completed work rather than list price per token. Several people replied with charts from Artificial Analysis (artificialanalysis.ai), which publishes an Intelligence Index plus a derived “Cost per Intelligence Index Task” metric. Their methodology takes the total spend required to run their full evaluation suite and divides by the number of tasks, so it reflects real token consumption (including reasoning tokens, agent steps, etc.). From the charts that were circulating: GPT-5.6 Luna (max) came in around $0.05 per Intelligence Index task. Claude Sonnet 5 (max) came in around $1.72. Intelligence Index scores were close (roughly 52–55 range in the screenshots). There were also agent/coding-oriented leaderboards (again from Artificial Analysis and related public dashboards updated around Aug 7) where Sonnet 5 sat noticeably higher on average cost-per-task than several models that scored similarly or higher on the same suites. I’m not treating any single benchmark as definitive but the gap on the public cost-per-task numbers is large enough that it made me question my earlier assumption about Sonnet being the cheaper everyday option. What I’m hoping the community can help with For people running significant volume on the Claude subscription (Pro / Max / Team): have you tracked approximate cost or usage limits relative to the amount of actual work completed? Does it still feel efficient compared with OpenAI’s equivalent tiers or other providers? For API users: are you seeing the same kind of token-volume difference (especially output + reasoning tokens) that shows up in the Artificial Analysis numbers? Has anyone done their own side-by-side cost-per-successful-task measurements on realistic agentic or multi-step coding / research workflows? I’m not claiming one model is universally better. I just want to update my mental model with real usage data rather than list prices. If the public cost-per-task numbers are directionally correct for most people, that changes how I allocate work between providers. If they’re not representative of normal Claude usage, that would also be useful to know. preview.redd.it/96xy4du7qqih1.png preview.redd.it/earp95b9qqih1.png preview.redd.it/khoze0gaqqih1.png preview.redd.it/bzhkff7bqqih1.png preview.redd.it/fkzd90dcqqih1.png preview.redd.it/nlusmscdqqih1.png x.com/kienbuilds/status/2086893283100553313 x.com/imnotchalk/status/2086914433775960518 x.com/angelbrodin/status/2086962855480308032 x.com/_wannabeabaddie/status/2086893874300232124 x.com/the_alex/status/2086899599810387970