GPT-5.6 Luna vs GPT-6 Astra: is a $1.20 model good enough for code review?
we benchmarked GPT-5.6 Luna vs GPT-6 Astra on 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana Astra found 92 confirmed bugs vs 69 for Luna, but cost $5.66 vs just $0.20 also added the full eval breakdown this time: cost, avg output tokens, latency, precision, and bug classes like data/logic, security, concurrency etc. we’re doing Astra vs Fable 5.1 this week, so would appreciate feedback on the methodology before we run the next one dropping the link in the comments if anyone wants to check it out preview.redd.it/190wg4zx8iph1.png
评论
?
参与讨论