I built a 1v1 nuclear strategy game to benchmark LLM reasoning (instead of just QCMs) — Age of LLM

In 2017, I watched OpenAI Five destroy pro players at Dota 2. That moment taught me something: games are the ultimate test of emergent intelligence. Traditional benchmarks (MMLU, HumanEval, etc.) mostly measure memorization and recitation. A model can pass a coding test by regurgitating its training data. But a game? A game forces you to adapt, plan under uncertainty, deal with hidden information, and bluff. You can't fake reasoning when someone is trying to nuke you. So I created Age of LLM — Benchmark . I

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论