I built a 1v1 nuclear strategy game to benchmark LLM reasoning (instead of just QCMs) — Age of LLM
In 2017, I watched OpenAI Five destroy pro players at Dota 2. That moment taught me something: games are the ultimate test of emergent intelligence. Traditional benchmarks (MMLU, HumanEval, etc.) mostly measure memorization and recitation. A model can pass a coding test by regurgitating its training data. But a game? A game forces you to adapt, plan under uncertainty, deal with hidden information, and bluff. You can't fake reasoning when someone is trying to nuke you. So I created Age of LLM — Benchmark . I
评论
?
参与讨论