Testing Fable 5, Opus 4.8, GPT-5.6, and more through playable 3D games

TL;DR at the end I wanted a way to evaluate models around something I care about and I think we’ll see more and more as we move to “world models“, which is spatial, temporal, and causal coherence in a 3D space. Meaning, does the model understand where things are, stay consistent over time, and when something happens, do the consequences make sense? Those qualities are hard to capture with static/benchmark questions, and I think games are the perfect vehicle for testing them. So I built WorldBuild Bench. In

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论