Distributed Local Agents Benchmark.

I wanted a setup where I can compare all the harnesses with all the local models. It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware. Everyone can do their share, but no one can pretend to do all possible tests extensively. To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. airbench.ai/leaderboard My hope is to turn this into a fully distributed Agent Benchmark. Let me know what you think.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论