Benchmark your custom Pi tools

A few people here mentioned interest in a way to test their custom Pi setups, so I figured I’d drop this here: RoastMyHarness The basic idea is a small engine that sets up an environment to run DeepSWE benchmark tasks using bare Pi as a control and a variant of your choice, your Pi harness, an extension, a skill, AGENTS.md file, etc. I used a Pi extension to have a wizard set it up for you so its easy. run the same coding tasks against bare Pi and your modified harness, then look at what actually got better, what broke, and what it cost. Since I like to play around with custom tools, I use this to get direction as to what is and isn't working. I know a lot here are making cool tools so I figured some might be interested in using it to help fine tune theirs. Any Pi users might be interested in figuring out if their tools are working like they think they should. You'd be surprised how hard it is to beat base Pi when it comes to task quality / token efficiency. Its a WIP. I do data analytics by trade but otherwise a vibe coder and I only really tested it on Linux.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论