DeepSWE-mini a 16 instance subset of DeepSWE that replicates the rankings of the leaderboard

I just released a new dataset huggingface.co/datasets/LocalLLaMA/deepswe-mini it is meant to be more efficent when testing and benchmarking local models on DeepSwe deepswe.datacurve.ai The full benchmark takes a while to run so we analyzed it and selected a subset that can faithfully replicate the relative ranking (not absolute scores) of models. You can use it to quicikly benchmark a new model locally.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论