DeepSWE-mini a 16 instance subset of DeepSWE that replicates the rankings of the leaderboard
I just released a new dataset huggingface.co/datasets/LocalLLaMA/deepswe-mini it is meant to be more efficent when testing and benchmarking local models on DeepSwe deepswe.datacurve.ai The full benchmark takes a while to run so we analyzed it and selected a subset that can faithfully replicate the relative ranking (not absolute scores) of models. You can use it to quicikly benchmark a new model locally.
评论
?
参与讨论