Embedding models benchmark for code duplication detection

For those who prefer to read the code instead of text, source code is available on GitHub . Relying only on model specification or common benchmarks, we can't predict how it would perform in a specific use case like detecting duplicated code. Focused evaluation revealed, for example, that a general-purpose model can be better than dedicated for code. Or that small model can outperform big providers.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论