Jev vs. Kev: open-source Jev alternative tested side by side

We hosted Kev 4B (Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B) and ran it side by side with Jev on the same endpoint to see how it compares. We built a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues), with answers taken from the source. A few findings: - Accuracy lands within 2 points on every task, inside the noise at this sample size - Jev is better calibrated and pulls ahead on paraphrase detection (PAWS 87.0% vs 74.5%) - Same list price, but Jev counts a fixed ~257 extra input tokens per request (same count calling TypeSafe directly), so short requests cost up to 12x more Benchmark code, test items and results are on GitHub if you want to run your own. Both models routed via my startup Opper. Happy to dig into specifics.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论