DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??

I ran the newly released DeepSeek-V4-Flash-0731 in my oneshot eval harness across 34 prompts and here are the results. oneshotlm.com/model/deepseek-deepseek-v4-flash-0731 The providers on openrouter were unstable and I had to retry generation multiple times. Surprisingly it costed $1.29 to go through all 34 prompts failing to produce 5 outputs whereas kimi k3 only costed $0.44 without any failures. DeepSeek V4 Flash 0731: 2.7/5 score, $1.29 cost, 753k tokens Kimi K3: 3.2/5 score, $0.44 cost, 233k tokens Am I doing something wrong?? How is your experience with this model compared to Kimi K3?

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论