Mistral Large 4 gets 38 on AA vs 39 for DeepSeek V4.1 Flash — but ProofBench is 10% vs 54%. Is AA missing too much formal math?
DeepSeek (Reddit)
,DeepSeek AI 非官方 Reddit 社区,讨论模型、研究与最新进展。
关注
添加评论
点赞
收藏
点踩
分享
查看原文
评论
?
参与讨论