Head to head: Kimi-K2.7-Code vs grok-4.3
This one is as close as the aggregate score suggests, but the task sheet tells a more useful story: Kimi-K2.7-Code was the steadier model across the set, while grok-4.3 won fewer categories and mostly on narrower formatting-faithfulness calls. The statistical edge belongs to Kimi-K2.7-Code, and the reasons are concrete
评论
?
参与讨论