Research|LLM: Kimi K3 - Scaling Still Works; An Expensive Model Competing at Front Tier

Research|LLM: Kimi K3 - Scaling Still Works; An Expensive Model Competing at Front Tier 图片 1

Kimi K3’s overall performance now sits in the global top tier. It is a large 2.8-trillion-parameter MoE model with a 1-million-token context window, and it stands out in tests spanning coding, agents, and long-horizon knowledge work. The first round of global user feedback has also broadly confirmed its capabilities, especially in complex coding and long-running task execution. However, the model has only just been released, and its real-world stability, success rates, and per-task cost still need further validation.

K3 shows once again that scaling laws have not broken down - larger models still generally deliver better absolute performance. Architectural optimization, MoE, distillation, and post-training can raise the efficiency with which parameters and compute are used, but for further gains in complex reasoning, long-horizon planning, and agent capability, increasing model size remains the most direct path. After K3 again scaled up substantially from the K2 series, its capabilities took a clear step up; this suggests architectural innovation is mostly improving scaling efficiency rather than replacing scaling itself.Subscribe nowRead more

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论