Gemma 4 is not your standard transformer

Gemma 4 makes five quiet departures from the standard transformer recipe. QK-norm instead of 1/√d, partial RoPE on global layers, per-layer input gating, KV sharing across layers, and an MoE that sits alongside the MLP rather than replacing it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论