I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

First of all, thank you for reading. I trained a small MoE model completely from scratch (no external base weights) and wanted to share the results + a couple of lessons. Apex-2 - Architecture: Decoder-only MoE, every layer is MoE (no dense layers) - Size: 3.87B total parameters, 1.45B active per token - 32 layers, d_model 2048, GQA 16Q/4KV, 16 experts, top-4 - Context: 4096 - Tokenizer: Qwen3 (151k) - Hugging Face: huggingface.co/YOON1v/Apex-2 (loads with transformers / vLLM via Qwen3MoeForCausalLM mapping) Training - Pretrain: 86.5B tokens (GH200 ×1 → ×2 with DiLoCo ) - SFT: ~2.5B tokens (code-heavy + math + instruction) - DPO: tried it, scores dropped, so I dropped the checkpoint Key numbers (SFT, greedy, chat template) Benchmark HumanEval 43.9 HumanEval+ 41.5 MBPP 56.3 MBPP+ 48.9 GSM8K (0-shot CoT) 32.4 MATH-500 21.0 IFEval (prompt strict) 44.7 MMLU (5-shot) 28.6 interesting comparison With only ~0.087T pretrain tokens, the base model’s HumanEval+ matched Qwen2.5-1.5B (which used 18T). Knowledge (MMLU) and math still lag far behind, as expected with the data gap. What didn’t work DPO (220k pairs, length-normalized) made answers much longer and hurt code / math / IFEval. I stopped it and kept the SFT checkpoint. Full log is in the benchmark write-up. Limitations (honest) - English-centric (almost no multilingual ability) - Weak knowledge → frequent hallucinations - LiveCodeBench medium/hard is near zero - 4k context only Happy to answer questions about the MoE setup, DiLoCo, or why DPO backfired.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论