Molt — a ~9K-line, PyTorch-native RL framework for agentic post-training that scales to hundred-B MoE

TL;DR — We built Molt , an agentic-first, PyTorch-native RL post-training framework. It's about 9K lines of core RL code you can actually read end to end , yet it trains MoE models from Qwen3.5-397B up to GLM-5.2 753B-scale . The whole stack is just Ray + vLLM + NVIDIA AutoModel/FSDP2 — no Megatron . Code: github.com/NVIDIA-NeMo/labs-molt Tech report (DOI): 10.13140/RG.2.2.23375.65447 License: Apache-2.0 The problem If you've done agentic RL on a large model, you know the drill: you open the framewo

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论