Ox Alpha was GLM-5.3-Flash all along 👀
Remember Ox Alpha? The anonymous model that showed up on OpenRouter/OpenCode on Aug 20 with a 1M context window, free pricing, and zero attribution just "Stealth" as the provider . In one week it processed 7T+ tokens across ~134K developers with nobody knowing who built it. Z.ai just claimed it: GLM-5.3-Flash, and dropped full open weights under MIT. Specs: 320B total params, only 18B active per token (MoE). Hybrid sparse + linear attention with mHC. (Manifold-Constrained Hyper-Connections), built for cheap long-context serving.1M context, up to 128K output tokens. First native multimodal model in the GLM-5 line — text, image, video, trained on a 30T-token multimodal corpus. Day-1 support for vLLM, SGLang, KTransformers. Reportedly ~1/10th the inference cost of GLM-5.2, with launch pricing at 1/20th of GLM-5.2