GLM 5.3 Flash now available on AI Gateway

GLM 5.3 Flash from Z.ai is now available on AI Gateway.

The model is a faster, cheaper sibling of GLM 5.3 built for coding and agent tasks that run across many steps.

GLM-5.3 Flash is a multimodal model that supports text and vision input, with a 1M token context window and a maximum output of 128K tokens. It supports function calling, structured output, and streaming.

To use it in a coding agent, run vercel ai-gateway coding-agents setup to connect agents like Claude Code, Codex, OpenCode, Cursor, Pi, and more, then select zai/glm-5.3-flash inside the agent.

Try GLM-5.3 Flash in the model playground.

AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more.

AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. View all language models on AI Gateway.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论