GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses.
The higher serving speed is useful for coding agents, tool loops, and interactive applications where users wait on generated output.
GLM 5.3 FlashX is now available on AI Gateway.
- 18 Sep 2026
- 1 min read
- Use zai/glm-5.3-flashx across API formats and in coding agents:
Chat Completions
5
prompt:'Name three checks to run after deploying a latency-sensitive API',
forawait(const text of result.textStream){
import{ streamText }from'ai';
const result =streamText({
model:'zai/glm-5.3-flashx',
To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway setup to create a key and configure your supported agents. Select zai/glm-5.3-flashx inside the agent.
AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, budgets for API keys, routing rules, and more.
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests.
Try GLM-5.3-FlashX in the model playground, or view all language models available on AI Gateway.