Z.ai: GLM 5.3 FlashX
Z-AI Developer Architecture Profile
Intelligence (ELO)1420Chatbot Arena Verified
Max Context Limit1,048,576Max Out: 131,072 tokens
Prompt Caching76% OFF$0.09 / 1M cached
Standard API / 1M$1.62Blended Standard Rate
Model Capabilities & Contract SLAs
- Drafting
- Classification
- Coding & Logic
- Fictional
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
Granular Pricing Matrix
Input Tokens (Standard Prompt)$0.37 / 1M
⚡ Cached Input (Prompt Cache Read)76% OFF$0.09 / 1M
Output Tokens (Completion)$1.25 / 1M
Pricing data via OpenRouter. Sync: 10/2/2026
Evaluate Competitors
VS Engine MatchupZ.ai: GLM 5.3 FlashX vs PrismML: Ternary Bonsai 2 27BVS Engine MatchupZ.ai: GLM 5.3 FlashX vs Z.ai: GLM 5.1VS Engine MatchupZ.ai: GLM 5.3 FlashX vs Mancer: Weaver (alpha)VS Engine MatchupZ.ai: GLM 5.3 FlashX vs AionLabs: Aion 3.5 MiniVS Engine MatchupZ.ai: GLM 5.3 FlashX vs Qwen: Qwen3.8 FlashVS Engine MatchupZ.ai: GLM 5.3 FlashX vs DeepSeek: DeepSeek V4 Flash Vision Exp