glm-5.3-flash
zhipumodel ID glm-5.3-flash-51%51% OFF
glm-5.3-flash costs $0.0735 per million input tokens and $0.245 per million output tokens through the KoRouter API — 51% below Zhipu GLM's official rate, pay as you go with credits.
Call it in one request
curl https://korouter.ai/v1/chat/completions \
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flash",
"messages": [{ "role": "user", "content": "Hello" }]
}'glm-5.3-flash answers on the OpenAI-compatible endpoint, so anything that accepts an OPENAI_BASE_URL — Codex, Cline, Continue, Aider, Open WebUI, the OpenAI SDKs — needs the base URL, a key, and this model ID. Setup guides per tool.
Replace sk-your-key with a key from the console. Model IDs are bare names — see the docs.
What a real coding session costs
Per-token price tables mislead for agentic coding. A focused two-hour Claude Code session re-sends its context on every request, so 88% of its tokens are cache reads rather than fresh input — and the cache rate, not the headline input price, decides most of the bill. Priced at glm-5.3-flash's live rates, that session costs $0.12 through KoRouter, against $0.24 at Zhipu GLM's official prices.
- Cache reads · 5M tokens$0.07
- Fresh input · 250K tokens$0.02
- Output · 100K tokens$0.02
- Cache writes · 350K tokensnot billed
- Session total$0.12
The token mix is a stated assumption, not a measurement — change it and re-price every model in the cost calculator, or read the full worked example in what a real Claude Code session costs. Prices come from the live billing catalog.
The same session elsewhere
The same reference session, priced at each model's live KoRouter rates — the nearest options on either side of glm-5.3-flash.
- deepseek-v4.1-flash$0.05
- gpt-6.1-sol$1.06
- glm-5.3$1.07
- claude-haiku-4-5$1.10
More models
Prices use the live billing catalog · live availability · all models