Zhipu

glm-5.3-flash

zhipumodel ID glm-5.3-flash-51%

glm-5.3-flash costs $0.0735 per million input tokens and $0.245 per million output tokens through the KoRouter API — 51% below Zhipu GLM's official rate, pay as you go with credits.

Call it in one request

curl · glm-5.3-flash
curl https://korouter.ai/v1/chat/completions \
  -H "Authorization: Bearer sk-your-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

glm-5.3-flash answers on the OpenAI-compatible endpoint, so anything that accepts an OPENAI_BASE_URL — Codex, Cline, Continue, Aider, Open WebUI, the OpenAI SDKs — needs the base URL, a key, and this model ID. Setup guides per tool.

Replace sk-your-key with a key from the console. Model IDs are bare names — see the docs.

What a real coding session costs

Per-token price tables mislead for agentic coding. A focused two-hour Claude Code session re-sends its context on every request, so 88% of its tokens are cache reads rather than fresh input — and the cache rate, not the headline input price, decides most of the bill. Priced at glm-5.3-flash's live rates, that session costs $0.12 through KoRouter, against $0.24 at Zhipu GLM's official prices.

  • Cache reads · 5M tokens$0.07
  • Fresh input · 250K tokens$0.02
  • Output · 100K tokens$0.02
  • Cache writes · 350K tokensnot billed
  • Session total$0.12

The token mix is a stated assumption, not a measurement — change it and re-price every model in the cost calculator, or read the full worked example in what a real Claude Code session costs. Prices come from the live billing catalog.

The same session elsewhere

The same reference session, priced at each model's live KoRouter rates — the nearest options on either side of glm-5.3-flash.

More models

Prices use the live billing catalog · live availability · all models