DeepSeek V4, GLM-5.3 and Kimi K3 in Claude Code: one config block, and what a 2-hour session costs on each
7 October 2026 · prices as listed on the KoRouter marketplace on that date
Claude Code doesn't have to run on Claude. It talks to any endpoint that speaks the Anthropic protocol, and a gateway that lists other vendors' models behind that protocol lets you point the same tool at DeepSeek V4, GLM-5.3 or Kimi K3 with one config block. This post covers the setup for Claude Code, Cline, Aider and Continue, what to expect from each model, and what the same two-hour session costs on all of them.
Prices below are USD per million tokens, with the vendor's official price first.
Why bother
Three reasons people do this:
- Cost. DeepSeek V4.1 Flash is $0.15 input / $0.60 output at DeepSeek's own off-peak list price (double that during weekday peak hours), and $0.066 / $0.264 through KoRouter at any hour. A two-hour Claude Code session that costs $10.97 on Claude Fable 5 costs about seven cents on V4.1 Flash (the numbers are below).
- One balance. The models are billed in USD on the same KoRouter balance as Claude and GPT — no separate vendor account, no second top-up.
- Task fit. Tests, refactors, docs, log triage and other high-volume, low-stakes work are where a cheap model with a 1M-token context earns its place. Keep the frontier model for the decisions that matter.
The models
| Model | Input | Output | Cache read | vs official |
|---|---|---|---|---|
| deepseek-v4.1-flash | $0.15 → $0.066 | $0.60 → $0.264 | $0.003 → $0.00132 | 56% |
| glm-5.3-flash | $0.15 → $0.0735 | $0.50 → $0.245 | $0.03 → $0.0147 | 51% |
| glm-5.3 | $1.40 → $0.714 | $4.40 → $2.24 | $0.26 → $0.1326 | 49% |
| kimi-k3 | $3.00 → $1.74 | $15.00 → $8.70 | $0.30 → $0.174 | 42% |
Notes that matter for agentic use:
- DeepSeek V4.1 Flash and GLM-5.3 Flash take a 1M-token input context and support tool calling (tools, tool_choice), response_format and reasoning parameters — everything Claude Code needs.
- None of the three vendors has a cache-write price: tokens that miss the cache are billed as ordinary input, at the input price. The session costs below count cache writes that way.
- The discount is one multiplier per model, applied to input, output and cache reads alike.
Claude Code
Add the block to ~/.claude/settings.json (Windows: %USERPROFILE%\.claude\settings.json). The same keys also work as environment variables.
{
"env": {
"ANTHROPIC_BASE_URL": "https://korouter.ai",
"ANTHROPIC_AUTH_TOKEN": "sk-your-key",
"ANTHROPIC_MODEL": "deepseek-v4.1-flash",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
"CLAUDE_CODE_ATTRIBUTION_HEADER": "0"
}
}Verify:
claude -p "Reply with OK"Two things people get wrong: the key goes in ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY; and ANTHROPIC_MODEL takes the bare id from the catalog — deepseek-v4.1-flash, glm-5.3-flash, glm-5.3, kimi-k3 — with no vendor prefix.
To switch models per project, keep a second settings file and launch with claude --settings path/to/settings.json, or just change ANTHROPIC_MODEL and restart.
Cline, Aider, Continue
These speak the OpenAI protocol, so they use the /v1 base URL and the same bare model ids.
export OPENAI_API_BASE="https://korouter.ai/v1"
export OPENAI_API_KEY="sk-your-key"
aider --model openai/deepseek-v4.1-flashmodels:
- name: KoRouter deepseek-v4.1-flash
provider: openai
model: deepseek-v4.1-flash
apiBase: https://korouter.ai/v1
apiKey: sk-your-keyCline — pick the OpenAI Compatible provider, set the base URL to https://korouter.ai/v1, paste the key, and enter deepseek-v4.1-flash as the model id.
Codex uses the Responses API and is documented with GPT models on the integrations page; it's not covered here.
What the same session costs
The reference workload is the one from our cost guide: a focused two-hour Claude Code session on a mid-size repository — 5M cache-read tokens, 250K fresh input, 100K output, 350K cache writes. Priced at KoRouter rates:
| Model | Session cost |
|---|---|
| deepseek-v4.1-flash | $0.07 |
| glm-5.3-flash | $0.14 |
| glm-5.3 | $1.32 |
| claude-sonnet-5 | $2.19 |
| kimi-k3 | $2.78 |
| claude-fable-5 | $10.97 |
A hundred and forty sessions of V4.1 Flash cost less than one session of Fable 5. That doesn't make V4.1 Flash the better model — it makes it the right model for the 80% of work where the cheaper answer is good enough, with the expensive one a config change away.
What to expect
- Tool calling works; Claude Code's file edits, shell commands and searches go through as tool calls the same way they do on Claude.
- Output style differs between vendors. Give the model the same CLAUDE.md you'd give Claude and expect to tighten instructions for the first hour.
- Long sessions still benefit from /clear between tasks. Cache reads are cheap on every model here, but a 1M context window doesn't mean you want to fill it.
- Availability history for each model is public on the status page — check it before you move a workload.
Disclosure
I run KoRouter, the gateway used in the config above. One key covers Claude, GPT, DeepSeek, GLM and Kimi; every model on the marketplace shows the vendor's official price next to ours, and one multiplier applies to all price components. Free to sign up, no monthly fee — pay as you go with credits. The config pattern works with any gateway that lists these models behind the Anthropic and OpenAI protocols; the prices are ours.