DeepSeek V4, GLM-5.3 and Kimi K3 in Claude Code: one config block, and what a 2-hour session costs on each

7 October 2026 · prices as listed on the KoRouter marketplace on that date

Claude Code doesn't have to run on Claude. It talks to any endpoint that speaks the Anthropic protocol, and a gateway that lists other vendors' models behind that protocol lets you point the same tool at DeepSeek V4, GLM-5.3 or Kimi K3 with one config block. This post covers the setup for Claude Code, Cline, Aider and Continue, what to expect from each model, and what the same two-hour session costs on all of them.

Prices below are USD per million tokens, with the vendor's official price first.

Why bother

Three reasons people do this:

  • Cost. DeepSeek V4.1 Flash is $0.15 input / $0.60 output at DeepSeek's own off-peak list price (double that during weekday peak hours), and $0.066 / $0.264 through KoRouter at any hour. A two-hour Claude Code session that costs $10.97 on Claude Fable 5 costs about seven cents on V4.1 Flash (the numbers are below).
  • One balance. The models are billed in USD on the same KoRouter balance as Claude and GPT — no separate vendor account, no second top-up.
  • Task fit. Tests, refactors, docs, log triage and other high-volume, low-stakes work are where a cheap model with a 1M-token context earns its place. Keep the frontier model for the decisions that matter.

The models

ModelInputOutputCache readvs official
deepseek-v4.1-flash$0.15 → $0.066$0.60 → $0.264$0.003 → $0.0013256%
glm-5.3-flash$0.15 → $0.0735$0.50 → $0.245$0.03 → $0.014751%
glm-5.3$1.40 → $0.714$4.40 → $2.24$0.26 → $0.132649%
kimi-k3$3.00 → $1.74$15.00 → $8.70$0.30 → $0.17442%

Notes that matter for agentic use:

  • DeepSeek V4.1 Flash and GLM-5.3 Flash take a 1M-token input context and support tool calling (tools, tool_choice), response_format and reasoning parameters — everything Claude Code needs.
  • None of the three vendors has a cache-write price: tokens that miss the cache are billed as ordinary input, at the input price. The session costs below count cache writes that way.
  • The discount is one multiplier per model, applied to input, output and cache reads alike.

Claude Code

Add the block to ~/.claude/settings.json (Windows: %USERPROFILE%\.claude\settings.json). The same keys also work as environment variables.

~/.claude/settings.json
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://korouter.ai",
    "ANTHROPIC_AUTH_TOKEN": "sk-your-key",
    "ANTHROPIC_MODEL": "deepseek-v4.1-flash",
    "CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
    "CLAUDE_CODE_ATTRIBUTION_HEADER": "0"
  }
}

Verify:

Verify
claude -p "Reply with OK"

Two things people get wrong: the key goes in ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY; and ANTHROPIC_MODEL takes the bare id from the catalog — deepseek-v4.1-flash, glm-5.3-flash, glm-5.3, kimi-k3 — with no vendor prefix.

To switch models per project, keep a second settings file and launch with claude --settings path/to/settings.json, or just change ANTHROPIC_MODEL and restart.

Cline, Aider, Continue

These speak the OpenAI protocol, so they use the /v1 base URL and the same bare model ids.

Aider
export OPENAI_API_BASE="https://korouter.ai/v1"
export OPENAI_API_KEY="sk-your-key"
aider --model openai/deepseek-v4.1-flash
~/.continue/config.yaml
models:
  - name: KoRouter deepseek-v4.1-flash
    provider: openai
    model: deepseek-v4.1-flash
    apiBase: https://korouter.ai/v1
    apiKey: sk-your-key

Cline — pick the OpenAI Compatible provider, set the base URL to https://korouter.ai/v1, paste the key, and enter deepseek-v4.1-flash as the model id.

Codex uses the Responses API and is documented with GPT models on the integrations page; it's not covered here.

What the same session costs

The reference workload is the one from our cost guide: a focused two-hour Claude Code session on a mid-size repository — 5M cache-read tokens, 250K fresh input, 100K output, 350K cache writes. Priced at KoRouter rates:

ModelSession cost
deepseek-v4.1-flash$0.07
glm-5.3-flash$0.14
glm-5.3$1.32
claude-sonnet-5$2.19
kimi-k3$2.78
claude-fable-5$10.97

A hundred and forty sessions of V4.1 Flash cost less than one session of Fable 5. That doesn't make V4.1 Flash the better model — it makes it the right model for the 80% of work where the cheaper answer is good enough, with the expensive one a config change away.

What to expect

  • Tool calling works; Claude Code's file edits, shell commands and searches go through as tool calls the same way they do on Claude.
  • Output style differs between vendors. Give the model the same CLAUDE.md you'd give Claude and expect to tighten instructions for the first hour.
  • Long sessions still benefit from /clear between tasks. Cache reads are cheap on every model here, but a 1M context window doesn't mean you want to fill it.
  • Availability history for each model is public on the status page — check it before you move a workload.

Disclosure

I run KoRouter, the gateway used in the config above. One key covers Claude, GPT, DeepSeek, GLM and Kimi; every model on the marketplace shows the vendor's official price next to ours, and one multiplier applies to all price components. Free to sign up, no monthly fee — pay as you go with credits. The config pattern works with any gateway that lists these models behind the Anthropic and OpenAI protocols; the prices are ours.