Product updates
Changelog
New features, behavior changes, model additions, and deprecations. Subscribe via RSS.
Clearer setup errors and a dashboard setup banner
When a request routes to a provider you have no key for, the error now explains why and how to fix it. The new reason field is plan_byok_only (your BYOK plan uses your own key for every provider you call) or provider_byok_only (the provider is BYOK-only for everyone), with a dashboard_url to add the key. These now return **400** instead of 503, so SDKs stop auto-retrying a request that can't succeed until the key is added. 503 remains only for no_managed_key, when our own managed key is missing. If a retired model was rewritten to a different provider, requested_model shows what you originally asked for. The dashboard also shows a banner on every page while a setup step is blocking all requests.
Sarvam 105B Conversations
sarvam/sarvam-105b-conversations is Sarvam's variant of 105B tuned for real-time dialogue and voice agents. Same rate card as sarvam-105b. Sarvam 30B remains retired; requests for it continue to be served by sarvam-105b with an x-gateway-model-deprecated header, as they have been since 31 July.
Pricing API: new replaced_by field for retired models
A request for a retired model id is served and billed as its replacement. GET /platform/v1/pricing and the MCP get_pricing tool now show that directly: retired rows quote the replacement's rates and carry a new replaced_by field. The MCP cheapest_for tool recommends live models only, and never a specialised OCR or translation-only model for a chat task. Billing is unchanged.
MCP requests always run their tools fresh
Requests carrying mcp_servers now skip both the exact-match and semantic caches, for reads and writes. The answer to an MCP request depends on what your tools return at that moment, so every MCP request now runs its tools live. Requests without mcp_servers are cached as before.
Guardrail library: seven built-in rules on every plan
Seven rules ship in the guardrail library and are included in every plan rather than sold as a plugin or a per-request add-on: PII redaction, secret and credential detection, prompt-injection detection, OpenAI moderation pre-check, data-exfiltration detection, profanity filtering, and a language/region gate. Each can redact rather than only block, which is usually what you actually want. Your own per-org patterns still layer on top.
Clear error when a reasoning model runs out of tokens
Reasoning models count their internal thinking against max_tokens. If the budget is too small, the model can use all of it thinking and return no answer. The gateway now returns 400 empty_completion in that case, naming the model, the finish reason and the output tokens billed, so you know to raise max_tokens. Tool-call turns, where empty content alongside tool_calls is normal, pass through as before.
DeepSeek peak-hour rates, GPT-6 Astra parameters, long-context tiers
**DeepSeek:** request costs now apply DeepSeek's weekday peak-hour rate (01:00-04:00 and 06:00-10:00 UTC, double the off-peak price), matching what DeepSeek invoices. Requests in those windows show a higher cost than before. **GPT-6 Astra:** max_tokens is now sent as max_completion_tokens, as OpenAI requires for this model. **Long context:** Gemini 3.1 Pro and xAI models now use their long-context rates for prompts over 200k tokens.
Point Claude Code at Leanroute with one environment variable
Leanroute now speaks the Anthropic Messages API at /anthropic/v1/messages, so anything built on the official Anthropic SDKs, Claude Code included, can route through the gateway without a code change:
``
export ANTHROPIC_BASE_URL="https://api.leanroute.dev/anthropic"
export ANTHROPIC_API_KEY="gw_live_..."
export CLAUDE_CODE_ATTRIBUTION_HEADER=0
`
This is a translation layer, not a passthrough. Requests are converted into the gateway's internal format, which means caching, spend caps, guardrails, failover, per-key limits and full request logging all apply to your Claude Code traffic. It also means you can pin a cheaper model: set ANTHROPIC_MODEL="deepseek/deepseek-v4-1-flash" and Claude Code keeps working while the bill drops roughly tenfold. Streaming, tool use, images and multi-turn tool round trips are all supported.
**Routing here defaults to explicit**, unlike /v1/chat/completions. The model you ask for is the model you get. Coding sessions are long, stateful and full of tool calls, and substituting a model at turn 12 of a refactor does not degrade gracefully, so automatic cost arbitrage is opt-in on this endpoint via ANTHROPIC_CUSTOM_HEADERS="X-Gateway-Routing: auto"`.
Requests also carry Claude Code's session and subagent identifiers into the request log, so spend can be attributed to an individual coding session rather than only to a day. Built against Anthropic's published gateway compatibility guide, including keep-alive pings during long thinking pauses and verbatim upstream error forwarding so Claude Code's own retry recovery keeps working.
MCP tools now work on every model, not just Claude
Leanroute now runs the MCP agent loop itself. Send mcp_servers on a non-streaming /v1/chat/completions request and the gateway connects to your MCP servers, discovers their tools, namespaces them, injects them as ordinary function tools, executes the calls, and loops until the model is done. Because MCP tools become plain function declarations on the way out, **any provider with function calling now supports MCP through Leanroute** — GPT, Gemini, Groq, DeepSeek, and the rest, not only Anthropic. Every tool call is logged with its server, duration, and payload sizes, and token usage is aggregated across all hops so you are billed for the whole run rather than the last turn. Response headers report the tools discovered, tool calls made, iterations used, and any MCP server that failed to connect. MCP requests run non-streaming.
MCP: one gateway-run path for every provider
The earlier Claude-only MCP path, which handed mcp_servers to Anthropic, has been replaced by the gateway-run loop above, so every tool call is now logged, costed and covered by guardrails. Three changes you may notice. **One:** mcp_servers with stream: true now returns 400 mcp_streaming_unsupported on every provider including Anthropic, where it previously streamed. Set stream: false and the gateway runs the loop for you. **Two:** the x-gateway-mcp-passthrough response header is gone; it had become a duplicate of x-gateway-mcp-allowlist, and x-gateway-mcp-mode is now always proxy. **Three:** the error code mcp_unsupported_provider_streaming is replaced by mcp_streaming_unsupported, since the restriction is no longer provider-specific.
Tool calling: full JSON Schema on Gemini, multi-turn Gemini 3, OpenAI reasoning models
**Gemini tool schemas:** Gemini accepts only a subset of JSON Schema. Schemas using additionalProperties, $ref, allOf or const are now converted to what Gemini accepts (const becomes a single-value enum), so the same tool definitions work on Gemini, OpenAI and Anthropic. Schemas that are already valid pass through unchanged. **Gemini 3 multi-turn tool use:** Gemini 3 requires each function call's thoughtSignature to be echoed back on the next turn. The gateway now preserves it and exposes it as tool_calls[].thought_signature in streaming and non-streaming responses, so you can round-trip it in your own agent loop. **OpenAI reasoning models with tools:** for gpt-5*, o1* and o3*, the gateway sets reasoning_effort: "none" when you send tools without setting it yourself, which OpenAI requires for tool calls on these models. If you set reasoning_effort explicitly, your value is sent as-is.
Added Claude Fable 5.1 and DeepSeek V4.1 Flash
Weekly catalog sweep. Two new SKUs added: **Claude Fable 5.1** (2026-09-01, $10/$50, cache reads dropped to $0.25 from $1.00 — a 75% cache-hit reduction on the same headline price) and **DeepSeek V4.1 Flash** (2026-09-10, $0.15/$0.60 with native multimodal, replaces V4 Flash and V4 Flash Vision Exp). DeepSeek's V4.1 launch also drops customer bills across the whole DeepSeek Flash line: old slugs (deepseek-chat, deepseek-v4-flash, deepseek-v4-flash-vision-exp) now auto-route to V4.1 Flash for a ~30% input / ~9% output price cut plus proper multimodal support (no more 384-token per-image cap). GPT-6 Astra's catalog entry now notes its 272K context limit.
Open-sourced opa-sidecar-a2a: chain-aware authorization reference implementation
Published opa-sidecar-a2a v0.1.0-alpha under MIT at github.com/leanroute/opa-sidecar-a2a. Single Go binary + OPA policy engine + reference Rego policies + worked planner-executor demo. Not a Leanroute-only feature; anyone building an A2A stack can adopt it. The announcement post covers the why and the code.
xAI Grok: model ids mapped to xAI's current naming
xAI's own model ids use dots (grok-build-0.1, grok-4.6, grok-4.20-*). Leanroute keeps its stable hyphenated ids (xai/grok-build-0-1, xai/grok-4-6) and now maps them to xAI's naming on the way out, across grok-build-0-1 and the full grok-4.x family. No change needed on your side.
`openai/gpt-6-astra` added (new OpenAI flagship)
OpenAI's new flagship above the GPT-5.6 line, launched 2026-09-03. Premium-reasoning tier at $10 in / $50 out per 1M tokens — matches Anthropic Fable 5's $-band. Cache-hit assumed 90% off; verify on first invoice.
`google/gemini-3.8-flash` added
Google's newest Flash, launched 2026-09-02. Same $0.75 / $3.75 intro rate as 3.6 and 3.7 Flash; all three revert to $1.50 / $7.50 on 2027-01-01. Recommended default for new integrations per Google.
`google/gemini-3.1-pro` added
Google's hard-reasoning / agentic-coding Pro SKU at $2 in / $12 out per 1M tokens. Flagship tier — competes with anthropic/claude-sonnet-5 and openai/gpt-5.6-terra in the $2-band contest.
`qwen/qwen3.8-flash` added (cheapest Qwen tier)
Alibaba's newest cheap_fast SKU, launched 2026-08-26. $0.14 in / $0.42 out per 1M tokens. Sits below qwen3.8-max ($2/$6.4) in the same generation.
`kimi/kimi-k2.5` retired (Moonshot 2026-08-31 sunset)
Moonshot retired K2.5 alongside the Moonshot V1 series on 2026-08-31. Requests now auto-rewrite to kimi/kimi-k2.6 (closest live general-purpose replacement, 256k context). Bill impact: K2.5 was $0.60 / $3.00, K2.6 is $0.95 / $4.00 — modest rate bump.
OpenAI catalog cleanup: gpt-4o, o1, gpt-5-nano, gpt-5-mini deprecated
OpenAI dropped gpt-4o, gpt-4o-mini, o1, gpt-5-nano, and gpt-5-mini from their public pricing page. The current catalog is gpt-5.6-sol / terra / luna / cyber + gpt-5.3-codex. Requests to the retired ids now auto-rewrite: gpt-4o → gpt-5.6-terra, gpt-4o-mini → gpt-5.6-luna, o1 → gpt-5.6-cyber, gpt-5-nano → gpt-5.6-luna, gpt-5-mini → gpt-5.6-luna. Watch for bill spikes: the old nano at $0.05/$0.40 was 3-4× cheaper than luna at $0.20/$1.20 — cost-sensitive workloads should switch to deepseek/deepseek-v4-flash ($0.22/$0.66).
`glm/glm-5.3` added (Z.AI flagship)
Z.AI's latest flagship, $1.40 in / $4.40 out, 81% cache-hit discount. Joins the 5.1/5.2 cluster at the same price point — 5.1/5.2 remain in the catalog for version pinning.
`glm/glm-5.3-flash` added (Z.AI cheap_fast)
New cheap_fast tier from Z.AI, list price $0.15 in / $0.50 out (80% cache-hit discount). Currently on a 50%-off promotion effective through 2026-09-09 24:00 SGT — during the promo effective rates are $0.075 / $0.25. Catalog encodes the list price; when the promo ends no code change is needed.
`glm/glm-5-turbo` retired (Z.AI catalog change)
Zhipu removed glm-5-turbo from docs.z.ai/guides/overview/pricing between our 2026-08-11 addition and today's sweep. Requests now auto-rewrite to glm/glm-5 (same flagship tier, next-generation sibling). Bill impact: glm-5 is $1/$3.20 vs glm-5-turbo was $1.20/$4.00 — customers see a small SAVING (~16%) on the rewrite. No proactive migration required.
AWS Bedrock (BYOK) — Anthropic Claude in your own AWS account
Route bedrock/anthropic.claude-sonnet-4-5 and bedrock/anthropic.claude-haiku-4-5 through your own AWS credentials. Inference runs inside your AWS boundary in the region you choose (Singapore ap-southeast-1 at launch, cross-region inference profiles for US/EU), so the data never leaves your account. Leanroute meters usage and applies routing, cache, and guardrails on the way through, but AWS bills you directly for the tokens. Non-streaming only in v1 — streaming lands in a follow-up. BYOK setup lives on /dashboard/providers.
`bedrock/anthropic.claude-sonnet-4-5` added
Claude Sonnet 4.5 on Bedrock, priced at Bedrock's public list rate ($3 / $15 per M tokens, 90% cache-hit discount). Requires a Bedrock BYOK row for your org; AWS bills you directly for inference.
`bedrock/anthropic.claude-haiku-4-5` added
Claude Haiku 4.5 on Bedrock at $1 / $5 per M tokens, 90% cache-hit discount. Same BYOK path as Sonnet — your AWS account, your region, your bill for the tokens.
OpenAI GPT-5.6 Sol price cut (~20% in, 33% out)
OpenAI dropped openai/gpt-5.6-sol from $5 / $30 to $4 / $20 per million tokens effective 2026-08-22. Cache-write and cached-input rates fell proportionally. Promo runs at least through 2026-11-21; if it ends, the standard rate returns to $5 / $30 and we'll refresh the catalog again. Blended 5:1 cost drops from $9.17 to $8.00 per 6M tokens.
Claude Sonnet 5 stays at $2 / $10 — Sept 1 bump cancelled
Anthropic made Sonnet 5's introductory pricing permanent on 2026-08-11 and cancelled the scheduled 2026-09-01 increase to $3 / $15. Sonnet 5 is now the permanent cheap-flagship option in the Anthropic lineup; Sonnet 4.6 stays at $3 / $15 as a higher-priced legacy pin.
Google Gemini 3.7 Flash added
google/gemini-3.7-flash launched 2026-08-13 with intro pricing $0.75 / $3.75 per M-tokens through 2026-12-31 (reverts to $1.50 / $7.50 in January). 1,048,576-token input window, 65,536 output. Google also cut gemini-3.6-flash to match — the two are at price parity for now, with 3.7 recommended for new integrations.
xAI Grok 4.6 added
xai/grok-4-6 launched 2026-08-12 as xAI's current flagship: $2 / $6 short context, $4 / $12 above 200K tokens, 500K-token window, cache-hit 75% off. Grok 4.5 stays in the catalog for customers who want to pin the previous flagship.
DeepSeek V4 Flash Vision (experimental)
deepseek/deepseek-v4-flash-vision-exp added — DeepSeek's experimental multimodal Flash variant. Same off-peak $0.22 / $0.66 as text v4-flash; images bill as tokens capped at 384 per image. Cheapest vision-capable model in the catalog by a wide margin.
DeepSeek peak-hour multiplier now Mon–Fri only
DeepSeek moved to weekday-only peak pricing effective 2026-08-16. Weekend requests (Sat + Sun UTC) always bill at off-peak rates now, regardless of hour. Peak windows remain 01–04 + 06–10 UTC on Mon–Fri. Billing calc and routing scorer updated to honor the new schedule.
Meta Muse Code added
meta/muse-code (2026-08-05 release) added at the standard Muse rate card: $1.25 / $4.25, cache-hit $0.15 (88% off). Meta's coding-agent SKU — persistent background workers, positioned as a general agent rather than an autocomplete completion. Contributor tier excluded per our no-data-sharing policy.
OpenAI GPT-5 Mini added
openai/gpt-5-mini at $0.25 / $2.00 per M-tokens — sits between gpt-5-nano ($0.05 / $0.40) and gpt-5.6-luna ($0.20 / $1.20) in the cheap_fast tier. Was missing from our catalog; added on customer request. Cache-hit rate assumed to follow OpenAI's Aug-2026 family-wide 90%-off policy.
Sarvam-105B repriced — 7× input, 4.6× output
⚠ Sarvam raised sarvam/sarvam-105b prices sharply effective 2026-08-05: input ₹4 → ₹29.28, output ₹16 → ₹73.2 per million (roughly $0.048 → $0.349 in and $0.19 → $0.871 out). Cache-hit discount is now 62.5% (was 37.5%), which softens the blow for cache-heavy workloads. Still cheaper than most Chinese flagship models on blended cost, but the gap has narrowed. Calculator + landing math now use the new numbers.
OpenAPI 3.1 spec + AI catalog + MCP server card
Three new agent-discovery endpoints: /openapi.json (also at /.well-known/openapi.json) documenting every stable API surface, /.well-known/ai-catalog.json with a machine-readable descriptor of the whole service (providers, models, endpoint URLs, auth schemes, capability tags), and /.well-known/mcp/server-card.json for MCP registries that need a static manifest fallback.
MCP over HTTP at api.leanroute.dev/mcp
The Leanroute MCP surface is now available via Streamable HTTP transport in addition to the stdio npm package. Point your MCP client at https://api.leanroute.dev/mcp with an Authorization: Bearer gw_admin_* header — zero install, no local subprocess. Same eight discovery/management tools as the stdio version (route_call remains stdio-only because it needs a runtime key alongside the admin bearer).
Groq Llama SKUs removed
Groq discontinued groq/llama-3.1-8b-instant and groq/llama-3.3-70b-versatile on their inference platform. Requests to these model ids will now 404 from the gateway. Migrate to groq/openai/gpt-oss-20b or groq/openai/gpt-oss-120b, both LPU-served and priced comparably.
Agent-first surfaces: MCP server, admin API, llms.txt
Full agent-discovery stack shipped in one release: the @leanroute/mcp-server npm package (also listed in the official MCP Registry as dev.leanroute/mcp-server), a new /platform/v1/* REST API for programmatic account control (list keys, check usage, get credit), a dashboard admin-keys page (owner-only) for minting gw_admin_* tokens, and /llms.txt + /llms-full.txt so agents doing web research find us structured-first.
Two-class API keys (runtime vs admin)
Every API key now has a class: gw_live_* (runtime; dispatches LLM traffic) and gw_admin_* (admin; manages the account). Runtime keys cannot hit /platform/v1/*; admin keys cannot dispatch LLM traffic. Neither class can top up credit or upload BYOK provider keys — those stay behind the dashboard. This is the invariant that makes it safe to hand an admin key to an agent.
Groq added as provider #13
LPU-served OpenAI-format inference for openai/gpt-oss-20b and openai/gpt-oss-120b. Sub-100ms first-token latency on both.
Meta Model API added as provider #12
Muse Spark 1.1 and 1.2 available at standard-tier pricing. Contributor tier not supported — it trains on customer prompts, incompatible with our privacy commitments.
Meta Muse Spark 1.1 and 1.2
Both checkpoints priced $1.25 / $4.25 per M-tokens (input / output), with 88% cached-input discount.
Workspace setup on first sign-in is now atomic
Your personal workspace is created exactly once on first sign-in, even if several tabs open at the same moment.
`/v1/models` now returns models from all 13 providers
A hardcoded reachability check was only scanning 5 providers, silently filtering out models from qwen, glm, doubao, kimi, sarvam, krutrim, meta, and groq even when their keys were configured. Now iterates the full provider set.
OpenAI gpt-5.6-cyber and gpt-5.3-codex
gpt-5.6-cyber for security-analysis workloads ($12.50 / $75). gpt-5.3-codex for code-generation ($1.75 / $14).
Anthropic Claude Fable 5, Opus 4.5, Sonnet 4.5
Fable 5 ($10 / $50) is Anthropic's new flagship long-context model. Opus 4.5 ($5 / $25) and Sonnet 4.5 ($3 / $15) are quality bumps to the existing tiers.
OpenAI price drops absorbed
gpt-5.6-luna dropped from $1.00 / $6.00 to $0.20 / $1.20 — a 5x reduction on input, 5x on output. gpt-5.6-terra now $2.00 / $12.00 (down from $2.50 / $15.00). No change to gateway markup; the savings flow through to your bill directly.
DeepSeek price hike absorbed (v4-flash +57%, v4-pro +52%)
DeepSeek raised published prices for v4-flash ($0.14 → $0.22 input, $0.28 → $0.66 output) and v4-pro ($0.435 → $0.66 input, $0.87 → $1.98 output). Our arbitrage router now routes around DeepSeek during their UTC peak-hour window (01–04 + 06–10 UTC, 2× surcharge) when a cheaper same-tier alternative exists.
Cross-provider failover shipped (#23)
When the primary provider returns 5xx, the gateway transparently retries on a same-tier alternative. Logs stamp failover=true + failover_from_provider so the dashboard can show which requests survived an outage. Cost math stays accurate — billing uses whichever provider actually answered.