Anthropic SDK · Python + TS

Anthropic SDK + Leanroute

Leanroute exposes an Anthropic-shaped endpoint at /anthropic/v1/messages (parallel to the OpenAI-shaped /v1/chat/completions). Point the official Anthropic SDK's base_url at https://api.leanroute.dev/anthropic and every SDK call routes through us without a wire-format translation. Same messages array, same tool_use / tool_result shapes, same content blocks.

Run Claude Code through Leanroute.

Claude Code speaks Messages API and nothing else, so this endpoint is what lets you put a gateway in front of an existing Claude Code session without touching any code. Three environment variables:

export ANTHROPIC_BASE_URL="https://api.leanroute.dev/anthropic"
export ANTHROPIC_AUTH_TOKEN="gw_live_YOUR_KEY"
export CLAUDE_CODE_ATTRIBUTION_HEADER=0

Use ANTHROPIC_AUTH_TOKEN rather than ANTHROPIC_API_KEY: the first overrides a saved Claude subscription login immediately, the second stops to ask for approval. The third variable is required because Claude Code prepends an attribution block that only survives an unmodified system array, and a gateway routing to multiple providers has to reshape it.

This is a translation layer, not a passthrough, so caching, spend caps, guardrails, failover and full request logging all apply to your Claude Code traffic. Set ANTHROPIC_MODEL="deepseek/deepseek-v4-1-flash" and the same session runs on DeepSeek at roughly a tenth the cost, with tools and streaming intact. Routing here defaults to explicit — the model you ask for is the model you get — because swapping providers midway through a long coding session does not degrade gracefully.

1. Get a gateway key

From /dashboard/keys create a key labeled "anthropic-sdk". Copy the gw_live_* value.

2. Python: chat

pip install anthropic
from anthropic import Anthropic

client = Anthropic(
    api_key="gw_live_YOUR_KEY",
    base_url="https://api.leanroute.dev/anthropic",
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=200,
    messages=[{"role": "user", "content": "Hello in one sentence."}],
)
print(msg.content[0].text)

Note the model name is claude-sonnet-5 (Anthropic's native form), NOT the canonical anthropic/claude-sonnet-5 form we use on the OpenAI-shaped endpoint. That's deliberate — the Anthropic wire is passed through unchanged, so upstream model strings apply.

3. Python: streaming

with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=200,
    messages=[{"role": "user", "content": "Write a haiku about caching."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

4. Python: serve the same code from a cheaper model

The Anthropic SDK only requires that the response is Messages-shaped, not that Anthropic produced it. Point model at a fully-qualified id from any provider and the same code keeps working.

# Pin a non-Claude model and the same code runs on it. The Anthropic
# SDK doesn't care what serves the request, only that the response is
# Messages-shaped — which is Leanroute's job.
#
# A bare id like "claude-sonnet-5" gets the anthropic/ prefix. A
# fully-qualified id is honoured as-is, which is the escape hatch:
# one string change moves the workload to a tenth of the cost.
msg = client.messages.create(
    model="deepseek/deepseek-v4-1-flash",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Refactor this function for readability."}],
)
print(msg.content[0].text)

# Check x-gateway-provider on the response to confirm who served it.
# The body deliberately echoes the model you ASKED for, because the
# Anthropic SDK asserts on that field.

For MCP tool servers, use the OpenAI-shaped /v1/chat/completions endpoint with an mcp_servers array — the gateway runs the tool loop there on any provider. Detailed guide at /docs/mcp.

5. TypeScript

npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: "gw_live_YOUR_KEY",
  baseURL: "https://api.leanroute.dev/anthropic",
});

const msg = await client.messages.create({
  model: "claude-sonnet-5",
  max_tokens: 200,
  messages: [{ role: "user", content: "Hello in one sentence." }],
});

console.log((msg.content[0] as { text: string }).text);

Same model-pinning trick applies: set model to a fully-qualified id such as deepseek/deepseek-v4-1-flash and the TS SDK keeps working unchanged.

Why route Anthropic through us at all?

Fair question — if you're only calling Claude, direct is fine. Reasons to route via Leanroute:

  • Cross-provider failover: when Anthropic 529-overloads (which they do), Leanroute retries on the cheapest same-tier alternative (e.g. Sonnet 4.6 → GPT-4o / DeepSeek v4-flash) transparently.
  • Per-org MCP allowlist: control which MCP servers your team can invoke centrally, without pushing config to every dev.
  • Unified billing: one Stripe subscription for all your model providers instead of a separate one per vendor.
  • Guardrails library: the same 7 curated rules (PII, secrets, prompt injection, moderation, data exfil, profanity, language) apply to Anthropic traffic without a separate integration.

Troubleshooting

  • 401 unauthenticated: use your gw_live_* Leanroute key, NOT your Anthropic sk-ant-* key.
  • 403 mcp_server_not_allowlisted: the MCP URL is not on your org's allowlist. Add it at /dashboard/mcp. The response header lists which URL failed the check.
  • 404 or wrong base URL: the Anthropic endpoint lives at /anthropic (no /v1 in the base URL — the SDK appends /v1/messages itself).
  • Model name rejected: use the Anthropic-native form (claude-sonnet-5, claude-opus-5), not the anthropic/* canonical form.