Put every AI call in your company on a budget
One endpoint for every model and coding agent, with a spending limit for each person, team and customer.
- OpenAI
- Claude
- Gemini
- ElevenLabs
- Codex
- AWS Bedrock
A limit for every person.
And every customer.
Set a daily, weekly or monthly allowance for anyone using AI in your company or your product.
Give each group member their own allowance, with resets on your billing cycle and your choice of warnings or blocked requests.
Explore the budget controlsBill your own customers for what they use
A great companion for Stripe: every call reaches your webhook with the customer, the model and the cost, ready to add to that customer’s bill.
{"type": "gateway.request.completed","data": {"gateway_request_id": "gwr_01J9M3T8Q2","virtual_key_id": "vk_01HZX9N","end_user_id": "customer_8841","model": "anthropic/claude-sonnet-4-5","usage": { "input_tokens": 1842, "output_tokens": 214 },"cost": { "total_usd": "0.031", "nano_usd": 31000000 },"metadata": { "plan": "pro", "feature": "summarize" },"status": "success"}}
Your whole AI stack can go through it
Give your apps and coding agents the same spending controls, across text, voice and images.
Models & agents
Chat, responses, tool calls
Voice
Speech, transcription, realtime
Images
Generation and editing
Coding agents
Claude Code, Codex, Cursor
- Chat completions (OpenAI)POST /v1/chat/completions
- Responses (OpenAI)POST /v1/responses
- Messages (Anthropic)POST /v1/messages
- EmbeddingsPOST /v1/embeddings
- SpeechPOST /v1/audio/speech
- TranscriptionPOST /v1/audio/transcriptions
- Image generationPOST /v1/images/generations
- Image editingPOST /v1/images/edits
- Realtime voicePOST /v1/realtime/client_secrets
- Gemini native/v1beta/*
- ElevenLabs nativePOST /v1/text-to-speech/{voice_id}
- Model listGET /v1/models
- OpenAI
- Anthropic
- Azure OpenAI
- AWS Bedrock
- Vertex AI
- Gemini
- xAI
- Groq
- Cerebras
- DeepSeek
- Voyage AI
- ElevenLabs
- OpenAI account
- Custom endpoint
11 microseconds of overhead, at 5,000 requests a second
Change one line in the SDK you already use, and the only wait left is the model’s own answer.
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://gateway.langwatch.ai/v1", apiKey: process.env.LANGWATCH_VIRTUAL_KEY,});
const res = await client.chat.completions.create({
model: "anthropic/claude-sonnet-4-5",
messages: [{ role: "user", content: "Say hi" }],
});One request, to scale
0 ms to 50 ms100% uptime.
Seriously.
The public status page has not recorded one gateway outage since it began watching on 29 July 2026.
See the live status pageFixed $6 per 100,000 calls. Zero markup on tokens.
The provider bills you at its own rate on your own keys, and the gateway adds nothing to the token price.
At 100,000 calls a month the gateway costs $6. A 5% markup on the same traffic costs $100.
See the pricingOpenRouter and other providers charge around 5% markup when you bring your own keys
Questions
Do I need my own provider keys?
Yes. You add your OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, Vertex AI or Gemini keys once, and the gateway calls the provider on your account at the provider’s price. LangWatch adds nothing to the token bill.
What does the AI gateway cost?
$6 per 100,000 calls, the same rate as LangWatch tracing, and every call comes with its trace. The tokens are billed by your providers, on your own keys, at their prices. Webhooks for billing your own customers are part of the Enterprise plan.
Do I have to change my code?
One line. Set the base URL of the OpenAI or Anthropic SDK to the gateway and use a LangWatch key as the API key. The request body, streaming and tool calls stay as they are.
What happens to a call at the limit?
A budget set to block refuses it with HTTP 402 and the code budget_exceeded, and the body names which budget ran out, so your app can tell a customer’s cap from a team’s. A budget set to warn serves the call and adds the X‑LangWatch‑Budget‑Warning header.
How much latency does it add?
About 11 microseconds of gateway time per call, measured at 5,000 sustained requests per second. The provider’s answer is the rest of what your user waits for.
What happens to my traffic if LangWatch has an incident?
The gateway keeps serving. It runs apart from the LangWatch app, remembers the keys it has seen and keeps accepting them for up to 6 hours while the app is unreachable. The public status page shows no gateway downtime since monitoring started on 29 July 2026.
What if a provider goes down?
Turn fallback on for a key, and a 5xx, a 429 or a timeout sends the same request to the next provider that serves the model, up to 3 attempts. A 400, 401 or 403 goes back to you unchanged, so a bad request fails once.
What happens to my prompts?
They go to your provider and come back. The gateway does not write request or response bodies to its logs. LangWatch keeps a trace of the call in your project, with the model, the tokens and the cost, for as long as your retention setting says.
Do Claude Code and Codex work through it?
Yes. Streams keep their tool call deltas, so Claude Code, Codex, Gemini CLI, OpenCode, Cursor and Aider run through the gateway unmodified, each engineer on a personal key with its own budget.
Can I run the AI Gateway in my own cluster?
Yes. The Helm chart runs the gateway next to a self‑hosted LangWatch, and the same keys, budgets and spend events apply.
Put your AI on a budget
Create a key with a budget in one command, then follow the quickstart.