Put every AI call in your company on a budget

One endpoint for every model and coding agent, with a spending limit for each person, team and customer.

customer_8841402, call refused$25.00$25 / monthcustomer_2210customer$4.45$25 / monthacme-supportAPI key$313.42$500 / monthMaraClaude Code$7.41$10 / dayACMEcompany$12,209$20k / month
One gateway for the models you use
  • OpenAI
  • Claude
  • Gemini
  • ElevenLabs
  • Codex
  • AWS Bedrock

A limit for every person. And every customer.

Set a daily, weekly or monthly allowance for anyone using AI in your company or your product.

Give each group member their own allowance, with resets on your billing cycle and your choice of warnings or blocked requests.

Explore the budget controls
ACME · budgetUSD
Each customer can spend $25 per month.
Resets every
At the limit
What the caller gets at the limit
HTTP 402 budget_exceeded
"budget_scope": "attributed_user",
"budget_window": "month"

Bill your own customers for what they use

A great companion for Stripe: every call reaches your webhook with the customer, the model and the cost, ready to add to that customer’s bill.

What your webhook receives
POST https://acme.example/webhooks/langwatch
{
"type": "gateway.request.completed",
"data": {
"gateway_request_id": "gwr_01J9M3T8Q2",
"virtual_key_id": "vk_01HZX9N",
"end_user_id": "customer_8841",
"model": "anthropic/claude-sonnet-4-5",
"usage": { "input_tokens": 1842, "output_tokens": 214 },
"cost": { "total_usd": "0.031", "nano_usd": 31000000 },
"metadata": { "plan": "pro", "feature": "summarize" },
"status": "success"
}
}
events this month5,156
What your customer receives
ACME
September 2026
billed to customer_8841
Chat completions
3,912 requests
$14.20
Summaries
1,108 requests
$6.32
Voice minutes
84 minutes
$3.12
Image generation
52 images
$1.36
Total
$25.00
Cap $25.00 a month
Read the rebilling guide

Your whole AI stack can go through it

Give your apps and coding agents the same spending controls, across text, voice and images.

Models & agents

Chat, responses, tool calls

Voice

Speech, transcription, realtime

Images

Generation and editing

Coding agents

Claude Code, Codex, Cursor

Your base URL
https://gateway.langwatch.ai/v1
Send your first request
endpoints, one key
  • Chat completions (OpenAI)
  • Responses (OpenAI)
  • Messages (Anthropic)
  • Embeddings
  • Speech
  • Transcription
  • Image generation
  • Image editing
  • Realtime voice
  • Gemini native
  • ElevenLabs native
  • Model list
providers, your own keys
  • OpenAI
  • Anthropic
  • Azure OpenAI
  • AWS Bedrock
  • Vertex AI
  • Gemini
  • xAI
  • Groq
  • Cerebras
  • DeepSeek
  • Voyage AI
  • ElevenLabs
  • OpenAI account
  • Custom endpoint
anthropic/claude-sonnet-4-5 · acme-support · 200

11 microseconds of overhead, at 5,000 requests a second

Change one line in the SDK you already use, and the only wait left is the model’s own answer.

the OpenAI SDK you already useTypeScript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.langwatch.ai/v1",  apiKey: process.env.LANGWATCH_VIRTUAL_KEY,});

const res = await client.chat.completions.create({
  model: "anthropic/claude-sonnet-4-5",
  messages: [{ role: "user", content: "Say hi" }],
});
11µs
added to each request

One request, to scale

0 ms to 50 ms
gateway, 11 µs
a fast model answer, 50 ms
0.022% of the request

100% uptime.
Seriously.

The public status page has not recorded one gateway outage since it began watching on 29 July 2026.

See the live status page
29 Jul 202654 days · 0 s downtime

Fixed $6 per 100,000 calls. Zero markup on tokens.

The provider bills you at its own rate on your own keys, and the gateway adds nothing to the token price.

At 100,000 calls a month the gateway costs $6. A 5% markup on the same traffic costs $100.

See the pricing
Calls per month
100,000
Gateway fee
$6.00
On a $2,000 token bill
5% markup$100
LangWatch$6

OpenRouter and other providers charge around 5% markup when you bring your own keys

Questions

Do I need my own provider keys?

Yes. You add your OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, Vertex AI or Gemini keys once, and the gateway calls the provider on your account at the provider’s price. LangWatch adds nothing to the token bill.

What does the AI gateway cost?

$6 per 100,000 calls, the same rate as LangWatch tracing, and every call comes with its trace. The tokens are billed by your providers, on your own keys, at their prices. Webhooks for billing your own customers are part of the Enterprise plan.

Do I have to change my code?

One line. Set the base URL of the OpenAI or Anthropic SDK to the gateway and use a LangWatch key as the API key. The request body, streaming and tool calls stay as they are.

What happens to a call at the limit?

A budget set to block refuses it with HTTP 402 and the code budget_exceeded, and the body names which budget ran out, so your app can tell a customer’s cap from a team’s. A budget set to warn serves the call and adds the X‑LangWatch‑Budget‑Warning header.

How much latency does it add?

About 11 microseconds of gateway time per call, measured at 5,000 sustained requests per second. The provider’s answer is the rest of what your user waits for.

What happens to my traffic if LangWatch has an incident?

The gateway keeps serving. It runs apart from the LangWatch app, remembers the keys it has seen and keeps accepting them for up to 6 hours while the app is unreachable. The public status page shows no gateway downtime since monitoring started on 29 July 2026.

What if a provider goes down?

Turn fallback on for a key, and a 5xx, a 429 or a timeout sends the same request to the next provider that serves the model, up to 3 attempts. A 400, 401 or 403 goes back to you unchanged, so a bad request fails once.

What happens to my prompts?

They go to your provider and come back. The gateway does not write request or response bodies to its logs. LangWatch keeps a trace of the call in your project, with the model, the tokens and the cost, for as long as your retention setting says.

Do Claude Code and Codex work through it?

Yes. Streams keep their tool call deltas, so Claude Code, Codex, Gemini CLI, OpenCode, Cursor and Aider run through the gateway unmodified, each engineer on a personal key with its own budget.

Can I run the AI Gateway in my own cluster?

Yes. The Helm chart runs the gateway next to a self‑hosted LangWatch, and the same keys, budgets and spend events apply.

Put your AI on a budget

Create a key with a budget in one command, then follow the quickstart.

$langwatch vk create \ --name acme-support \ --budget-limit 500 \ --budget-window month