Webhook endpoints are an Enterprise feature (the
webhookEndpointsEnabled plan flag). On other plans the API answers 403 and the settings page shows an upgrade state. Self-hosted enterprise licenses include it.The model
An endpoint belongs to your organization and carries:
Manage endpoints under AI Gateway > Webhooks (
/settings/gateway/webhooks) or over REST with an organization API key:
201 response carries the endpoint plus secret, returned only this once. Store it where your receiver can verify signatures. Send an Idempotency-Key and a retry that lost its response returns the original secret instead of stranding it.
What a receiver URL must look like
A webhook URL is a destination our workers dial on your behalf, so it is held to one policy across the product:
Anything else is refused at create and update time with
400 and error.code = "webhook_endpoint_invalid", naming the rule it broke:
The envelope
Every delivery is one POST with a JSON body of the form{"batch": [envelope, ...]}. Each envelope is one real-world occurrence:
idis the idempotency key: stable across retries and replays. Dedup on it. For spend events it is the gateway request id with a type suffix (<gateway_request_id>:completed,<gateway_request_id>:settled), so the settled and completed events for one request never collide in your dedup layer whiledata.gateway_request_idjoins them.createdis when the event occurred, never when it was delivered. Late deliveries land in the right period.schema_versionversions thedatapayload per type.
max_batch_size envelopes, possibly of mixed types within your subscription. Ordering is best-effort: rely on created and your own dedup, not arrival order.
Event catalog
GET /api/webhooks/v1/event-types returns the live catalog. Types are grouped by family (the first dotted segment), which is also what the settings UI renders as checkbox groups and what the gateway.* wildcard matches.
Family: gateway
The spend payloads (
gateway.request.*) are documented field by field on Billing and spend events. The governance payloads:
gateway.budget.threshold_crossed and gateway.budget.breached (data):
bucket_scope_id is <anchor_id>:<end_user_id> and end_user_id is set, so your platform can tell “this user’s cap” from “the tenant cap”.
virtual_key_id and anchor_project_id name the key and the project the budget targets, as their own fields. virtual_key_id is set for a virtual-key budget and for an attributed-user template (where it is the anchor); anchor_project_id is set for a project-scoped budget. Read those rather than splitting bucket_scope_id, which you cannot split reliably when an end-user id contains a colon.
Every enum on these payloads is
lower_snake_case (scope_type, window, on_breach), matching the rest of the wire.gateway.virtual_key.* (data):
Verifying signatures
Every delivery carries:.code when you want to tell a clock-skew problem from a wrong secret. Timestamps outside a 5-minute tolerance are rejected; pass toleranceSeconds / tolerance_seconds if your receiver needs a different window.
Working receivers built on both snippets live in the agent-billing-demo reference repo.
Rotating the signing secret
roll-secret returns a new secret and keeps the old one valid for 24 hours. For that window every delivery is signed with both, so there is no coordinated deploy and no gap:
- Call
roll-secretand store the new value. - Deploy it to your receiver any time in the next 24 hours. Deliveries keep verifying under the old secret until you do.
- After the window the old secret stops being signed with and stops verifying.
Signing automation webhooks
An automation trigger can also call a webhook when it fires. Those requests go out unsigned unless you give the trigger a signing secret, in the webhook action’s Signing secret field. Set one and every fire from that trigger carries the sameX-LangWatch-Signature: t=...,v1=... header, computed the same way, so the verifier above validates it without a single change. The trigger’s own id header is X-LangWatch-Event-Id rather than X-LangWatch-Delivery-Id, and it groups the attempts of one fire.
- Opt-in per trigger. A trigger with no secret keeps sending exactly what it sent before, so adding this breaks no existing receiver.
- Rotation works the same way. Replacing the secret keeps the previous one signing for 24 hours, so you deploy the new value on your own schedule. During the window the header carries a
v1for each, which is why a verifier has to accept any match. - Secrets are encrypted at rest, and, like an endpoint’s, are not readable again after you save them.
Delivery, retries, and auto-disable
Your receiver acks with any2xx. The delivery classifier:
5xx,429, and408are retryable. ARetry-Afterheader on those is honored as a floor on the next attempt.- Any other non-2xx status is terminal for that batch: retrying a misconfigured endpoint just spams it. Redirects are never followed, so a
3xxis terminal too.
disabled with disabled_reason: "auto_failures_72h" and delivery stops. Sources keep accruing the whole time: disabling delivery never loses events.
To recover:
- Fix the receiver.
- Re-enable:
PATCH /endpoints/:id {"status": "active"}. - Re-enabling does not re-send the gap. Replay it explicitly: for spend events,
POST /api/gateway/v1/spend-events/replaywith the gap window and this endpoint’s id.
Delivering to an Amazon SQS queue
An endpoint can put each batch on your own Amazon SQS queue instead of POSTing it to a URL. Only the last hop changes. The batching, the retry ladder, the delivery log, and the signature over the same bytes all work as they do for an HTTPS endpoint. The body is byte-identical to the HTTP body.MessageBody is the exact same {"batch": [...]} JSON an HTTPS receiver would be POSTed. There is no outer wrapper. The signature, the delivery id and the attempt ride as message attributes under the same names they use as HTTP headers:
So verification is the same call over the same bytes, and moving an integration from HTTP to a queue does not change your verification code.
The trust policy on your role, with the
external_id the create response returned:
On LangWatch Cloud, ask support for the principal to trust; it is a property of the deployment, not of your endpoint. On a self-hosted install it is whatever identity your workers run as.
- Standard queues only. A
.fifoqueue is refused at save time. Delivery is at-least-once and consumers deduplicate on the envelopeid, which is what a standard queue asks of them; ordering is not part of the delivery contract, and a FIFO queue would add a throughput ceiling for a guarantee this path does not need. - A canonical Amazon SQS queue URL,
https://sqs.<region>.amazonaws.com/<account id>/<queue name>. The region and the owning account are read from it, so they can never disagree with the queue. Anything else is refused. - One batch is one message, and LangWatch caps it at 256 KiB including the attributes. Amazon SQS itself accepts up to 1 MiB; the lower cap is ours, and it keeps one slow consumer from having to hold a megabyte in memory per message. A batch over the cap fails terminally and says so; lower
max_batch_sizeso each delivery carries fewer events.
id. A standard queue is at-least-once and can redeliver a message on its own, and it has no queue-level deduplication of its own to lean on. X-LangWatch-Delivery-Id names the batch delivery, not an event, so deduplicating on it would drop every envelope in the batch but one. The id inside each envelope is the idempotency key.
Asking for a retry. Delete the message only after the batch is durably ingested. Leaving it alone is the retry: it reappears after the visibility timeout, and after your queue’s maxReceiveCount it lands in your dead letter queue. Set a redrive policy, so a message your consumer can never accept lands somewhere you can inspect it.
Changing where an endpoint delivers is not an update. destination_kind is fixed once created, because batches already planned against the old transport are in flight. Create a second endpoint, let the first drain, then archive it.
Delivery controls
Three per-endpoint knobs, validated server-side against fixed bounds; out-of-range values are rejected with the bound named in the error:Health and the lag number
GET /endpoints/:id/health:
oldest_undelivered_age_ms: the age of the oldest envelope still buffered or retrying for this endpoint. It is your feed’s staleness, the number that tells a billing operator how far behind their invoice data could be. Zero or small is healthy; growing means your receiver is failing or slow. sends_per_minute and success_rate aggregate the last hour; p95_latency_ms is sampled over the same window. dlq_depth counts dead-lettered batches awaiting manual attention.
The delivery log
GET /endpoints/:id/deliveries lists recent attempts, newest first: the attempt number, envelope count, outcome, your endpoint’s HTTP status, latency, and the error text when transport failed. Your receiver’s responses are recorded, so debugging a rejecting endpoint starts here rather than in your own logs.
The list is cursor-paginated on (fired_at, id) at 25 rows by default (max 200), and the response carries next_cursor when more remain. Cursors rather than offsets because a busy endpoint writes new attempts while you page: an offset walk would show you the same row twice and skip another. Delivery log rows are retained for 30 days.
Endpoint deliveries and the webhooks fired by your automation triggers share one log, each row tagged with which of the two sent it. This endpoint’s list shows only its own rows; the trigger drawer shows a trigger’s. Nothing about our request is stored, only the outcome: no URL, no headers, no body, plus your receiver’s status, the latency, and a truncated copy of a failure response.
Test fire
POST /endpoints/:id/test sends one signed test.ping envelope through the full delivery path, including SSRF checks and signing, plus an X-LangWatch-Test-Fire: true header so your receiver can tell it apart. The route answers 200 whenever the test itself ran; read data.delivered and data.response_status for what your receiver did. Test fires appear in the delivery log too.
The events log
GET /api/webhooks/v1/events?from=&to=&type=&cursor=&limit= is the organization’s emitted-events log: cursor-paged, newest first, filterable by type. Webhooks are push over this log, never the only copy of it, so a consumer that missed deliveries can list what was emitted and recover.
from and to bound the created range in unix milliseconds and are required, and from must be less than or equal to to. The log reads the same 13-month spend table the reconciliation pull reads, so an unbounded listing sorts every month the organization has, cold storage included, to serve one page. A read with no range, with only one of the two bounds, or with a range that ends before it starts is rejected with the canonical 400 this surface answers every validation failure with, naming the offending parameter.
GET /api/webhooks/v1/events/{id} reads one event back by the id its envelope carried. A 404 means the log cannot answer for that id, and deliberately does not say which reason applies.
Which families the log holds
The log serves the request families only:
The governance families are delivered to your endpoints exactly like the request families, but they are not retained in a queryable log, so they cannot be listed, read by id, or replayed. If you need a durable record of budget and virtual-key events, persist them from your receiver when they arrive. A request for a type this log does not hold returns an empty page rather than an error, so a client can probe for new families without breaking.
What events never carry
No prompts, no completions, no conversation content, ever. Envelopes carry ids, quantities, classes, states, and your own metadata echo. If you need content downstream, that is the traces API, a separate surface with its own access control.See also
- Billing and spend events: the spend payloads, the money contract, reconciliation, replay.
- Metering and rebilling your customers: the end-to-end integration cookbook.
- Self-hosting webhooks: egress policy, local receivers, license flag.
- Budgets: what threshold_crossed and breached mean and when they fire.