How the broker works
The gateway mints the vendor’s own session credential. Your client asks the gateway for a session, the gateway admits it against the key’s budgets and caps, calls the vendor’s mint endpoint with your stored provider key, and hands back what the vendor answered plus a LangWatch session id. The audio never passes through the gateway. The media socket runs from your client straight to the vendor, so the gateway does not add to turn latency and cannot see, record or end a call in progress. One session is one spend record: admitted at the mint, settled when the call’s usage arrives.Because the gateway is not in the audio path, a voice budget admits at session start and reconciles when the call ends. Overshoot on voice is bounded by the length of a call.
ElevenLabs Conversational AI
Mint a signed URL for one hosted agent.agent_id is required: a signed URL is bound to one agent, and a request without it answers 400 bad_request.
The answer is the vendor’s own body with one added key:
X-LangWatch-Session-Id response header, so a client that parses the vendor shape strictly reads it there instead.
Open the socket with signed_url exactly as returned. The gateway asks ElevenLabs for the conversation id at mint time, so the call is joined to its spend record before the socket exists.
Closing an ElevenLabs session
ElevenLabs reports cost and duration after the call, not over the socket. LangWatch reads that report two ways, and either one settles the session:- The post-call webhook. Point your ElevenLabs workspace webhook at
https://gateway.langwatch.ai/v1/convai/webhook/<model-provider-id>and store the webhook secret on that provider row. See ElevenLabs for the route, the signature check and the status codes. - The poller. Two minutes after a mint, LangWatch reads the conversation back from ElevenLabs by its own id.
post_call_transcription event is applied. post_call_audio and call_initiation_failure are acknowledged and not applied: they name the conversation and carry no duration.
OpenAI Realtime
Mint an ephemeral client secret.session.model as the model and forwards your declaration as you wrote it, with two edits: the resolved model is written back into session.model, and expires_after.seconds is held between 10 and 7200. A body that names no expiry is forwarded untouched, so OpenAI’s own default applies. The voice, the tools, the turn detection and the audio formats reach OpenAI unchanged. The request body is capped at 256 KiB.
The answer is OpenAI’s own body plus langwatch.session_id, and the same id is on X-LangWatch-Session-Id.
Open the socket with the returned secret. Which form you use depends on where the code runs, because a browser WebSocket takes only a URL and a subprotocol list and cannot set request headers.
From a server, with the ws package or any Node websocket client that accepts headers:
Closing an OpenAI session
OpenAI reports usage over the socket, inresponse.done, and that socket does not pass through the gateway. Post what you read back:
202. SESSION_ID is the X-LangWatch-Session-Id from the mint. Three body shapes are accepted, so you can forward the frame unchanged: the bare usage object, an object with a usage key, and the whole response.done event. The rest of the event is not read, and no transcript content is looked at. The body is capped at 64 KiB.
Audio tokens are taken out of the text totals before pricing, because audio tokens cost several times what text tokens cost.
A session that never reports settles as cost unknown when the settlement grace expires, and a report arriving after that replaces the settled record.
Limiting how many calls a key runs at once
Set max open sessions on the virtual key. The request-rate limits do not bound voice: one mint is a single request that opens a call billing for as long as it runs, so a key at 60 requests per minute can hold sixty ten-minute calls without tripping anything. A mint over the cap answers429 realtime_session_limit, and the refusal is recorded in the spend surface. A slot frees when the call’s report arrives. A session older than the longest call a vendor allows stops holding a slot, which matters for OpenAI, whose socket never signals that it closed.
What the broker does not do
- It does not proxy the voice socket. There is no frame-level tracing, and no way to end a call in progress.
- It does not enforce the session’s tools. A hosted agent’s tools live at the vendor, and an OpenAI session declares its own.
- Guardrails do not run on a mint. The body is a session declaration, not a prompt, and the conversation never reaches the gateway. The response says so on
X-LangWatch-Guardrails-Not-Applied: realtime_sessioninstead of reporting a check that could not mean anything.
Errors
Each mint route serves one vendor:
/v1/realtime/client_secrets needs an OpenAI credential, and /v1/convai/conversation/get-signed-url needs an ElevenLabs one. A mint never falls back to a second credential, because a session credential is bound to one account and a fallback would sign for an agent that does not exist there.
Also check: ElevenLabs, Budgets, Errors.