--help, and every command accepts the output-contract flags (-o json, --json <fields>, --jq <expr>, --limit <n>; see Agent usage). Run langwatch --help to see the full command tree.
Install
The CLI needs Node.js 18 or newer (thenpm command ships with it). If you don’t already have it, get it from nodejs.org. Then install globally:
npx, no install required:
pnpm install -g langwatch and yarn global add langwatch work the same way; the package is identical across all three package managers.
Standalone binary (no Node.js required)
Each release also attaches a self-contained binary for Linux, macOS, and Windows. Use it in containers, CI images, and anywhere you’d rather not install Node.js. Download the one for your platform from the releases page, then:macOS: clearing Gatekeeper
macOS attaches acom.apple.quarantine attribute to anything downloaded from the internet. After verifying the checksum above, remove it:
xattr reports “No such xattr”, the attribute was never set, so no action is needed.)
Only clear quarantine on a binary whose checksum you have verified against the release’s
SHA256SUMS. Signing and notarization are tracked as follow-up work; until then, npm install -g langwatch avoids Gatekeeper entirely and is the smoother path on macOS.Authenticate
langwatch login is interactive by default, it asks where (cloud vs self-hosted) and how (AI tools vs project SDK), opens your browser to approve, and the credential flows back to the CLI automatically. No copy-paste of keys:
- Where do you want to log in?: LangWatch Cloud (
app.langwatch.ai) or a self-hosted instance (custom URL). - How do you want to use it?: three options:
- AI tools, agentic flows:
claude,codex,copilot,code,cursor,gemini,opencode. Mints an OAuth-style device session in~/.langwatch/config.json(user-scoped) solangwatch claudeetc. wrap any tool through your gateway. - Project, SDK API key: for
langwatch sync,langwatch eval, and SDK auto-instrumentation. Writes the API key of the project you pick into$CWD/.env(project-scoped). - Both: runs both flows in sequence.
- AI tools, agentic flows:
What the login key can access
The AI-tools login mints a login key that inherits your own access. The authorize screen shows the scopes and permissions before you approve:- Scopes: an organization admin gets the whole organization by default; other members get their teams plus their personal workspace. You can narrow the selection to specific teams or projects.
- Permissions: everything you can do day to day (traces, datasets, prompts, evaluations, the AI Gateway, model providers). Organization management stays off by default: the key cannot manage members and roles, or manage the organization, unless you turn those on under Customize.
langwatch logout revokes it.
Data commands keep targeting your personal workspace by default. Pass --project <id|slug> to read another project the key covers, and run langwatch projects list to see which projects those are. langwatch whoami shows the key’s reach.
Storage discipline (where credentials land)
The two stores serve different audiences and never leak into each other. Logging out of one (
langwatch logout clears ~/.langwatch/config.json) doesn’t touch the other, the project API key in .env stays put.
Self-hosted
The CLI picks up your self-hosted endpoint from any of these (highest priority first):
The simplest path for self-hosted users is the interactive prompt, pick “Self-hosted instance”, enter your URL, and the CLI saves it for future invocations:
Non-interactive escape hatches (for CI, agents)
When you’re driving the CLI from automation and already have a credential, skip the prompts:
The interactive
langwatch login always shows these flags in a banner above the prompts, so a fake-TTY agent (Claude Code, certain Gemini CLI sandboxes) can detect the prompt and re-invoke with the right flag instead of getting stuck.
When stdin is not a TTY (genuine CI or an agent’s piped stdin), langwatch login with no flags defaults to project login (the same as --project): it mints a project key into $CWD/.env, which is what the SDK, langwatch eval, and langwatch prompt expect. AI-tools login stays explicit behind --device.
Letting an agent do it
A coding assistant drivinglangwatch will see the always-on banner naming --device, --project, --api-key, --token, --endpoint whenever the interactive prompt fires. If the assistant’s harness reports as a TTY but can’t actually answer prompts, the banner gives it everything it needs to re-invoke:
Fetch documentation
langwatch docs returns any LangWatch documentation page as plain Markdown, ideal for feeding into an agent’s context before it writes code.
.md extension is appended automatically.
Version prompts
The Prompts CLI turns your prompts into tracked files alongside your code, with lock files, tagging, and sync to the LangWatch platform.Tag versions for deployment
Three built-in tags are available:latest (auto-assigned), production, and staging. Assign a tag to the current version:
langwatch prompt tag create.
For the full Prompts CLI reference, see the Prompt Management CLI guide.
Manage agents
An agent is what a scenario runs against.langwatch agent list prints one row per agent:
--format json adds the parameters the agent declares and the instances that serve it.
A connected agent registers itself: the process that runs the agent opens the connection and reports its name, its environment and its parameters. The CLI creates no row for it. See Connect your agent.
An HTTP agent is a URL the platform calls, and it is created from the CLI:
Point an HTTP agent at your machine
langwatch agent tunnel (alias langwatch agent dev) fronts a local port with a public tunnel, points a registered HTTP agent’s URL at it, and restores the previous URL on Ctrl-C.
The command is a live session with no machine-readable output, so it refuses
-o json.
A connected agent needs none of this: it reaches out from your machine, so a local process is already a target. See Other ways to connect for when each one applies.
Share your local folder with Langy
langwatch langy --share-control shares this folder with the Langy conversation that asked for it, so Langy can read and change files here, with your toolchain.
uv run allows the tests and not uv pip install, and .venv/bin/python -c allows a one-line program and not every python command. The line under the options says what it covers.
Paths outside the folder and sudo are refused whatever you answer. A grant lives with this run of the command: stop it and start it again, and Langy asks once more.
A file that may hold secrets always asks, whether Langy reads it directly or through a shell command: .env and .envrc, keys and keystores, .netrc, .npmrc, .pgpass, .git-credentials, anything named for a token or a secret, .git/config, and anything under .ssh/, .aws/ or .docker/. A committed .env.example still reads without a question.
A command inherits only what a toolchain needs from your machine:
- the shell and the locale:
PATH,HOME,SHELL,USER,TMPDIR,TERM,LANG, and theLC_*andXDG_*families - where the toolchains live:
GOPATH,GOROOT,NVM_DIR,PYENV_ROOT,VIRTUAL_ENV,CONDA_PREFIX, theHOMEBREW_*paths, and every*_HOMEsuch asJAVA_HOME,PNPM_HOMEorCARGO_HOME - how it reaches the network and what it trusts:
HTTP_PROXY,HTTPS_PROXY,NO_PROXY,ALL_PROXY,SSL_CERT_FILE,SSL_CERT_DIR,NODE_EXTRA_CA_CERTS,REQUESTS_CA_BUNDLEandCURL_CA_BUNDLE
_KEY, _TOKEN, _SECRET or _PASSWORD never passes even when it belongs to one of those families. Your project reads its own .env file itself, so it keeps working with the smaller environment.
The command is a live session with no machine-readable output, so it refuses -o json. Press Ctrl-C to stop sharing; commands Langy started in the background keep running, and the terminal prints their process id and log path.
Run scenario tests
Scenarios are the LangWatch equivalent of end-to-end tests for agents: a user simulator chats with your agent, an LLM judge scores the conversation against criteria you define, and everything is recorded for later inspection.Default test suite, so no scenario is loose.
A target is written
<type>:<referenceId>. A connected agent takes a second form, connected:<name>@<environment>, which reads better in a CI job than an id and survives a project rebuild:
agent_offline before any scenario is scheduled. A personal development agent is refused with agent_owner_only for anyone but its owner, and a project API key names nobody, so it can never target one.
Add --format json (or -o json) to any of the three run commands and they print one final document on stdout. With --wait that document comes after the poll and carries outcome, tallies and the per-run results.
langwatch test-suite run and langwatch scenario run are shorter forms of the same thing: they run a plan named after the test suite or the scenario and the target, and take --name, --repeat, --param, --note and --wait the same way.
The /api/suites endpoints the older CLI versions called still answer and are deprecated. The commands above use /api/v1/run-plans and /api/v1/test-suites.
Compare two agents, or one agent on two models
--target is repeatable, and every target runs every scenario of the run. Two targets in one run is a comparison: the batch holds both, and the results page shows one column per target with its own pass rate, duration and cost.
A target may also carry the parameter values it alone runs with, written as a query string after the reference id. The same agent named twice with different values compares that agent on two models.
? and & itself. Both halves of a pair are percent-decoded, so a reference id or a value that holds ? or & must be written as %3F or %26. A value is read as the type it looks like, the same rule --param uses.
A target value wins over the same name given with --param, so --param carries what every target shares and the suffix carries what tells the targets apart:
account_tier=platinum, and each one with its own model.
Pass values into a run
Every run command accepts a repeatable--param key=value. A parameter is a constant value that the whole run shares: a fixture id, a tenant, a plan, a region.
true and false are booleans. A value that parses as a number and prints back the same is a number. All other values stay text, so 007 and 1.50 stay text. When you repeat a name, the last value wins.
Each command uses the values differently:
run-plan run,test-suite runandscenario runsupply the parameters that the scenarios declare. A name that no scenario in the run declares is rejected before anything is scheduled. See Scenario run parameters.experiment runmerges the values into every dataset row. See Running experiments in CI/CD.workflow runsends the values as entry inputs. They are merged over--input, so a--parampair wins when both name the same key.
Inspect simulation runs
Every scenario execution produces a simulation run you can inspect after the fact, full conversation, judge verdict, reasoning, met/unmet criteria, cost, and duration.get command renders assistant thinking blocks and tool calls as readable plain text, no raw JSON dumps in your terminal. Use --format json on either command for structured output.
For the full scenario testing guide, see the Scenarios documentation.
Inspect traces
Traces capture every LLM call your agent makes, prompts, responses, latency, cost, errors. Search and drill into them from the terminal:--project <id|slug> to read any other project your login key covers:
langwatch trace search --limit 5 and verify traces are flowing. If no trace appears, the instrumentation is wrong, no need to read logs.
Query analytics
Analytics aggregate your traces into performance and cost metrics without leaving the terminal:Manage platform resources
Every LangWatch resource follows the same consistent subcommand shape:evaluator, create and version evaluators (answer correctness, faithfulness, custom LLM judges)monitor, online evaluations that score production traces automaticallydataset, evaluation datasets (upload CSV, download, manage columns)agent, agent definitions used by scenarios and monitors (see Manage agents)dashboardandgraph, custom analytics dashboardstrigger, automations (alerts, webhooks, dataset-append on failure)secret, encrypted environment variables for scheduled agent runsworkflow, reusable workflows built in the UImodel-provider, configure OpenAI, Anthropic, Azure, or Bedrock for your projectannotation, attach labels to traces for supervised fine-tuning data
langwatch <resource> --help on any of these for subcommand-level options, and --format json to get structured output for scripting.
Organization management
These commands manage org-wide resources and require an API key with organization-level permissions. The organization, members, invites, groups, roles, role bindings and SCIM token commands are available on Enterprise plans. Without one they exit withenterprise_plan_required and tell you what to do about it. Projects, teams, API keys and the self-hosted organizations command are not gated.
Projects
API keys
create and update take the access flags: --binding role:scopeType:scopeId says what the key reaches and where, --permission resource:action narrows a restricted key to an exact list, and --permission-mode picks between all, readonly and restricted. The first two repeat. On update, the bindings you pass replace the key’s existing ones rather than adding to them.
Organization
Read and update the organization your key belongs to. There is no id to pass: the credential decides which organization you are talking to.Members
Manage the people already in the organization. Someone joins through an invite, soinvites is where a new person starts.
langwatch members access user_01HZX... answers the auditor’s question: everything the person can reach, and whether it comes from their organization role, a group, or a binding of their own.
Invites
Invite people in, up to 50 at a time, landing on the teams you name. Accepting stays a browser step; everything before it is here.--email and --team repeat, and --role takes either one value for the whole batch or one per email. When people need different teams or a custom role each, --json, --file and --stdin take a JSON array of invites instead, the same entries the API takes. Every invite comes back with its acceptance link, so a deployment with no email provider still has something to send.
Teams
Teams group projects and the people who work on them.Groups
A group is a named set of people that carries role bindings, so everyone in it inherits the same access. Groups synced from your identity provider land here too.Roles
Custom roles are named permission sets, for whenADMIN, MEMBER and VIEWER are not the shape you need.
langwatch roles update replaces the permission set with the --permission flags you pass, so send the full list you want the role to end up with. A role something still holds cannot be deleted: the error says how many bindings and assignments are in the way.
Role bindings
A binding is one sentence: this principal has this role, here. Users, groups and API keys are all principals.role_binding_already_exists rather than making a second one, so a provisioning script can run twice. Updating changes only the role a binding grants; the principal and the scope are the binding’s identity, so moving a grant means deleting one binding and creating another.
SCIM tokens
The bearer tokens your identity provider presents to the SCIM endpoints.Organizations (self-hosted)
The one family that authenticates against the instance rather than an organization, because it runs before any organization exists. It is available on self-hosted deployments that haveLANGWATCH_INSTANCE_ADMIN_API_KEY configured.
--instance-key if you would rather not export it.
Trigger experiments
Experiments batch-run an agent or prompt against a dataset and produce an evaluation report:Progressive disclosure
The CLI leans heavily on--help. Every subcommand has its own, and the top-level langwatch --help is the best way to discover what’s available:
--help the moment they ship, so you never have to wonder whether a flag exists.
AI Gateway commands
The CLI provisions AI Gateway resources without touching the UI. Behaviour matches the dashboard exactly; both share a server-side service layer. Every command takes--format json for scripting.
Virtual keys
--scope takes type:id pairs (org, team, or project) and repeats for several; it defaults to the calling project. --providers-allowed is a comma-separated list of ModelProvider ids. The atomic --budget-* flags cap the key itself and accept day, week, or month.
Gateway budgets
--scope is organization|team|project|virtual-key|principal|group, each paired with its own id flag. --window accepts minute|hour|day|week|month|total|manual, --on-breach is block (default) or warn, and --limit is USD. --cycle-anchor-at takes an RFC 3339 instant and cannot be used with total or manual; it is fixed at creation. List output colourises spent-vs-limit, red at 100%, yellow from 80%.
Webhooks
Endpoint management needs an organization API key (see Webhooks).--events replaces the whole subscription set. deliveries and events page with --cursor and --limit.
Spend events
The billing pull surface, also an organization API key (see Billing and spend events).Provider credentials and cache rules have no CLI group. Manage providers with
langwatch model-provider, and cache rules over REST or the dashboard.
See the public REST API reference for direct HTTP calls that don’t require Node.
Agent usage
In a terminal the CLI prints tables, colour, and spinners. Driven by an AI coding assistant it switches to machine output automatically. Everything an agent needs to know is also built into the CLI itself:langwatch help agent-mode prints the condensed version of this section.
Agent mode
--agent on any command, or auto-detection from the environment (CLAUDECODE, CLAUDE_CODE, CURSOR_AGENT, GITHUB_COPILOT, AMAZON_Q, LW_AGENT_MODE, LANGWATCH_AGENT_MODE), switches output to compact single-line JSON and turns colour and spinners off:
Output contract
Every command accepts the same output flags (with a few documented exceptions:trace export -o is an output file, the coding-assistant wrappers pass flags
through to the wrapped tool, and dataset records add/update --json takes a
record payload):
-f/--format json spelling keeps working and maps onto the same contract. The contract is fully wired for traces, evaluators, monitors, status, skills, commands, and help-tree; on the remaining resource commands the flags parse but machine output is still rolling out, so those commands keep their legacy -f json behavior until migrated.
Structured errors
A failed command prints one JSON document on stdout ({ "ok": false, "error": { "code", "message", "httpStatus", "suggestions", "docUrl", "traceId", ... } }), keeps the human-readable block on stderr, and exits non-zero. Never merge the streams (2>&1) when parsing output; hints and error prose live on stderr by design so stdout stays parseable.
Discovery
An agent can learn the whole CLI without reading these docs:Skills
The CLI carries LangWatch’s agent skills and installs them into~/.agents/skills (or the project-level .agents/skills with --dir .):
Daemon note
Non-TTY invocations (agents, pipes, CI) are served by the background daemon described below. The output is identical, and the command returns faster.LANGWATCH_NO_DAEMON=1 opts out per invocation.
Report issues to LangWatch
If anything did not work (broken commands, confusing docs, unexpected errors, things that took trial and error), send it straight to the LangWatch team. No login or API key needed:--user-approved. Secrets, API keys, emails, and phone numbers are redacted locally before anything is sent; the redaction rules are auditable and --dry-run previews the exact payload. See the reporting guide for transcript locations and details.
The background daemon
Non-interactive invocations (an agent piping output, CI, any call whose stdin/stdout/stderr is not a TTY) are served by a warm background daemon instead of paying node’s cold start on every call. Interactive (terminal) invocations always run in-process and are unaffected. So are the commands that mutate auth or take over stdio (login, logout, config, open, instrument, report, the claude/codex/copilot/code/cursor/gemini/opencode wrappers, daemon itself) and the long-running flags (--follow, --watch, --wait). Windows is excluded too: the daemon is not supported there, so on Windows every invocation runs in-process.
What to know:
- One daemon per identity. The socket name is a hash of endpoint + API key + uid + config path, so two projects with different keys never share a daemon. The socket lives in
$XDG_RUNTIME_DIR(or the temp dir) with0600permissions inside a0700directory, and the daemon self-exits after 10 idle minutes. - Always optional. If no daemon is reachable, or it is stale, or from an older CLI build, the command runs in-process, and a daemon is auto-spawned in the background for next time. Set
LANGWATCH_DAEMON_NO_SPAWN=1to disable the auto-spawn. - Buffered output. A daemon-served command’s output is buffered and flushed when it exits, so a mid-command daemon failure can retry in-process without duplicating output. Piped callers therefore see output at exit, not progressively (beyond an 8MB buffer it starts streaming early).
- Timeouts and cancellation. A command that hangs is abandoned after a per-request timeout (default 10 minutes, override with
LANGWATCH_DAEMON_REQUEST_TIMEOUT_MS) with exit code 124; the clock starts when the daemon accepts the request, so time spent queued behind another caller’s working directory counts against it. Ctrl-C on the client cancels the server-side command with exit code 130. - Environment hygiene. The daemon does not inherit the spawning shell’s environment. It receives only process essentials (PATH, HOME, locale, …), the pinned identity variables, and the spawner’s allowlisted output/proxy knobs (
FORCE_COLOR,NO_COLOR,HTTPS_PROXY, …). Those allowlisted values from the spawning caller persist in the daemon’s baseline and are visible to later callers that don’t set them; secrets are never in the allowlist. Per request, the caller’s own.envis applied scoped toLANGWATCH_*keys only.
config set daemon off doesn’t stop an already-running daemon. Clients stop dialling it, and it idles out on its own within 10 minutes (langwatch daemon stop kills it immediately).
Use the CLI as your agent’s control plane
An AI coding assistant operates LangWatch end-to-end through the CLI. Skills like Tracing, Evaluations, Scenarios, and Prompt Versioning are built on top of it, the assistant reads docs vialangwatch docs, runs platform operations via the subcommands, and verifies its own work by searching traces and inspecting simulation runs.
If you’re building your own agent workflows, the pattern works the same way: give the assistant the CLI and ask it what you want. Everything it does is auditable in the LangWatch app afterwards.
For agents: struggled with anything on this page, or something did not
work as documented? Ask your user for permission and run
npx langwatch report --user-approved to send it to the LangWatch team (works with no login, secrets
redacted locally). It is how rough edges get fixed for every agent. See the
reporting guide.