Authentication & Authorization
LangWatch handles user authentication with support for:- Email/password (default)
- SSO providers: Azure AD, Okta, Auth0, Google, GitHub, GitLab
- Organization-level roles (owner, admin, member)
- Project-level permissions
API_TOKEN_JWT_SECRET) for SDK authentication.
Encryption
At Rest
The
CREDENTIALS_SECRET environment variable is used to encrypt API keys and credentials stored in PostgreSQL (e.g., LLM provider keys configured in the UI). This is application-level encryption on top of database-level encryption.
In Transit
Secrets Management
Development (Auto-Generated)
For development, enableautogen.enabled: true in the Helm chart. This generates random secrets automatically. Not suitable for production, secrets change on reinstall.
Production (Kubernetes Secrets)
Create secrets manually and reference them in the Helm chart:values.yaml:
Production (External Secret Managers)
For tighter security, use an external secrets operator to sync secrets from your cloud provider:- AWS Secrets Manager: via External Secrets Operator
- HashiCorp Vault: via Vault Secrets Operator
- Azure Key Vault: via Azure Key Vault Provider
secretKeyRef pattern works with any Kubernetes Secret, regardless of how it was created.
Network Security
Recommended Network Architecture
- Only the LangWatch App should be exposed externally via Ingress or Load Balancer
- All other components (Workers, NLP, LangEvals, PostgreSQL, ClickHouse, Redis) should be on internal networks only (ClusterIP services)
- Place databases in private subnets with no internet access
- Use VPC endpoints, PrivateLink for S3 access
Private control-plane paths are blocked at the ingress
Everything under/api/internal/* is the app’s private control plane: the Langy
agent’s callbacks and the AI gateway’s usage callbacks. Those routes authenticate
with a shared secret, but they are never meant to be reachable from the internet.
The chart blocks them at the ingress by default, via ingress.blockedPaths
(default ["/api/internal"]). Each blocked prefix is routed to a Service with no
endpoints, so the controller has nowhere to forward the request:
- Blocked requests answer 503 on ingress-nginx. The exact code is controller-dependent (the Ingress API does not specify what a backend with no endpoints returns), and this is verified against ingress-nginx only. If you run a different controller, confirm it answers 5xx rather than falling through to another rule before relying on the block. Expect a low background rate in your ingress 5xx metrics from scanners: that is the block working, not the app faulting.
- In-cluster callers are unaffected. The agent, gateway and CronJobs use the internal Service.
- It covers this ingress only. If you publish the app another way (NodePort, LoadBalancer, your own Ingress), restrict these paths where that route ends.
- The chart refuses configurations that would defeat the block. Each of
these renders a manifest that looks protected and is not, so the chart fails
at render time rather than shipping it:
- an
ingress.hostspath nested under a blocked prefix (it out-matches the blackhole) - a blocked prefix with a trailing slash, a repeated slash, or surrounding whitespace (none of them match any real request path)
- a regex
pathType: ImplementationSpecificpath, which ingress-nginx ranks above prefix rules - the
nginx.ingress.kubernetes.io/default-backendannotation, which is defined as the handler for a backend with no endpoints (exactly what the blackhole is), so it converts the block into a proxy - a host whose
http.pathslist is empty, which would leave an Ingress whose only rules are blackholes
- an
- The block depends on no principal being able to create an EndpointSlice named
for the blackhole Service. It is selector-less, so endpoints can only arrive by
hand. Restrict
discovery.k8s.io/endpointslices: createin this namespace if that is not already the case. - AWS LBC (default
target-type: instance) and GKE without NEG reject a ClusterIP backend and stop reconciling the Ingress. Setalb.ingress.kubernetes.io/target-type: ip(or the NEG equivalent), or emptyblockedPathsand block at the load balancer instead. - To disable, set
blockedPaths: []in a values file or use--set-json; plain--setassigns a string and is rejected. - Gateway or Langy agent running outside the cluster? Their callbacks arrive over the public ingress and are blocked. Narrow the list rather than emptying it.
Kubernetes Network Policies
Restrict traffic between pods:Firewall Rules
Pod Security
Every LangWatch pod (app, workers, NLP, LangEvals, gateway, and the cron pods) and every bundled datastore (PostgreSQL, Redis, ClickHouse, Keeper) runs hardened. The defaults the first-party services inherit:MustRunAs constraint, allow all of those, not only 1000: a constraint that misses the init container’s 65534 denies the whole Keeper pod. (The bundled Prometheus also runs 65534, though strict-admission removes it.)
runAsNonRoot is deliberately set at both pod and container level. Kubernetes inherits the pod-level value, but some Gatekeeper constraints read the container-level field directly and deny pods that only carry it on the pod.
No LangWatch pod mounts a ServiceAccount token (global.automountServiceAccountToken: false); none of these services talk to the Kubernetes API. The bundled Prometheus does mount one. It is an upstream subchart, and strict-admission removes it.
The root filesystem is read-only on every one of those pods. Anything a process writes at runtime (the app’s /tmp, ClickHouse’s server logs and /tmp, Postgres’ socket directory) lands on an emptyDir mounted over that path; persistent data stays on its PVC. The image layer is never writable.
Strict admission control
On a cluster enforcing Pod Security Admissionrestricted or an equivalent Gatekeeper / Kyverno bundle, apply the strict-admission overlay:
RuntimeDefault seccomp, no privilege escalation, no automounted token, CPU + memory requests and limits), but three bundled components cannot, and the overlay turns all three off:
- Prometheus (
prometheus.chartManaged) is an upstream subchart LangWatch doesn’t control. Its pods carry noreadOnlyRootFilesystem, no seccomp profile, and no resource limits on the config-reload sidecar. Bring your own, or disable metrics. - The Langy assistant (
langyagent.chartManaged) runs its manager as root withCHOWN,DAC_OVERRIDE,FOWNER,SETUIDandSETGID. That is by design: the manager gives every assistant worker a distinct UID, and without those capabilities sibling workers share a UID and can read each other’s credentials. Don’t force it non-root. Deploy it on a cluster that allows it, or run without the assistant. - The ClickHouse preflight Job (
clickhouse.preflight.enabled) shells out tokubectlto check your Secret’s keys, so it needs both an automounted token and a writable root for kubectl’s discovery cache. It only renders when you supply the ClickHouse Secret yourself rather than using autogen, which is the production path, so the overlay pins it off.
- Gateway HPA (
gateway.autoscaling.enabled) scales on a custom metric viaprometheus-adapter. Set it tofalseif you don’t run one; the overlay pinsgateway.replicaCount: 2instead.
app.telemetry.metrics.enabled) is off by default too, so nothing exposes or scrapes a /metrics endpoint; the overlay pins it off to keep the two coupled. Anonymous usage analytics (app.telemetry.usage.enabled) is a separate setting, on by default; set it to false for an air-gapped or no-egress install.
In-cluster datastores satisfy read-only-root policies, but managed databases (the postgres-external, redis-external, and clickhouse-external overlays) are still the better production choice: you get backups, HA, and patching, and there are no datastore pods to admit.
PII Redaction
LangWatch includes a built-in PII redaction pipeline step that automatically detects and masks personally identifiable information in traces before storage.- Enabled by default in the Helm chart
- Disable with
app.features.disablePiiRedaction: true(not recommended) - Runs as part of the event sourcing pipeline in workers
Multitenancy
LangWatch enforces tenant isolation at the application level:- Every ClickHouse query includes
WHERE TenantId = ...as the first predicate - PostgreSQL queries include
projectIdin WHERE clauses - API tokens are scoped to a specific project
- Cross-tenant data access is prevented at the query layer
Supply Chain
LangWatch container images and CLI packages are published with verifiable supply-chain attestations so operators can confirm an artifact was built by LangWatch CI from a specific source commit.Container images (Docker Hub)
Every release oflangwatch/langwatch, langwatch/langwatch_nlp, langwatch/langevals, and langwatch/ai-gateway is signed with Sigstore cosign using keyless OIDC. Both the multi-arch index manifest and each per-platform manifest (linux/amd64, linux/arm64) are signed by digest. A CycloneDX SBOM is generated per platform and attached as a cosign attestation against the matching platform manifest digest, so the SBOM you verify always corresponds to the architecture you actually pulled.
Verify a signature with cosign:
*.cdx.json files (e.g. langwatch-linux-amd64.cdx.json, langwatch-linux-arm64.cdx.json) are also attached to each langwatch@vX.Y.Z GitHub release.
npm CLI
Thelangwatch npm package is published with npm provenance attestations via GitHub Actions OIDC, also backed by Sigstore. The provenance link is visible on the package page and can be verified with npm audit signatures.
Production Hardening Checklist
-
autogen.enabled: false, use manually created secrets - All secrets stored in a secrets manager (not inline in values.yaml)
- TLS enabled on Ingress (HTTPS only)
- Database connections use TLS (
?sslmode=require) - PostgreSQL, ClickHouse, Redis in private subnets (no public access)
- Network policies restrict pod-to-pod traffic
- S3 buckets have public access blocked
- ClickHouse backups enabled and tested
- Monitoring and alerting configured
- Secret rotation procedure documented
- Pod security contexts verified (non-root, read-only filesystem,
RuntimeDefaultseccomp, no automounted token) - On strict admission clusters: apply
examples/overlays/strict-admission.yaml(prometheus.chartManaged: false,langyagent.chartManaged: false, and, without a custom metrics API,gateway.autoscaling.enabled: false) - Ingress rate limiting configured
- Audit logs enabled (Enterprise)
- Image signatures verified at pull time (admission controller or
cosign verifyin CI)