Skip to main content
Let your agent set this up. Copy the evaluations prompt into your coding agent to get started automatically.
LangWatch provides comprehensive evaluations tools for your LLM applications. Whether you’re evaluating before deployment or monitoring in production, we have you covered.

The Agent Evaluation Lifecycle

Core Concepts

Experiments

Online Evaluation

Guardrails

Evaluators

When to Use What

Quick Start

1. Run Your First Experiment

Test your LLM on a dataset using the Experiments via UI or via code:
Go to Experiments and click “New Experiment” to get started with the UI.

2. Set Up Online Evaluation

Monitor your production traffic with evaluators that run on every trace:
  1. Go to Online Evaluations
  2. Create a new online evaluation with the “When a message arrives” trigger
  3. Select evaluators (e.g., PII Detection, Faithfulness)
  4. Enable monitoring

3. Add Guardrails

Protect your users by blocking harmful content in real-time:

Supporting Resources

Datasets

Annotations