The Agent Evaluation Lifecycle
Core Concepts
Experiments
Online Evaluation
Guardrails
Evaluators
When to Use What
Quick Start
1. Run Your First Experiment
Test your LLM on a dataset using the Experiments via UI or via code:- Platform
- Python
- TypeScript
Go to Experiments and click “New Experiment” to get started with the UI.
2. Set Up Online Evaluation
Monitor your production traffic with evaluators that run on every trace:- Go to Online Evaluations
- Create a new online evaluation with the “When a message arrives” trigger
- Select evaluators (e.g., PII Detection, Faithfulness)
- Enable monitoring