← Changelog
Thread-level Evaluations
Evaluations now run across a whole thread, not just one message, so you can score context retention and see where a conversation drifts.
LangWatch Team · November 6, 2025 · 1.6.0
Evaluate LLM performance at the thread level to understand end-to-end outcomes, context retention, and where conversations drift. Add thread-based mapping to real-time evaluations.
