The LangWatch Blog
Engineering deep-dives, product updates, and field lessons on evaluating, testing, and observing AI agents in production.
Product Releases
Things you can ask Langy
The real questions teams ask Langy about their agents, and how it answers them from your traces and code, then opens…
Rogerio Chaves · July 23, 2026
Developer findings
We Red-Teamed Our Own AI Agent and Found 14 Real Bugs
Langy is the AI agent LangWatch runs on its own product. A teammate distracted it with an unrelated coding question,…
Aryan Sharma · July 22, 2026
Developer findings
The PR Hound: how we fixed PR review assignment with a daily agent
We leaned into agentic coding and started shipping far more code than our review process could keep up with, so PRs…
Andrew Joia · July 9, 2026
Article
Claude vs Codex: which is the better background agent?
Some of our preferences and stories on Claude and Codex at LangWatch.
Rogerio Chaves · July 8, 2026
Developer findings
Background Agents on Slack: How we built our own Claude Tag before it was cool
A fleet of background agents runs a chunk of our engineering, each one scoped to one job and living in its own Slack…
Rogerio Chaves · July 5, 2026
More from the blog
Getting to value with LangWatch, faster than ever - how to migrate from Langfuse to LangWatch with Skills.
LLM Evaluations Explained: Experiments, Online Evaluations, Guardrails, and when to use each in 2026
October 27, 2025Governance AIManouk Draisma & FlagSmith
How LangWatch helps enterprises test, evaluate, and trust their AI before release
Build vs Buy - Should you build your own LLMOps stack or leverage a purpose-built platform designed for enterprise scale?
The 6 Best LLM Evaluation Platforms in 2025: Why LangWatch redefines the category with Agent Testing (with Simulations)
Introducing the Evaluations Wizard: How to evaluate your LLM: Building an LLM evaluation framework that actually works
LangWatch vs. LangSmith vs. Braintrust vs. Langfuse: Choosing the Best LLM Evaluation & Monitoring Tool in 2025
LangWatch.ai - Announcing - €1M funding round to bring the power of Evaluations and Auto-Optimizations to AI teams.
OpenAI, Anthropic, Deepseek and other LLM Providers keep dropping prices: Should you host your own model?
December 20, 2024LLM EvalsCEO of HolidayHero - redated by Manouk



