Claude Code usage tracking by LangWatch

Check it out →

The LangWatch Blog

Engineering deep-dives, product updates, and field lessons on evaluating, testing, and observing AI agents in production.

Product Releases

Things you can ask Langy

The real questions teams ask Langy about their agents, and how it answers them from your traces and code, then opens…

Rogerio Chaves · July 23, 2026
Introducing Langy: Your Automated AI Engineer
Product Releases

Introducing Langy: Your Automated AI Engineer

Rogerio Chaves · July 22, 2026
Developer findings

We Red-Teamed Our Own AI Agent and Found 14 Real Bugs

Langy is the AI agent LangWatch runs on its own product. A teammate distracted it with an unrelated coding question,…

Aryan Sharma · July 22, 2026
Developer findings

The PR Hound: how we fixed PR review assignment with a daily agent

We leaned into agentic coding and started shipping far more code than our review process could keep up with, so PRs…

Andrew Joia · July 9, 2026
Article

Claude vs Codex: which is the better background agent?

Some of our preferences and stories on Claude and Codex at LangWatch.

Rogerio Chaves · July 8, 2026
Developer findings

Background Agents on Slack: How we built our own Claude Tag before it was cool

A fleet of background agents runs a chunk of our engineering, each one scoped to one job and living in its own Slack…

Rogerio Chaves · July 5, 2026
More from the blog
July 5, 2026Governance AIRogerio Chaves

EU AI Act compliance: are you affected?

July 3, 2026LLM EvalsManouk Draisma

Cost per successful task: the metric that will decide your AI stack when model subsidies end

June 2, 2026Voice AIManouk Draisma

Introducing: Testing voice agents like you test your chat agents

May 31, 2026Product ReleasesManouk Draisma

The Whole Platform Is Now Open Source: LangWatch May 2026 Update

May 19, 2026Developer findingsManouk Draisma

What happens when two engineering teams just... talk

April 30, 2026Product ReleasesManouk Draisma

LangWatch v3.0 and the April 2026 Product Drop

April 20, 2026Developer findingsAlex Forbes-Reed

Eat Sleep Append Repeat…

April 20, 2026Developer findingsAlex Forbes-Reed

Four Refactors and a Funeral: Migrating a Live System to Event Sourcing

April 20, 2026Developer findingsAlex Forbes-Reed

Internal Product vs Internalised Trauma: Supporting Event Sourced Systems

April 15, 2026Governance AIAryan

Every way your AI agent can be broken (and how attackers actually do it)

April 14, 2026Governance AIRogerio Chaves

Why AI Red teaming is broken (and how we fixed it)

March 27, 2026AgentsSergio Cardenas

How we test Agent Skills with Scenario simulations

March 26, 2026IntegrationsManouk Draisma

Getting to value with LangWatch, faster than ever - how to migrate from Langfuse to LangWatch with Skills.

March 25, 2026Governance AIRogerio

A Note on the LiteLLM Vulnerability

March 25, 2026AgentsSergio Cardenas

Product Managers and leaders are running agent simulations now, and it changing how AI ships

March 24, 2026LLM EvalsSergio Cardenas

Making your AI Agent reliable: Adding Evaluations to your multi-modal agent with LangWatch Skills

March 23, 2026AgentsManouk Draisma

LangWatch Skills: Your coding agent already knows how to test your agent

March 12, 2026Product ReleasesManouk Draisma

Introducing LangWatch MCP: Test and evaluate AI Agents without leaving your workflow

March 6, 2026AgentsManouk Draisma

The Agent Development Lifecycle: Why shipping is the easy part

February 28, 2026Product ReleasesManouk Draisma

The LangWatch February Drop: Cheaper Events, Claude Code, and Multimodal Evals

February 20, 2026Product ReleasesManouk Draisma

New Pricing: AI growth shouldn’t increase your bill

February 10, 2026LLM EvalsManouk Draisma

What is LLM monitoring? (Quality, cost, latency, and drift in production)

February 10, 2026LLM EvalsManouk Draisma

What is Prompt Management? And how to version, control & deploy prompts in productions

February 3, 2026IntegrationsRogerio Chaves

How OpenClaw / ClawBot works behind the scenes - and why agent observability matter

February 3, 2026IntegrationsRogerio Chaves

Instrumenting Your OpenClaw Agent with LangWatch via OpenTelemetry

February 3, 2026IntegrationsRogerio Chaves

How to Use Clawdbot + LangWatch to Monitor Your Agents in Production

February 2, 2026LLM EvalsRogerio Chaves

LLM Evaluations Explained: Experiments, Online Evaluations, Guardrails, and when to use each in 2026

January 31, 2026Product ReleasesManouk Draisma

The LangWatch Monthly Drop: January 2026

January 30, 2026LLM EvalsBram P

4 best tools for monitoring LLM & agent applications in 2026

January 30, 2026LLM EvalsBram P

Arize AI alternatives: Top 5 Arize competitors compared (2026)

January 30, 2026LLM EvalsBram P

Top 10 LLM Observability Tools: Complete Guide for 2026

January 30, 2026LLM EvalsManouk Draisma

Top 5 AI evaluation tools for AI agents & products in production (2026)

January 29, 2026IntegrationsSergio Cardenas

How to test AI Agents with LangWatch & Mastra / Google ADK and ship them reliably

December 30, 2025Voice AIBram P

Top Tools for Evaluating Voice Agents in 2025

December 29, 2025AgentsManouk Draisma

What are the AI Agent Events in 2026: The must-attend conferences for Agentic AI Builders

December 24, 2025Product ReleasesManouk Draisma

Closing the year Strong: December Product Updates

December 23, 2025IntegrationsManouk Draisma

How to do Tracing, Evaluation, and Observability for Google ADK

December 23, 2025LLM EvalsManouk Draisma

Top 5 AI Prompt Management Tools of 2025

December 23, 2025LLM EvalsManouk Draisma

Writing Effective AI Evaluations, that hold up in production

December 12, 2025AgentsManouk Draisma

Why Agentic AI needs a new layer of testing

November 26, 2025Product ReleasesRogerio Chaves

Launch Week Day 5: Better Agents CLI: The reliability layer for the next wave of agent development

November 25, 2025Product ReleasesAryan

Scenario MCP: Automatic Agent Test Generation inside your editor

November 24, 2025Voice AIAndrew Joia

Testing Voice Agents with LangWatch Scenario in Real Time

November 20, 2025AgentsManouk Draisma

A Systematic way of Testing of AI Agents

November 20, 2025Product ReleasesAndrew Garde Joia

Introducing: LangWatch newest Prompt Playground

October 27, 2025Governance AIManouk Draisma & FlagSmith

How LangWatch helps enterprises test, evaluate, and trust their AI before release

October 17, 2025LLM EvalsManouk Draisma

Build vs Buy - Should you build your own LLMOps stack or leverage a purpose-built platform designed for enterprise scale?

October 17, 2025LLM EvalsManouk Draisma

The 6 Best LLM Evaluation Platforms in 2025: Why LangWatch redefines the category with Agent Testing (with Simulations)

October 15, 2025AgentsAndrew Joia

Need-based Context Engineering: Let tests tell you what your AI agent actually needs

October 6, 2025IntegrationsRogerio

The Ultimate RAG Blueprint: Everything you need to know about RAG in 2025/2026

September 26, 2025AgentsAndrew Joia

From Scenario to Finished: How to Test AI Agents with Domain-Driven TDD

September 25, 2025LLM EvalsManouk Draisma

Building Reliable AI Applications: Why Evals (and Scenarios) Are the backbone of trustworthy AI

September 7, 2025LLM EvalsRogerio Chaves

Are evals dead?

September 3, 2025LLM EvalsRogerio Chaves

Essential LLM evaluation metrics for AI quality control: From error analysis to binary checks

August 22, 2025IntegrationsManouk Draisma

Trace IDs in AI: LLM Observability and Distributed Tracing

August 19, 2025AgentsManouk Draisma

The 6 context engineering challenges stopping AI from scaling in production

August 18, 2025LLM EvalsManouk Draisma

LLMOps is the new DevOps, here’s what every developer must know

August 14, 2025IntegrationsManouk Draisma

LLM observability: What is it and why it matters

August 8, 2025LLM EvalsManouk Draisma

GPT-5 Release: From Benchmarks to production reality

August 7, 2025LLM EvalsRogerio Chaves

LLM-as-a-Judge: Using the Panel of Judges Approach to Approximate Human Preference

August 1, 2025IntegrationsManouk Draisma

Observability Framework Design for LLM Apps - The Complete LangWatch Guide

July 18, 2025LLM EvalsManouk Draisma

Top 4 Humanloop Alternatives in 2025

Why Agent Simulations are the new Unit Tests for AI

June 27, 2025AgentsRogerio Chaves

Real-time simulation visualization and debug mode

June 26, 2025AgentsRogerio Chaves

Scripted simulations, evaluations, and guardrails

June 25, 2025AgentsManouk Draisma

Customer Story: How Roojoom automates AI Agent Quality Control with LangWatch Scenario

June 25, 2025IntegrationsRogerio Chaves

Test agents on Mastra, Agno, and 10+ other frameworks

June 24, 2025Product ReleasesRogerio Chaves

Introducing simulation-based agent testing

June 24, 2025AgentsRogerio Chaves

Why LangWatch Scenarios represents the future of AI agent testing

June 21, 2025AgentsRogerio Chaves

Best AI Agent Frameworks in 2025: Comparing LangGraph, DSPy, CrewAI, Agno, and More

Multilingual AI Agent Testing: Using Scenario to Simulate, Break, and Improve LLMs

June 18, 2025LLM EvalsManouk Draisma

LangSmith Alternatives: What to use if you need more security and control

Intro to Scenario (Testing AI agents)

Simulations from First Principles (How to test your agents)

Agent Evaluation: Framework for Testing AI Agents

Simulation Based Eval Framework

Introduction: The Real Issue isn’t RL

Simulations to Test My Agent

May 15, 2025Developer findingsAlex Forbes-Reed

New Python SDK Brings Native OpenTelemetry to GenAI Observability

May 5, 2025Product ReleasesManouk Draisma

April Product Recap: Selene Integration, Eval Wizard Upgrades, Prompt Studio & More

May 5, 2025LLM EvalsManouk Draisma

LLM Monitoring & Evaluation for Real-World Production Use

April 24, 2025AgentsTahmid Tapadar

Systematically Improving RAG Agents

April 22, 2025Product ReleasesRogerio

Introducing the Evaluations Wizard: How to evaluate your LLM: Building an LLM evaluation framework that actually works

April 18, 2025IntegrationsManouk Draisma

Function Calling vs. MCP: Why You Need Both - and How LangWatch Makes It Click

April 18, 2025IntegrationsManouk Draisma

Why LLM Observability is Now Table Stakes

April 17, 2025LLM EvalsManouk Draisma

LangWatch vs. LangSmith vs. Braintrust vs. Langfuse: Choosing the Best LLM Evaluation & Monitoring Tool in 2025

April 8, 2025Product ReleasesRogerio Chaves

Introducing Scenario: Use an Agent to Test Your Agent

April 4, 2025LLM EvalsManouk Draisma

Tackling LLM Hallucinations with LangWatch: Why Monitoring and Evaluation Matter

April 3, 2025LLM EvalsManouk Draisma

LLM evaluations at Swis for Dutch government projects by LangWatch

April 2, 2025LLM EvalsManouk Draisma

Why Your AI Team Needs an AI PM (Quality) Lead

March 27, 2025Governance AIManouk Draisma

LangWatch and adesso join forces: Accelerating Secure LLM Adoption for Enterprises

March 25, 2025LLM EvalsManouk Draisma

LLMOps Is Still About People: How to Build AI Teams That Don’t Implode

March 20, 2025LLM EvalsManouk Draisma

Practical LLM Evaluation Framework for AI Development Teams

March 16, 2025IntegrationsManouk Draisma

What is Model Context Protocol (MCP)? And how's LangWatch involved?

March 14, 2025IntegrationsManouk Draisma

How PHWL.ai uses LLM Observability and Optimization to Improve AI Coaching with LangWatch

February 25, 2025Product ReleasesManouk Draisma

LangWatch.ai - Announcing - €1M funding round to bring the power of Evaluations and Auto-Optimizations to AI teams.

February 20, 2025Governance AIManouk Draisma

OpenAI, Anthropic, Deepseek and other LLM Providers keep dropping prices: Should you host your own model?

January 1, 2025AgentsRogerio

7 Predictions for AI in 2025: A CTO's, Rogerio Chaves Perspective

December 20, 2024LLM EvalsCEO of HolidayHero - redated by Manouk

Customer Stories: HolidayHero AI start-up <> LangWatch

December 10, 2024Product ReleasesRogerio

LangWatch Optimization Studio - Built for AI Engineers, by AI Engineers

November 10, 2024Product ReleasesManouk Draisma

The power of MIPROv2 (DSPy) in a Low-Code environment with LangWatch’s Optimization Studio

November 7, 2024LLM EvalsManouk Draisma

What is Prompt Optimization? An Introduction to DSPy and Optimization Studio

July 27, 2024IntegrationsZhenya

Deploying an OpenAI RAG Application to AWS ElasticBeanstalk

July 3, 2024AgentsRogerio - CTO

The complete guide for TDD with LLMs

June 27, 2024LLM EvalsRogerio - CTO

Data Flywheel: Using your production data to build better LLM products

June 11, 2024LLM EvalsManouk Draisma

How Algomo reduced AI hallucinations with LangWatch

June 10, 2024LLM EvalsManouk Draisma

The AI Team: Integrating User and Domain Expert Feedback to Enhance LLM-Powered Applications

June 10, 2024LLM EvalsRogerio Chaves - CTO

Unit Testing Your LLM: The Power of Datasets

June 3, 2024Product ReleasesRogerio - CTO

Introducing DSPy Visualizer

May 20, 2024Product ReleasesManouk Draisma

New Dutch Startup, LangWatch, brings much-needed quality control to GenAI

May 14, 2024IntegrationsZhenya

How to build a RAG application from scratch with the least possible AI Hallucinations

May 13, 2024IntegrationsManouk Draisma

LLM Reliability with Retrieval-Augmented Generation

May 13, 2024Governance AIManouk Draisma

Safeguarding Your First LLM-Powered Innovation: Essential Practices for Security

May 10, 2024LLM EvalsManouk Draisma

What is User Analytics for LLMs, The Difference With Traditional Analytics, And Why is it Important?

May 8, 2024LLM EvalsManouk Draisma

Unlocking the Potential of Large Language Models: The LLM's Beyond the Hype

May 6, 2024LLM EvalsManouk Draisma

The 8 Types of LLM Hallucinations

May 1, 2024LLM EvalsManouk Draisma

5 Things You Must Consider Before Putting Your Chatbot Live in Production

May 1, 2024LLM EvalsManouk Draisma

Navigating the Complexities of AI-Powered Products

April 29, 2024LLM EvalsManouk Draisma

Understanding Hallucinations: What are they?

April 18, 2024LLM EvalsManouk Draisma

Mastering the GenAI Wave: Strategies for Success in AI Adoption

April 18, 2024LLM EvalsManouk Draisma

Successfully building an AI Startup in the current booming industry

April 17, 2024LLM EvalsManouk Draisma

How Struck.build improved AI Performance with LangWatch

April 8, 2024LLM EvalsManouk Draisma

Journey Through Innovation: The LLM Adventure

Date coming soonLLM EvalsManouk Draisma

Webinar recap: LLM Evaluations: Best Practices, LLM Eval types & real-world insights