Skip to main content

The Results tab

Test Runs shows the number of scenario executions in the period. The period is Last 7 days, Last 15 days or Last 30 days; runs older than 30 days are in cold storage and do not appear here.
The Results tab grouped by run plan, with the pass rate and the trend of each plan
Group by sets what a row is: The filters narrow every grouping: Scenario, Label, Target and All statuses, Passed or Failed. Reset filters clears them. Pass is green at 100%, amber from 40%, red below, and grey with no result. Trend shows the last runs as bars, oldest first, at most 14. Charts opens the tiles Executions, Pass rate, Failing scenarios and Cost for the period, and Pass rate over time.
The Charts block with the four tiles and the pass rate over time

A run plan and its runs

Press a run plan. The sidebar lists its runs as Run #n with the note, the time and the pass rate; Load More… loads older ones. The header of the selected run shows Pass, Completed, Avg Agent Latency, Avg Agent Cost, Total Duration and Total Cost.
A run plan with its runs in the sidebar, the summary pills and the table of results
Show run settings opens the configuration the run used: Started, Targets, Parameters, Repeat, Simulator model and Judge model. The Table view has one row per scenario execution: Result, Scenario, Evaluators and Time · cost. The row menu offers Open the conversation, Rerun this scenario and Edit scenario. The Grid view shows the same executions as cards, one colour per result, each ending in Completed or Failed.
The Grid view of a run, one card per scenario execution
Export as CSV downloads the run. Run again starts a new run with the plan’s configuration, and Stop all cancels the conversations still running.

A conversation

Press a row to open the drawer. The header shows the Status, the criteria count, the Duration, the Cost and when it Ran. The transcript follows, with View trace on every turn whose trace has arrived.
The conversation drawer of a failed run, with the failed criteria and the judge reasoning
Under the transcript: PASSED CRITERIA, FAILED CRITERIA and JUDGE REASONING, then Parameters with the values the run used, a secret parameter masked. Open Scenario opens the scenario in the editor. While the run is going, the drawer shows The conversation is running… and then The judge is reading the conversation…, and the verdict appears in place when it is in.

What a status means

The Result column of the table shows Passed, Failed or Running. The Status chip of the drawer is more precise:
The conversation drawer of a run whose target could not be reached, with the error
A criterion the judge could not decide is inconclusive. A criterion about an internal action goes inconclusive when the agent’s traces did not arrive; see Linking your traces. Also check: Run plans, Compare agents, Run from CI for the batch API behind the same runs.
Last modified on August 30, 2026