AUTOMATED EVALUATION

Test Runs

Run a reliability test collection against a built-in demo or external agent, inspect the deterministic evidence behind every result, and compare it against a baseline.

LAST 20 RUNS
01

Configure run

Select a pack or suite, then run it against a compatible agent target.

Agent target
Agent mode
02

NO RUN SELECTED

Evidence begins with a run.

Choose the deterministic demo baseline or connect an external HTTP agent using the same scenario evidence workflow.

  • Multi-turn Turkish reliability scenarios
  • Structured tool-contract evaluation
  • No external API key required