AUTOMATED EVALUATION
Test Runs
Run a reliability test collection against a built-in demo or external agent, inspect the deterministic evidence behind every result, and compare it against a baseline.
01
Configure run
Select a pack or suite, then run it against a compatible agent target.
Reliability trend
Version-by-version movement using the same deterministic Goal Completion score and regression threshold as run comparison. New critical failures still force regression.
no packNo completed or errored runs for this pack yet.
02
NO RUN SELECTED
Evidence begins with a run.
Choose the deterministic demo baseline or connect an external HTTP agent using the same scenario evidence workflow.
- Multi-turn Turkish reliability scenarios
- Structured tool-contract evaluation
- No external API key required
