Prompt Regression Runner
Same tickets. Different instructions. Visible evidence.
This report compares a vague baseline with a structured candidate across 15 fixed support tickets. Each response is checked for syntax, schema, task correctness, missing information, grounding, and actionability.
Comparison
Candidate Scorecards
Case Explorer
Inspect Why A Candidate Passed Or Failed
Teaching Point
What The Score Means
- Every candidate receives the same tickets.
- Valid JSON does not prove task correctness.
- Aggregate scores point to cases worth inspecting.
- Critical regressions can outweigh an improved average.
Reproducibility
Recorded And Live Are Separate
CI scores reviewed recordings without credentials. An optional live adapter creates new recordings with an explicit model id for separate review.
prompt-regression compare ...
prompt-regression record-live ...