Engineering Notebook

Prompt Regression Runner

Same tickets. Different instructions. Visible evidence.

This report compares a vague baseline with a structured candidate across 15 fixed support tickets. Each response is checked for syntax, schema, task correctness, missing information, grounding, and actionability.

loading comparison

Comparison

Candidate Scorecards

Loading...

Case Explorer

Inspect Why A Candidate Passed Or Failed

Teaching Point

What The Score Means

  • Every candidate receives the same tickets.
  • Valid JSON does not prove task correctness.
  • Aggregate scores point to cases worth inspecting.
  • Critical regressions can outweigh an improved average.

Reproducibility

Recorded And Live Are Separate

CI scores reviewed recordings without credentials. An optional live adapter creates new recordings with an explicit model id for separate review.

prompt-regression compare ... prompt-regression record-live ...