Human feedback to evaluation
Corrections become candidates, not automatic truth.
This recorded evidence validates a sanitized Review Console export, summarizes reviewer decisions, removes duplicate corrections, and prepares regression cases. Every candidate remains outside the permanent golden set until explicitly reviewed.
Review queue
Candidate Evaluation Cases
Learning signal
Override Reasons
Safety boundary
Promotion Is Deliberate
- Only the three published synthetic sample ids are supported.
- Malformed, non-synthetic, and duplicate review records fail clearly.
- An approval artifact must name the reviewer and selected candidate ids.
- Permanent eval data changes only through a separate reviewed code change.
Reproduce it
From Browser Export To Reviewed Evidence
prompt-regression prepare-feedback support-triage-synthetic-reviews.json --output candidates.json --report review.md
prompt-regression approve-feedback candidates.json --reviewer "Name" --approve CANDIDATE_ID --output approval.json
The second command records approval; it deliberately does not edit
data/cases.jsonl.