Decision #008
Nightly eval harness on golden conversations
Prompt and model changes currently rely on spot checks. Twice this quarter a "harmless" prompt tweak degraded reroute confidence calibration and we found out from dispatcher complaints.
A nightly job replays golden conversation suites (curated from real,
anonymized transcripts) against the active prompt versions and scores
grounding, refusal correctness, and suggestion acceptance-rate proxies into
eval_runs. A score drop beyond threshold pages the on-call and blocks
prompt activation.
Prompt regressions surface overnight instead of via customers. Golden suites need curation — stale suites will drift; ownership and refresh cadence are open questions to settle before acceptance.