Decision #008

Nightly eval harness on golden conversations

proposed decide by 2026-08-15 by bruno

Prompt and model changes currently rely on spot checks. Twice this quarter a "harmless" prompt tweak degraded reroute confidence calibration and we found out from dispatcher complaints.

A nightly job replays golden conversation suites (curated from real, anonymized transcripts) against the active prompt versions and scores grounding, refusal correctness, and suggestion acceptance-rate proxies into eval_runs. A score drop beyond threshold pages the on-call and blocks prompt activation.

Prompt regressions surface overnight instead of via customers. Golden suites need curation — stale suites will drift; ownership and refresh cadence are open questions to settle before acceptance.