Compare outputs from different model versions for quality regression or tuning.
Replays a fixed set of prompts against two model or prompt versions and compares outputs, scores and behaviour. The mechanism by which a model upgrade becomes a decision rather than a leap.
Any model, prompt or retrieval change reaching production. Provider model updates make this recurring rather than occasional.
Where no stable evaluation set exists to replay, in which case the comparison has nothing meaningful to say.
Aggregate improvement must not be allowed to mask regression on high-consequence cases. Those cases need their own gate, passed independently of the average.
A new version scoring better overall while failing a specific important category, which ships because the headline number improved.
A provider model upgrade evaluated against a curated corpus of prior reconciliation cases, where a 3% aggregate improvement is accompanied by a regression on multi-currency explanations.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.