Measure how well responses align with user intent and contextual needs.
Scores whether the response actually answers what was asked, as opposed to answering something adjacent. Catches the common failure where a model produces good content about the wrong question.
Conversational and query interfaces where users ask open-ended questions and the failure mode is a well-written non-answer.
Constrained tasks with a defined output shape. If the task is 'extract these six fields', relevance is not the risk.
Where a low relevance score is detected, the honest response is to ask a clarifying question rather than to answer more confidently.
Ambiguous queries scored as low relevance when the real problem was the question, not the answer. The metric blames the model for the user's imprecision.
A natural-language query interface where 'why did the balance move' could mean period-over-period, versus budget, or versus the same period last year — and the system asks rather than picking.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.