Use standardized or custom rubrics (e.g., BLEU, G-Eval) to score model performance.
Applies a defined scoring rubric across multiple criteria, either classical metrics or model-as-judge evaluation against written criteria. The workhorse of systematic evaluation, and the mechanism behind most evaluation suites.
Comparing prompt or model changes systematically, and anywhere quality must be tracked over time rather than assessed once.
Novel tasks where nobody yet knows what good looks like. A rubric encodes an existing judgement; it cannot substitute for forming one.
Where a model judges model output, the judge's own biases and failure modes must be understood and periodically checked against human scoring.
Rubric drift, where the criteria stop matching what users actually value and scores improve while satisfaction does not.
A reconciliation narrative rubric scoring whether the explanation names the amount, the cause and the correcting entry — the three things an auditor will look for.