Run academic/industry benchmarks (e.g., MMLU, TruthfulQA, MT-Bench) to evaluate performance.
Runs published benchmarks to position a model's general capability. Useful for model selection, largely uninformative about performance on a specific task.
Initial model shortlisting, and communicating capability to stakeholders who recognise the benchmark names.
Deciding whether a model works for your task. Benchmark performance and task performance correlate loosely and sometimes not at all.
Benchmark results should never be presented as evidence of fitness for a specific purpose. That requires a domain evaluation set.
Benchmark contamination — models trained on data including the benchmark, producing scores that measure memorisation.
Shortlisting three candidate models on published benchmarks, then deciding between them entirely on a domain corpus of real cases.
The source material did not expand this solution into components, and the row says so in its own notes. Expanding it here would be authorship, not research.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.