AI solutionsShared across all subject areas

Standardized Benchmark Executors

Run academic/industry benchmarks (e.g., MMLU, TruthfulQA, MT-Bench) to evaluate performance.

Description

Runs published benchmarks to position a model's general capability. Useful for model selection, largely uninformative about performance on a specific task.

When it fits

Initial model shortlisting, and communicating capability to stakeholders who recognise the benchmark names.

When it does not fit

Deciding whether a model works for your task. Benchmark performance and task performance correlate loosely and sometimes not at all.

Governance requirement

Benchmark results should never be presented as evidence of fitness for a specific purpose. That requires a domain evaluation set.

Characteristic failure

Benchmark contamination — models trained on data including the benchmark, producing scores that measure memorisation.

Example

Shortlisting three candidate models on published benchmarks, then deciding between them entirely on a domain corpus of real cases.

AI solution components

Deliberately empty

The source material did not expand this solution into components, and the row says so in its own notes. Expanding it here would be authorship, not research.

AI opportunity solutions

Deliberately empty

Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.