Share of AI outputs judged correct against a held-out evaluation set.
Correct outputs ÷ evaluated outputs × 100