Detect social, gender, racial, and cultural biases in LLM outputs.
Tests whether outputs differ systematically across demographic groups when they should not. Usually implemented by perturbing prompts across group markers and measuring whether the response changes in ways that cannot be justified.
Any system whose output affects individuals differentially — hiring, lending, identity verification, claims assessment, content moderation.
Systems with no human subject. Reconciling a bank statement has no demographic dimension, and applying the machinery there is theatre.
Testing must cover the groups actually present in the population served, not a standard list. Results must be reported by subgroup rather than in aggregate, since aggregate accuracy is precisely what conceals concentrated harm.
Aggregate metrics hiding subgroup disparity. A 98% pass rate can contain a 70% rate for one population, and the headline figure will never reveal it.
An identity verification system tested across skin tones and name origins, where face-match and name-extraction accuracy are reported separately for each rather than pooled.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.