Monitor for emergent behavior shifts, degraded outputs, or risky trends in LLMs over time.
Watches production behaviour over time for gradual change — rising bias scores, drifting output style, degrading trust signals, emerging behaviours not present at launch. Continuous monitoring rather than point-in-time evaluation.
Long-running deployments, particularly on provider-hosted models that change without notice underneath you.
Short-lived or pinned-version deployments where nothing is changing.
Drift alerts need a defined response path. An alert nobody owns is a log entry.
Drift so gradual that each period's change falls below threshold while the cumulative shift is substantial. Baselines must be absolute, not rolling.
A monitored deployment where the proportion of generated explanations citing specific evidence declines month over month, none of the individual drops significant, the trend clear.