Instrumentation of latency, token consumption, error rates, cache performance and quality signals in production.
Operational monitoring of model usage as distinct from evaluation of output quality: end-to-end latency including gateway and retrieval, token consumption by component, error and retry rates, cache hit rate.
Every production deployment.
Not applicable, though what is captured must respect the sensitivity of prompt content.
Telemetry capturing prompt or response content inherits the sensitivity of the source data. Redaction belongs in the capture path, not the dashboard.
Measuring only the provider's reported generation time, which excludes gateway, retrieval and queueing — most of what the user actually experiences.
A p95 latency of nine seconds where the provider reports two, the remainder being retrieval and a cold gateway connection.
Two different absences share this shape. Foundational solutions get built whatever the domain, so no domain links them; the rest are solutions this domain genuinely does not reach for. v_ai_solutions_unlinked separates the two.