Proportion of merged changes where an AI agent produced most of the code.
Share of AI-proposed changes a developer accepts without substantial rework.
Proportion of automated tests passing on a given run.
Share of scheduled pipeline runs completing successfully and on schedule.
Proportion of inference volume run asynchronously at batch rather than interactive rates.
Time work spends unable to proceed, and why.
Elapsed time from an incident to the customer's business process being correct again, not merely the service being available.
Share of deployments causing degraded service requiring remediation.
Proportion of post-incident actions completed and verified as effective in a later period.
Total inference spend divided by tasks that produced an accepted result.
Infrastructure, model inference and support cost attributable to one tenant.
Proportion of identified high-risk scenarios with explicit test coverage.
Occurrences where a request returned or could return data belonging to another tenant.
Elapsed time from work starting to being complete.
How long open decisions have blocked progress.
Defects per merged change, compared between AI-authored and human-authored code.
How long cross-team or external dependencies have been outstanding.
How often changes reach production.
Share of customers who could use a capability and actually do.
Proportion of the permitted unreliability used so far in the window.
Defects reaching production that verification should have caught.
Forecast total effort or cost to finish, based on progress and consumption so far rather than on the original plan.
Share of evaluation cases that previously passed and now fail after a model, prompt or retrieval change.
How long service degradation persists after a failed change.
Proportion of tests producing inconsistent results on identical code.
Revenue retained excluding expansion — pure churn and contraction.
Share of generated statements traceable to a cited source record.
Share of AI proposals a reviewer rejects or materially edits.
Count of production incidents, weighted by severity rather than pooled.
Model spend attributable to a single tenant.
Elapsed time from code commit to that code running in production.
Reduction in human hours on a process attributable to a capability.
Elapsed time from an incident beginning to it being detected.
Revenue from the existing customer base this period against the same base a year ago, including expansion, contraction and churn.
Pages received per on-call engineer per rotation, and how many arrive outside working hours.
Elapsed time from contract signature to the tenant operating in production.
The response time 95% of model calls fall under.
The response time 95% of requests fall under.
Physical progress against effort spent, reported as a pair rather than as a single figure.
Share of input tokens served from cache rather than processed fresh.
Proportion of incidents that are repeats of a previously resolved cause.
How much requirements change after work has entered delivery.
Reviewer time spent per AI-authored change against human-authored.
Time changes spend waiting for review rather than being worked on.
Proportion of completed work that returns for further change.
Work introduced into a sprint after commitment.
Work added to an initiative after its baseline was agreed.
Share of time or requests where the service met its objective.
How reliably a team delivers what it committed to at sprint start.
Work completed per sprint by a given team, used to forecast that same team's near-term capacity.
Proportion of work entering delivery that met agreed readiness criteria.
Engineering capacity consumed answering support and operational queries.
Share of AI outputs judged correct against a held-out evaluation set.
Accumulated deferred work, weighted by the risk it carries rather than counted.
Deferred work actually resolved, against new debt accrued.
Proportion of code exercised by automated tests.
Proportion of time test environments are usable when needed.
Number of work items completed per period.
Elapsed time from a customer starting onboarding to first achieving a meaningful outcome.
Elapsed time from a critical vulnerability being disclosed to it being remediated in production.
Input, output and cached tokens consumed by one completed task.
Proportion of engineering capacity consumed by interrupts, incidents and unplanned fixes.
Proportion of automated outcomes users reject, override or abandon mid-flow.
How much started work remains unfinished at a point in time.
How long currently in-progress items have been open.
Proportion of a customer's relevant workflow actually running through the capability.