Aurora Hill Advisors
Menu
← All insights

Measurement

How to measure AI adoption when usage is not the goal

A workflow-level scorecard for measuring AI adoption through behavior, quality, risk, and operating results, with baselines and owners.

By

Aurora Hill Advisors

Reading time

9 min

Published

Measure AI adoption at the level where the work changes.

Licenses, active users, training completion, and prompt volume are useful measures of access and activity. They cannot tell a leader whether a customer request was resolved faster, a research brief became more reliable, or a review process became safer. Those results belong to workflows, not tools.

A practical adoption scorecard gives each AI-enabled workflow seven things:

  1. an accountable owner;
  2. a defined population of qualifying work;
  3. a baseline;
  4. a behavior signal;
  5. an outcome measure;
  6. a quality and risk threshold; and
  7. a decision rule.

The scorecard should help a team decide whether to scale, revise, or stop the workflow. If it cannot support one of those decisions, it is probably reporting activity rather than adoption.

Usage answers a narrower question

Usage metrics answer useful questions:

  • Did people receive access?
  • Did they try the product?
  • Are they returning?
  • Which features do they use?
  • Where does activity drop?

Microsoft’s AI adoption category in Adoption Score illustrates this scope clearly. It measures Copilot use as a daily habit, based on active days across selected Microsoft 365 applications during a 28-day period. That is a concrete product-usage measure. It does not claim to establish workflow quality or business value.

Trouble starts when a narrow activity metric is made to carry a broad operating claim. A team can have frequent use and weak outcomes. A specialist group can create strong outcomes with modest usage. Low use can also be rational for work that occurs infrequently or has little suitable AI exposure.

Use activity as an early signal, then connect it to the workflow.

A Gallup study fielded February 4–19, 2026 among 23,717 US employees found workplace AI use was associated with task fit and support from leaders and managers. The survey establishes associations rather than causal effects. It reinforces the need to define which work is genuinely eligible before interpreting a low usage rate.

Define the denominator before choosing the metric

“Sixty percent adoption” is meaningless without the denominator.

Possible denominators include licensed people, trained people, eligible people, active days, qualifying cases, completed workflows, or teams. Each creates a different result. Choose the denominator that matches the operating question.

For example, imagine that 40 people can access a system, 20 have a role that should use it, and those 20 handled 100 qualifying cases this month. The system appeared in 65 cases, and people completed the approved review path in 52.

  • License adoption: 20 eligible users out of 40 licensed users.
  • Case adoption: 65 AI-assisted cases out of 100 qualifying cases.
  • Correct workflow adoption: 52 correctly completed cases out of 100 qualifying cases.

All three figures are true. Only the third directly answers whether the approved method became part of the work.

Write the denominator beside every percentage on the dashboard. This small rule prevents many false comparisons.

Build one scorecard per workflow

The workflow is specific enough to measure and stable enough to manage. Start by naming its trigger and outcome.

“Use AI for communications” is too broad. “Turn an approved research packet into a first draft, verify every cited claim, route legal exceptions, and secure final editorial approval” can be observed from beginning to end.

Use the following scorecard:

Field Question Example evidence
Owner Who can change, expand, or stop the workflow? Named business owner and review authority
Population When and for whom should this method apply? Teams, case types, volume, exclusions
Baseline How did the workflow perform before the change? Cycle time, rework, quality, cost, risk events
Behavior Is the intended method being used correctly and repeatedly? Qualifying cases, correct handoffs, sustained use
Outcome Did a result important to the business improve? Faster resolution, lower rework, better conversion, backlog reduced
Quality and risk What cannot deteriorate while the outcome improves? Factual accuracy, customer harm, policy compliance, exception rate
Decision rule What evidence leads to scale, revision, or shutdown? Threshold, review date, decision owner

The owner matters because measurement without decision authority becomes a reporting ritual. The population matters because a pilot with six enthusiasts should not be described as adoption by a 600-person function. The baseline matters because an impressive number after launch may be worse than the process it replaced.

Separate four layers of evidence

Different metrics mature at different speeds. Organize them into four layers rather than forcing them into one blended score.

1. Access and habit

Track licenses, permissions, trained users, trial, active days, and repeat use. These measures help identify basic access and awareness problems.

They are leading signals. They should trigger questions, not settle the case.

2. Workflow use

Measure what share of qualifying cases used the intended method, whether required human reviews occurred, and whether people followed the exception path. Also record where people bypassed, abandoned, or modified the workflow.

This layer distinguishes product activity from operating adoption. It may require event data, a case tag, a small sample, or a supervisor review. Perfect instrumentation is rarely necessary at the start.

3. Outcome and quality

Choose one primary outcome tied to the workflow’s purpose and pair it with at least one quality threshold.

Examples include:

  • median resolution time, with a customer-correction threshold;
  • research cycle time, with a source-verification threshold;
  • proposal throughput, with an approval and win-quality threshold;
  • invoice processing cost, with duplicate-payment and exception thresholds;
  • claims-review time, with accuracy and escalation thresholds.

Avoid reporting only averages when a tail matters. A faster median can coexist with a small number of severe failures. Where consequences vary, inspect the distribution and the exceptions.

4. Governance and system health

Record policy exceptions, access failures, data-quality problems, model or workflow version, review completion, incidents, and unresolved risks. NIST’s AI Risk Management Framework treats governance as cross-cutting and risk work as continuous across the system lifecycle. The framework is voluntary and focused on trustworthiness and risk, so it complements rather than replaces business-outcome measurement.

A workflow can meet its speed target and still fail a governance threshold. Keep those facts visible rather than netting them into a composite score.

Use one primary outcome and a small set of guardrails

A dashboard with more measures can still leave the decision undefined. Give each workflow one primary outcome that reflects its reason for existing.

Then add only the guardrails needed to prevent a misleading win:

  • one quality measure;
  • one risk or policy measure;
  • one adoption measure; and
  • the full operating cost.

Suppose the primary outcome is time from customer question to approved answer. A workable scorecard might use:

  • median approval time as the outcome;
  • factual-correction rate as the quality guardrail;
  • required-review completion as the risk guardrail;
  • correctly completed cases divided by qualifying cases as adoption; and
  • licenses, model use, support, review time, and maintenance as cost.

That set can support a decision. Twelve loosely related productivity indicators usually cannot.

A communications example

Consider a fictional corporate communications team that prepares weekly issue briefs for executives. The existing workflow takes a median of 12 staff hours from request to approval. Editors return 22 percent of briefs for missing or weak source support. During high-volume weeks, the queue forces some briefs to miss the requested delivery window.

The team introduces an AI-assisted workflow for source triage, outline development, and first-draft preparation. A researcher selects the source set. The system drafts only from that set and attaches citations. An editor verifies every material claim, and legal or reputational exceptions follow a separate path.

The workflow scorecard could read:

  • Owner: Director of corporate communications.
  • Population: Recurring internal issue briefs that use public sources; crisis response and privileged material excluded.
  • Baseline: 12 median staff hours, 22 percent returned for source problems, 18 percent delivered late.
  • Behavior: Share of qualifying briefs that follow the research, draft, verification, and approval path without an undocumented shortcut.
  • Outcome: Median staff hours from request to approved brief.
  • Quality threshold: No increase in material corrections after approval; fewer than 10 percent returned for source support.
  • Risk threshold: Every material claim has a checked source, and every required review is recorded.
  • Decision rule: Expand after two full review cycles only if the outcome improves and both thresholds hold.

The numbers above are illustrative. The design matters: one workflow, one baseline, one decision, and explicit boundaries around what the method should not handle.

The same record can feed a three-ledger transformation review when a central team needs to reconcile workflow, adoption, and value evidence across a portfolio.

Set a cadence that matches the decision

Review early signals frequently and durable outcomes less often.

During a pilot, a team might check access failures, support requests, bypass reasons, and quality samples weekly. It may review outcome and cost monthly after the workflow has enough volume. A rare, high-consequence workflow may need case-by-case review rather than a monthly average.

At each review, ask:

  1. Did the workflow boundary or population change?
  2. Did people use the intended method for qualifying work?
  3. Which bypasses were rational, and which show a design or capability problem?
  4. Did the primary outcome move relative to baseline?
  5. Did quality, risk, and cost remain inside the agreed bounds?
  6. What decision follows from the evidence?

The final question should produce one of four answers: continue the test, revise a named part of the workflow, expand to a defined population, or stop.

If behavior is weak and the cause is unknown, use a stalled-rollout diagnostic before choosing an intervention. If decision rights are unclear, resolve ownership of the AI-enabled workflow before adding more metrics.

Avoid five common measurement errors

Comparing people who do different work. Usage differences may reflect role exposure rather than willingness or skill. Compare qualifying work and relevant cohorts.

Calling time saved a financial return. Explain whether the time reduced overtime, absorbed more volume, cleared a backlog, improved service, or changed staffing. Otherwise report capacity released, not dollars saved.

Ignoring the cost of review. Human oversight is part of the operating model. Count it, then improve the workflow without hiding it.

Moving the baseline after launch. Version the baseline and record material changes in volume, staffing, policy, or demand.

Letting a composite score hide a failure. A quality or risk breach should remain visible even when adoption and speed are strong.

The executive test

For every material AI workflow, an executive team should be able to answer:

  • What work changed?
  • Who owns the result?
  • Which cases should use the new method?
  • What was the baseline?
  • How do we know the method is being followed?
  • Which outcome improved?
  • What quality or risk threshold must hold?
  • What did the workflow cost to operate?
  • What evidence would make us scale, revise, or stop it?

If those questions can be answered, usage data becomes useful context. If they cannot, a larger usage number will not close the gap.

Sources