Most AI dashboards answer the easiest questions.
How many people logged in? How many prompts did they send? How many tokens did the company buy? Which model consumed the most credits?
Those numbers help manage access and cost. They do not prove that work improved.
OpenAI’s new admin analytics for ChatGPT Work and Codex reflect an important shift. The product can group AI activity into tasks, show usage and spend, surface tool and skill adoption, and track Codex contributions to merged code. OpenAI also says business owners still need to compare that activity with measures such as review time, defects, rework, delivery time, profitability, and sales outcomes.[1]
That distinction matters.
AI usage is an input. Changed work is the result.
An SME can have high adoption and weak returns. Staff may generate more documents but spend longer correcting them. A sales team may produce account briefs faster without improving customer conversations. A service agent may close simple tickets quickly while creating a larger exception queue for experienced staff.
If the dashboard stops at activity, leaders can mistake motion for progress.
Start with the unit of work
Do not begin with the model. Begin with a repeatable unit of work.
For a sales team, that unit might be one qualified opportunity. For finance, one reconciled invoice. For customer service, one resolved case. For operations, one supplier onboarded with all required checks completed.
The unit must be specific enough to count and inspect. “Improve productivity” is not a unit of work. “Prepare a complete account brief before a sales call” is.
Once the unit is clear, map the current workflow:
- What triggers the work?
- Which systems and data does it require?
- Who performs each step?
- Where do delays, corrections, and escalations occur?
- What marks a completed, acceptable result?
This gives the AI implementation something concrete to improve. It also exposes unnecessary handoffs that should be removed before they are automated.
Establish the baseline before the pilot
A pilot without a baseline produces anecdotes.
Before adding AI, record how the work performs today. Use the measures already available in the ERP, CRM, ticketing system, or operating log. If the business does not have perfect data, start with a small sample and document the limitations.
A useful baseline usually includes:
- Cycle time: How long does one unit take from trigger to completion?
- Active effort: How much human time is spent doing the work?
- Throughput: How many acceptable units are completed in a period?
- Quality: How often does the result pass its first review?
- Exceptions: How often does the standard path fail or require escalation?
- Rework: How much time is spent correcting completed work?
- Business result: What happens after the work is completed?
- Total cost: What labour, software, integration, and oversight does the process consume?
The purpose is not to build a perfect measurement system. It is to prevent the organisation from declaring victory because an AI feature was switched on.
Count the human work that AI creates
Every AI workflow has a review burden. Good implementations reduce or concentrate that burden. Poor ones hide it.
Suppose an agent drafts customer proposals. The generation step may take seconds, but the commercial manager still checks pricing, scope, delivery assumptions, and contract language. If the draft is frequently wrong, the agent has shifted work rather than removed it.
Measure:
- minutes spent reviewing each output;
- percentage approved without changes;
- percentage requiring minor edits;
- percentage rejected or rebuilt;
- recurring error categories;
- seniority of the reviewer required;
- time spent investigating uncertain claims or missing context.
Review is not automatically waste. A five-minute approval may be the correct control for a financial, legal, customer-facing, or production action. The mistake is excluding that five minutes from the ROI calculation.
Human-in-the-loop works only when the human receives enough context, has authority to intervene, and can see what the agent did. Otherwise the approval becomes a ceremonial click.
Track exceptions, not just averages
Averages make unstable workflows look healthy.
An agent may handle most cases quickly but fail badly on a small number of high-value or high-risk cases. Those failures can erase the savings from routine work.
Track the exception rate and classify why the standard path failed. Common categories include missing data, conflicting records, unclear policy, insufficient permission, tool failure, low confidence, and customer-specific judgment.
Then decide what should happen next:
- the agent retries with better context;
- the case moves to a specialist;
- a human approves a proposed action;
- the workflow stops safely;
- the underlying process or data is redesigned.
This is where domain experts become AI architects. They define the boundaries, escalation rules, and evidence needed for a safe decision. The model handles execution inside those boundaries.
Connect activity to an outcome
The final measure must sit outside the AI tool.
For a CRM research agent, do not stop at briefs generated. Compare preparation time, brief quality, meetings held, qualified opportunities, and contribution margin where attribution is credible.
For an invoice-reconciliation agent, look at close time, unresolved exceptions, duplicate payments, late fees, and finance hours released for analysis.
For a customer-service agent, look at resolution time, reopen rate, escalation quality, customer satisfaction, and the complexity of the cases handed to people.
Analytics do not prove causation by themselves. Use a defined test period, compare against the baseline, record other operational changes, and review the result with the process owner. The aim is a defensible operating decision, not a marketing claim.
Use a seven-part operator scorecard
A practical AI workflow review can fit on one page:
| Field | Operator question |
|---|---|
| Work unit | What specific piece of work are we improving? |
| Baseline | How did it perform before AI? |
| AI contribution | Which steps does the agent complete or support? |
| Review burden | What must a person inspect, correct, or approve? |
| Exceptions | Where does the workflow fail, stop, or escalate? |
| Outcome | Which operational or commercial result changed? |
| Total cost | What do models, tools, integration, support, and human oversight cost? |
Review the scorecard on a fixed cadence with the process owner. Do not let the AI team review it alone. The person accountable for the business outcome must challenge the numbers and decide what happens next.
There are only four useful decisions:
- Expand when quality and outcomes improve at an acceptable total cost.
- Redesign when the concept works but data, steps, skills, or handoffs create friction.
- Constrain when the agent is useful only within narrower permissions or case types.
- Stop when review burden, exceptions, risk, or cost outweighs the benefit.
Stopping a weak workflow is not failure. It is evidence that the measurement system works.
Measure the operation, not the excitement
The next phase of AI adoption will create more agents, skills, integrations, and usage data. That can make dashboards look impressive while the underlying operation stays the same.
Leaders need a different question.
Not: “How much AI are we using?”
Ask: “Which unit of work changed, what improved, what new burden appeared, and is the result worth the total cost?”
That is the difference between buying AI and operating an AI-enabled business.
Sources
[1] https://openai.com/index/how-to-connect-ai-usage-to-business-value — How to connect AI usage to business value