Most AI buying conversations start with the wrong unit.
How many users need access? How many prompts will they send? How many agents can the platform run? How many tasks can it process?
Those numbers may help a vendor package software. They do not tell an operator whether useful work was completed.
A recent MatrixLabX announcement offers a useful signal. The company launched a financial-services agent platform that it says uses human approval gates, records workflow activity in an audit ledger, and charges per executed, approved workflow rather than per seat.[1]
That is a vendor announcement, not independent proof of customer ROI. It does not establish a universal pricing trend. But it raises the right question for any SME evaluating agentic AI:
What exactly counts as one completed, approved unit of work?
If you cannot answer that before buying the platform, you will struggle to govern the agent, measure value, or challenge the invoice.
Activity is not an outcome
Seats made sense when software value depended heavily on how many people logged in. Agentic systems change the operating model. A digital coworker may run in the background, move across several systems, and involve one human only when judgment is required.
That makes familiar usage metrics weak proxies for value:
- A login does not mean work was completed.
- A prompt does not mean the answer was used.
- A tool call does not mean the action was correct.
- A finished run does not mean the business accepted the result.
- An approval does not mean the work produced a useful outcome.
The distinction matters because four different events can occur:
- Work attempted: The agent started the workflow.
- Work technically completed: The agent reached the end of its programmed path.
- Work approved: An authorised person accepted the proposed action or result.
- Useful outcome produced: The work changed the business result in the intended way.
These are not equivalent. They should not automatically become equivalent billable events.
A failed invoice match, a duplicate supplier record, and an approved but unnecessary purchase order may all look like “completed workflows” inside a dashboard. An operator should see three exceptions, not three units of value.
Define the work before you automate it
Consider a recurring accounts-payable workflow in an SME.
An invoice arrives. The agent extracts the supplier, amount, purchase-order number and payment terms. It checks the ERP for the purchase order and goods receipt. If the documents match within the company’s tolerance, it prepares the invoice for posting. If they do not match, it routes the exception to the right owner.
That description is still too loose. An approved unit of work needs a proper operating definition.
1. Trigger
What starts the work?
For example: a new supplier invoice received through the approved finance inbox, with a unique invoice number and supported file type.
A forwarded duplicate should not create another billable unit.
2. Required inputs
What must exist before the agent can proceed?
This could include the invoice, supplier master record, purchase order, goods receipt, tax fields and approval matrix. Missing inputs should create a visible exception, not encourage the agent to guess.
3. Permitted actions
What may the agent do without asking?
It may read approved records, extract fields, compare documents, prepare a posting and request approval. It may not create a new supplier, change bank details, release payment or alter approval limits unless those actions have separate authority.
4. Completion criteria
What makes the workflow technically complete?
The invoice data must be captured, validated against the required records, checked for duplication, classified as matched or exceptional, and routed to the correct next step.
“Agent finished” is not a completion criterion.
5. Named approver
Who has the decision right?
Approval should belong to a role with the authority, context and accountability to accept the action. The agent may recommend. The named human decides when the consequence crosses the agreed boundary.
Human-in-the-loop is not a button added at the end. It is an operating decision about who owns risk.
6. Evidence receipt
What proof remains after the work?
The receipt should show the source records used, checks performed, exceptions found, proposed action, approver, decision, timestamps and final system response. It should be possible to reconstruct what happened without relying on the agent’s summary.
7. Exception path
What happens when the normal path fails?
Define where missing documents, mismatched quantities, unusual tax treatment, duplicate invoices and unavailable systems go. Name the owner. Set the response expectation. Do not allow difficult cases to disappear into a generic “needs review” queue.
8. Business outcome
What result are you trying to improve?
For this workflow, it could be shorter invoice cycle time, fewer duplicate payments, less manual matching, or faster exception resolution. Choose the measure before deployment so the project cannot quietly redefine success as usage.
Approval is necessary, but not sufficient
Approval can establish accountability. It cannot rescue a badly designed workflow.
A person may approve the wrong recommendation because the evidence is incomplete. They may approve automatically because the queue is too large. They may lack the authority to make the decision. The workflow may be compliant but economically pointless.
This is why the approval step must be designed together with the inputs, boundaries, evidence and outcome measure.
The aim is not to keep a human clicking forever. The aim is to place human judgment where consequence, ambiguity or policy requires it—and make routine work safe enough to execute consistently.
A buyer checklist for outcome-based agent pricing
Before accepting “per workflow,” “per resolution” or “per outcome” pricing, ask:
- Where does the workflow begin and end?
- What exact event creates a billable unit?
- Are rejected results billed?
- Are duplicates, retries and failed runs billed?
- Who decides that the work is approved?
- Can the customer reverse or dispute an approval?
- What evidence is retained for each unit?
- How are exceptions and partial completions classified?
- Can we separate agent activity from accepted work?
- Can we connect accepted work to a business outcome?
- What is the fully loaded unit cost, including human review and integration?
- Who owns remediation when the agent changes the wrong record or triggers the wrong action?
If the answers are vague, the pricing is not outcome-based. It is activity pricing with better language.
Start with one definition
Do not begin by negotiating how many agents you can deploy.
Choose one recurring workflow. Write a one-page definition of its trigger, inputs, permitted actions, completion criteria, approver, evidence receipt, exception path and intended business outcome.
Then test whether the platform can execute that definition and prove each completed unit.
Orchestrate the work before you price the automation.
Sources
[1] https://vermontbiz.com/news/2026/september/17/matrixlabx-launches-governed-agentic-ai-built-mid-market-banks-lenders-asset — MatrixLabX launches governed agentic AI for mid-market finance