Release readiness
Stop selling voice AI on cost per call. Model cost per resolved outcome.
Build a voice AI ROI model around verified resolution, repeat contacts, human recovery, vendor spend, risk, and retained revenue instead of deflection.

Voice AI resolution economics model
A CSV worksheet for entering call volume, automation cost, human labor, repeat contacts, handoff recovery, errors, and verified resolution.
voice-ai-resolution-economics-model.csv
Your AI call cost 38 cents. Then the customer called back, spent nine minutes with a human, received a fee reversal, and left a one-star review. The first contact was cheap; the journey was not.
This is for contact-centre leaders and founders building the business case for voice AI. You will get an ROI model that counts verified resolution, repeat effort, human recovery, risk, and revenue impact. No universal savings percentage. Your own operating data goes in.
Deflection is the easiest number to inflate
A call that never reaches a human can mean several things:
- The AI solved it.
- The caller gave up.
- The caller used another channel.
- The caller plans to call again.
- The AI changed the wrong thing and the failure has not surfaced yet.
If all five become “contained,” the dashboard is telling a finance story, not an outcome story.
Choose a follow-up window that matches the workflow. Delivery-status contacts may reveal a failure within a day. A billing adjustment may not surface until the next statement. Document the window, the channels included, and how you connect related contacts without exposing customer data in analytics. If the window is too short, unresolved work appears successful; if it is too long, unrelated contacts can be counted as repeats.
A June 2026 customer-experience thread described bot-first cases that can poison the interaction before a human sees it. The claims in a Reddit thread are anecdotal and often unverifiable. The measurement idea is worth keeping: tag the whole journey, not only the last touch.
Scroll diagram horizontally on smaller screens.
Use cost per verified resolution
Define a resolved outcome for each workflow:
| Workflow | Verified outcome |
|---|---|
| Delivery status | Correct current status communicated, with no avoidable repeat contact within the selected window. |
| Appointment change | Intended appointment updated in the system of record, old slot released, confirmation issued, and no conflicting booking. |
| Restaurant order | Confirmed cart matches the committed POS ticket and payment state. |
Then calculate:
cost per verified resolution =
total operating cost for the lane
÷
verified resolved outcomes
The denominator must come from authoritative evidence. For a read-only status call, compare what the agent said with the system state at that time. For a write, verify the durable record, any required notification, and the absence of a conflicting state. For a handoff, verify that the intended queue or callback path accepted the case and that the human outcome is known.
Unknown outcomes should remain unknown. Do not assign them to success because no complaint arrived. Track why verification failed: missing event, unsupported system, incomplete retention window, or caller disconnect. A falling unknown rate is often as valuable as an early improvement in the resolution number because it tells the team whether the business case rests on observable work.
The cost stack people forget
The numerator must include more than model tokens and telephony. Group the full cost stack so reviewers can see what drives it:
- Technology: telephony and media; speech recognition, model, and synthesis; orchestration fees; tool and data services; monitoring, storage, and evaluation.
- People: human review, handoff labor, and engineering and operations support.
- Failure and recovery: repeat contacts, refunds, credits, chargebacks, and rework caused by errors.
- Governance: compliance and security work allocated to the lane.
Separate fixed costs from costs that vary with volume so scale assumptions stay honest.
Use a loaded human cost that reflects the operating model, not only hourly wages. Include the allocation method for supervision, training, occupancy, facilities or vendor fees, and after-call work. The right method differs by company, but it should be applied consistently to the baseline and the AI-assisted path.
Separate one-time build and integration work from recurring run cost. Buyers need both a payback view and a steady-state view. A pilot can have unattractive unit economics while the team is building evaluation and integrations, yet a proposal that simply omits that work is not honest either. State the amortization period and show the result with and without it.
Give recovery work back to the AI lane
This accounting choice changes decisions. Suppose a bot mishandles a billing dispute and a human spends twelve minutes repairing it. If the human team absorbs that time, the AI lane looks cheap and the human team looks slow. Tag automation_touched=true from first contact through final resolution, then attribute the recovery effort to the full journey.
The containment tax includes human minutes after AI escalation or error, supervisor involvement, repeat contacts, time to repair the business state, and any credits or refunds.
Create journey IDs or privacy-safe linkage that follow the issue across channels. A failed voice interaction may continue in chat, email, a branch visit, or another phone number. If the analysis sees only the original call, the recovery cost disappears. Define how closely the reason and customer state must match before contacts are linked, and sample the result for false matches.
Human recovery time should be split into normal escalation and error repair. A planned handoff for a policy decision is part of the lane design. Time spent reversing a wrong address change or explaining a false promise is a quality cost. Combining them can make a healthy hybrid workflow look inefficient or make a harmful automated path look ordinary.
A practical model
Microsoft’s current agent value blueprints point teams toward first-contact resolution, handle time for escalated cases, loaded representative cost, quality, and revenue effects. That is a better shape than “AI handled 10,000 calls.” Build one weekly record per workflow around four groups:
Demand covers offered calls, answered calls, abandonment before a useful response, and call reason. Keeping them together stops a shift in call mix from disappearing inside one volume number.
Outcome records verified AI resolutions, human resolutions after handoff, failures, unknown results, and repeat contacts. A call that ended after a refund promise is not resolved until the authoritative system shows the approved result.
Cost combines AI charges, allocated platform cost, human minutes, engineering and review allocation, and error remediation. Keep one-time implementation work separate from recurring operations.
Business effect captures orders, appointments, payments, retention events, refunds, credits, and losses. Store the attribution rule beside each value. A call before a purchase does not prove the agent caused it.
That weekly record lets finance inspect the assumptions and operations trace a bad week to a route or call reason.
Build a baseline from the same call reasons, hours, regions, and customer mix. Comparing a voice agent handling simple overnight status calls with humans handling daytime disputes will overstate the benefit. If random assignment is possible and appropriate, a controlled pilot gives a stronger comparison. If it is not, use matched cohorts and state the remaining differences.
Watch for demand that automation creates or uncovers. Better answer rates may capture orders that were previously abandoned, while an easier phone channel can also increase low-value contacts. Report offered demand, completed work, and repeat demand separately. The volume after launch is not automatically the volume the human operation would have seen.
Use ranges for uncertain inputs. Calculate conservative, expected, and optimistic cases for verification rate, repeat contacts, human recovery, and variable vendor cost. Do not change every assumption in the favorable direction at once without showing the combined risk. A simple sensitivity table often reveals that one field, such as repeat-contact rate, matters more than model price.
External claims can seed those scenarios, but they cannot approve them. Avaya’s 2026 ROI and TCO article says strong self-service can deflect 30 to 60 percent of contacts and lower cost per contact by 30 to 50 percent against a human baseline. Those are vendor-published ranges, not a law of contact centres, and one range cannot describe both simple order-status calls and angry insurance disputes. Replace it with observed pilot data before approval.
Keep an evidence field beside every assumption. Mark whether it came from finance, workforce management, provider billing, a pilot observation, or a vendor estimate, then add a date and owner. This prevents a range copied into an early business case from surviving unchanged after the company has better data.
A fictional pilot with honest math
This illustrative example did not happen. A dental group pilots appointment rescheduling on 1,000 first contacts and assigns each one to a single primary path:
| Primary path | Calls | Rate |
|---|---|---|
| Verified appointment change by AI | 620 | 62% |
| Transfer to a human and resolve | 180 | 18% |
| Caller abandons | 90 | 9% |
| First contact remains unresolved and produces a repeat contact | 60 | 6% |
| Outcome remains unknown after the observation window | 30 | 3% |
| Verified failed outcome | 20 | 2% |
In this simplified cohort, the repeat-contact row is mutually exclusive so the six paths sum to 100 percent. If repeat contacts in your reporting can follow any initial path, treat repeat contact as a separate overlay and do not add it into the outcome distribution.
Only 62 percent count as verified AI resolutions. If the question is pure automation economics, the resolution denominator is 620. If the question is whether the complete hybrid rescheduling lane resolved the work, the denominator is 800 and the cost numerator must include the AI and human spend across both successful paths. The 180 calls resolved after transfer are hybrid resolutions, not automated ones. Attach costs to all six paths and the result may still be excellent, but it will also be believable.
Suppose the AI-only path costs less per attempt but the repeat-contact group requires a long human repair. The model will show whether improving that one failure path creates more value than negotiating a slightly lower per-minute model price. This is why journey economics changes the engineering backlog: it points to the defects that consume real money and customer effort.
Repeat the calculation by cohort over several weeks. One launch week can be distorted by novelty, careful staffing, or a narrow invitation list. Track whether verified resolution, transfer recovery, and unknown outcomes remain stable as volume and call diversity grow.
Add quality and risk gates
Some errors are too expensive to price as an average, so create critical gates for:
- Wrong-account action.
- Unauthorized payment or refund.
- Missed mandatory disclosure.
- Unsafe medical or financial claim.
- Lost human handoff.
- Sensitive-data exposure.
NIST’s AI RMF Playbook encourages production monitoring, error documentation, feedback, and attention to error propagation. It is voluntary guidance, not a voice AI ROI formula. The useful connection is governance: a lane with an unresolved critical risk should not win budget approval because its per-call cost is low.
Give critical events a separate incident cost and decision rule. Finance may still need an estimated exposure range, but product should not treat one unauthorized payment as a few dollars added to the weekly average. The release record should show the event, containment action, affected scope, and whether the lane remains open.
Privacy, security, and compliance work also changes with scope. Retaining audio for evaluation, expanding to a new jurisdiction, or adding payment data can add controls and review. Include those costs in the lane where they arise rather than burying them in a company-wide overhead bucket that makes every automation proposal look equally safe.
Compare lanes, not one blended average
Order status may be cheap and safe to automate, while account closure may need a human. A refund under a small threshold may work with deterministic controls, while a disputed fraud case may not.
Report economics by workflow, language, channel, customer segment where lawful, and outcome severity. Blend only after the operators can inspect each lane.
Set a decision cadence before the pilot. A weekly operating view can find failures quickly; a monthly or quarterly finance view can confirm whether savings reached the budget. Name the owner who can narrow a lane when repeat effort rises and the owner who can approve expansion when quality and economics both hold.
The best dashboard connects unit economics with the release evidence behind it. An operator should be able to move from “cost per verified reschedule increased” to the relevant repeat-contact cohort, handoff failures, and regression cases. That connection keeps ROI from becoming a sales slide detached from how the agent behaves.
Define those boundaries with the hybrid automation lane method. For ordering, feed verified results from the restaurant order-state engine into the model instead of counting a polite goodbye as revenue. Voxeval’s voice AI resolution economics model is deliberately blank where your numbers belong, with fields for assumptions, observed values, confidence, and source. The best business case is not “we removed people.” It is “we resolved more of the right work, at a lower full-journey cost, without hiding customer effort.” That claim is harder to calculate and much easier to defend.
Subscribe to Voxeval for grounded voice AI buying and release-readiness guides.
Reference list