The Hidden Cost in Your AI Agent’s P&L: Bad Conversation Design
AI leaders are spending a lot of time thinking about tokens.
Which model should we use? How can we reduce inference costs? Can we route simpler tasks to cheaper models? How do we optimise prompts and negotiate better API rates?
Those are legitimate questions. But a recent McKinsey analysis suggests they may not be where some of the biggest costs are hiding.
In “Where AI agents pay off: A practical guide to the economics of agentic workflows,” McKinsey examines the emerging unit economics of AI agents and makes a striking observation: for an AI agent performing a customer service task in banking, token costs can represent just 20–25% of variable run costs, while human oversight can account for 70–75%.
That oversight includes functional and risk experts reviewing agent activity, handling exceptions and ensuring that automated work is completed appropriately. McKinsey’s broader point is important: organisations need to look beyond the visible cost of AI and understand the fully loaded economics of getting work done.
We agree.
But there is another question worth asking:
What causes avoidable exceptions and human intervention in the first place?
For customer-facing AI agents, one answer is surprisingly familiar.
Conversation design.
We’ve seen this problem before
Anyone who worked with automated voice systems twenty-five years ago will recognise some of what is happening today.
The voice industry learned through experience that simply asking a caller:
How can I help you?
wasn’t always the most effective way to begin an automated interaction.
It sounds natural. It gives the customer complete freedom. And technically, it appears elegant.
But humans don’t always behave the way system designers expect.
Give callers an entirely open-ended prompt with no examples, context or conversational scaffolding and some will hesitate. Others will provide answers that are unnecessarily complex or ambiguous. Some will say nothing useful at all.
The industry learned to use directed and mixed-initiative prompting: retaining natural interaction while providing enough guidance to help the caller understand what the system can do and what sort of response is expected.
It was a usability lesson learned through years of real-world VUI design, call-in testing and production tuning.
Then generative AI arrived.
Large language models dramatically improved the ability of machines to interpret natural language, and rightly created excitement about moving beyond rigid menus and traditional IVR structures.
But somewhere along the way, an important distinction became blurred:
Being able to understand almost anything a customer says doesn’t necessarily mean that asking them anything, in any way, produces a good conversation.
The technology has changed enormously.
Human beings haven’t.
We’re seeing the evidence in production interactions
This isn’t just a historical observation.
In recent production interaction data we’ve analysed, the majority of conversations reaching one open-ended intent-gathering experience contained fewer than 60 words, with many ending around the initial exchange — before the AI agent reached a meaningful resolution or handoff point.
That doesn’t mean every short interaction represents failure.
But when you analyse the conversations themselves, patterns start to become visible: customers hesitating, failing to engage with the opening prompt, producing responses the experience doesn’t handle effectively, or abandoning the interaction before any useful outcome is achieved.
At a dashboard level, these can simply appear as:
abandonment, failed intent, exception, escalation or incomplete interaction.
At the conversation level, you can begin to understand why.
And that distinction matters economically.
Bad conversation design is no longer just a UX problem
In the traditional IVR world, a poorly designed interaction was primarily viewed as a customer-experience problem.
Customers became frustrated. They pressed zero. They asked for an operator. They abandoned the call.
Those things certainly had a cost, but it was often difficult to connect an individual design decision to a specific financial outcome.
AI agents change the equation.
McKinsey estimates that customer-facing AI agents in some banking environments can cost $20,000–$30,000 annually for a single-agent workflow and $100,000–$200,000 for a multiagent team. (McKinsey & Company)
More importantly, McKinsey argues that organisations shouldn’t optimise solely around model and token costs. They should redesign workflows to reduce exception rates and simplify review processes that require expensive human intervention.
McKinsey does not suggest that conversation design accounts for all — or even most — of those oversight costs. Risk review, governance, compliance and functional oversight will remain necessary in many agentic environments.
But its analysis exposes an important economic truth:
Every avoidable exception matters.
If poor conversation design creates ambiguity, abandonment, unnecessary escalation, repeat contact or human intervention, it adds cost to an operating model in which human involvement is already expensive.
Bad VUI design used to be primarily a UX problem.
In the agentic era, it can also become a unit-economics problem.
Containment isn’t completion
There is another important idea in McKinsey’s analysis that we think deserves more attention.
McKinsey argues that the metric that ultimately matters is “completed work ROI.”
In other words, don’t simply measure what it costs the AI agent to perform its part of a process. Measure the fully loaded cost of actually completing the customer’s job — whether that’s opening an account, resolving a claim or closing a sale — across AI agents, deterministic systems and humans.
For customer-facing AI, that distinction is critical.
Containment isn’t necessarily completion.
An AI agent can technically contain an interaction without resolving the customer’s underlying requirement.
It can successfully classify an intent but ask the wrong follow-up question.
It can complete its workflow but leave the customer confused.
It can avoid an immediate escalation only for the customer to call again tomorrow.
And it can report a successful interaction while creating downstream human work elsewhere in the organisation.
So the real question shouldn’t simply be:
Did the AI handle the interaction?
It should be:
“Was the customer’s job actually completed — and what did it take to complete it?”
That is a much harder question.
But it’s also where the real economics live.
AgentOps can tell you what happened. Conversation intelligence can help tell you why.
McKinsey argues that organisations will increasingly need an AgentOps discipline analogous to FinOps — continuously monitoring AI economics, performance and the allocation of work rather than treating an agent as something that is built once and left alone.
We think there needs to be a corresponding discipline at the interaction layer.
Traditional operational monitoring can tell you that:
escalation increased, completion deteriorated, abandonment rose or human intervention became more expensive.
Conversation intelligence can help explain why.
- Was the customer confused by the opening prompt?
- Did the agent ask an ambiguous question?
- Did it fail to recognise a perfectly reasonable customer response?
- Was the right knowledge unavailable?
- Did authentication introduce unnecessary friction?
- Did the conversation repeatedly loop?
- Was the human escalation appropriate?
- Did the interaction appear successful, only for the customer to contact again because their underlying issue wasn’t actually resolved?
Those aren’t primarily infrastructure questions.
They’re interaction questions.
And answering them requires analysing what actually happened inside the conversations.
The lesson isn’t new. The economics are.
- The traditional voice industry had a discipline for this.
- Design the experience.
- Call into it.
- Observe how real people behave.
- Analyse where interactions fail.
- Tune the prompts and dialogue.
- Test again.
- Then continue monitoring after deployment.
The difference today is that AI allows that discipline to operate at a scale that wasn’t previously practical.
Instead of manually listening to a small sample of failed calls, organisations can analyse every AI interaction and continuously identify patterns in intent, dialogue, abandonment, escalation, resolution, repeat contact and customer outcome.
That creates a feedback loop:
Measure → Diagnose → Improve → Re-measure
And that loop becomes increasingly important as AI agents themselves continue to evolve.
McKinsey makes exactly this broader point about agent economics: workflows that make economic sense today may not tomorrow, and vice versa, as model capability, costs and operating practices change. Organisations therefore need an ongoing review process rather than a one-off implementation exercise.
The same should apply to the conversations those agents are having.
The takeaway
Tokenomics is having its moment, and rightly so.
Model selection, inference costs, orchestration and infrastructure all matter.
But the economics of a customer-facing AI agent aren’t determined by token costs alone.
They are ultimately determined by what it takes to successfully complete the customer’s job.
If poor conversation design creates confusion, abandonment, unnecessary escalation, repeat contact or human intervention, then conversation design isn’t merely a UX consideration.
It’s part of the unit economics of AI.
And that means it deserves the same continuous measurement and optimisation discipline organisations are beginning to apply to models, infrastructure and agent orchestration.
The industry has forgotten this lesson before.
The difference this time is that it’s finally showing up on a spreadsheet.
________________________________________
CXEX AutoInsights provides the intelligence layer to measure, understand, optimise and assure customer interactions across human, automated and AI-driven experiences — helping organisations understand not just whether an interaction completed, but why it succeeded or failed and what should improve next.
