AI in customer service is the safe bet, going further holds the real upside

AI in customer service is the safe bet, going further holds the real upside

Benjamin Guy Saunders

·

Vendite

In breve

Customer support is where UK businesses often first deploy AI agents. This makes sense, it’s low risk, and there’s a strong cost optimisation angle. But there's a catch. Traditional metrics and containment strategies that flow from them consistently fail to measure real customer experience, particularly in the AI channel. Furthermore, stopping at support leaves a valuable asset untapped: conversational AI provides an unmediated record of what customers want, in their own words. The best teams are going beyond support, fully leveraging this pure source of insights, and building differentiated personalised customer experiences. (Part 1 of 3)

It's time to talk about how UK businesses are deploying AI

Launching in a new market, you come up against the prevailing conditions. How are companies experiencing the problem your product solves, and do their established habits help or hinder them in seeing a better way? Nebuly is a broad church by design; any agentic conversational product in almost any sector, with clients from farming equipment manufacturers to luxury car makers and health tech in between. A few hundred conversations later, here's what I think is holding UK businesses back.

This is the first of three short pieces. Plenty has been said about the challenges getting from pilot-to-production, governance and the importance of talent, all of it valid; I want to explore a more structural perspective; that putting AI in your customer support function or contact centre is simultaneously the safest bet and the most limited application of the technology. The first part below will unpack this view, before highlighting companies going further - a sort of Hall of Fame, from which the anti-examples should be clear. Part two goes deeper on our client Oura as a standout leader in merging the worlds of customer support, CX and AI; and part three will offer some starter thoughts on strategy.

Part one: twenty years of learning to listen

For more than two decades the contact centre has dominated how UK businesses invest in serving customers. Hardly surprising in an economy where services have grown from around 70% of output to more than 80%. Alongside this trend, marketing evolved from telling to listening as aspirational, product-led gave way to consumer-led and insights driven messaging. With digital and mobile the discipline continued to reorganise around consumer insight, and I watched the tail end of that at Qualtrics, working with names like Samsung, Unilever and Vodafone. Having the best product was no longer enough, "Experience" became the thing that won market share.

An enormous software category grew up to serve the shift: Qualtrics, Medallia, Sprinklr, NICE, ServiceNow, Inmoment, Intercom, now Fin (so many) plus the smaller companies with one valuable capability who got acquired and bolted on. All of them promising to connect insights to action in one way or another. An honest caveat, there are obviously differences in how these providers and others in the space emphasize the way they drive value – from contact centre optimisation to more granular experiential insights – but really, it’s all the same.

An ironic outcome

Here's the part I find ironic. A movement founded on the premise that customer experience is felt, subjective and qualitative ended up encoding itself almost entirely in hard metrics. Not through laziness: a score can be owned, targeted and improved, and a description of how someone felt can't. Sentiment detection doesn't rescue this, it just compresses a different variable into the same shape. No single vendor is to blame either; it's the problem with category building. Efforts to build an ever-better system-of-action compound, and action trumps insight. Twenty years on, CX measures operational performance more closely than it measures customers' actual experience.

That's not to say it isn't valuable. I've seen what Qualtrics and its peers deliver in that operational context, and I know the limitations too (happy to talk about those offline). But the value comes from optimising an operation you can already see: your channels, your queues, your teams. A score is a lossy summary of an experience, and lossy works when journeys are predictable enough that one number stands in for thousands of near-identical interactions, and when there's a convenient moment at the end to ask "how would you rate your experience?" and check your working.

Almost none of that carries over to AI. A good AI experience isn't navigation, it's delegation, and ideally for your customer it's end-to-end. Interrupt them to ask how it's going and you break the magic. And once every conversation is its own journey, an aggregate score has nothing left to aggregate over. So teams fall back on adoption and token counts, which sit downstream of experience quality: they'll tell you something went wrong, eventually, and never what.

Worse, many think they've solved it, and will tell you about their outcome-based framework, which turns out to mean, drumroll, CSAT, as measured by a survey. If the only tool you have is a hammer, everything looks like a nail.

What that conditioning does to your roadmap

Buyers haven’t only built up a toolkit over these two decades, they’ve built a preference for evidence that arrives with a number attached. Business cases get funded when they do, and the cleanest numbers in the business live in the contact centre. Drop AI in there and the case writes itself.

Which is how a cost metric ends up passing as a CX metric. CSAT at least tries to ask about the experience. Containment doesn't pretend to: it counts throughput and gets reported as if it tells you something about how customers were served. Hence strategies named Deflection and Containment, terms you'd sooner associate with a zombie horde than with your paying customers.

The number isn’t even useful for ranking companies. When one celebrates 50% of cases handled by its AI agent and another claims 80%, you'd be forgiven for assuming the second is doing better. Check Trustpilot and you'll find the same pattern either way: doom loops, no route to a human, and "I'm sorry, I don't have that information, please check our website"; all persistent signs of poor AI products serving as the face of your brand.

What the number can't tell you

Three things get blurred here. There's the service being discussed. There's the AI product: the agent itself, whether it grasps the request and knows when it's out of its depth. And there's the interaction data, what the user asked for in their own words and what happened next.

Containment and CSAT distinguish none of them. An 80% rate is equally consistent with a good agent papering over a broken service as with a bad agent trapping people who'd have been better off speaking with a human. So whether any of this is valuable to the user goes unmeasured almost everywhere. The answer lives inside the conversations, and the inherited apparatus was built to listen to something else.

What you're missing out on by only applying the support lens

  • The product signal in the failures. Every conversation your agent can't resolve is a customer telling you, unprompted, what they wanted and couldn't get. A support frame files that as a containment miss. Read properly, it's a signal for your roadmap.

  • The only unmediated record of intent you'll ever own. Surveys tell you what people say when asked. Conversations tell you what they wanted at the moment they wanted it.

  • The ability to personalise anything later. The transcripts sit in your platform, but nobody has worked out what people were asking for or whether they got it, so there's nothing to build on. Real personalisation needs that groundwork, and it takes time to accumulate.

And there's a limit on what support can tell you however well you instrument it. People arrive there because something went wrong, which makes it a good record of friction and a poor one of desire. You'll learn what your customers can't do. What they actually want shows up on the surfaces where they come to you willingly.

There's a final point, the defensive one: if customers get their answers about your category from a general assistant rather than yours, generic and occasionally questionable advice gets given in your name. That's really a strategy question, so I'll save it for part three.

The Hall of Fame

IKEA's Billie is always cited for its containment rate, 47% of queries at first and around 74% now. The interesting number is the other one. Ingka looked at what Billie couldn't resolve, found customers asking for help with home planning, and retrained 8,500 contact centre staff as remote interior design advisers. Those sales centres reportedly turned over €1.25bn last year, up from €1.08bn.

Octopus Energy shows the low-stakes path done honestly. Arlo handled around 8,000 emails a week in trial, roughly 4% of UK customer emails, scoring 76% satisfaction against 72% for comparable human replies. Never vulnerable customers, a person always available. Four percent, published on purpose, by a company that competes on service.

Virgin Atlantic worked backwards from their best human conversation rather than their cheapest. Their highest-value holiday baskets came from physical stores, where staff could sit and talk a customer through their plans. So the digital team sat with those reps, identified the seven contextual preferences the good ones surfaced through casual chat, and set their agent the job of discovering those seven by talking. No if-then-else.

Nebuly

Nebuly is the ROI platform for enterprise AI. It connects to the AI agents your business runs on, the assistants your customers interact with, and the tools your employees use every day, including Claude, ChatGPT, and Copilot, and translates that activity into business value: How much time is being saved across teams. What revenue your AI is influencing. What adoption and AI proficiency look like in practice, across departments and geographies.

If you need clarity on what your AI investment is actually delivering, book a demo.

Iscriviti alla nostra newsletter

Iscriviti alla nostra newsletter

Resta aggiornato su ciò che stiamo imparando, costruendo e osservando mentre i team enterprise distribuiscono e misurano gli agenti AI in produzione.