Oura, and the standard worth holding your AI to (part 2 of 3)
Oura, and the standard worth holding your AI to (part 2 of 3)

Benjamin Guy Saunders
·
Director UK&I

TLDR
Oura runs two customer-facing AI agents: one for support, and Advisor, the in-app health companion members use to make sense of their own data. On the support side they built a metric that reads nearly every conversation rather than surveying a small sample and holds the AI agent to the same bar as a human. Nobody set a containment target, and the operational numbers improved anyway. (Part 2 of 3)
One bar for the agent and the human
Part one ended with Oura as the outlier worth a piece of its own. Here's why.
Most companies running an AI support agent measure it against the contact centre's own numbers. Oura published something more interesting: a metric that reads the conversations rather than counting them and evaluates each one against the same scorecard regardless of whether the member was talking to the AI agent, or a human being.
It's called TCX, or Total Customer Experience. Four classifiers read the actual conversation and each answers one question: was the issue understood, did the member express frustration, was it resolved or properly escalated, and was their sentiment positive afterwards. An interaction only passes if all four come back positive. Solving the problem but leaving the member unhappy doesn’t count as job well done.
Two decisions in there are worth more than the score.
It deliberately ignores internal notes and backend data, because the member never saw any of it. And it makes no distinction between channels, so a conversation that starts with their AI agent is held to the same bar as one that starts with a Member Care Representative. No separate bot scorecard sitting in its own dashboard where it can look flattering.
This is not the easy path. Roughly a year to build, over a dozen candidate classifiers narrowed to four, and quarterly calibration against a golden set of at least a thousand cases Oura's own experts rated. When the model and the humans disagree, the humans set the standard.
Nobody set a containment target
Since rollout, their TCX score has climbed nine points, repeated troubleshooting has dropped from 40% to around 20%, and median handle time on live chat is down by more than a quarter.
Those are exactly the operational numbers challenged in part one, but critically in Oura’s case they improved because the lens was explicitly qualitative. Nobody set a containment target. Nobody deflected anybody. They read what members experienced, found where it broke, and fixed it, and the experience improved as a consequence.
That's the whole argument, and it's the answer to anyone who thinks measuring experience properly is a luxury you fund after you've hit your cost numbers.
What it surfaces is more interesting than what it scores. Oura say root-causing an emerging issue used to take days or weeks and now takes under 30 minutes. In one case TCX flagged a gap between their automated battery diagnostics and what members were seeing on their own ring & app, which prompted them to update an outdated diagnostic model. A conversation exposed a product model that was wrong, and the product changed.
The surface people come to willingly
Which brings us to the other AI agent, and the reason this isn't a support story.
Oura Advisor is their in-app health companion described as the operating interface for the Oura experience. Members aren't going there because something broke. They're asking why they slept badly, what their readiness score means, whether to train today.
Oura's published numbers on it: members most often discuss sleep and recovery (75%), then stress and resilience (57%), then activity and workouts (45%). 83% find the answers reliable, 60% say Advisor helped them understand something about their own body they hadn't grasped before, and over half report turning an insight into an action with a tangible benefit.
That's a record of what customers want, in their own words, at the moment they wanted it. No survey produces it. No containment rate captures it. And it only exists because Oura built a surface worth talking to and committed to keeping it that way.
Why this is a brand argument
Writing about Oura this month, Scott Galloway put them alongside Apple, Google and Amazon on a specific attribute: companies that control the entire customer experience and own the relationship, turning their brand experience into a moat.
Owning the relationship increasingly means owning the conversation. When a member asks Advisor about their health, it answers in Oura's voice, with Oura's science, based on the member’s own data. When that same question goes to a ChatGPT or Gemini, someone else answers on Oura’s behalf, using generic information and without understanding that member’s unique situation.
A brand that reads its own conversations gets better at the thing that distinguishes it. A brand that doesn't is renting its own expertise back from a model it has no say over.
Three practical thoughts to takeaway
Judge the conversation from the user's side. Read what they experienced and keep the internal metadata out of it. It forces the question everyone's been avoiding: did this person get what they came for?
One bar for the AI and the humans. The moment your agent has its own scorecard, you've created a number that can improve in isolation of the user’s actual experience – these can move in opposite directions, without you knowing.
Calibrate against your own people, and let them lead. A golden set your experts rated, re-run on a schedule. Without it, you've automated the flattery.
And whatever you build, point it at the surface people come to willingly, not only the one they reach for when something's gone wrong.
Part three gets into what that means for competitive advantage, whether you pursue Agent-of-choice, or choice-of-agents or a mix, and what the main questions will be.


Stay up to date on what we're learning, building, and seeing as enterprise teams deploy and measure AI agents in production.



