How to measure the ROI of ChatGPT, Copilot, and the other AI tools your team already uses

How to measure the ROI of ChatGPT, Copilot, and the other AI tools your team already uses

TLDR

-> Enterprises have shifted decisively from building their own AI to buying off-the-shelf tools like ChatGPT, Copilot, and Claude. -> Those tools report usage, not business value, so most of that spend is measured by adoption alone. -> AI ROI becomes a real number when each interaction is translated into hours saved and the money those hours represent. -> Measured on the same terms across every tool, usage data turns into a business case you can compare and act on.

For a few years, the ambitious move in enterprise AI was to build your own. Train a model, own the stack, keep the advantage in-house. That instinct has reversed.

In 2025, 76% of enterprise AI use cases were bought rather than built, up from 53% a year earlier, according to Menlo Ventures' State of Generative AI in the Enterprise. Companies now run day-to-day work on commercial tools like ChatGPT Enterprise, Microsoft 365 Copilot, and Claude.

Buying is the rational call. Off-the-shelf tools reach production in weeks instead of years, and they improve on their own. The tradeoff arrives afterward. You are now spending real money on software you did not build, and the feedback you get back is a usage dashboard.

That dashboard tells you a rollout took hold. It does not tell you what the business got in return.

Usage analytics answer "is anyone using this," not "is it worth it"

The reporting inside ChatGPT Enterprise, Microsoft 365 Copilot, and Claude is built to show adoption. Active users, prompts sent, features touched, all trending over time. That information has a job. It confirms people are showing up.

It stops one step short of the question finance actually has to answer. Adoption is an activity metric. A high prompt count tells you people are typing. It does not tell you whether those prompts removed an hour of manual work, replaced a slower process, or produced output the business valued.

AI ROI is time first, then money

AI ROI, in business terms, is simple to define. It is the time an AI interaction removes from a task, multiplied across users and tasks, converted into money using a cost of time. Hours saved on one side, the value of those hours on the other.

None of that lives in a usage dashboard, because usage dashboards were built to prove a rollout happened. Turning the same activity into hours and value takes a different measurement, one that estimates the time each action saves rather than only counting the actions.

With several tools, the numbers stop lining up

Most enterprises do not run one AI tool. They run ChatGPT in some teams, Copilot inside Microsoft 365, Claude in others, and often an internal assistant or two on top. Each arrives with its own dashboard and its own definition of engagement.

That leaves leaders with reports that do not reconcile and no way to answer a basic portfolio question. Which of these tools is actually producing value, and where should the next dollar of budget go? Answering it takes a measurement layer that sits above any single vendor and reads all of them the same way.

What changes when usage becomes hours and value

Once every interaction carries an estimated time cost, the same data that filled the adoption dashboard becomes a business case. You can see total hours saved, hours saved per user and per task, task completion rate, and the money that time represents. You can put one tool next to another on identical terms. You can back the spend with a return instead of a volume chart.

Because Nebuly sits as a measurement layer between your users and whichever models they use, it translates raw AI usage into hours and business value automatically, across ChatGPT, Copilot, Claude, and your internal assistants at once. A proprietary model estimates the time each action removes from a task, then rolls it up into hours saved and value created that finance can read. The question stops being how much a tool was used and becomes how much it returned.

Want to see this in practice? Book a demo with us.

FAQs

What is AI ROI, and how is it measured?

AI ROI is the business value an AI tool returns relative to its cost. In practice it is measured by estimating the time each AI interaction removes from a task, adding that up across users and tasks to get hours saved, then converting those hours into money using an employee cost of time. Usage counts alone, like active users or prompts sent, are not ROI, because they measure activity rather than value.

Why isn't the usage data from ChatGPT or Copilot enough to prove ROI?

Those admin dashboards are built to show adoption: how many people use the tool and how often. That confirms a rollout took hold, but it never translates into hours saved or money returned. Two companies with identical prompt counts can get very different value from the same tool, and usage data cannot tell them apart.

How does Nebuly calculate hours saved and value saved?

Nebuly uses a proprietary model to estimate the time each action removes from a task, then multiplies that across usage and applies an employee cost of time to produce a monetary value. While the model computes estimates specific to your workspace, it starts from sensible defaults of 15 minutes per action and $75 per hour. Both the hourly cost and the per-action time estimates are adjustable at any time in ROI Settings, and changes apply across all reports immediately.

Can you measure ROI across ChatGPT, Copilot, and Claude in one place?

Yes. Nebuly is model agnostic, so it reads usage across different AI tools and internal assistants and applies the same measurement to all of them. That means you can compare hours saved and value created tool by tool, instead of stitching together separate vendor dashboards that each define engagement differently.

Is Nebuly's ROI reporting available now?

ROI reporting is available in beta on both v1 and v2 of the platform. You can turn it on from Settings, under Feature Flags, by enabling ROI Analysis. After you enable it, the model needs a short period to compute time-per-action estimates for your workspace, and uses default values in the meantime.

Subscribe to our newsletter

Subscribe to our newsletter

Stay up to date on what we're learning, building, and seeing as enterprise teams deploy and measure AI agents in production.