What executives need to see from their AI investments
What executives need to see from their AI investments

TLDR
→ McKinsey found 92% of executives plan to increase AI spending. Gartner found only 39% are confident their current AI investments will have positive financial impact. The gap between investment intent and confidence in returns is a measurement problem. → The questions executives actually ask - is this generating return, which teams are getting value, are governance controls working, where should we scale - are not answered by uptime figures or session counts. → Effective executive AI reporting requires four views: segmented adoption by department and role, business impact by use case, risk and governance signals from behavioral patterns, and trend data over time. → McKinsey found AI high performers are three times more likely to have strong senior leadership ownership. That ownership requires the right data - adoption depth, hours saved, governance signals, and business impact - not aggregate activity reports. → The reporting infrastructure that makes this possible requires baselines before deployment, segmentation built into the data structure from day one, and a regular review cadence connected to business data rather than IT metrics. Updated on 6th July 2026
According to McKinsey, 92% of executives expect to increase AI spending over the next three years. The same research notes that they can no longer simply spend without expecting results. The question that is increasingly landing in board meetings and leadership reviews is not whether the organization has deployed AI. It is what the AI is actually doing for the business.
This is a harder question than it sounds. Most enterprises have reasonable visibility into whether their AI systems are running. Uptime is tracked. API calls are logged. Usage counts are recorded. What most lack is the view that connects AI activity to business outcomes — the data that tells a CFO whether the AI investment is generating return, tells a COO which departments are actually embedding AI into their workflows, and tells a CISO whether the governance controls are working in practice.
According to Gartner, only 39% of technology leaders are confident that their enterprise's current AI investments will have a positive impact on financial performance. The confidence gap is not about the technology. It is about the measurement.
What executives are actually asking
The questions that enterprise leaders ask when reviewing AI performance are specific and operational. They are not answered by session counts or system uptime.
Is our AI investment generating return? This requires connecting AI activity to business outcomes: hours saved per department, task completion rates, resolution rates for customer-facing agents, and commercial signals from AI conversations. Without this connection, the investment sits on the balance sheet as a cost with uncertain benefit.
Which parts of the organization are getting real value? Enterprise AI adoption is rarely uniform. Some departments embed AI into daily workflows within weeks of deployment. Others show nominal access with minimal real use. Knowing which is which at the department and role level is what makes investment decisions defensible. Without segmented visibility, leadership is making decisions based on company-wide averages that mask the variation that actually matters.
Are our governance controls working? In regulated environments, this is not a hypothetical question. Gartner found that only 23% of IT leaders are very confident in their organizations' ability to manage security and governance when deploying AI tools. Behavioral patterns in AI conversations — employees including sensitive data, queries approaching policy limits, outputs that approach compliance boundaries — are only visible through conversation analytics, not infrastructure monitoring.
Where should we invest next? Whether to expand a deployment, fix a specific failure pattern, or pause a use case that is not generating return requires data at the task and department level. Without it, scaling decisions are made on intuition rather than evidence.
The four views that answer these questions
Effective AI reporting to leadership is built from four distinct views, each answering a different executive question.
Adoption by department and role. Not total session count, but segmented adoption showing which teams have genuinely embedded AI agents into their workflows and which have nominal access. The meaningful measure is return rate — whether users come back after their first interaction — and session depth, whether usage is deepening over time toward more complex tasks. A department with 90% of employees registered but a 20% return rate is a very different situation from one with 60% registered and a 75% return rate. Both can produce the same aggregate usage number.
Business impact by use case. For internal AI agents: hours saved per department, measured against a pre-deployment baseline. Task completion rate for the specific workflows the agent was deployed to support. AI proficiency growth by team, showing whether employees are using the AI for increasingly sophisticated work or keeping it for simple queries. For customer-facing agents: intent resolution rate, escalation timing, churn signals by interaction category, and revenue influence from commercial signals in customer conversations.
Risk and governance signals. The behavioral patterns that precede compliance incidents: PII appearing in employee prompts, queries approaching policy boundaries, output patterns that suggest the AI is being used outside its intended scope. These signals appear in conversation data before they manifest as incidents. Tracking them continuously, segmented by department and by interaction type, is the governance layer that converts theoretical policy into operational oversight.
Trend data over time. Single-point data tells you the current state. Trend data tells you whether the investment is working. Adoption rates climbing in month two and three indicate genuine embedding. Declining session depth after an initial adoption period is a leading indicator of silent churn. Governance risk signals increasing in a specific department indicate either a training gap or a policy clarity issue. Leadership decisions about where to scale, where to intervene, and where to pause require trend data — not snapshots.
The difference between reporting and accountability
There is a meaningful difference between producing an AI report and creating AI accountability.
An AI report describes what the systems are doing. Uptime percentage, session volume, user counts. This is useful for operational oversight and is the data that most AI dashboards currently provide.
AI accountability connects what the systems are doing to what the organization is trying to achieve. It answers whether the investment is generating the productivity gains that justified it, whether the governance controls are sufficient for the regulatory environment, and whether the scaling decisions being made are based on evidence or assumption.
McKinsey's research found that AI high performers are three times more likely to have strong senior leadership ownership and engagement than organizations seeing limited returns. That ownership requires the right information. Leaders who can see adoption depth by department, hours saved by function, governance signal trends, and business impact by use case are in a position to own the AI programme rather than monitor it from a distance.
Iveco Group built this kind of reporting from the start of their AI copilot deployment across 35,000 employees. The result was more than 100 times more feedback data than their previous manual review process had generated — not because more data was collected, but because the data was structured around the questions that mattered: where is the AI being used, for what tasks, how effectively, and what is it worth.
Building the reporting infrastructure
Three conditions make executive-level AI reporting practical rather than aspirational.
Baselines established before deployment. Hours saved can only be measured against how long the same tasks took before the AI agent existed. Adoption depth can only be assessed against who was expected to use the tool and for what. Without pre-deployment baselines, post-deployment reporting describes activity rather than change.
Segmentation built into the data structure from day one. Aggregated company-wide metrics are easy to produce and easy to report. Segmented data by department, role, geography, and use case requires deliberate instrumentation. Organizations that build this segmentation in from the start have the data structure needed for meaningful leadership reporting at 90 days. Those that retrofit it later are always behind.
Regular cadence connected to business data. An AI performance review that runs quarterly and is disconnected from business metrics produces interesting data. An AI performance review that runs monthly, alongside CRM data, support cost data, and productivity baselines, produces decisions. The cadence and the data integration are what convert measurement from a reporting exercise into a management tool.
Nebuly
Nebuly is the ROI platform for enterprise AI. It connects to the AI agents your business runs on, the assistants your customers interact with, and the tools your employees use every day, including Claude, ChatGPT, and Copilot, and translates that activity into business value. How much time is being saved across teams. What revenue your AI is influencing. What adoption and AI proficiency look like in practice, across departments and geographies. All aggregated at the organizational level, never tied to individuals.
If you need clarity on what your AI investment is actually delivering, book a demo.
FAQs
What metrics do enterprise executives need to assess AI ROI?
The metrics that connect AI activity to business outcomes vary by deployment type. For internal productivity agents: hours saved per department measured against pre-deployment baselines, task completion rate, return rate and session depth by team, and AI proficiency growth showing whether employees are using the AI for increasingly complex work. For customer-facing agents: intent resolution rate, escalation timing, churn signals by interaction category, and revenue influence from commercial signals in customer conversations. Aggregate metrics like total session volume and system uptime describe activity, not return.
Why does segmented AI adoption data matter more than company-wide averages?
Company-wide averages mask the variation that determines where AI investment is actually working. Two organizations can show identical company-wide adoption rates while having completely different distributions underneath: one with deep, productive use concentrated in a few departments and nominal engagement everywhere else, and one with genuine embedding across functions. Investment decisions about where to scale and where to intervene require seeing the variation, not the average. Segmentation by department, role, geography, and use case is what makes AI reporting actionable rather than descriptive.
How should executives think about AI governance reporting?
Governance reporting for AI should surface behavioral patterns in employee AI interactions, not just technical security alerts. The patterns that precede compliance incidents — employees including sensitive data in prompts, queries approaching policy limits, output patterns outside intended scope — appear in conversation data before they manifest as incidents. Gartner found only 23% of IT leaders are very confident in their organizations' ability to manage AI governance. The organizations in that 23% are those with behavioral visibility into how employees actually use AI tools, not only technical monitoring of system performance.
What is the difference between AI activity reporting and AI accountability?
Activity reporting describes what AI systems are doing: sessions, uptime, usage counts. It is useful for operational oversight and is what most AI dashboards currently provide. Accountability connects what AI systems are doing to what the organization is trying to achieve. It answers whether the investment is generating the productivity gains that justified it, whether governance controls are sufficient, and whether scaling decisions are based on evidence. McKinsey's research found that AI high performers are three times more likely to have strong senior leadership ownership — which requires the second type of reporting, not the first.
What infrastructure is needed to produce reliable executive-level AI reporting?
Three conditions make executive-level AI reporting practical. Pre-deployment baselines: hours saved and adoption depth can only be measured against what came before. Segmentation built in from day one: retrofitting department and role-level segmentation after deployment produces incomplete data. A review cadence connected to business data: AI performance data reviewed alongside CRM outcomes, support costs, and productivity baselines produces decisions. The same data reviewed in isolation produces interesting observations. The cadence and data integration are what convert measurement from a reporting exercise into a management tool.


Stay up to date on what we're learning, building, and seeing as enterprise teams deploy and measure AI agents in production.



