How to do effective AI upskilling

How to do effective AI upskilling

Anton Freyberg

·

Solutions Lead

TLDR

Tracking quantitative adoption is a solved problem. Understanding the quality of conversations, is a problem starting to arise with discussions around ROI and effectiveness of AI. Surveys are not enough for the AI price tag. What you need to understand to ensure your workforce becomes truly AI native: 1. The tasks users (try to) do with AI, and what they never think to use it for. 2. Proficiency visibility by task, department, seniority and region. 3. A root cause: the specific proficiency gaps. Sometimes the answer is a skills file, not a training. 4. Continuous visibility, to rate and improve upskilling measures over time.

Four steps to make AI training specific instead of generic.

Enterprises easily spend six figures for AI enablement, often in the millions. Workshops, a learning path, maybe an external partner. Given this budget, I was surprised to see the difficulties they have in measuring the ROI of the enablement investments.

The pain point is not being able to answer these simple questions:

  • Am I training the right people?

  • Is my training efficient? If so, how and why?

How are we trying to answer these questions today? In my experience: surveys. I hate surveys. They are biased, unreliable, infrequent and cost everyone time.

Among other things, 91% of HR professionals believe employees overstate their own skill proficiency, and only 18% of organisations measure skills regularly at all. [1]

Let’s look at how we can answer these questions automatically, reliably and economically.

Step 1: Start from what people actually do with AI

Not what the pilot committee thinks they do. What is in the conversations.

You need to understand the tasks your users solve and how. Engineering and Product might be using Claude skills, but defaulting to very advanced models. Legal might review and generate documents, while sales summarizes their meeting notes, without using any skills, and does research with vague prompts on expensive models, burning immense amounts of tokens.

Then look at the inverse, which is the part almost everyone skips: what people are not using AI for. If your customer service team never uses AI for anything but rewording emails, that is not a proficiency problem you can see in a score. It is a whitespace problem. The highest-return training is often not “get better at what you do” but “here is a task you did not know to delegate.”

Step 2: Measure proficiency for your organization

Adoption is not proficiency. Lots of people might log in. But are they using AI correctly?

People default to questionnaires to answer this question, but the data sits right there in the conversations. That is what Nebuly’s AI Proficiency score does. It reads how people actually work with AI and scores conversations on the things that define good AI interactions; for example whether the task was stated clearly, whether the person supplied enough context, whether they specified the format or other guidelines, or how many turns it took to get there. All of this is harder to capture than you might think. More on this in a follow-up post. Those signals combine into a single score, grouping users by proficiency, somewhere between novice and expert.

No survey, no test, no self-assessment. Just continuous analytics.

See it by department, seniority, role and region, and the plan writes itself. You stop buying one training for 4,000 people and start buying three for the three groups that actually differ. Maybe no training is needed? Let’s look into it.

Step 3: Diagnose the specific failure, not the general one

“Low proficiency” is not actionable. And a generic course on AI won’t help much. Generic courses are for the average employee who does not exist. We need to understand in more detail what to improve.

Take two departments doing completely different work. Sales spends its day summarizing meeting notes and pushing them into the CRM. Turns out they do not need training at all. We look at how their best people already do it, turn that into a skills file, and hand it to the rest of the team. Same for every other recurring task we find: map what people already do into reusable skills and ship the templates. You just moved a whole department’s fluency and efficiency without losing anyone’s time. Trainer or trainee.

Legal is the opposite case. Nobody ever told them about half the connectors they have access to, so they were asking questions the AI had no chance of answering, missing a lot of context without ever noticing. That one is a real training, and now we know exactly what it has to cover.

Same data, two completely different interventions, and neither of them is a curriculum. One is a template file for one team. The other is one focused session for another.

Step 4: Measure the impact

If you’re responsible for AI upskilling, you need to know what works and what does not. With continuous proficiency scoring, it’s simply a line on a chart. Below is how it looks in Nebuly: AI proficiency for managers across different offices, before and after an upskilling day. The lift shows up in the weeks that follow, per office, without anyone filling in a form. I knew what they were doing (and what they were not), I knew how well they were doing it, and I knew what to fix. The rest was communication. Now I can tell the Chief People Officer exactly how much our training moved, and when.

Sources

[1] Skillsoft, Global Skills Intelligence Survey, September 2025. Survey of 1,000 HR and L&D professionals in the US, UK, Germany and Australia. Both figures are from this survey: 91% of HR professionals believe employees overstate their skill proficiency, particularly in AI, leadership and technical domains; 18% regularly measure skills throughout the talent development journey.

Nebuly

Nebuly is the ROI platform for enterprise AI. It connects to the AI agents your business runs on, the assistants your customers interact with, and the tools your employees use every day, including Claude, ChatGPT, and Copilot, and translates that activity into business value. How much time is being saved across teams. What revenue your AI is influencing. What adoption and AI proficiency look like in practice, across departments and geographies. All aggregated at the organizational level, never tied to individuals.

If you need clarity on what your AI investment is actually delivering, book a demo.

FAQs

How do you measure AI proficiency across an organization?

You measure it directly from the conversations employees already have with AI, not from what they report about themselves. Nebuly's AI Proficiency score reads how people work with AI and scores each conversation on the things that define a good interaction, then combines those signals into a single score per user. You can view it by department, seniority, role, and region, and it updates continuously instead of at survey intervals.

Can you measure AI skills without surveys?

Yes. The data needed to assess proficiency already sits in the conversations, so continuous analytics can replace self-assessment entirely. This matters because surveys tend to capture belief rather than behavior: in a September 2025 Skillsoft survey of 1,000 HR and L&D professionals, 91% said employees overstate their own skill proficiency, and only 18% of organizations measure skills regularly at all.

Is AI adoption the same as AI proficiency?

No. Adoption tells you people logged in, while proficiency tells you whether they get real value from AI. Teams usually reach high adoption quickly, but proficiency varies widely by department and role, which is why usage counts alone do not tell you what to teach or whom to teach it to.

How do you decide who needs AI training?

Start from what people actually do with AI, broken down by department, seniority, role, and region. Proficiency scoring shows which groups genuinely differ, so you can replace one generic course for thousands of people with focused interventions for the few groups that need them. In some cases the data shows a team needs no formal training at all.

When is a skills file better than a training course?

A skills file works when a team already has people doing a recurring task well. You map how the strongest performers do it, turn that into a reusable template, and hand it to the rest of the team, which raises fluency without anyone sitting through a course. Training is the better call when a team is missing capabilities it never knew it had, such as connectors or context it was never shown how to use.

How do you measure the ROI of an AI upskilling program?

Track the proficiency score over time, by group, before and after each intervention. Because the score comes from real conversations rather than forms, the lift from a training day shows up in the following weeks per department or office, giving you a defensible figure to report on what the training actually moved.

How do you find what employees are not using AI for?

Look at the tasks absent from the conversation data, not only the ones present. If a team uses AI for one narrow task and nothing else, that gap is a whitespace opportunity, and it is often where the highest-return training sits, because it points to work people did not know they could delegate.

Subscribe to our newsletter

Subscribe to our newsletter

Stay up to date on what we're learning, building, and seeing as enterprise teams deploy and measure AI agents in production.