Scale, fix or retire your AI agents: how to justify continued AI investment

Scale, fix or retire your AI agents: how to justify continued AI investment

Scale, fix or retire your AI agents, Nebuly title card

TLDR

Once budgets tighten, a program-wide ROI average gives the AI team nothing to protect, so AI agents have to be justified one at a time. Each agent's verdict (scale, hold, fix or retire) comes from two questions: does it return more than it costs, and do people want it? There are no universal benchmarks for those thresholds, so agree your own with finance before the review. Strong demand with weak returns usually means fix, not retire.

It's budget season, and the AI team has twelve agents in production across HR, IT, sales and legal. This is an illustrative example, not a client. Usage is up and, on average, the program looks healthy. Then finance asks for a 20% cut, and the team realizes it can't say which agents produce the value.

That's the question this article answers: how to decide, agent by agent, where AI budget should go next.

How do enterprises justify continued investment in AI agents?

One agent at a time. For each agent, compare the value it produces with what it costs, and check whether people actually use it. That gives a clear verdict for every agent: scale it, fix it or retire it. Finance can support a set of evidenced decisions far more easily than one blended ROI figure for the whole program.

Why a program-wide ROI number stops being enough

A single figure for the whole program averages strong agents and weak ones together. When budgets tighten, an average gives you nothing to protect. A flat cut lands on every agent, including the ones producing most of the value.

Per-agent evidence changes the conversation. Instead of defending the program, the AI team proposes where money should move.

What metrics show an AI agent is delivering value?

Four signals, collected the same way for every agent:

  • Demand. How many of the intended users come back to the agent each month. Repeat use is the measure, because logins and seat counts overstate it.

  • Task success. The share of conversations where the agent actually completed the job. Implicit feedback is the best guide here: behavior such as rephrasing, correcting or abandoning that shows how an interaction went without the user rating it.

  • Value. For employee agents, time saved on successful tasks, converted at net hourly cost (the employee's hourly cost, net of what the AI costs to run). Customer-facing agents need their own value rule agreed with finance, because revenue influenced is not margin and can't be set against cost directly.

  • Cost of ownership. What the agent costs to own each month beyond running it: licenses, upkeep, and the people who support it. Running cost is already netted out of value, so it isn't counted again here.

How to decide: scale, fix or retire

The decision rests on two questions. Does the agent return more than it costs? And do people want it?

Verdict

Value compared with cost

Demand

What to do

Scale

Above cost and rising for two quarters

At or above target

Roll out to more teams or use cases, and fund it

Hold

Above cost, but the scale conditions aren't met

Any

Keep funding at the current level and watch the trend

Fix

Below cost, or falling for two quarters

At or above target

Find what is failing (content, integrations, scope) and review next quarter

Retire

Below cost

Below target, after a round of fixes

Shut it down and move the budget to an agent marked scale

Demand is what separates fix from retire. A weak agent that people keep trying to use is telling you there is a real need, and showing you where it falls short. A weak agent nobody uses offers neither.

One case sits outside the table: an agent that returns well above its cost but serves a small specialist team. Low usage doesn't make it a bad investment. Keep it, and judge it on value per user rather than total usage.

There are no universal benchmarks, so set your own

No industry standard says an agent should reach a certain success rate or return a certain multiple of its cost. Use cases differ too much for that. So the rules have to rest on lines you can defend without a benchmark.

Two lines do that. Break-even is the first: an agent either returns more than it costs to own, or it doesn't. The second is a demand target you set when the agent launches, expressed as the share of its intended users you expect to come back each month. Checked in this order, they give every agent exactly one verdict:

  1. Fix when value is below cost, or has fallen for two quarters, and demand is at or above target.

  2. Retire when value is below cost, demand is below target, and one round of fixes hasn't moved either. If no fix has been tried yet, treat it as fix.

  3. Scale when value has been above cost and rising for two consecutive quarters, with demand at or above target.

  4. Hold in every other case where value is above cost.

Write the demand targets down at launch, agree the rules with finance before the first review, and apply them the same way to every agent.

Give new agents a realistic ramp

An agent's first months rarely reflect what it will return. Gartner's September 2026 survey of 160 finance leaders found that even simple finance AI use cases, such as data extraction, accounts payable and receivable automation, and report creation, generally deliver returns within nine to 10 months. More complex use cases take longer.

Agents still inside their ramp aren't eligible for retire. Until they have had that time, the question is whether demand and task success are moving in the right direction, not whether value already covers cost.

An example portfolio review

The figures below are illustrative. Value is time saved multiplied by a net hourly cost of $75. Each agent launched with a demand target of 30% of its intended users returning each month.

Agent

Returning users a month (share of intended)

Task success

Time saved (hours a month)

Value a month

Licenses, upkeep and support a month

Verdict

IT ticket triage

1,900 (63%)

81%

940

$70,500

$9,000

Scale

Sales proposal drafting

450 (56%)

84%

610

$45,750

$7,000

Scale

HR policy assistant

3,200 (40%)

58%

180

$13,500

$16,000

Fix

Legal clause checker

60 (8%)

49%

15

$1,125

$5,000

Retire

IT triage and proposal drafting both return several times their cost, have grown for two quarters, and beat their demand targets, so they get more budget. The HR assistant has the most returning users but returns less than it costs. Retiring it would waste clear demand, and its low success rate points to the fix: the policy content it answers from. The legal clause checker is past its ramp, has already had a round of fixes, and still has neither demand nor value, so its budget moves to ticket triage.

Who owns the review, and how often

  • The AI leader owns the portfolio and proposes each verdict.

  • Finance agrees the thresholds and the value assumptions, ideally before the first review.

  • Use case owners explain the context behind each agent's numbers and own the fix plans.

Run the review every quarter. Monthly is too noisy for most agents, and once a year is too slow to move money to where it works.

How to prove your AI investment is working in the budget review

Proof is evidence finance can check, agent by agent. Bring:

  1. One line per agent with its verdict.

  2. The four signals behind each verdict: demand, task success, value and cost of ownership.

  3. The thresholds and assumptions you agreed in advance.

  4. The trend for each agent across the last few quarters.

  5. The reallocation: what moves from retired agents to the ones you are scaling.

This turns a defensive conversation about the whole program into decisions finance can support, agent by agent. For how to lay the numbers out, see building an AI agents value dashboard.

Evidence for every agent

A scale, fix or retire decision is only as good as the evidence behind it. Nebuly gives AI leaders that evidence for every agent, from the conversations themselves: who keeps coming back, which tasks get completed, and what each agent returns in time saved or revenue influenced. The portfolio review becomes a set of decisions you can defend, one agent at a time.

Know which agents to scale, fix, or retire before your next budget review: book a demo.

FAQs

How do you justify continued investment in AI agents?

Evaluate each agent on its own evidence instead of the program average. Show demand, task success, value and cost per agent, then recommend scaling, fixing or retiring it. Finance can support a set of clear decisions far more easily than one blended ROI figure.

How do I prove our AI investment is actually working?

Measure outcomes from real usage: whether tasks were completed, time saved for employee agents, and revenue influenced for customer-facing ones. Usage and seat counts show activity, which is a different thing from value.

When should you retire an AI agent?

When its value stays below its cost, returning users stay below the target set at launch, and a round of fixes hasn't changed either. Agree these thresholds with finance before the review, and give new agents the ramp their use case realistically needs.

How long should an AI agent get before you judge its ROI?

Longer than most teams expect. Gartner found that even simple finance AI use cases generally deliver returns within nine to 10 months, and complex ones take longer. Review quarterly, but judge against the timeline the use case realistically needs.

Are there benchmarks for when to scale an AI agent?

No universal ones. Success rates and returns vary too much by use case. Build your rules on two lines you can defend: break-even against the cost of ownership, and a demand target set at launch. For example, scale only when value has been above cost and rising for two quarters, with demand on target.

What is implicit feedback?

Implicit feedback is behavior in a conversation that shows how it went without anyone clicking a rating. Rephrasing, correcting an answer or abandoning a task are all examples.

Subscribe to our newsletter

Subscribe to our newsletter

Stay up to date on what we're learning, building, and seeing as enterprise teams deploy and measure AI agents in production.