An executive with a coffee by an office window, looking out over the city

Five metrics that prove AI is paying.

The board's AI question for 2027.

Boards should judge AI on five metrics: AI-sourced pipeline and its conversion, resolution rate on agent-owned work, cost per resolved outcome, revenue per customer, and adoption by the teams who work with the agents. Each needs a baseline before launch, an owner and a weekly or monthly report. Activity counts do not prove value.

By Aaron Goh, CEO, Azend Group · 2 October 2026 · 4 min read

In 2026, most boards in Southeast Asia asked whether the company was doing enough with AI. In 2027, the question will be sharper: is it paying? Answering that with a list of pilots, licences and usage statistics will not survive the meeting.

The gap between activity and value is wide. McKinsey's 2026 State of AI survey finds only 37 percent of respondents attribute at least some EBIT impact to AI, and that high performers are twice as likely to have defined processes to measure the impact of their AI initiatives (McKinsey). Measurement is not a reporting nicety. It is part of what separates the two groups.

Why are activity metrics not enough?

Conversations handled, emails drafted, prompts run and seats activated tell you the tools are switched on. They do not tell you whether the business is better off. An agent can touch every conversation and resolve very few. A sales team can send twice the emails and book the same number of meetings.

Info-Tech's SoftwareReviews makes the same point about HubSpot's new agents: buyers should separate activity from results by tracking “conversion, cycle time, resolution quality, customer response, and employee effort” (SoftwareReviews).

Framework 01

Five metrics on one board page.

AI scorecard · one page, once a quarterBaseline · Owner · Cadence
  1. Metric 01

    AI-sourced pipeline

    Is pipeline from AI work converting to revenue?

    MonthlyBaseline first
  2. Metric 02

    Resolution rate

    What share of agent-owned work is fully resolved?

    WeeklyBaseline first
  3. Metric 03

    Cost per resolved outcome

    Fully loaded, including credits and review time.

    MonthlyBaseline first
  4. Metric 04

    Revenue per customer

    Retention and cross-sell, with AI support against without.

    QuarterlyBaseline first
  5. Metric 05

    Adoption

    Are people using drafts, taking handoffs and updating content?

    WeeklyBaseline first

A model, not a result: no values are shown because yours start from your own baseline.

Metric 1: Is AI-sourced pipeline converting?

For sales and marketing agents, measure pipeline created from AI-sourced or AI-assisted work, and its conversion to revenue against the rest of the pipeline. If AI-sourced opportunities convert worse, you are adding volume, not value. This requires one thing many CRMs lack: a reliable record of where each opportunity came from. Fix attribution first.

Metric 2: What share of work does the agent fully resolve?

For service, the core measure is resolution rate on the intents the agent owns, alongside the handoff rate and the reasons for handoff. HubSpot's reporting for Customer Agent includes resolutions, deflections, human handoffs and handoff rate over time (HubSpot Knowledge Base). Report resolution on agent-owned intents, not across all conversations, or the number will look worse than reality early and better than reality later.

Metric 3: What does each resolved outcome cost?

Cost per resolved conversation, per qualified meeting or per collected invoice is the number finance will ask for. Include everything: licences, HubSpot Credits, integration, and the time people spend reviewing and improving agents. HubSpot says Customer Agent consumes credits only when a conversation is resolved (HubSpot Knowledge Base), which makes cost per resolution easier to calculate than most. Compare it with the fully loaded cost of the same outcome handled by people.

Metric 4: Is revenue per customer moving?

Efficiency is the first gain. Growth is the bigger one. Track revenue per customer, retention and cross-sell for customers served with AI support against those without. For regional groups, measure across brands and markets, which requires one customer record across the group. This is a slower metric. Report it quarterly, and expect it to lag the others.

Metric 5: Are people actually working with the agents?

Every agent depends on people: reps who review drafts, service teams who handle handoffs, owners who improve the content. Measure adoption where it counts. What share of agent drafts are used? How quickly are handoffs picked up? Is the knowledge base being updated? Is there a named owner reviewing each agent monthly?

Low adoption is the most common reason a technically working agent shows no business result. It is also the earliest warning.

How do you set this up?

  1. Baseline before launch. Record each metric for the process as it runs today. Without a baseline, every later number is an opinion.
  2. One owner per metric, from the business, not the technology team.
  3. Same definitions everywhere. What counts as resolved, qualified or sourced must be agreed once, across markets.
  4. A cadence. Weekly for resolution, handoff and adoption in the first quarter; monthly for pipeline and cost; quarterly for revenue per customer.
  5. One dashboard in the CRM, not a slide assembled by hand.

Keep the set small. Five metrics that are trusted beat twenty that are argued over. Where a use case genuinely needs a different measure, such as days sales outstanding for invoice follow-up, swap it in, but keep the same discipline of baseline, owner and cadence.

How should you read vendor numbers?

Vendors publish averages. HubSpot, for example, reports that teams using Customer Agent close 77% more tickets per month on average, and that teams using Prospecting Agent create 65% more sales leads per month on average (HubSpot). These show direction across many customers. They are not a forecast for your business, and a board should not be shown them as one. Your five metrics, against your baseline, are the only numbers that answer the question.

What are the common traps?

  • Moving the baseline. Re-defining “resolved” or “qualified” after launch makes every comparison meaningless.
  • Counting gross, not net. Pipeline the agent touched is not pipeline the agent created. Be strict about attribution.
  • Ignoring hidden cost. Review time, content upkeep and integration support are real costs and belong in cost per outcome.
  • Averaging across markets. A strong result in one market can hide a poor one in another. Report by market.
  • Reporting too early. Agents improve as context improves. Judge trend over a quarter, not a single week.

What should the board see?

One page, once a quarter: the five metrics for each funded AI use case, against baseline, with a short note on what changed and what is next. Use cases that are not moving their number after two quarters are reviewed or stopped. That discipline is what turns an AI programme into an AI portfolio.

We build these measures into every engagement, and run them on an ongoing basis through managed operations. To see which use cases will pay first, start with an AI readiness assessment.

Questions.

How often should AI metrics be reported?

Weekly for resolution, handoff and adoption in the first quarter, monthly for pipeline and cost, and quarterly for revenue per customer, with a one-page board summary each quarter.

Should boards rely on vendor-reported AI results?

No. Vendor figures are averages across many customers and show direction only. Decisions should rest on your own metrics measured against your own baseline.

Measure what the board will ask.

A strategy call sets the five metrics and baselines for your first AI use cases.