Metric 1: Is AI-sourced pipeline converting?
For sales and marketing agents, measure pipeline created from AI-sourced or AI-assisted work, and its conversion to revenue against the rest of the pipeline. If AI-sourced opportunities convert worse, you are adding volume, not value. This requires one thing many CRMs lack: a reliable record of where each opportunity came from. Fix attribution first.
Metric 2: What share of work does the agent fully resolve?
For service, the core measure is resolution rate on the intents the agent owns, alongside the handoff rate and the reasons for handoff. HubSpot's reporting for Customer Agent includes resolutions, deflections, human handoffs and handoff rate over time (HubSpot Knowledge Base). Report resolution on agent-owned intents, not across all conversations, or the number will look worse than reality early and better than reality later.
Metric 3: What does each resolved outcome cost?
Cost per resolved conversation, per qualified meeting or per collected invoice is the number finance will ask for. Include everything: licences, HubSpot Credits, integration, and the time people spend reviewing and improving agents. HubSpot says Customer Agent consumes credits only when a conversation is resolved (HubSpot Knowledge Base), which makes cost per resolution easier to calculate than most. Compare it with the fully loaded cost of the same outcome handled by people.
Metric 4: Is revenue per customer moving?
Efficiency is the first gain. Growth is the bigger one. Track revenue per customer, retention and cross-sell for customers served with AI support against those without. For regional groups, measure across brands and markets, which requires one customer record across the group. This is a slower metric. Report it quarterly, and expect it to lag the others.
Metric 5: Are people actually working with the agents?
Every agent depends on people: reps who review drafts, service teams who handle handoffs, owners who improve the content. Measure adoption where it counts. What share of agent drafts are used? How quickly are handoffs picked up? Is the knowledge base being updated? Is there a named owner reviewing each agent monthly?
Low adoption is the most common reason a technically working agent shows no business result. It is also the earliest warning.
How do you set this up?
- Baseline before launch. Record each metric for the process as it runs today. Without a baseline, every later number is an opinion.
- One owner per metric, from the business, not the technology team.
- Same definitions everywhere. What counts as resolved, qualified or sourced must be agreed once, across markets.
- A cadence. Weekly for resolution, handoff and adoption in the first quarter; monthly for pipeline and cost; quarterly for revenue per customer.
- One dashboard in the CRM, not a slide assembled by hand.
Keep the set small. Five metrics that are trusted beat twenty that are argued over. Where a use case genuinely needs a different measure, such as days sales outstanding for invoice follow-up, swap it in, but keep the same discipline of baseline, owner and cadence.
How should you read vendor numbers?
Vendors publish averages. HubSpot, for example, reports that teams using Customer Agent close 77% more tickets per month on average, and that teams using Prospecting Agent create 65% more sales leads per month on average (HubSpot). These show direction across many customers. They are not a forecast for your business, and a board should not be shown them as one. Your five metrics, against your baseline, are the only numbers that answer the question.
What are the common traps?
- Moving the baseline. Re-defining “resolved” or “qualified” after launch makes every comparison meaningless.
- Counting gross, not net. Pipeline the agent touched is not pipeline the agent created. Be strict about attribution.
- Ignoring hidden cost. Review time, content upkeep and integration support are real costs and belong in cost per outcome.
- Averaging across markets. A strong result in one market can hide a poor one in another. Report by market.
- Reporting too early. Agents improve as context improves. Judge trend over a quarter, not a single week.
What should the board see?
One page, once a quarter: the five metrics for each funded AI use case, against baseline, with a short note on what changed and what is next. Use cases that are not moving their number after two quarters are reviewed or stopped. That discipline is what turns an AI programme into an AI portfolio.
We build these measures into every engagement, and run them on an ongoing basis through managed operations. To see which use cases will pay first, start with an AI readiness assessment.