5 dimensions for analyzing human and AI agent performance
Your AI Agents need a performance review — just like the humans next to them. Dave Rennyson, CEO, SuccessKPI sets out five dimensions you should measure to determine which calls your human or AI agents should be handling
Add bookmark
Most contact centers have spent two years buying AI agents and little time deciding how to supervise them. That gap is less an oversight than an inherited habit. For two decades the automated layer in a contact center was an IVR, and an IVR did not need much supervision. But your new AI agents do.
Your callers want an outcome when they call. Regardless of who, or what, handles the call, there's an expectation of an outcome, and hopefully a successful resolution after a good experience.
But what you need to understand is not simply how to get an outcome, but who can give the best outcome based on each caller's need, or intent.
If you are going to govern two workforces, you need one definition of a good outcome that applies to both. At SuccessKPI we use five dimensions, and call it our CARES framework.
This article explains
- Why AI contact center agents need active governance, not just deployment
- The CARES framework that can measure AI and human agents on a unified standard
- Why containment and cost-per-contact hide the real customer outcome
Don't miss any news, updates or insider tips from CX Network by getting them delivered to your inbox. Sign up to our newsletter and join our community of experts.
5 dimensions that define a successful outcome for human and AI agents
The five dimensions of the CARES framework are:
- Compliance: Did the agent stay within its intended scope?
- Accuracy: Was the judgment and the action correct, and did the agent recognize when to escalate?
- Resolution: Was the issue correctly solved? The simplest test: did the customer come back within 72 hours?
- Effort: How hard did the customer have to work to get their outcome?
- Sentiment: How did the customer feel across the whole interaction, not just at the end?
Cost matters, but it belongs after these five, not in front of them. A cheap contained interaction that generates a callback, a complaint and a churn risk is not cheap. It is a delayed cost to your business.
None of these five dimensions are new. They are what a quality program already asks about a human agent.
Let's unpack each of the five dimensions in detail.
Compliance: Did the agent stay within its intended scope?
This can be various things, and we're not just talking about AI hallucinations. With human agents this is staying on script, general disclosures, and verification steps.
We ran across an example where an AI agent asked for a credit card number to verify an employee for an internal password reset. Completely off script for human common sense, but the AI bot knows asking for the last four digits of a credit card is used in the general world of verifying that a caller is who they say they are.
AI runs the risk of using general knowledge at scale until someone notices.
In regulated industries the liability is the same as a human breach. Regardless, it can make for very confusing interactions for callers.
Accuracy: Was the judgment correct and did the agent know when to escalate?
Accuracy has two halves, and the second one is where AI agents may quietly fail.
A competent human agent knows when they are out of their depth and hands off to a supervisor. An AI agent sounds equally assured whether it is right or wrong, which means recognizing the limits of its own competence is a skill you have to measure directly rather than assume.
It's important to note, escalation is not a failure condition. An AI agent that transfers the right eight percent of calls is performing better than one that contains all of them badly, and your current containment dashboard will tell you the opposite.
Resolution: Was the issue correctly and fully solved?
This is where most AI measurements show cracks.
Containment, automation rate, and cost per contact show what happened to the individual interaction, not what happened to the customer. A health plan member calls about an unexpected bill, the AI agent confirms the claim was processed correctly, emails the explanation of benefits, and closes the interaction.
No transfer or further human touch occurred, and the dashboard logs a contained self-service resolution.
Three days later the member calls back about the same bill. Nothing in the vendor's report was inaccurate; it simply was not measuring customer resolution.
The simplest test applies to both workforces equally: did the customer contact you again within 72 hours about the same thing?
Effort: How hard did the customer have to work?
Effort is the dimension that really exposes the customer experience. Callers have historical frustrations with IVRs and being contained. This goes back to callers wanting an outcome, and most importantly a quick and easy outcome.
Until your AI agent can prove helpful, callers will be skeptical and give little leeway.
AI may provide a password reset seamlessly, but create a loop trying to solve a billing question that should be handed off. The more you analyze and understand both human and agentic calls, the better you can route accordingly, and the quicker you alleviate customer concern and frustration.
Sentiment: How did the customer feel across the interaction, not just at the end?
AI agents are consistently polite, which is not the same as satisfying, and end-of-call survey scores are the least reliable place to look for the difference. Sentiment as a governance measure means tracking the trajectory of the conversation: where frustration entered, whether it was recognized, and whether anything changed after it was. A human agent who hears frustration adjusts. Whether your AI agent does, or plows through its path while the customer escalates, is observable in the transcript and can be invisible on a dashboard.
Measuring intent success as an independent governance layer
Scoring both workforces the same way is what turns governance into a decision you can act on. The table below gives an example of how you can get started:
| Intent | Human Resolution | AI Resolution | Sentiment | Accuracy | Cost of Success | Best Workforce |
| Password reset | 91% | 96% | AI better | 99% | AI lower | AI |
| Billing dispute | 88% | 69% | Human better | 74% | Human higher | HUMAN |
| Order status | 94% | 97% | Similar | 98% | AI lower | AI |
| Cancellation | 87% | 81% | Human better | 85% | Mixed | HYBRID |
The question executives keep asking – 'is our AI working?' – has no useful answer because it is asked without context. There is no universal verdict on an AI agent any more than on a contact center team.
Asked one intent at a time, as in the table above, it becomes meaningful: who is better at this intent, measured the same way, on the same scale?
This requires a grade book neither the CCaaS nor the AI vendor holds. Most organizations get their AI performance data from the company that built the AI and their platform data from the CCaaS vendor. Neither arrangement is giving you true, independent oversight. And unfortunately, neither vendor does both, unless you buy their entire tech stack.
Either way, this is like having the home builder perform their own inspections.
SuccessKPI builds neither but measures both, which is exactly why we can score both workforces independently, against one standard, across every interaction.
Start with one intent. Pick something high-volume and unambiguous that your AI handles today, and check whether you can currently see if those customers came back within 72 hours.
If you cannot, containment is the only thing you are measuring.
Then score the AI and human interactions for that intent against all five dimensions and see who wins it. Repeat on the next two. That is how an AI deployment becomes a managed one.
Quick links
- Top 30 contact center leaders to follow in 2026
- Before you cut your contact center workforce for AI agents, read the fine print
- Using human-led omnichannel to address the irony in AI