Why 2% call sampling costs government agencies more than it saves
Delayed quality assurance was built to score representatives. Real-time QA can do something far more valuable: tell you whether what your people are saying actually resolves the call.

Government contact centers have increasingly invested in artificial intelligence over the past three years. The returns are not showing up in the cost line.
In our latest research, not one public sector leader surveyed reported that AI investments had reduced overall CX costs, and 51.5% said their costs had actually gone up.
That finding is easy to misread as a verdict on the technology, and it is not one. Agent assist tools genuinely do navigate complex agency knowledge bases in seconds, keeping representatives from putting callers on hold while they hunt down a rule.
The problem sits somewhere less interesting and far more fixable. Agencies are running 2026 artificial intelligence on a quality assurance model built in the 1990s, and that model cannot tell them whether any of it is working until weeks after the fact.
The limits of sampling: Built to score representatives, not improve service
Most contact centers still evaluate a random 2 to 5% of calls, days or weeks after the interaction happened. The usual criticism is that the sample is too small. That is true, but it is not the real problem. Even a perfect sample produces a score, and a score answers only one question: did the representative follow the guidance they were given? It says nothing at all about whether that guidance was any good.
The distinction matters more in public service than almost anywhere else, because a caller who cannot get a straight answer about a lien or a benefit dispute has nowhere else to take their business. If the knowledge article a representative is working from is ambiguous, every representative reading it will give the same unhelpful answer, and every one of them will pass their quality review for doing so. The 95 to 98% of calls nobody listens to is precisely where that pattern would have been visible.
The cost of finding out weeks later
Gartner's customer service benchmarks put the median cost per contact at $1.84 for self-service against $13.50 for live, assisted channels, and in high-complexity federal environments those assisted costs run considerably higher. When a voice bot reads a policy, or a representative repeats a poorly worded procedure, the error compounds quietly across thousands of calls before anyone happens to sample one of them.
It is tempting to read those two numbers as an argument for pushing every contact into the cheaper channel. That is the wrong lesson. The expensive outcome is not the assisted call. It is paying $1.84 and then $13.50 for the same unresolved issue, frequently a third time when the caller rings back a week later. Automation that contains a contact without resolving it has not saved anyone money; it has deferred the cost and added to it.
What real-time quality assurance changes: Evaluating every call live
The alternative is to stop sampling altogether. Enterprise speech-to-text engines can transcribe and evaluate every interaction as it occurs, matching spoken dialogue against form entries and system logic. At TTEC we recommeautomated quality assurance in real time, and the value of it shows up in four distinct places.
- Discrepancies surface during the call rather than after it: logic errors, sentiment spikes, or a mismatch between what the caller said and what was entered on the form.
- Supervisors gain complete observability and receive immediate alerts, so they can step in or guide a representative before a frustrated caller hangs up.
- Recurring misstatements point back at their source. When forty representatives give the same wrong answer, the defect is in the knowledge article, not in forty people.
- Resolution becomes the measured outcome, in place of checklist adherence.
The last two are the ones agencies consistently underuse. Real-time QA tends to get bought as supervision, and supervision is the least interesting thing it does. Its more valuable output is a continuous read on whether what we are telling people to say is actually resolving the call, and a specific, evidence-backed answer about where it is not. That turns quality assurance from a monthly compliance ritual into the mechanism by which the service itself gets better.
What we learned moving the IRS 1040 line to conversational voice
When I was with the IRS and we moved the 1040 product line to conversational voice, we learned quickly that canned, formatted responses were not the most effective way to interact with a taxpayer. How we speak is very different from how we format text for a website.
The same lesson applies on the assisted side of the operation: a knowledge article written to be read on a screen often lands badly when a representative has to deliver it aloud, under time pressure, to someone who is already anxious. You do not discover that from a monthly scorecard. You discover it by listening to every call and noticing where the same explanation keeps having to be repeated.
Doing it responsibly: Security designed in, not added on
Real-time feedback can never come at the expense of data stewardship.
Public sector interaction data falls under strict statutory frameworks, including Title 26 protections for federal tax information and Title 13 confidentiality requirements, and none of those obligations relax because a QA engine is doing the listening rather than a person. The tooling has to operate inside FedRAMP-certified and StateRAMP or Impact Level 4 compliant architectures, with automated masking that redacts sensitive attributes at the ingestion layer, before any guided tool or speech engine touches the representative's desktop.
Where human representatives still belong
There are players in this market pushing to automate every public interaction end to end. I advocate strongly for 100% automated quality assurance, and I am equally firm that human representatives have to stay in the loop for complex, high-stakes service.
Artificial intelligence can model empathy. It cannot supply genuine reassurance. When a taxpayer is working through a property lien, or settling an estate after losing someone they love, most people want a person on the other end of the line, and forcing them into an inescapable automated loop only adds to their distress.
That is the strongest argument for investing in quality assurance on the assisted side rather than treating it as the legacy channel. The calls that stay with human representatives are the hardest ones an agency handles. They are the calls where getting the guidance right matters most, and they are the least likely to turn up in a 2% sample.
Getting there without a rip and replace
None of this requires a risky overhaul of legacy infrastructure.
Agile AI gateways let agencies layer real-time automated quality assurance, intelligent routing and agent assist onto the architecture they already run. Pair that full operational visibility with a well-supported human workforce and you close the loop that sampling never could. You stop merely confirming that representatives said what we asked them to say and start learning, continuously, whether what we asked them to say was right.
[cta-1]
Stop paying twice for the same unresolved constituent issue.
Talk with our public sector experts to learn how a real-time AI quality assurance model can reduce repeat call volume and optimize your agent desktop — without a risky "rip and replace."

Teresa manages strategic public sector client relationships and drives long-term account value across complex digital and CX engagements.