Key Takeaways
- Comm100's 2026 benchmark covers more than 220 million chats and reports a 75.3% AI handling rate, but only 44.8% of AI-handled chats were resolved without a person.
- An NBER field study of 5,179 support agents found that AI assistance increased issues resolved per hour by 13.8%.
- The same NBER study found a 1.3% increase in successful resolutions and no statistically significant change in customer satisfaction.
- Comm100 reports 92.6% CSAT for chatbot-to-agent handoffs, evidence that a measured handoff can outperform a containment-only design.
- No source reviewed publishes a comparable error rate across AI-only, human-only, and blended support, so teams must audit all three modes with the same rubric.
The strongest 2026 evidence does not show that AI-only support beats human support across every contact. It shows three narrower results. Autonomous AI can contain a meaningful share of selected chats. Human agents remain the path for contacts the system cannot or should not finish. A blended model can help those agents resolve more work without a measured drop in satisfaction.
Those findings come from different datasets, so they should not be averaged into one score. Comm100 reports platform-wide chat benchmarks. Intercom publishes results from its own product. The best causal evidence comes from field studies in which human agents gained access to AI assistance. Each source answers a different question.
For companies choosing an operating model, our customer support specialist service explains the human delivery option. The AI versus human virtual assistant comparison covers the broader staffing decision, while our AI customer service statistics research tracks adoption and market data.
AI and human support statistics at a glance
| Measure | AI-only evidence | Human-only evidence | Blended evidence |
|---|---|---|---|
| Resolution and containment | AI handled 75.3% of chats in Comm100's 2026 dataset, but resolved 44.8% of the chats it handled without a person (Comm100, 2026) | Public human benchmarks use different channels and FCR definitions, so there is no valid like-for-like rate in the reviewed sources | AI-assisted agents resolved 13.8% more issues per hour in an NBER field study (NBER, June 2023) |
| First-contact resolution | Vendor resolution rates range from 44.8% in Comm100's cross-industry data to 76% in Intercom's product disclosure (Comm100, 2026; Intercom, April 2026) | The human control group is the baseline in the NBER study; the public summary does not publish its absolute FCR | AI assistance raised successful resolutions by 1.3%, relative to the human-only baseline (NBER, June 2023) |
| Escalation | Comm100's figures imply a 30.5 percentage-point gap between AI handling and AI resolution, but that gap is not an escalation rate (Comm100, 2026) | Human agents receive bot escalations as well as human-first contacts; the reviewed sources do not publish one comparable escalation rate | Chatbot-to-agent handoff CSAT reached 92.6% in Comm100's benchmark (Comm100, 2026) |
| Handle time | Comm100 says overall chat duration stayed flat in its benchmark, without publishing an AI-only causal change (Comm100, 2026) | The NBER human-only cohort supplies a control, but its public summary does not give the absolute time per chat | AI-assisted agents spent about 9% less time per chat (NBER, June 2023) |
| CSAT | Comm100 reports a 9.1% jump in chatbot satisfaction and an overall score of 4.1 out of 5 (Comm100, 2026) | Absolute human-only CSAT is not separated on the public benchmark page | The NBER study found no statistically significant satisfaction change after assistance was introduced (NBER, June 2023) |
| Error or repeat contact | No comparable AI-only error rate is disclosed in the reviewed benchmark pages | No comparable human-only error rate is disclosed in the reviewed benchmark pages | A 2026 Alibaba field experiment found no significant change in three-day customer retrial rates on average (Ni et al., February 2026) |
The blank cells matter. A percentage becomes misleading when its denominator changes. AI resolution usually measures conversations the bot had an opportunity to answer. Human FCR may measure all eligible calls. Blended productivity counts work completed by an agent who can accept, edit, or ignore a suggestion.
AI-only support: handling is not resolution
Comm100's 2026 Live Chat Benchmark Report analyzes more than 220 million interactions across 18 industries (Comm100, 2026). AI agents handled 75.3% of chats, yet resolved 44.8% of the chats they handled without human involvement (Comm100, 2026; Comm100, August 24, 2026).
That distinction prevents a common reporting error. Subtracting the resolution rate from 100% produces a non-resolution share, not a verified escalation share. A non-resolved conversation may transfer, become a ticket, time out, or end when the customer leaves.
The range inside the benchmark is also wide. Comm100 reports AI resolution rates from 41.2% to 89.0% by team size and from 38.1% to 97.7% by industry (Comm100, August 24, 2026). The company says small teams with one to five agents routed 54.3% of chats to AI and resolved 89.0% of those handled chats. Teams with 11 to 25 agents routed 92.5% to AI but resolved 47.8% (Comm100, August 24, 2026). Broader coverage exposes the system to harder contacts, so a lower resolution rate can coexist with more total work removed from the human queue.
Intercom reports a higher product result. In April 2026, it said Fin served almost 8,000 customers, averaged a 76% resolution rate, and resolved close to 2 million queries each week (Intercom, April 2026). An earlier October 2025 disclosure reported 66% across more than 6,000 customers (Intercom, October 15, 2025). These are different reporting dates and customer populations, not conflicting measurements of one fixed market average.
Intercom's July 2026 metric change shows why comparisons need definitions. In its worked example, 250 resolutions among 750 active conversations produced a 33% legacy resolution rate. Excluding conversations in which the bot had no opportunity to answer changed the denominator to 500 and the displayed resolution rate to 50%. Automation stayed at 25% because the system still resolved 250 of 1,000 total conversations (Intercom, updated July 2026).
For an AI-only queue, report both resolution and automation:
AI resolution rate = AI-resolved contacts / contacts AI attempted
AI automation rate = AI-resolved contacts / all eligible incoming contacts
Human-only support: the public comparison gap
Human support is often presented as the benchmark, yet recent public sources rarely publish an absolute human-only rate beside the AI-only rate for the same intents, channel, period, and customer population. The absence of a matched figure does not mean human resolution is unknowable. It means companies should calculate it from their own human-first queue.
The NBER field study provides a better comparison design. Researchers followed 5,179 customer support agents at a Fortune 500 enterprise software company while it introduced a conversational assistant, mostly between November 2020 and February 2021 (NBER Working Paper 31161, April 2023, revised November 2023). Agents without the tool formed the human-only comparison group. The public results report changes from that baseline rather than one universal human FCR.
This approach controls for the employer and work setting more effectively than comparing a vendor's chatbot portfolio with an unrelated phone-support survey. A company can reproduce the logic by matching contacts on intent, channel, customer type, policy complexity, and arrival period.
Human-only support also needs an explicit denominator. Exclude contacts that cannot reasonably be completed on first contact only if the same exclusion applies to AI and blended queues. Count a callback, reopened ticket, or same-intent repeat contact as unresolved under the same time window for every mode.
Blended support: the strongest causal evidence
The NBER study found that access to AI increased issues resolved per hour by 13.8%. Agents spent about 9% less time per chat, handled about 14% more chats per hour, and improved their successful resolution rate by 1.3% (NBER, June 2023). Customer satisfaction did not change significantly (NBER, June 2023).
The gains were uneven. Less experienced and lower-skilled agents improved by 35%, while the most experienced agents saw little or no benefit (NBER, June 2023). That result argues for targeted assistance and coaching rather than one productivity assumption for the whole team.
A February 2026 Alibaba field experiment adds a useful caution. Human agents in ecommerce after-sales chat were randomly assigned access to AI diagnosis and response suggestions. The study found faster issue identification and shorter chats, better customer ratings, and lower dissatisfaction, but no significant average change in three-day retrial rates (Ni et al., February 2026). Top-performing agents experienced declines in subjective and objective service quality, which the researchers associate with more multitasking and slower responses (Ni et al., February 2026).
Blended support therefore has two separate levers: a handoff from automation to a person, and AI assistance during the person's work. Comm100's 92.6% chatbot-to-agent handoff CSAT suggests that customers can rate a transfer highly when the receiving agent gets the right context (Comm100, 2026). It does not prove that every handoff works or that the score applies outside Comm100's customer base.
What the evidence says about each outcome
First-contact resolution
Do not compare Comm100's 44.8% resolution rate with Intercom's 76% as if one system beat another by 31.2 percentage points. Comm100 spans 18 industries and defines the figure among chats handled by AI, while Intercom reports its own customer portfolio (Comm100, 2026; Intercom, April 2026). Use each as a vendor reference, then run a matched evaluation on the intents your operation receives.
Escalation and containment
Containment should require confirmed resolution with no same-intent repeat contact during a declared window. Escalation should count transfers and tickets routed to a person. Comm100's 30.5 percentage-point handling-to-resolution gap shows that contact is not completion, but it cannot be relabeled as an escalation rate without outcome data (Comm100, 2026).
Handle time
The NBER result supports a blended speed benefit: about 9% less time per chat in one enterprise software setting (NBER, June 2023). Comm100 says chat duration was flat across its broader 2026 benchmark even as agent workload fell by 5.8% (Comm100, 2026). These findings are compatible because one is a causal rollout study and the other is an aggregate platform benchmark.
CSAT
CSAT depends on which customers answer the survey. Comm100 reports overall satisfaction of 4.1 out of 5, a 9.1% increase in chatbot satisfaction, and 92.6% handoff CSAT (Comm100, 2026). The NBER study found no significant change after agents gained AI assistance (NBER, June 2023). Keep AI-only, human-first, and transferred-contact response rates beside their CSAT scores.
Error and repeat-contact rates
No source in this review publishes one error rubric across all three support modes. The Alibaba study's three-day retrial result is the closest durable-resolution check, and it found no significant average change with AI assistance (Ni et al., February 2026). That is not an error rate. It is evidence that faster chats and better ratings did not produce a measured reduction in customers trying again within three days.
An internal quality audit should sample resolved contacts from every mode and label wrong answers, incomplete actions, policy violations, unnecessary escalations, missed escalations, and repeat contacts. Report each error category per audited resolved contact. One blended total can conceal a bot that escalates too often or a human queue that accepts bad suggestions.
A fair scorecard for AI, human, and blended support
Use the same contact population and definitions for all three columns:
| Metric | Formula |
|---|---|
| Durable first-contact resolution | Contacts resolved on the first interaction with no same-intent repeat inside the declared window / eligible contacts |
| AI automation | Contacts resolved by AI without human work / all eligible contacts |
| Escalation rate | Contacts routed from AI to a person / AI-started contacts |
| Handoff acceptance | Transfers accepted by the correct human queue with transcript and customer context / attempted transfers |
| Average handle time | Total active handling time / completed contacts |
| Mode-specific CSAT | Positive or average survey result for one mode, shown with its survey response rate |
| Audited error rate | Resolved contacts with at least one defined quality defect / audited resolved contacts |
Publish the channel, intent mix, customer segment, sample count, test dates, and repeat-contact window. Keep abandonment separate from successful containment. If the AI vendor changes a denominator, preserve the old series or restate history before declaring an improvement.
Which support model has the best resolution performance?
There is no universal winner in the public data. AI-only support can close routine contacts at scale, but reported resolution depends heavily on what the bot attempts and how the vendor defines the denominator. Human-only support is the necessary baseline for matched tests and the destination for work that requires judgment or authority. Blended support has the best causal evidence for higher throughput: 13.8% more issues resolved per hour in the NBER study, with no significant CSAT decline (NBER, June 2023).
The practical choice is to route predictable, reversible contacts to automation; preserve a clear human path; and assist people where retrieval or drafting can shorten the work. Judge all three modes by durable resolution, not by how many conversations the AI touched.
Sources
- National Bureau of Economic Research, Generative AI at Work, Working Paper 31161, published April 2023 and revised November 2023. Field data from 5,179 customer support agents at one Fortune 500 enterprise software company.
- National Bureau of Economic Research, "Measuring the Productivity Impact of Generative AI", published June 1, 2023. Reports the study's handle-time, throughput, resolution, experience-level, and satisfaction results.
- Comm100, AI Live Chat Benchmark Report 2026, published 2026. Covers more than 220 million live chat interactions across 18 industries.
- Comm100, "What Is a Good AI Resolution Rate for Player Support?", published August 24, 2026. Gives resolution results by team size and industry from the 2026 benchmark.
- Intercom, "Announcing Monitors: Opening the AI black box", published April 2026. First-party Fin customer, weekly volume, and resolution disclosure.
- Intercom, "What's new with Fin 3", published October 15, 2025. Earlier first-party resolution disclosure across more than 6,000 customers.
- Intercom, "Update to Fin performance metrics", updated July 2026. Documents the resolution denominator change and worked example.
- Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, and Congyi Zhou, Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations, submitted February 8, 2026. Randomized field experiment in ecommerce after-sales chat.
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →