Key Takeaways
- An NBER field study found that AI assistance raised customer support resolutions per hour by 13.8%, with a 35% gain for less experienced and lower skill agents.
- Salesforce found that 69% of service organizations used at least one form of AI in 2025, while service professionals expected AI to resolve 50% of cases by 2027.
- Gartner found that 64% of customers would prefer companies not use AI for customer service, and 53% would consider switching if a company planned to use it.
- Zendesk found that 95% of consumers expect an explanation for decisions made by AI, making transparency part of support quality rather than an optional disclosure.
- Human handoffs need context: 81% of consumers want representatives to continue from the prior interaction, and 74% are frustrated when they must repeat information.
Human in the loop AI customer support statistics point to a practical division of work. Automation can answer routine questions, retrieve account details, and draft responses. People still handle exceptions, emotional conversations, high-risk decisions, and cases where the system lacks enough evidence to answer safely.
That division matters because automation and human service are not interchangeable. A bot's resolution rate measures what it finishes without a person. An agent-assist study measures how much more a person completes with AI support. A consumer preference survey measures willingness, not operational performance. This report keeps those measures separate and identifies the sample, date, and source behind each figure.
Human-in-the-loop AI customer support statistics at a glance
| Measure | Finding | Study context |
|---|---|---|
| Agent productivity | 13.8% more issues resolved per hour | NBER field study of about 5,000 support agents, published June 2023 |
| Newer-agent productivity | 35% improvement | Same NBER study; effect concentrated among less experienced and lower skill agents |
| Service AI adoption | 69% use at least one form of AI | Salesforce survey of 6,500 service professionals, fielded April to June 2025 |
| Current agentic AI use | 39% | Same Salesforce survey |
| Cases resolved by AI | 30% in 2025, expected to reach 50% in 2027 | Salesforce respondent estimate and forecast, not an independently audited resolution rate |
| Preference against service AI | 64% | Gartner survey of 5,728 customers, conducted December 2023 |
| Switching consideration | 53% | Same Gartner survey |
| Explanation expected for AI decisions | 95% | Zendesk survey of more than 11,000 consumers and business leaders, published November 2025 |
| Desire for context-preserving handoff | 81% | Same Zendesk study |
| Frustration with repeating information | 74% | Same Zendesk study |
These results do not support a simple claim that customers either like or dislike AI. Customers accept automation for some jobs and remain wary of losing access to a person. The strongest operating model gives the automated system a defined scope and makes escalation easy when that scope ends.
What human in the loop means in customer support
A human-in-the-loop system gives a person authority at a defined point in an AI workflow. The person may review a drafted response, approve a refund, take over an uncertain conversation, or audit a sample after completion. The term should not be used for any support operation that merely employs people somewhere in the company.
Four common designs put the human at different points:
| Design | AI role | Human role | Useful measure |
|---|---|---|---|
| Agent assist | Retrieves knowledge, summarizes, or drafts | Reviews and sends the response | Resolutions per hour and quality |
| Automated front line | Resolves bounded requests | Receives exceptions or customer-requested handoffs | Confirmed resolution and escalation rate |
| Approval workflow | Recommends an action | Approves high-risk or high-value actions | Error rate and approval time |
| Quality review | Handles eligible cases | Audits conversations and corrects policy or content | Defect rate and recurrence |
These designs answer different questions. A team assessing a customer service virtual assistant should decide which decisions belong to the assistant, which require human approval, and which conditions trigger an immediate handoff.
Automation and resolution metrics
Salesforce's State of Service, Seventh Edition offers a broad view of adoption. The company surveyed 6,500 service professionals across 40 countries between April 25 and June 6, 2025. Sixty-nine percent said their organization used at least one form of AI. Current use was 53% for generative AI, 44% for predictive AI, and 39% for agentic AI.
The report also asked about the share of cases resolved by AI. Respondents put the 2025 share at 30% and expected it to reach 50% by 2027. That figure is a reported current estimate paired with a forecast. It should not be treated as proof that half of all support cases will reach verified resolution without human involvement.
Vendor operating data shows why definitions matter. Intercom reported in October 2024 that Fin 2 achieved a 51% average resolution rate across customers. The vendor defined resolution around conversations that its AI answered and closed. This is product data rather than an independent cross-industry benchmark, but it gives operators a concrete example of a measured autonomous-resolution metric. Intercom's reporting documentation separately tracks AI involvement, confirmed or assumed resolution, and routing to a team. Those fields prevent a handoff from being counted as an AI resolution.
Teams should report at least four separate outcomes:
- Confirmed AI resolution, where the customer verifies that the issue is solved
- Assumed AI resolution, where the conversation ends without an explicit confirmation
- Human handoff, where the AI routes the conversation to a person
- Abandonment, where the conversation ends without evidence of resolution or a completed handoff
Combining these outcomes inflates the apparent success of automation. A high deflection rate can coexist with poor resolution if customers leave after an unhelpful answer.
What controlled research says about assisted agents
The strongest causal evidence in this topic comes from Generative AI at Work by Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond. The NBER summary, published in June 2023, describes a staggered rollout to roughly 5,000 customer support agents at a Fortune 500 software company. The system suggested responses during live text conversations, and agents could accept or ignore them.
AI-assisted agents resolved 13.8% more issues per hour. The study broke that gain into about 9% less time per chat, roughly 14% more chats handled per hour, and a 1.3% increase in the share of chats successfully resolved. Customer satisfaction did not change significantly.
The average hides a large experience effect. Less experienced and lower skill agents improved by 35%, while the most experienced and highest-performing agents saw little benefit or small negative effects. This is evidence for agent assist in one software support environment, not evidence that the same gain will appear in every channel or industry.
The study also shows why human oversight can improve deployment safety. Agents retained control of the final response. They could use a helpful suggestion and reject one that did not fit the conversation. The AI made knowledge from stronger workers easier to apply, while the person remained responsible for the customer interaction.
Escalation is part of the product, not a failure state
Customer resistance often centers on access to a person. Gartner surveyed 5,728 customers in December 2023 and published the results on July 9, 2024. Sixty-four percent said they would prefer companies not use AI for customer service, and 53% would consider switching to a competitor if they learned a company planned to use it.
Gartner reported that the leading concern was increased difficulty reaching a person. Other concerns included job displacement and incorrect answers. The result does not mean 64% refuse every automated interaction. It does mean companies risk losing trust when automation acts as a barrier instead of a quick route to the right resolution.
A useful escalation policy includes:
- Customer choice: a clear request for a person triggers a handoff
- Confidence: low confidence or conflicting source material stops an automated answer
- Risk: account security, legal threats, safety, regulated advice, and material financial actions go to authorized staff
- Repetition: repeated questions or failed steps trigger review rather than another version of the same answer
- Sentiment and vulnerability: distress, bereavement, accessibility needs, or a sensitive complaint receive human attention
Escalation rate alone cannot show whether this policy works. A very low rate may mean the AI resolves most eligible cases, or it may mean customers cannot escape the bot. Pair escalation rate with repeat contact, abandonment, complaint, and confirmed-resolution measures.
Trust depends on disclosure, explanation, and control
Zendesk's 2026 CX Trends research, published November 18, 2025, included more than 11,000 consumers and business leaders across 22 countries. It found that 95% of consumers expected clear explanations for decisions made by AI. Sixty-three percent said their demand for transparency had increased from the prior year.
The gap on the company side was substantial. Eighty percent of CX leaders said transparency would become a requirement for customer-facing AI, but the report's public findings said only 37% currently offered reasoning behind AI decisions.
For support leaders, an explanation needs to be useful. A message that says an answer was generated by AI identifies the tool but does not tell the customer what happened. A better notice states which account fact, policy, or customer instruction led to the result. It also gives the customer a path to challenge the result and reach a person.
Trust also depends on limiting the AI's authority. A system may be allowed to report an order status but not cancel the order. It may draft a retention offer but require a person to approve the credit. Those boundaries should follow the cost of a wrong action, not the novelty of the technology.
Customer preference changes with the job
The available datasets show preference is conditional. Salesforce's October 2024 State of the AI Connected Customer found that 34% of consumers would choose an AI agent over a person to avoid repeating themselves. Thirty percent would make the same choice for faster service, rising to 37% among Gen Z and millennial respondents.
Zendesk's 2026 research found broader demand for continuity. Eighty-one percent of consumers wanted representatives to pick up where the prior interaction ended, while 74% were frustrated when they had to repeat their information. This is a strong case for shared context between AI and people. It is not a case for removing people from the journey.
The customer preference pattern is straightforward:
- For a routine status check, speed and availability can matter more than who provides the answer.
- For a complex or disputed outcome, access to a person and a clear explanation carry more weight.
- During a handoff, customers expect the next agent to receive the transcript, attempted steps, identity state, and reason for escalation.
Customer support teams can apply this pattern without forcing every request through the same channel. The Stealth Agents blog includes broader service and workforce research, while the services overview explains where managed human support can fit around automated workflows.
A measurement framework for 2026
A human-in-the-loop dashboard should show whether customers receive correct resolutions and whether people intervene at the right moments. Volume reduction by itself does not answer either question.
| Area | Metric | What to check |
|---|---|---|
| Automation | Confirmed AI resolution rate | Customer verified the outcome, with no repeat contact inside the chosen window |
| Escalation | Appropriate handoff rate | Eligible risk, confidence, and customer-choice triggers reached a person |
| Continuity | Context-preserved handoff rate | The person received the transcript, intent, and completed steps |
| Resolution | First-contact resolution | The issue stayed solved after the first contact across both AI and human channels |
| Trust | Disclosure and explanation coverage | Customers could identify AI use and understand material decisions |
| Quality | Unsupported-answer rate | Audits found claims or actions without an approved source |
| Workforce | Resolutions per paid hour | Productivity changed without a fall in satisfaction or resolution quality |
| Learning | Repeat-defect rate | Corrected failure types stopped recurring after content or workflow changes |
Segment these metrics by intent, channel, customer type, language, and risk level. A single average can hide a system that performs well on order tracking and badly on billing disputes. Review the segments often enough to pause an unsafe workflow before it affects a large number of customers.
What the statistics support
The evidence supports AI assistance for bounded work and clear human authority for exceptions. In a large field study, agent assist raised output most for newer workers without reducing customer satisfaction. Industry surveys show rapid AI adoption and high expectations for autonomous resolution. Consumer research adds a firm constraint: people want access to a person, continuity during handoff, and explanations for AI-made decisions.
The sensible target is not the highest possible automation rate. It is the highest safe confirmed-resolution rate, with a fast and context-rich path to a qualified person when automation stops being useful.
Source notes
| Source | Publication date | Sample or dataset | Statistics used |
|---|---|---|---|
| NBER, Measuring the Productivity Impact of Generative AI | June 2023 | Roughly 5,000 customer support agents | 13.8% more resolutions per hour; 35% gain for less experienced and lower skill workers |
| Gartner, Survey Finds 64% of Customers Would Prefer That Companies Didn't Use AI for Customer Service | July 9, 2024 | 5,728 customers surveyed in December 2023 | 64% preference against service AI; 53% would consider switching |
| Salesforce, State of Service, Seventh Edition | 2025 | 6,500 service professionals in 40 countries, surveyed April 25 to June 6, 2025 | 69% AI adoption; 39% agentic AI use; 30% current and 50% expected AI case resolution |
| Salesforce, All Stats: State of the AI Connected Customer | October 2024 | Global double-blind, product-agnostic consumer research | 34% would use AI to avoid repetition; 30% would use AI for faster service |
| Zendesk, AI Ushers In Era of Contextual Intelligence | November 18, 2025 | More than 11,000 consumers and business leaders in 22 countries | 95% expect explanations; 81% want context continuity; 74% dislike repetition |
| Intercom, Fin 2 launch and customer-base results | October 2024 | Intercom Fin 2 customer base | 51% average resolution rate, presented as vendor operating data |
Frequently asked questions
What is human-in-the-loop AI in customer service?
It is a support workflow in which AI performs a defined task while a person retains authority to review, approve, correct, or take over. The human checkpoint may occur before a response, during an escalation, before a high-risk action, or during quality review.
Does AI improve customer support agent productivity?
In the NBER field study of roughly 5,000 support agents, AI assistance increased resolutions per hour by 13.8%. Less experienced and lower skill agents improved by 35%. The result came from one company's text-support environment, so teams should validate the effect in their own channels and case mix.
Do customers prefer AI or human customer support?
Preference depends on the task. Gartner found broad concern about companies using AI in service, especially when it makes people harder to reach. Salesforce found that some consumers would choose AI for speed or to avoid repeating information. Both findings support customer choice and a visible human handoff.
What should trigger escalation from AI to a person?
Common triggers include a customer request, low model confidence, repeated failed steps, an unsupported answer, a sensitive complaint, account security, regulated advice, and actions with material financial or legal consequences.
How should a company measure AI support resolution?
Separate confirmed resolution, assumed resolution, human handoff, and abandonment. Then check repeat contact and complaint rates. Counting every conversation that does not reach an agent as a resolution will overstate performance.
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →