Key Takeaways
- AI assistance raised customer support issues resolved per hour by 14%, but the gain reached 34% for novice and lower-skilled agents.
- In one clinical study, good AI raised clinician accuracy by 4.4 percentage points, while systematically biased AI reduced it by 9.1 points even when explanations were shown.
- A 2026 mammography trial cut radiologist reading workload by 63.6% and raised cancer detection by 15.2%, but it also raised the recall rate by 14.8%.
- Specialized legal research tools hallucinated in 17% to 33% of tested responses, so professional review remains necessary.
- Escalation thresholds should follow the consequence of an error and measured system performance, not a universal confidence score.
Human review does not automatically make an AI system accurate. It can catch errors, but it can also add delay, pass along a confident mistake, or consume more labor than the automation saves. The useful question is narrower: which work can the system complete, which work needs review, and when must a person take control?
The strongest human in the loop AI quality statistics come from studies that compare real workflows. They show gains in customer support and medical screening, but they also show a stubborn problem. People do not always correct bad AI advice, even when the system supplies an explanation.
Teams planning these workflows may also need AI and machine learning specialists, a consistent set of virtual assistant performance metrics, or managed services for the work that remains with people.
Human-in-the-loop AI statistics at a glance
| Setting | Human and AI result | Comparison | What the study does not prove |
|---|---|---|---|
| Customer support | 14% more issues resolved per hour | Agents with an AI assistant versus agents without access | It does not measure an autonomous bot or a universal quality gain |
| Less experienced support agents | 34% more issues resolved per hour | Novice and lower-skilled agents with versus without the assistant | The result came from one company |
| Clinical diagnosis with standard AI and explanations | 4.4 percentage-point accuracy gain | Clinician performance versus baseline | Vignettes are not live patient care |
| Clinical diagnosis with biased AI and explanations | 9.1 percentage-point accuracy loss | Clinician performance versus baseline | The experiment tested a deliberately biased model |
| Mammography screening | 63.6% lower reading workload | Partly autonomous AI workflow versus standard double reading | The recall rate was higher in the AI strategy |
| Mammography cancer detection | 15.2% higher | 7.3 versus 6.3 cancers detected per 1,000 screens | One screening design should not be copied without local validation |
| Specialized legal research AI | 17% to 33% hallucination rate | Four commercial research tools on a preregistered legal benchmark | The rates do not apply to every legal task or model |
These results use different definitions and cannot be averaged into one human-in-the-loop accuracy rate. A resolved support conversation, a correct diagnosis, and a verified legal citation are separate outcomes. Each needs its own denominator and review rule.
Automated-only and human-reviewed outcomes
The customer support evidence shows where assistance can work. An NBER study followed 5,179 support agents during a staggered rollout of a generative AI assistant. The agent stayed responsible for the customer conversation and could accept or ignore the model's suggested response.
Access to the assistant increased issues resolved per hour by 14% on average. The increase was 34% for novice and lower-skilled workers, while the effect on experienced and highly skilled workers was small. The paper also reports that two months of tenure with the tool produced performance comparable to more than six months of tenure without it. This is evidence for AI-assisted human work, not autonomous resolution.
Health care studies give a more complicated comparison. A randomized diagnostic reasoning trial of 50 physicians found median scores of 76% with access to GPT-4 and 74% with conventional resources. The adjusted two-point difference was not statistically significant. In an exploratory comparison, the LLM alone scored 16 percentage points higher than the conventional-resources group.
The result is uncomfortable but useful. Adding a person to a capable model did not capture the model's full measured performance. Workflow design, training, and a person's willingness to use or challenge a suggestion all affect the final result.
Bad AI can make expert review worse
A separate JAMA study tested 457 clinicians across 13 US states. Baseline diagnostic accuracy was 73.0%. Standard AI predictions with explanations raised accuracy by 4.4 percentage points. When the researchers supplied systematically biased AI predictions with explanations, clinician accuracy fell by 9.1 percentage points from baseline.
The explanation did not protect the reviewer. Biased predictions without explanations reduced accuracy by 11.3 points, and the 2.3-point difference between biased AI with and without explanations was not statistically significant. A review process that merely displays a rationale can still produce automation bias.
This changes what quality control should measure. A team should test whether reviewers identify planted errors instead of relying on approval clicks. Useful review measures include:
| Review measure | Formula |
|---|---|
| Defect escape rate | Incorrect outputs approved by reviewers / all incorrect outputs sent for review |
| Reviewer correction rate | Incorrect outputs corrected or rejected / all incorrect outputs sent for review |
| False escalation rate | Acceptable outputs sent to human review / all acceptable outputs |
| Review yield | Outputs materially changed or rejected / all reviewed outputs |
| Time to disposition | Median time from AI output to approval, correction, or escalation |
The defect set should contain known errors and ordinary production samples. Without known defects, a high approval rate may mean that the model is accurate, or it may mean that reviewers are not catching mistakes.
Hallucination rates depend on the task
There is no defensible universal hallucination rate for generative AI. NIST defines confabulation as confidently presented false or erroneous content and notes that the risk varies with the prompt, domain, and use context. Its Generative AI Profile calls for monitoring, human intervention alerts, moderation where models perform poorly, and retraining or decommissioning when performance falls outside defined limits.
Legal research provides a measured example. Stanford researchers ran a preregistered evaluation of Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI, and Bloomberg AI Assistant. The tools hallucinated in 17% to 33% of responses, despite using retrieval systems intended to ground answers in legal sources.
That range belongs to the tested products, prompts, and scoring rules. It should not be applied to support chat, medical imaging, document classification, or a newer model. The operational lesson is simpler: any output that contains a citation, legal proposition, amount, date, or named party needs verification against the underlying record before it is used.
Human review workload is a design variable
Reviewing every output can erase the capacity gained from automation. A 2026 prospective mammography trial offers unusually clear workload evidence. It enrolled 31,301 women and compared standard double reading with an AI workflow. The AI classified low-risk exams as normal without human reading, while higher-risk exams received double reading with AI support.
Radiologists read only 36.4% of exams in the AI strategy. That reduced reading workload by 63.6%. Cancer detection rose from 6.3 to 7.3 per 1,000 screens, a 15.2% relative increase. Yet the recall rate also rose from 4.8% to 5.5%, a 14.8% relative increase, and failed the study's noninferiority test.
The system saved review work by routing cases according to risk. It also created more downstream recalls. Capacity planning therefore needs both stages of work:
net review hours saved = avoided initial review hours - added escalation and correction hours
This calculation is an operating formula, not a published benchmark. Teams should supply observed volumes and handling times from their own workflow.
Set escalation thresholds by consequence
A confidence score is not a universal safety threshold. Scores may be poorly calibrated, and a 90% score from one system is not equivalent to 90% from another. NIST says organizations should use their own risk tolerance and performance measures to define limits. The US Government Accountability Office also describes human-in-the-loop control as active human oversight in which the person retains full control.
A practical policy can use three routes:
| Route | Appropriate work | Required control |
|---|---|---|
| Automatic completion | Reversible, low-consequence tasks with validated performance | Sample audits and rollback logging |
| Human review before action | Outputs with financial, contractual, privacy, medical, or reputational consequences | Source evidence, named reviewer, and recorded disposition |
| Immediate escalation | Suspected fraud, safety threats, regulated decisions, conflicts in source records, or model behavior outside tested limits | Qualified owner takes control and the AI cannot complete the action |
The threshold should tighten when the harm from a false negative or false positive increases. In screening, teams may tolerate more review to avoid missed disease. In customer support, a low-risk address-formatting task may be audited through sampling, while an account takeover report should go straight to a trained person.
A measurement plan for AI quality review
Begin with one bounded task and a labeled evaluation set. Record the AI output, source material, model version, confidence or policy signal, reviewer decision, edit, escalation reason, final outcome, and any later correction. Compare four groups when the workflow allows it: human only, AI only, AI with mandatory review, and AI with risk-based review.
Report accuracy and workload together. At minimum, track defect escape, correction rate, false escalation, review time, and the final business or safety outcome. Break results out by task type and risk tier. An overall average can hide a strong routine-task result and a dangerous failure rate on rare cases.
Recalibrate after model, prompt, data, or policy changes. A threshold validated on one version is not evidence for the next. If performance falls outside the documented limit, stop automatic completion for that task until a fresh evaluation passes.
Sources and limitations
- NBER, Generative AI at Work, April 2023, revised November 2023. Field study of 5,179 customer support agents at one company.
- JAMA Network Open, Large Language Model Influence on Diagnostic Reasoning, October 2024. Randomized vignette trial with 50 physicians.
- JAMA, Measuring the Impact of AI in the Diagnosis of Hospitalized Patients, December 2023. Randomized vignette study with 457 clinicians.
- Nature Medicine, AI-based triage and decision support in mammography, 2026. Prospective paired trial of 31,301 women.
- Journal of Empirical Legal Studies, Hallucination-Free?, published in 2025 from a 2024 preprint. Preregistered benchmark of four legal AI research tools.
- NIST AI 600-1, Generative Artificial Intelligence Profile, July 2024, updated April 2026. Risk management guidance rather than an outcome study.
- GAO, Artificial Intelligence: An Accountability Framework, June 2021. Governance framework rather than a performance benchmark.
The studies use different systems, dates, populations, and definitions. Their percentages are not a league table for AI products. They show why each organization needs local tests, risk-based routing, and evidence that reviewers catch errors rather than simply confirming them.
Related reading
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →