Key Takeaways
- A confidence threshold sets a workload and risk tradeoff, not a universal quality guarantee.
- Higher thresholds usually increase human review volume while improving the quality of the cases left to automation.
- Precision, recall, coverage, and exception volume must be reported together because any one metric can hide important failures.
- Confidence scores require calibration and monitoring on the current population, task, and model version.
- Human review time should be measured from queue entry through final disposition, including corrections and repeat work.
An AI confidence threshold decides which cases a system handles and which cases people review. Moving that cutoff changes both risk and workload. A stricter threshold can improve the accuracy of automated decisions, but it also sends more cases to a review queue. A looser threshold can reduce review demand while allowing more model errors to pass without inspection.
The research does not support one confidence cutoff for every workflow. Recent studies show that the useful threshold depends on the confidence signal, model, task, population, and cost of an error. These AI confidence threshold human review statistics separate observed results from planning examples and explain how to measure the tradeoff.
Teams that need people to manage AI review queues can compare AI-enabled virtual assistant services and the practical uses of an AI virtual assistant.
AI confidence threshold statistics at a glance
| Measure | Finding | Population and date |
|---|---|---|
| Autonomous coverage at a 0.90 consistency threshold | 49.4% | 2026 Nature Medicine study; 272 of 551 MIRA-v2 medical cases |
| Accuracy among retained cases | 98.9% | Same study, model, dataset, and threshold |
| Autonomous errors | 3 cases | Same 272 retained cases |
| Cases deferred for human review | 279 of 551, or 50.6% | Calculated from the published study counts; 230 deferred cases were correct and 49 were errors |
| Coverage with an internal probability threshold of at least 0.98 | 22.5% | Same medical benchmark; retained-case accuracy was 97.6% |
| External benchmark performance | 89.9% accuracy at 32.0% coverage | 2026 VivaBench evaluation at a consistency threshold of at least 0.85 |
| Explicit-threshold LLM experiment | 11,000 model queries | 2026 Nature Machine Intelligence study; 1,000 SimpleQA questions at each of 11 thresholds |
| Manual review in a screening simulation | 2% to 3% postponed | 2026 conference study; held-out test set of 100 Cochrane reviews under realistic simulated human errors |
| Recall in the same simulation | At least 95% | Same simulated conditions and test set |
| Review volume under model collapse | 12.6% deferred | Same study's stress test; recall remained 97.5% |
What a confidence threshold controls
A binary classifier assigns a score and compares it with a threshold. Google explains that raising a positive-class threshold generally reduces both true positives and false positives. It also increases false negatives. That is why a higher cutoff can improve the precision of accepted positive predictions while reducing recall.
A human-review system often uses two boundaries instead of one. Cases above an upper threshold are accepted, cases below a lower threshold are rejected, and cases between the two go to a person. This middle band is the exception queue. Widening the band increases review volume. Narrowing it increases automation but gives people fewer chances to catch mistakes.
The score itself needs scrutiny. A value of 0.90 is not automatically a 90% probability of correctness. Calibration tests whether outcomes assigned a given probability occur at about that rate. A threshold based on an uncalibrated score can look precise while routing the wrong cases.
Precision and recall move with the cutoff
Precision is the share of predicted positive cases that are truly positive. Recall is the share of all actual positive cases the system finds. Both depend on a fixed threshold, and their importance depends on the cost of each type of error.
Consider an invoice-fraud screen. A false positive creates an unnecessary review. A false negative allows a risky payment to continue. Lowering the positive threshold may catch more fraud and raise recall, but it can also create more false alarms and lower precision. Raising it can produce a smaller, more precise alert queue while missing more fraud.
Accuracy alone can obscure this tradeoff. If only 1% of cases are positive, a model that labels every case negative is 99% accurate and detects none of the positives. Threshold reports should include the confusion matrix, precision, recall, class prevalence, and case counts.
A 2026 medical study quantifies the coverage tradeoff
A 2026 Nature Medicine study of on-premise medical AI agents tested confidence-based routing on 551 MIRA-v2 cases. The researchers measured behavioral consistency by running the system repeatedly and checking whether its diagnosis remained stable.
At a consistency threshold of 0.90, the system retained 272 cases for autonomous handling. That was 49.4% coverage, with 98.9% accuracy in the retained set. Three autonomous cases were wrong. The other 279 cases went to review, including 49 of the 52 total errors and 230 cases that the model had answered correctly.
The result makes the workload tradeoff visible. Review captured most errors, but reviewers also had to process many correct cases. Of the 279 deferred cases, 82.4% were correct model outputs. That percentage is calculated from the study's published counts and is not a separate reported endpoint.
The confidence signal mattered as much as the numeric cutoff. An internal probability score of at least 0.98 retained only 22.5% of cases and reached 97.6% accuracy. Behavioral consistency produced more coverage and higher retained-set accuracy in that experiment. On the external VivaBench dataset, the best reported point within the evaluated range was 89.9% retained accuracy at 32.0% coverage with a threshold of at least 0.85.
These findings cover benchmark cases, not live clinical outcomes. The authors also found that stricter routing reduced but did not eliminate autonomous errors. A high threshold did not create a zero-error system.
Language models respond to explicit abstention thresholds
A 2026 Nature Machine Intelligence experiment tested whether language models used confidence when deciding to answer or abstain. Its explicit-threshold phase ran 11,000 independent GPT-4o queries: 1,000 SimpleQA questions at each of 11 instructed confidence thresholds.
Higher instructed thresholds increased abstention. Accuracy among answered questions also increased, producing the expected accuracy-versus-coverage tradeoff. The statistical model reported a positive accuracy effect as thresholds rose, with a coefficient of 0.0051, a z score of 3.97, and a P value of 7.27 × 10^-5.
The study also found that verbal confidence was less effective at distinguishing correct from incorrect answers than log-probability-based confidence. This matters for operations. Asking a model to state that it is 90% confident does not prove that the score is calibrated well enough to control a production review queue.
Exception volume becomes human workload
Threshold testing should convert the deferred share into cases and hours. The basic calculation is:
review cases = eligible cases × deferred share
review hours = review cases × average handling minutes ÷ 60
The next table is a planning example, not a published benchmark. It assumes 20,000 eligible cases per week and four minutes of active review per deferred case.
| Deferred share | Weekly review cases | Active review hours | Review FTE at 30 productive queue hours |
|---|---|---|---|
| 10% | 2,000 | 133.3 | 4.4 |
| 25% | 5,000 | 333.3 | 11.1 |
| 50.6% | 10,120 | 674.7 | 22.5 |
| 75% | 15,000 | 1,000.0 | 33.3 |
The 50.6% row uses the deferred share calculated from the 2026 MIRA-v2 experiment only to show the arithmetic. It is not a staffing benchmark for another organization. Actual handling time varies with the evidence reviewers must inspect, the decision's consequence, and whether the person must correct the record or only approve it.
Queue time should be measured separately from active handling time. A case may need four minutes of work but wait two hours before a reviewer accepts it. Staffing reports need both figures because a short handling time does not guarantee a short customer or business delay.
Threshold failure can reveal itself through review demand
A 2026 ISPOR conference study simulated AI-assisted title and abstract screening across 160 Cochrane reviews. The researchers tuned their approach on 60 reviews and evaluated it on 100 held-out reviews.
Under the study's realistic simulated human-error conditions, the two-threshold strategy maintained at least 95% recall and postponed 2% to 3% of records for manual review. Under a model-collapse stress test, the deferred share rose to 12.6% while recall remained 97.5%.
That jump in exceptions acted as a warning. Yet the study also showed a more difficult failure mode. Under simulated automation bias, recall fell below 90% while conflict and abstention stayed low. Agreement between people and AI can therefore hide shared errors. A quiet review queue is not enough evidence that a threshold is safe.
The study was a simulation presented at a conference, and it used a weaker non-LLM classifier by design. Its percentages should not be applied as universal operating targets.
Review time needs more than one average
Average handling time is useful for budgeting, but it can hide slow and consequential cases. Report at least the median and 90th percentile for each exception type. Separate simple approvals from corrections, escalations, appeals, and incidents.
Review time should start when a case enters the queue and end when the final action is recorded. Useful intervals include:
- queue wait before a reviewer accepts the case
- active inspection and evidence-checking time
- correction or escalation time
- time waiting for a specialist or customer response
- repeat work after a case is reopened
Sampling matters too. Time studies should cover ordinary work instead of focusing on reviewers who know they are being observed or the easiest cases selected for a pilot.
Quality controls for a threshold-based workflow
NIST's AI Risk Management Framework Core calls for documented human-AI roles, ongoing monitoring, feedback, appeal, and override processes. Those controls translate into specific checks for confidence routing.
- Validate the threshold on a held-out sample that matches the intended population. Report dates, case counts, prevalence, and subgroup results.
- Test calibration, precision, recall, coverage, and residual errors. Do not approve a cutoff from accuracy alone.
- Run a shadow period in which people review every case. Compare what the proposed routing policy would have automated with the known outcomes.
- Audit a random sample above and below each threshold. High-confidence errors and low-confidence correct cases both affect the operating decision.
- Give reviewers the evidence, time, and authority to disagree. Record changes and reason codes instead of collecting approval clicks only.
- Monitor drift by model version, input source, language, customer group, and task type. Revalidate after material changes.
- Define a stop rule. A rise in errors, appeals, exception volume, or queue delay should reduce automation until the cause is understood.
Human review is itself a control that needs measurement. Double-review a sample, compare reviewer agreement, and adjudicate disagreements. Otherwise, the evaluation treats the reviewer as perfect even when the human labels are inconsistent.
Metrics to publish for every confidence threshold
| Metric | Definition | Why it matters |
|---|---|---|
| Coverage | Cases handled automatically divided by eligible cases | Shows how much work remains automated |
| Deferred share | Cases sent to people divided by eligible cases | Converts the cutoff into review demand |
| Precision | True positives divided by all predicted positives | Shows how reliable positive decisions are |
| Recall | True positives divided by all actual positives | Shows how many positive cases the system finds |
| Residual autonomous error rate | Wrong automated decisions divided by automated decisions | Measures errors that bypass review |
| Review correction rate | Deferred outputs materially changed divided by reviewed outputs | Shows the useful error-catching yield of review |
| Queue wait time | Time from deferral to reviewer acceptance | Reveals staffing and service-level pressure |
| Active review time | Minutes spent inspecting and resolving the case | Converts exceptions into labor demand |
| Appeal reversal rate | Appealed decisions changed after review divided by appeals | Finds failures that passed the first control |
| Calibration error | Difference between predicted confidence and observed outcome rate | Tests whether confidence values mean what users assume |
Report counts with percentages. A 1% error rate means 10 errors in 1,000 cases and 10,000 errors in one million cases. The operational consequence depends on volume.
Frequently asked questions
What is a good AI confidence threshold for human review?
No universal cutoff is supported by the evidence. Choose a threshold with a representative validation set and the costs of false positives, false negatives, review labor, and delay. Recheck it after the model or input population changes.
Does a 90% confidence score mean the AI is correct 90% of the time?
Only if the score is calibrated for that task and population. The 2026 medical study found that behavioral consistency and internal probability produced different accuracy-and-coverage results. The source of the score matters.
Does raising the threshold always improve quality?
It generally makes the automated set smaller and can improve its accuracy, but it can also miss more positive cases and increase human workload. The 2026 medical study still found three errors in the autonomous stream at its reported 0.90 consistency threshold.
How do teams estimate human review volume?
Multiply eligible case volume by the deferred share at the chosen threshold. Then multiply the result by measured handling time. Add queue coverage for peaks, audits, training, leave, escalations, and reopened cases.
Should reviewers see the AI confidence score?
The score can help with routing, but showing it can influence the reviewer. Test whether the display improves joint accuracy and watch for automation bias. Keep blind audits in the quality program.
How often should a confidence threshold be changed?
Review it on a scheduled cadence and after a material model, prompt, data, policy, or population change. Use control limits or defined stop rules rather than changing the cutoff in response to every small fluctuation.
Sources and limitations
- Google for Developers, Thresholds and the Confusion Matrix, accessed October 7, 2026. Official educational guidance on how thresholds change classification outcomes.
- Google for Developers, Accuracy, Recall, Precision, and Related Metrics, accessed October 7, 2026. Official metric definitions and cautions about class imbalance.
- Gu and Hopkins, On the Evaluation of Neural Selective Prediction Methods for Natural Language Processing, ACL 2023. Peer-reviewed comparison across six GLUE classification tasks; it evaluates confidence ranking methods rather than human staffing.
- Gao and colleagues, Causal Evidence That Language Models Use Confidence to Drive Behaviour, Nature Machine Intelligence, 2026. Controlled multiple-choice and abstention experiments; results do not establish production review thresholds.
- On-premise Medical AI Agents for Reliable Clinical Decision-making, Nature Medicine, 2026. Benchmark evaluation of confidence-based medical routing; not a clinical staffing trial or live patient-outcome study.
- Nowak and colleagues, A Tale of Two Thresholds, ISPOR 2026. Conference simulation across Cochrane review datasets; not a universal workload benchmark.
- NIST, AI Risk Management Framework Core, 2023. Voluntary risk-management guidance; it does not prescribe a confidence cutoff.
- NIST, Artificial Intelligence Risk Management Framework 1.0, 2023. Governance and monitoring framework; it does not supply staffing ratios or threshold targets.
All calculated percentages and staffing examples are labeled in the text. Published findings remain tied to their original populations, dates, tasks, and study designs.
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →