Key Takeaways
- Audit coverage and autonomous resolution measure different parts of the workflow.
- The public evidence does not establish one universal human exception rate.
- A useful benchmark separates policy exceptions, suspected fraud, missing evidence, and model uncertainty.
- Human workload should be calculated from exception volume, handling time, quality sampling, and appeals.
AI expense audit human exception benchmarks are easy to misread. A system can inspect every transaction while still sending a meaningful share to people. It can also auto-approve most reports without proving that every approval was correct. Coverage, autonomous resolution, exception rate, correction rate, and confirmed loss are separate measures.
Public data in 2026 gives finance teams useful reference points, but no independent source publishes a universal exception target. The strongest operational figures come from audited platform activity and named customer programs. Government and fraud research explain why human judgment remains necessary.
For the manual control steps behind these measures, use the employee expense report review checklist. Teams can connect review capacity to the virtual assistant performance metrics guide and compare the broader labor model in AI virtual assistant versus human virtual assistant statistics for 2026.
2026 benchmarks at a glance
| Measure | Published result | What it measures |
|---|---|---|
| Expense lines audited | 258 million in 2025 | Activity across AppZen customers, not an independent market total |
| Autonomous resolution | 61% to 78% across disclosed industries | Lines approved, rejected, or routed without human input |
| Global bank auto-approval | 76% of about 250,000 annual reports | One vendor customer deployment |
| Manual review after automation | Under 24 hours, down from as much as 22 days | Turnaround for the bank's remaining human review |
| Occupational fraud cases studied | 1,921 cases in 138 countries and territories | ACFE's 2024 case sample, not expense reports processed |
| Federal improper payments | $162 billion in fiscal year 2024 | Government payment errors, including but not limited to fraud |
These figures do not belong in one performance ranking. AppZen reports how its software and customers operated. ACFE studies detected occupational fraud cases. GAO measures improper payments across selected federal programs. Each source has a different population and denominator.
What an exception rate actually counts
An expense audit exception is a transaction or report that the automated process does not close on its own. That definition can cover several events:
| Exception type | Human decision required |
|---|---|
| Missing or unreadable evidence | Request a receipt, confirm an allowable substitute, or deny the claim |
| Policy exception | Apply a documented exception, seek approval, or reject the expense |
| Possible duplicate | Compare reports, card feeds, dates, merchants, and prior reimbursements |
| Suspected manipulation | Examine the original file, transaction evidence, and employee explanation |
| Tax or regulatory question | Apply the relevant recordkeeping and reimbursement rule |
| Model uncertainty | Review the evidence when the system cannot support a confident action |
Companies set different routing rules. One may send every out-of-policy meal to a manager. Another may reject the same claim automatically. Their human exception rates are not comparable until the action categories and thresholds match.
AppZen defines autonomous resolution as the share of expense lines its agents approve, reject, or route without human input. Its 2026 report gives industry rates from 61% in education to 78% in information, based on customer activity during 2025. The implied non-autonomous share ranges from 22% to 39%, but that subtraction is a planning illustration. AppZen does not label the remainder as a standardized human exception rate.
A bank case gives a clearer human-workload reference
AppZen describes a global bank that processes about 250,000 expense reports a year. The deployment audited all reports before payment and auto-approved 76%. The remaining share went through human review in an average of 0.83 days, compared with an outsourced cycle that had taken as long as 22 days.
If the published percentages are applied to the disclosed annual volume, the workload illustration is:
250,000 reports x 24% not auto-approved = 60,000 reports for another action
That calculation does not prove that 60,000 reports were errors or fraud. A non-auto-approved report may need evidence, an employee response, a policy decision, or a reviewer confirmation. It does show why an automation business case needs a human queue model rather than an auto-approval percentage alone.
The case study comes from the software provider and does not disclose the bank, sampling method, false-positive rate, or correction rate. Treat it as an operating reference, not an independent accuracy study.
Fraud data sets the risk context, not the exception target
ACFE's Occupational Fraud 2024: A Report to the Nations, published in 2024, analyzed 1,921 cases across 138 countries and territories. The cases caused more than $3.1 billion in total losses, and the median loss was $145,000. Asset misappropriation appeared in 89% of cases, while tips detected 43% of the cases.
Those results support layered controls. They do not show what percentage of ordinary expense reports should be flagged. ACFE's sample contains investigated fraud cases rather than a denominator of all transactions. It also shows why finance teams should keep reporting channels and investigative review even when software checks every submitted line.
GAO provides a second warning about labels. On March 11, 2025, it reported $162 billion in estimated federal improper payments for fiscal year 2024. About $135 billion, or roughly 84%, consisted of overpayments. GAO explicitly notes that improper payments can result from inaccurate records, errors, or fraud, and that not every improper payment is fraudulent.
An expense program should follow the same discipline. A duplicate, missing receipt, policy violation, and deliberate false claim should not share one undifferentiated fraud label.
IRS rules explain why documentation exceptions need people
IRS Publication 463 for 2025 says taxpayers must substantiate travel, gift, and transportation expenses with records covering the required elements. The publication states that written evidence is generally more reliable than oral evidence and that a computer record can qualify as an adequate record. It also says documentary evidence is generally required, with exceptions that include a non-lodging expense under $75.
That $75 rule is not a universal company auto-approval threshold. Publication 463 also requires evidence for amount, time, place or description, and business purpose as applicable. Company policy, accountable-plan rules, card controls, and local tax requirements may demand more.
Automation can check whether fields and files exist. A person may still need to decide whether other evidence is sufficient, whether an explanation fits the business purpose, or whether an exception is allowed. Those decisions belong in the exception taxonomy so documentation work does not get mixed with suspected fraud.
Research supports human review of anomaly alerts
Peer-reviewed accounting research has tested machine learning on journal-entry data, where the operating problem resembles expense exception triage. A 2022 study in the Journal of Information Systems evaluated an autoencoder on real-world accounting data and synthetic fraud cases. The authors frame anomaly detection as decision support for auditors, not proof that a flagged entry is fraudulent.
That distinction matters in expense review. An anomaly score ranks unusual records. It does not establish intent, policy treatment, or recoverable loss. Human reviewers need the underlying receipt, transaction, policy rule, employee response, and prior history to close a case.
Public vendor benchmarks rarely disclose precision, recall, or a human overturn rate. Without those measures, a lower exception rate could mean better automation, looser controls, or more automated rejection. Finance leaders should require a labeled evaluation set before accepting accuracy claims.
Build a benchmark that measures both AI and human work
Start with a stable population such as submitted expense lines or complete reports. Then record the first automated decision and final disposition separately.
| Metric | Calculation |
|---|---|
| Audit coverage | Items checked by the system divided by eligible items |
| Autonomous closure rate | Items finally approved or rejected without human action divided by checked items |
| Human exception rate | Items requiring a person to act divided by checked items |
| Confirmed exception rate | Human-reviewed items with a policy, evidence, duplicate, or fraud finding divided by reviewed items |
| Overturn rate | Automated decisions changed by a reviewer or appeal divided by reviewed automated decisions |
| False-positive rate | Compliant items incorrectly flagged divided by all known compliant items |
| Review time | Productive reviewer minutes divided by completed human exceptions |
| Recovery rate | Money prevented or recovered divided by confirmed recoverable value |
Use mutually exclusive final dispositions. An item can trigger several alerts, but it should have one closing result for reporting. Keep the individual alerts for model evaluation.
Calculate weekly reviewer demand with observed values:
human review hours = checked items x human exception rate x average review minutes / 60
Add time for employee follow-up, appeals, quality sampling, policy maintenance, investigations, and model monitoring. Divide by productive review hours per person rather than paid hours. This is an internal staffing method, not a published industry ratio.
Controls to keep when automation rises
A high autonomous resolution rate should not remove human accountability. Keep approval authority separate from model administration. Sample auto-approved and auto-rejected items. Track reviewer disagreement by policy rule and model version. Require a documented path for employee explanations and appeals. Preserve the evidence used for each decision.
The employee expense report review checklist can anchor the case-level process. For staffing, assign service measures such as queue age, completion time, rework, and documented accuracy using the virtual assistant performance metrics.
Source record and limitations
| Source | Publication date | Claim used | Limitation |
|---|---|---|---|
| The rise of the autonomous CFO, AppZen | 2026 | 258 million lines audited in 2025; 61% to 78% disclosed industry autonomous rates | Vendor analysis of its customers |
| Agentic AI expense report audit for financial services and banking, AppZen | 2026 | Global bank volume, 76% auto-approval, and review turnaround | Vendor case study; customer not named |
| Occupational Fraud 2024: A Report to the Nations, ACFE | 2024 | Case count, countries, losses, scheme prevalence, and detection methods | Detected fraud cases, not all expenses |
| GAO reports estimated $162 billion in improper payments, GAO | March 11, 2025 | Fiscal 2024 improper payments and overpayments | Federal programs; improper payment is broader than fraud |
| Publication 463 (2025), IRS | March 2, 2026 revision posting | Expense substantiation and documentary-evidence rules | U.S. federal tax guidance, not an AI benchmark |
| Deep Learning for Anomaly Detection in Accounting Data, Journal of Information Systems | 2022 | Anomaly detection tested as audit decision support | Journal-entry setting, not employee expense reports |
Frequently asked questions
What is a good human exception rate for AI expense audit?
No authoritative source provides one universal rate. AppZen's disclosed autonomous rates imply that a material remainder is not closed autonomously, but company policy and routing design determine the human queue. Benchmark against your own stable categories and risk thresholds.
Does 100% audit coverage mean no manual review?
No. Coverage means the system checked every eligible item. Exceptions, uncertain evidence, investigations, quality samples, and appeals can still require people.
Is every flagged expense fraudulent?
No. A flag may identify missing evidence, a policy breach, a duplicate, unusual behavior, or model uncertainty. Fraud requires investigation and evidence of intentional deception.
Which number matters most?
Use a balanced set: coverage, autonomous closure, human exception volume, overturn rate, confirmed findings, review time, and value prevented or recovered. A single automation percentage cannot show control quality.
Final benchmark takeaway
The public evidence supports a hybrid operating model. AI can inspect the full expense population and close many routine items. People remain responsible for ambiguous evidence, policy judgment, investigation, appeals, and quality control.
The most defensible AI expense audit human exception benchmark is local and reproducible. Define the denominator, preserve the final disposition, measure reviewer effort, and test automated decisions against labeled outcomes. That gives finance leaders a workload forecast and a control-quality measure instead of a marketing percentage.
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →