Key Takeaways
- The traditional four-fifths screen flags a group selection rate below 80% of the highest group's rate, but federal guidance calls it a rule of thumb rather than a legal definition
- In the federal agencies' worked example, a 45% American Indian selection rate divided by a 60% White selection rate produces a 75% impact ratio and an adverse-impact signal
- A 2022 meta-analysis ranked structured interviews highest among reviewed selection methods, with an estimated validity of .42, compared with .31 for cognitive ability tests
- In an EEOC case against Dial, 38% of women passed a strength test compared with 97% of men, and the final award was approximately $3.3 million for 52 rejected applicants
- The 2026 federal policy shift makes legal advice essential, but employers still need job-related assessments, consistent administration, applicant-level records, and checks for less discriminatory alternatives
A hiring test can predict performance and still produce sharply different outcomes across applicant groups. Validity asks whether an assessment predicts or represents important work. Adverse-impact analysis asks who advances after the assessment is applied. A favorable vendor validity report does not reveal the selection rates for a particular job, location, or applicant pool.
The legal context changed in 2026. On June 9, the U.S. Department of Justice's Office of Legal Counsel concluded that the federal Uniform Guidelines on Employee Selection Procedures, or UGESP, embody an unconstitutional interpretation of Title VII disparate-impact liability. OPM subsequently moved to remove UGESP references from federal civil-service regulations. This federal policy change is not a court decision that erases every other obligation. New York City's automated-employment law, for example, still requires annual bias audits for covered tools. Employers should use the figures below as risk and process measures, then obtain advice for the jurisdictions where they hire.
The selection-rate calculation
The traditional four-fifths analysis uses four steps:
- Divide the number selected in each group by the number of applicants in that group.
- Find the group with the highest selection rate.
- Divide every other group's selection rate by that highest rate.
- Review any impact ratio below 0.80.
The federal agencies' official UGESP questions and answers include this example:
| Applicant group | Selection rate | Impact ratio versus highest rate | Traditional screen |
|---|---|---|---|
| White | 60% | 100% | Reference group |
| American Indian | 45% | 75% | Below four-fifths |
| Hispanic | 48% | 80% | At four-fifths |
| Black | 51% | 85% | Above four-fifths |
The same guidance gives a simpler example: 48 of 80 White applicants are hired, a 60% rate, while 12 of 40 Black applicants are hired, a 30% rate. The impact ratio is 30% divided by 60%, or 50%. That is well below 80%.
The agencies expressly call four-fifths a rule of thumb. A ratio below 80% is not an automatic finding of unlawful discrimination. A ratio above 80% is not a safe harbor either. Smaller differences can matter when they are statistically and practically significant, while small applicant pools can make a ratio unstable.
What the evidence says about assessment types
No assessment format has a fixed adverse-impact rate. The outcome changes with the construct being measured, cut score, applicant population, job, scoring method, and how the assessment is combined with later stages. Published research is still useful for choosing where to look first.
| Assessment type | What the evidence shows | Practical implication |
|---|---|---|
| Structured interview | A 2022 review placed structured interviews first among the selection procedures studied, with estimated validity of .42. | Use the same job-linked questions and scoring anchors for every candidate, then monitor interviewer and stage-level results. |
| Cognitive ability test | The same review estimated validity at .31 and paired the method with evidence of Black-White subgroup differences. | Do not assume predictive value resolves impact. Check the cut score, job relevance, and available alternatives. |
| Short-term memory test | A meta-analysis covering 27,973 people in 31 samples found an average Black-White mean difference of 0.42 standard deviations, described by the authors as less than half the difference typical of general cognitive tests. It reported corrected validity of .41 for job performance. | A narrower cognitive construct may reduce group differences in some settings, but the employer still needs local outcome data. |
| Work sample | Research warns that the common claim of low adverse impact often rests on incumbent samples. Two public-sector applicant datasets showed that work samples can produce more impact than expected. | A realistic task is not automatically low impact. Audit actual applicants, not only current employees or vendor pools. |
| Personality assessment | The EEOC found probable cause that Best Buy personality assessments used from 2003 through 2010 adversely affected applicants based on race and national origin. Best Buy stopped using the assessments and reached a conciliation agreement without admitting liability. | Confirm the measured trait is tied to the job and inspect item, score, and advancement patterns. |
| Physical ability test | Public enforcement cases show large sex differences when a test does not match essential job demands. | Recreate necessary work demands, document them, and avoid cutoffs based on convenience or incumbent tradition. |
| Automated scoring tool | New York City requires a bias audit within one year before a covered automated employment decision tool is used, publication of a summary, and candidate or employee notices. | A vendor report does not replace the employer's own inventory, notice process, and local selection-rate review. |
The .42 and .31 figures are validity coefficients, not selection rates. They describe the relationship between assessment results and later performance in the reviewed research. They cannot be converted into an impact ratio without applicant scores, group counts, and the employer's decision rule.
Documented demographic outcomes and employer costs
Enforcement records show what weak assessment design can cost. They also show why stage-level data matters.
Dial strength test: women fell from 46% to 15% of hires
Dial Corporation introduced a pre-employment strength test at a sausage plant. Across the tested applicants, 38% of women passed compared with 97% of men. The Eighth Circuit opinion also reports that women had made up 46% of new hires before the test. After the test was introduced, that share fell to 15%. The court found that the test had a disparate impact on women and that Dial had not shown the test was an effective measure of the strength required for the job. The final award was approximately $3.3 million for 52 rejected women.
This case is particularly useful because the record included both an outcome shift and a validation problem. The employer argued that the test would reduce injuries. The court found the injury evidence did not establish that the test delivered that result.
Ford cognitive test: $1.6 million for nearly 700 Black workers
Ford and related defendants agreed to pay $1.6 million to settle an EEOC lawsuit involving a written apprenticeship test. The test measured verbal, numerical, spatial-reasoning, and mechanical concepts. The EEOC said it disproportionately excluded Black workers. The settlement covered nearly 700 people and required a jointly selected industrial-organizational psychologist to design a replacement procedure that would predict success while reducing adverse impact.
The case makes an important point about validation. EEOC materials state that the test had been validated in 1991. The employer still faced exposure because later, less discriminatory procedures could serve its needs and it did not change the process.
Pennsylvania physical test: $2.2 million and up to 65 priority hires
The Justice Department alleged that Pennsylvania State Police physical tests assessed skills that were not required for entry-level trooper work and disproportionately excluded women. A court-approved settlement required $2.2 million in compensation, priority hiring relief for up to 65 qualified women, and a new test for future candidates.
Written public-safety exams: $145,000 to $700,000 settlements
A Justice Department summary lists several test cases. Portsmouth paid $145,000 and offered 10 priority hires over an entry-level firefighter written examination. Dayton paid $450,000 and offered 14 priority hires after challenges to a police written exam and firefighter minimum requirements. Another physical-abilities case produced $700,000 in back pay and 18 priority hires for women.
These are settlement or judgment amounts, not an estimate of total employer cost. Legal fees, expert analysis, staff time, test replacement, record reconstruction, monitoring, and delayed hiring sit outside the announced relief.
Validation practices that stand up to scrutiny
The classic UGESP framework recognizes three strategies:
- Criterion-related validation tests whether scores are statistically related to a relevant performance measure.
- Content validation shows that assessment content represents important job behaviors or work products.
- Construct validation shows that the assessment measures a defined characteristic and that the characteristic matters to job performance.
A validation label does not substitute for evidence. Start with a current job analysis. Define the work behavior or characteristic being measured, explain the scoring rule, and record why the cutoff is appropriate. Keep enough applicant-level data to calculate results for each stage rather than only for the final hiring decision.
Employers should ask vendors for the sample size, jobs, applicant population, demographic results, criterion measures, confidence intervals, and date of each study. A pooled vendor audit can conceal a poor result for one employer or job. Local monitoring tests whether the tool behaves as expected in the environment where it is actually used.
When a stage shows a material difference, examine less discriminatory ways to measure the same requirement. Options may include a better structured interview, a more representative work sample, a changed cutoff, accessible administration, or a different weighting across predictors. The goal is not to manipulate group outcomes. It is to remove requirements that do not improve job-related decisions.
A practical audit table for every hiring stage
| Field | Why it belongs in the audit |
|---|---|
| Job, location, and requisition | Aggregating unlike roles can hide a problem. |
| Assessment version and date | Vendor models, items, and scoring rules change. |
| Applicants entering and advancing | These counts produce the selection rate. |
| Race or ethnicity and sex groups | These support the traditional UGESP calculation and many local audit rules. |
| Intersectional groups where sample size permits | A broad category can hide a sharper disparity. |
| Pass score and override | Overrides can create impact even when the test itself does not. |
| Disability accommodation request and outcome | The ADA separately governs disability-related inquiries, medical exams, and reasonable accommodation. |
| Validity evidence and job analysis date | Old evidence may not match the present job or current assessment. |
| Alternative reviewed and decision | This records whether an effective option with less impact was available. |
Small samples need context. Report the raw counts beside every percentage, extend the review window when the same procedure is used consistently, and use qualified statistical and legal review before drawing a conclusion. A ratio based on one hire out of two applicants should not be presented with the certainty of a result based on thousands.
How recruiting operations affect the data
Adverse-impact work depends on clean process records. Candidate sourcing, assessment invitations, accommodation messages, interview scheduling, disposition codes, and version tracking all need consistent handling. A recruiting virtual assistant can support those administrative steps, but the employer must retain responsibility for assessment design and hiring decisions.
Teams building the operating process can use this guide on how to hire a virtual assistant to define access and review controls. Stealth Agents also provides recruitment support services for sourcing, scheduling, candidate communication, and recruiting administration.
2026 compliance note
Federal policy is unsettled. The DOJ's June 2026 opinion rejects the prior federal interpretation, and OPM's subsequent rule concerns federal civil-service regulations. At the same time, the text of Title VII, court precedent, other federal statutes, and state and local rules still shape employer obligations. New York City's audit requirement is one clear example of an outcome-based rule that remains operational.
Selection-rate monitoring is still useful even where a specific four-fifths requirement does not apply. It can reveal a broken cutoff, an inaccessible test, inconsistent administration, a vendor-model change, or a recruiting funnel that is losing qualified applicants. Employers should treat 80% as a diagnostic threshold, not permission to discriminate and not proof that a process is lawful.
Sources
- U.S. Equal Employment Opportunity Commission, Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures.
- Electronic Code of Federal Regulations, 29 CFR Part 1607, Uniform Guidelines on Employee Selection Procedures.
- U.S. Equal Employment Opportunity Commission, Employment Tests and Selection Procedures.
- U.S. Department of Justice, Office of Legal Counsel, Constitutionality of Disparate-Impact Liability Under Title VII, June 9, 2026.
- U.S. Office of Personnel Management, Removal of References to the Uniform Guidelines on Employee Selection Procedures in Federal Personnel Regulations, 2026.
- Sackett, Zhang, Berry, and Lievens, Revisiting meta-analytic estimates of validity in personnel selection, Journal of Applied Psychology, 2022.
- Roth, Bobko, McFarland, and Buster, Short-term memory tests in personnel selection: Low adverse impact and high validity, Intelligence, 1996.
- Bobko, Roth, and Buster, Work Sample Selection Tests and Expected Reduction in Adverse Impact: A Cautionary Note, 2005.
- U.S. Equal Employment Opportunity Commission, Appeals Court Upholds EEOC Sex Discrimination Claim Against Dial, 2006.
- U.S. Equal Employment Opportunity Commission, Ford, affiliates, and UAW agree to pay $1.6 million, 2007.
- U.S. Department of Justice, United States v. Commonwealth of Pennsylvania, settlement approved 2021.
- New York City Department of Consumer and Worker Protection, Automated Employment Decision Tools.
- U.S. Department of Justice, Civil Rights Division, Accomplishments, 2009 to 2012: Challenging Unlawful and Ineffective Employment Tests.
Tags
Ready to put this into practice?
Book a free 15-min match call
Tell us what role you're filling. We'll match you with a pre-vetted virtual assistant - or tell you honestly if we're not the right fit.
Book a free call →