- Blog
- AI Interview Bias Audit: The 2026 5-Step Framework
AI Interview Bias Audit: The 2026 5-Step Framework

The regulators changed. Your obligation didn't.
In January 2025, the EEOC pulled its AI hiring guidance from its website. That didn't repeal Title VII or erase the four-fifths rule. It handed the enforcement job to the states, and they're moving fast.
An AI interview bias audit checks whether your screening tool passes or rejects protected groups at different rates, using impact ratios from the EEOC Uniform Guidelines (29 CFR 1607.4(D), 1978). It's the document regulators, plaintiffs' counsel, and hiring managers will ask for.
The gap is already visible. New York State Comptroller Thomas DiNapoli's office reviewed the same 32 companies NYC's regulator had checked under Local Law 144. The city flagged one instance of non-compliance. The state auditors found at least 17. The regulator has committed to tougher enforcement.
Audit yourself before someone else does.
TL;DR
- Federal AI guidance is gone; Illinois (Jan. 1, 2026), Colorado (delayed to Jan. 1, 2027), NYC, and California (Oct. 1, 2025) now set the standard.
- The four-fifths rule: any group's selection rate below 80% of the highest group's rate is evidence of adverse impact.
- Five steps: define groups, compute impact ratios, handle small samples, split vendor vs. employer duties, remediate and re-test.
- Step 5 is the one most teams skip and the one that actually protects you.
Step 1: Define your groups and collect the right data
Under NYC Local Law 144, you need to cover sex, race/ethnicity, and intersectional sex-by-race groups (Hispanic women or Black men as distinct cells, not just "women" and "Black candidates" separately).
Three prerequisites:
- Pass-through outcomes by stage. How many candidates entered and cleared each stage (screen, interview, advance), broken out by group.
- Demographic self-ID rates. What share of candidates disclosed sex and race/ethnicity.
- A denominator plan. Incomplete self-ID means partial populations. A "62% pass rate for women" built on 30% self-ID tells you less than it sounds.
If you can't produce stage-by-stage counts by group, you've got a dashboard, not an audit. Fix data collection first. For how your tool generates these outcomes, see our explainer on how AI interview scoring works.
Executive takeaway: Your audit is only as trustworthy as your self-ID rate. Collect stage-level outcomes by protected class, including intersectional cells, before you run any numbers.
Step 2: Compute selection rates and impact ratios, the worked example
Here's the four-fifths rule: if any group's selection rate falls below 80% of the highest group's rate, that's evidence of adverse impact (29 CFR 1607.4(D), 1978). Three operations: calculate the selection rate, divide by the top group's rate, compare to 0.80.
A passing case on 1,000 candidates, analyzed by sex:
- Men: 520 screened, 234 passed, 45.0% selection rate, impact ratio = 45.0 / 48.1 = 0.94
- Women: 480 screened, 231 passed, 48.1% selection rate (reference group)
Women are the reference group (48.1%). Men's ratio is 0.94. Passes.
Now a failing case on race/ethnicity:
- White: 500 screened, 275 passed, 55.0% selection rate (reference group)
- Hispanic: 250 screened, 120 passed, 48.0% selection rate, impact ratio = 48.0 / 55.0 = 0.87
- Black: 250 screened, 103 passed, 41.2% selection rate, impact ratio = 41.2 / 55.0 = 0.75
White candidates set the reference (55.0%). Hispanic candidates clear at 0.87. Black candidates land at 0.75, below the 0.80 threshold. That's evidence of adverse impact, and it triggers Step 5.
The same dataset passes on sex and fails on race. Run every required cut and report the failing one, not the flattering one.
Step 3: Handle the statistics honestly, small samples and intersectionality
The four-fifths rule is a rule of thumb, not a statistical test. It ignores sample size, which means it can flag noise or miss a real gap. The Uniform Guidelines note that smaller differences can still constitute adverse impact when they're statistically and practically significant (29 CFR 1607.4(D)).
A 12-person intersectional cell can swing from 0.60 to 1.10 on a single candidate. Pair the four-fifths ratio with a significance check on larger cells and flag any cell that's too small to interpret. When numbers are too small, aggregate over a longer period or across similar uses of the tool.
Local Law 144 still requires intersectional categories even when cells are small. Small means annotated, not omitted. Show the cell, note that N is too low, and aggregate across a longer window.
Buyer test: if your vendor's report shows a clean ratio for a 9-person subgroup and calls it a pass, ask how they handled significance. No answer means the report is decoration.
Step 4: What your vendor owes you vs. what you run yourself
You need two separate audits.
The vendor's audit is what Local Law 144 requires: independent, no more than one year old, published publicly, with candidates notified at least 10 business days before the tool is used. It runs on aggregate data. Your audit runs on your applicant flow: your job titles, markets, and candidate mix. The vendor's aggregate can pass while your deployment fails.
The liability math changed in 2025. In Mobley v. Workday, the court let discrimination claims proceed against the vendor on an agent theory, meaning a tool provider screening on your behalf can be treated as an "employer" under Title VII. On May 16, 2025, the court granted conditional certification of a nationwide ADEA collective covering applicants 40 and older rejected through the system since 2020.
Vendors can now be sued directly, yet according to TermScout data published through Stanford CodeX, 88% of AI vendors cap their liability at monthly subscription fees and only 17% provide a regulatory-compliance warranty. Ask:
- Does the liability cap carve out discrimination and compliance claims, or does it stop at your subscription fee?
- Will the vendor put compliance with AI hiring laws in writing?
- Will they indemnify you for adverse-impact claims from their model?
- Will they deliver audit-ready output on your data, not just their aggregate?
For the full vendor-selection framework, see our RFP checklist for AI recruiting software and the guide to first-round interview automation.
Executive takeaway: Post-Mobley, get discrimination and compliance carve-outs in writing. A subscription-fee cap won't protect you.
Step 5: Remediate, document, and re-test
A ratio failed. That doesn't automatically mean you broke the law. It's a signal that requires a decision, root-cause investigation, and documentation.
Pause or proceed?
- Pause if the ratio is below 0.70, the sample is robust, and the tool gates advancement.
- Proceed with monitoring if the ratio is marginal (0.78 to 0.80), the sample is small, or a human reviews every decision. Document why and set a re-test date.
If the tool alone can end a candidacy and it's failing, pause it. If a human owns the decision, investigate while you monitor.
Find the root cause
Work these root-cause paths in order:
- Rubric wording. Criteria that reward a specific communication style or credential can disadvantage groups. Most common, most fixable.
- Scoring model. Look for features that correlate with a protected class. Ask the vendor what drives the score and whether they tested it on a population like yours.
- Question set. Questions that assume a specific background narrow the funnel before scoring starts.
- Modality effects. Voice screens can penalize accents; video screens can penalize certain disabilities. Compare pass rates across modalities.
Document everything
For every failed audit, keep the pass-through data, impact-ratio calculations, root-cause finding, remediation, and re-test result. California's FEHA regulations require employers to retain automated-decision-system data for four years as of October 1, 2025. Your documentation window is measured in years, not cycles.
Set a re-test cadence and know when to call counsel
Re-test after any remediation and on a schedule: quarterly for high-volume roles, annually at minimum for Local Law 144. Loop in counsel when a failing ratio is large and persistent, when it appears on a tool that gates advancement without human review, or when you're about to pause a tool mid-cycle.
The 30-day bias-audit sprint
If you're starting from zero, here's a realistic timeline:
- Week 1, scope and data. Identify tools, pull stage-level pass-through data by group, check self-ID coverage. Request the vendor's most recent independent audit.
- Week 2, compute. Run impact ratios for sex, race/ethnicity, and intersectional cells. Flag small samples. Add significance testing where possible.
- Week 3, diagnose. Work the root-cause paths for every failing ratio. Draft remediations and log needed contract carve-outs.
- Week 4, remediate and document. Apply fixes, re-test, assemble documentation, publish your Local Law 144 summary if required, set the next re-test date.
Executive takeaway: A failed ratio is a decision point, not a verdict. Pause tools that gate advancement, document the chain, and re-test on a cadence.
Continuous monitoring beats annual panic
An annual audit tells you whether your tool was fair last January, with last January's model and applicant mix. But a tool that passed in Q1 can drift by Q3, and you won't know until the next check, or until a plaintiff's expert runs the numbers first.
Humanly's screening across chat, SMS, and voice, with native Workday and iCIMS integrations, includes a compliance layer that monitors adverse impact continuously. A drifting ratio surfaces the week it happens, not eleven months later. For the broader framework, see our defensible hiring playbook.
Want to see what continuous, audit-ready monitoring looks like in practice? Book a demo.
FAQs
What is an AI interview bias audit?
It's a structured review that checks whether an automated screening tool passes or rejects protected groups at different rates. It uses selection rates and impact ratios by sex, race/ethnicity, and intersectional sex-by-race categories, compared against the four-fifths rule (29 CFR 1607.4(D)).
What is the four-fifths rule for adverse impact analysis?
If any group's selection rate falls below 80% of the highest group's rate, federal enforcement agencies generally treat that as evidence of adverse impact. The rule comes from the EEOC Uniform Guidelines (29 CFR 1607.4(D), 1978). It's a rule of thumb, not a formal statistical test, so you should pair it with significance testing on larger samples.
Does NYC Local Law 144 require both a vendor audit and my own audit?
Local Law 144 requires an independent bias audit dated within one year, published publicly, with candidates notified at least 10 business days before use. That audit runs on aggregate data. To know whether the tool is fair on your specific applicant flow, you also need an adverse-impact analysis on your own hiring data. The two answer different questions.
Is a failing four-fifths ratio automatically illegal?
No. A rate below 80% of the top group's is evidence of adverse impact, not proof of a violation. It means you need to investigate the root cause and either fix the tool or document a business-necessity justification. Failing to act on a known disparity is the larger risk.
Which AI hiring laws take effect in 2026?
Illinois HB 3773 amends the Illinois Human Rights Act to regulate AI in employment effective January 1, 2026. Colorado's AI Act has been delayed to January 1, 2027 and significantly scaled back after the governor signed SB 189 in May 2026. California's FEHA regulations on automated-decision systems took effect October 1, 2025, including a four-year data-retention requirement. With federal AI hiring guidance rescinded in January 2025, these state laws now carry the enforcement weight.