← AI hiring compliance

AI Hiring Bias Audits: Who Needs One, Who Can Perform It, and What It Costs

8 min read · Last reviewed 1 Aug 2026

General information, not legal advice. Laws in this area change; verify against the official sources at the end of this guide and confirm specifics with employment counsel.

"Bias audit" gets used loosely, but in AI hiring it has a precise legal meaning in exactly one place: New York City's Local Law 144, which requires an annual, independent, published audit of any automated employment decision tool used for NYC jobs. Everywhere else, auditing your screening for bias is not (yet) a formal mandate, but California makes your testing admissible evidence, Colorado expects impact assessments from larger deployers, and any disparate-impact claim anywhere will start with the same math an audit runs.

This guide separates the official audit from the self-check, explains who is allowed to perform the official one, what the auditors actually compute, what it realistically costs a small employer, and how to show up prepared. General information, not legal advice.

When an audit is legally required, and when it is just smart

Required: if you use an AEDT to screen candidates for NYC-based jobs, the tool must have had a bias audit by an independent auditor within the year before use, refreshed annually, with a results summary published on your careers site. That is the only US jurisdiction with a hard audit mandate for employers as of this guide's last review.

Dropped: Colorado's original AI Act expected impact assessments from larger deployers, but the state repealed and reenacted that law in May 2026 (SB 26-189) before it took effect; from January 1, 2027 Colorado requires notice, adverse-decision explanations, and records rather than audits.

Smart everywhere: California's FEHA regulations make anti-bias testing (or its absence) relevant evidence in discrimination claims, Illinois attaches liability to discriminatory effect regardless of intent, and Title VII disparate-impact claims run on selection-rate statistics. In every one of those postures, an employer who periodically checks their funnel and documents the results is in a categorically better position than one who never looked. The difference between the mandated audit and the smart practice is who performs it and whether it is published, not the underlying math.

Who counts as an independent auditor

For the official LL144 audit, independence is defined by exclusion. The auditor cannot be someone who was involved in using, developing, or distributing the tool; cannot have an employment relationship with you or with the tool's vendor; and cannot have a direct or material indirect financial interest in either, beyond the audit fee itself.

That rules out the three shortcuts everyone asks about first. You cannot audit yourself. Your screening vendor cannot audit their own product for you. And your regular consultant who also helped you configure the tool has an independence problem. What remains is a genuine third party: specialist algorithmic-audit firms, some accounting and consulting practices that built AI-audit lines, and academic or boutique statistical shops.

One structural relief for small employers: the audit attaches to the tool, and the rules allow an audit built on historical data pooled across multiple employers using the same tool. In practice, responsible vendors commission an independent audit of their product annually and hand customers the summary and distribution date. Your job then shrinks to verifying the audit exists, is current, was genuinely independent, and covers the tool as you use it, and publishing the summary. If a vendor cannot produce this, price in commissioning your own or switching.

What the audit actually computes

The core of an LL144 bias audit is the same selection-rate analysis the four-fifths rule uses, formalized. For tools that produce a binary outcome (advance / do not advance), the auditor computes selection rates by sex, by race/ethnicity, and for intersectional combinations of the two, then the impact ratio: each category's rate divided by the most-selected category's rate.

For tools that output scores rather than decisions, which is most modern screening software, the rules use a scoring-rate variant: the share of each group scoring above the median score, compared across groups the same way. The audit must also disclose the number of individuals in each category and can exclude categories below a small share of the data with justification, which keeps one-applicant categories from producing absurd ratios.

What the audit needs as input is therefore concrete: per-candidate tool outputs (scores or classifications), per-candidate outcomes at the audited step, and per-candidate demographic categories. The demographics come from voluntary self-identification or the employer's EEO records, never from inference. Note what the audit is not: it is not a code review, not a validation that your criteria predict job performance, and not a certification of overall fairness. It is a statistical snapshot of outcomes by group, published so outsiders can see it.

What it costs and how long it takes

Prices vary widely with scope, and the market is young, so treat ranges as orientation rather than quotes. A single-tool audit riding on clean, well-structured data has been quoted in the low-to-mid four figures by boutique firms; complex engagements covering multiple tools, messy data, and bespoke methodology run into five figures and beyond. The dominant cost driver is not the statistics, which are a spreadsheet afternoon, but the data preparation: assembling per-candidate outputs, outcomes, and demographics into one joinable dataset, and documenting where each column came from.

Timeline follows the same logic: with prepared data, weeks; with data archaeology, months. Since LL144 requires the audit within a year before use and annually thereafter, the sane pattern is a standing yearly engagement with the data pipeline built once.

For most small employers the realistic cost is lower than any of this, because the vendor-level pooled audit carries the load: your expense is verifying and publishing, plus maintaining your own records in case your usage pattern ever needs an employer-specific look. The self-check tier costs nothing at all, which is the subject of the next section.

Preparing: the audit-ready export and the free self-check

Whether you face a mandated audit or just want the smart-practice version, preparation is identical: be able to produce, on demand, a structured per-candidate record of what the tool scored, what was decided, and when.

SiftFirst builds this in. Every screening stores the human-set rubric, each candidate's per-criterion scores with quoted evidence, and the shortlist outcome; the audit-ready export emits it as a clean table with no demographics in it, formatted so an auditor (or you) can join it against separately held self-ID data and compute selection and scoring rates immediately. That join-and-compute step is also available free: the bias audit self-check at /tools/bias-audit-check runs the impact-ratio math entirely in your browser, flags groups under the four-fifths line, and never uploads or stores anything, so the sensitive demographic data never leaves your machine.

Be precise about what the self-check is: a pre-audit and monitoring practice, explicitly not the official LL144 audit, which requires the independent auditor described above. Used quarterly, it means the official audit never surprises you, and in the jurisdictions where testing is evidence rather than mandate, it is the evidence. Confirm audit specifics, especially auditor independence and publication details, with counsel.

Key takeaways

  • Only NYC LL144 mandates an independent, annual, published bias audit; California and Illinois make testing your practical defense, and Colorado's replacement ADMT law (2027) asks for records and explanations, not audits.
  • Independent means genuinely third-party: not you, not your vendor, no financial ties beyond the audit fee. Vendor-commissioned pooled audits of the tool are the realistic path for small employers.
  • The audit computes selection rates (or above-median scoring rates) and impact ratios by sex, race/ethnicity, and intersections, from tool outputs joined with voluntarily self-identified demographics.
  • Costs range from low four figures to five-plus depending on scope, and data preparation, not statistics, drives the bill; a standing yearly engagement with a built-once pipeline is the sane pattern.
  • Prepare with a structured per-candidate export of scores and outcomes, and run a free browser-based four-fifths self-check quarterly so the official audit never surprises you.

Screening built for these rules

SiftFirst scores candidates against criteria you set, quotes the resume line behind every score, never auto-rejects, and exports the records these laws expect. The candidate notice generator and bias audit self-check are free.

FAQ

Can my screening vendor audit their own tool for me?

Not for LL144: the auditor must be independent of both you and the vendor. What vendors legitimately do is commission an independent third-party audit of their tool, often on historical data pooled across customers, and give you the summary and distribution date to publish. That satisfies the requirement if the audit is current and genuinely independent, so the question to ask your vendor is who performed it and when, not whether they ran it themselves.

Can I run the audit math myself?

For monitoring and for evidence of testing, yes, and you should: the four-fifths computation on your own funnel takes minutes with a browser-based tool and no data leaves your machine. For the official NYC audit, no: self-computed results do not satisfy the independence requirement no matter how correct the math is.

What if I have too few candidates for the statistics to mean anything?

Small samples are a real limitation: one candidate flipping outcome can swing a ratio across the 0.80 line. The LL144 framework lets audits exclude very small categories with justification and lean on pooled or test data where history is thin, and any honest self-check should annotate group sizes. For a small employer this cuts kindly: rely on the vendor's pooled audit for the official requirement, and treat your own small-sample checks as trend awareness rather than verdicts.

Related

Official sources

More guides