AI resume screening: how it works, where it fails, and how to use it safely

AI resume screening uses a language model to read applications and rate them against the requirements for a role, instead of matching keywords. The practical difference is that it can read 'ran the month-end close and cut days-to-close from nine to four' as strong accounting evidence even though the resume never uses the word accounting, and it can tell that a candidate listing a tool in a skills block has not actually shown using it.

The gain is real and so is the risk, and they have the same source. A system flexible enough to recognise experience described in unfamiliar words is also flexible enough to pick up patterns nobody chose and nobody can see. That is why US regulators moved on this category first among AI applications: New York City requires an annual bias audit and candidate notice, Illinois HB 3773 took effect in 2026, and Colorado's AI Act treats hiring as high risk. In every one of them the employer carries the duty, not the vendor.

What separates a defensible setup from a risky one turns out not to be the model. It is three design choices: whether a human wrote the criteria, whether every score carries the evidence it came from, and whether the tool is permitted to reject anyone on its own. A tool that ranks and explains, with a person deciding, is a very different legal object from one that filters people out before anyone looks.

The four approaches, and how each one fails

Keyword and boolean matching

Suits: Hard filters where the requirement is a fact: a licence number, a certification, the right to work.

Trade-off: Not AI at all, and now actively misleading. Applicants using language models produce near-perfect keyword overlap with your job post, so the filter passes the best-optimised resumes rather than the best candidates.

A model trained on your past hiring decisions

Suits: Very high volume in one repeated role, where there is enough history to learn from.

Trade-off: It learns your history including the parts you would not defend, and the pattern is invisible once it is inside the weights. This is the setup regulators examine hardest, and the one where a bias audit is most likely to find something.

A language model scoring against a rubric a human wrote

Suits: Teams that can state what matters for the role and want the reasoning visible per candidate.

Trade-off: Someone has to write the rubric, and a vague rubric produces vague scores. It costs per candidate rather than per seat, and it will not tell you anything you did not ask about.

AI-written-resume detectors used as a filter

Suits: Nothing, as a filter. As background context for a human reader, treated sceptically, they have some use.

Trade-off: Detection accuracy on short professional text is poor, and the false positives concentrate on non-native English writers and anyone who used a grammar tool. Rejecting on this signal creates exactly the disparate-impact exposure the rest of your process is trying to avoid.

The questions that decide whether your setup is defensible

Did a person write the criteria, and can you produce them?

Human-set criteria are both a quality control and the artefact you show if anyone asks what the tool was judging. If the criteria live inside a vendor's model, you cannot produce them, and you are answering for a decision process you never saw.

Does every score cite the specific line it came from?

Evidence is what turns a score into something a manager can disagree with. Without it there is no way to distinguish a model that read carefully from one that pattern-matched a job title, and no way to correct it when it is wrong.

Can the tool remove a candidate without a person seeing them?

This is the line that matters most in NYC LL144 and in the state laws that followed. A tool that ranks leaves the decision with a human; a tool that filters makes an employment decision automatically, with all the duties that attach to it. Keep the tool on the ranking side of that line.

Can you reconstruct a past screening months later?

A bias audit and a discrimination claim both work backwards from records. If your criteria, scores and evidence are not retrievable for a screening you ran in March, you will be reconstructing your defence from memory.

Does the vendor train on your candidate data?

Resumes are personal data belonging to people who applied for a job, not material you are free to donate. Get the answer in writing, and note that 'we may use data to improve our services' is not a no.

Have the candidates been told?

NYC requires notice before an automated tool is used, with specific timing and content, and other jurisdictions are converging on the same expectation. Notice is cheap to send and expensive to have skipped, and a template costs nothing.

Where SiftFirst fits

SiftFirst is the third approach on that list, built deliberately on the ranking side of the automatic-decision line.

  • Scores each candidate against a rubric you wrote and can edit, criterion by criterion
  • Quotes the line from the resume that produced each score, so the reasoning is auditable rather than asserted
  • Keeps the criteria, the weights and the evidence together, which is the record a bias audit or a challenge works from
  • Lets you change weights and watch the ranking move, so a rubric that is scoring the wrong thing surfaces immediately
  • Runs on your own pile free without a signup, which is the only honest way to judge screening quality

What it does not do

  • Reject, filter, or auto-advance any candidate at any score
  • Train on your candidate data
  • Claim to identify AI-written resumes reliably, or treat any such signal as a reason to score someone down

See the output on your own applications

Paste your job description and the applications, get a ranked shortlist with a quote from each resume behind every score. You set the criteria; you decide every hire. Free to try, no signup.

Try a free screening →

Frequently asked questions

Is AI resume screening legal in the United States?

Yes, subject to rules that depend on where the job is. New York City requires an annual independent bias audit and candidate notice for automated employment decision tools. Illinois HB 3773 and the Colorado AI Act impose their own duties. Federal discrimination law applies throughout, and the EEOC has been explicit that using a vendor's tool does not move responsibility to the vendor. The compliance check on this site works out which of these reach you.

Is AI resume screening biased?

It can be, and the mechanism depends on the approach. A model trained on past hiring decisions reproduces whatever those decisions encoded. A model scoring against explicit human-written criteria is bounded by those criteria, which makes the failure mode visible: a biased rubric produces biased scores, and you can read the rubric. Neither approach is automatically safe, which is why the audit duty exists.

Does human review make an AI hiring tool legally safe?

No, and this is a common and expensive misreading. Mobley v. Workday established that a human in the loop is not a shield when the tool has already shaped who the human sees. What actually helps is narrower: the tool never removes anyone, the criteria are yours and recorded, every score is evidenced, and candidates were notified. Those are auditable facts rather than a posture.

Can AI tell if a resume was written by AI?

Not reliably, on text this short and this formulaic. Detection tools produce false positives that fall hardest on people writing in a second language and anyone who ran their draft through a grammar checker, which makes acting on the signal a discrimination risk rather than a screening improvement. The better response to AI-written applications is to score against specific evidence of having done the work, which is harder to fabricate convincingly than vocabulary.

How accurate is AI resume screening?

Accuracy is the wrong frame, because there is no ground truth about who would have been the best hire. The useful questions are whether the tool is consistent across candidates, whether its reasoning survives a manager reading it, and whether it surfaces people a keyword filter would have dropped. All three are things you can check on your own applications in an afternoon, which is worth more than any vendor's accuracy figure.

Keep reading