AI in HR

Screening 10,000 Applicants: What We Learned

Our head of talent shares the reality of AI-assisted hiring at scale, including the bias traps we almost fell into.

January 6, 2026
6 min read
Nina Patel

The Scale Problem

When you're a growing company that posts on a few popular job boards, you get a lot of applications. A lot. Our engineering roles averaged 800 applicants. Our marketing roles hit 1,200. Our customer success posting got 2,400 applications in two weeks.

We had three recruiters. The math doesn't work. At five minutes per application (a generous average for a thoughtful review), screening 2,400 applicants takes 200 hours. That's five weeks of full-time work for one person, for one role.

Something had to give. Usually what gives is quality—recruiters start scanning for keywords, prestigious company names, and recognizable schools. This is fast but terrible. It systematically disadvantages non-traditional candidates who are often exactly the kind of creative, resilient people you want on your team.

Setting Up AI Screening

We built our screening pipeline using Billix to connect our ATS with our evaluation criteria. Instead of keyword matching, we trained the system to evaluate applications against structured rubrics.

For our engineering roles, the rubric included: evidence of problem-solving (from any context, not just tech), communication clarity, relevant technical skills (weighted by proficiency, not years), and project complexity. Notice what's not on that list: school name, years of experience, current employer prestige.

The AI reads each application and scores it against the rubric, providing specific evidence for each score. A recruiter then reviews the top-scored applications and the AI's reasoning. This catches both great candidates the AI might have missed and good reasoning the recruiter might not have considered.

The rubric for our engineering roles looked like this:

  1. Problem-solving evidence (any context, not just tech) — 30%
  2. Communication clarity — 20%
  3. Technical skill relevance (weighted by proficiency, not years) — 25%
  4. Project complexity — 25%

What's deliberately absent: school name, years of experience, employer prestige, keyword density. We wanted to evaluate what people did, not where they were.

Bias Detection

This is where things get uncomfortable. Our first version of the screener had a bias problem. It scored candidates who described their experience using confident, assertive language higher than equally qualified candidates who used more modest phrasing. This is a known issue in NLP models—they inherit the biases present in their training data, which overrepresents certain communication styles.

We caught this by running a bias audit: we anonymized a set of applications, had the AI score them, then checked whether scores correlated with demographic proxies. They did, slightly but measurably.

The fix wasn't simple. We couldn't just "debias" the model. Instead, we restructured the evaluation to focus on specific, verifiable claims rather than overall impressions. "Built a system that processed 1M daily transactions" is evaluated the same regardless of whether it's stated confidently or modestly. This reduced the bias significantly, though I won't claim we eliminated it entirely.

What Surprised Us

Three things we didn't expect:

Career changers scored well. The rubric didn't penalize non-traditional backgrounds, and the AI evaluated transferable skills fairly. A former teacher who'd learned to code and built classroom management tools scored higher than some candidates with five years of industry experience but generic project descriptions. That felt right.

The AI was better at consistency than humans. When we had recruiters and the AI both score the same batch of 200 applications, the AI's scores were far more internally consistent. Human scores varied significantly based on time of day, which application they reviewed right before, and how many they'd already reviewed that session.

Candidates appreciated the transparency. We told applicants that AI was part of our screening process. Some were uneasy, but most appreciated that we were trying to reduce bias in hiring. Several candidates specifically mentioned it as a reason they applied.

Recommendations for HR Teams

If you're considering AI-assisted screening:

  • Define your rubric before deploying AI. If you can't articulate what you're looking for in structured terms, the AI will invent its own criteria—and they might not be good ones.
  • Audit for bias regularly. Not once. Regularly. Bias can creep in as your candidate pool changes.
  • Keep humans in the loop. AI should narrow the funnel, not make final decisions. Every candidate who reaches the interview stage should have been reviewed by a person.
  • Be transparent. Tell candidates AI is involved. The ones who object probably aren't the right fit for a company that values innovation anyway.

Let's Make Life Easier

Try Billix for
free right now

Start for free

Enterprise-
Grade Security

Effortless
For Everyone

Automation
Made Natural

Unified
Workflow Sync

Start for free, right now