Skip to content

The AI Bias Audit Industry Has a Fake Candidate Problem

Happy Friday Job Board Doctor friends!

If you haven’t checked it out yet: Maki People answered every question I sent them, on the record and in full, and that exchange is on the site. Last week, I wrote about disability bias in AI assessments, and I hope you will give it a read.

This week, I wrap up my thoughts on the Maki People launch with Recruitics by sharing what I have learned about AI bias audit testing at the vendor level, which, to be fully transparent, was a lot.

What Maki People has done this month is open my eyes. They are the case study, not the scandal. The scandal is industry-wide. Buyers must be aware of it, and must require more from their vendors and from their own data sharing practices.

Here’s what we’re covering this week:

Sponsor content

Talroo brings clarity to improve new-hire retention

You filled the seat, but not with someone who plans to stay. The candidate made it through your entire interview process while quietly interviewing with other companies that pay 25% more.

When expectations aren't clear up front, retention suffers.

Get clarity up front with Talroo's SmartQualify.

Get a demo today

The audits running on invented candidates

After taking a deeper dive into bias auditing practices, I realized I was no longer writing about one vendor. I was looking at an industry practice that creates greater peril for all of us and a greater likelihood of bias against our candidates.

The practice? Bias audits built on candidates who do not exist at all.

New York’s Local Law 144 (NY LL144) is an easy example we can use to talk through current AI tech vendor bias auditing practices.

NY LL144 requires an annual bias audit for automated hiring tools, known in the law as automated employment decision tools (AEDTs), and it permits two kinds of data:

    •   Historical data, meaning real candidates, or
    •   Test data, when historical data is insufficient.

While “test data” sounds procedural, it means the auditor never evaluated a single human interaction with the technology.

An ACLU-led research team published a comprehensive study at FAccT 2025 analyzing every publicly available audit from the law’s first sixteen months, and roughly one in six audit reports ran on test data.

Their dataset names the cases, all covering audits published through late 2024:

    •   Landed’s 2024 audit is the one the researchers cite for generating fake applicant scores with a random-number function.
    •   HackerRank’s image analysis screening system was audited two years running on test data (2023 audit, 2024 audit) that, in 2024, produced identical selection rates for every demographic group, a result no real applicant pool has ever produced.
    •   Covey used test data two years straight, under two different auditors (2023 audit, 2024 audit).

For the worst of these, the researchers expressed serious concern whether the results may have no relationship to the tool’s real-world impact.

 Everyone gets a gold star

Here is what the more sophisticated version of test data looks like, and why it guarantees the outcome:

    •   Thousands of interviews generated by an AI model.
    •   Every demographic group is the exact same size.
    •   Every group assigned the exact same distribution of response quality, by instruction.
    •   Demographics reduced to a single tidy phrase inserted into the transcript.

Now remember what the audit is measuring: whether selection rates differ across groups. Adverse impact in the real world comes from two places. It comes from small groups, where a handful of decisions swings a rate, and it comes from the real-world tangle between demographics and everything else: education access, first language, dialect, disability, nerves, circumstance. A perfectly balanced synthetic pool deletes both.

Equal group sizes eliminate the small-sample problem. Identical quality distributions sever every correlation that produces disparate outcomes in actual hiring. The design does not just make the exam easy. It removes, by construction, the two mechanisms through which this exam can be failed.

Let me be clear: nobody cheated. The methodology is disclosed, the auditor is independent, the math is correct. The exam itself is the cheat code. A pass was close to guaranteed the moment the dataset was designed, and a gold star certifies nothing.

This is not fraud. It is something a buyer should treat with the same weight: a result that is legal, disclosed, and, in my opinion, not worth the paper it is printed on.

THE BOTTOM LINE

A perfectly balanced synthetic audit deletes the two conditions under which bias tests fail: small groups and real-world correlation. It cannot flunk, so it cannot certify. Treat it as no evidence about your applicant pool, and say so in procurement, because the law permits it and only market pressure will end it.

If you must simulate, simulate reality

Here is the part that removes the last excuse, and it comes from data this industry already uses every day. The demographic composition of the actual labor market is not a mystery.

The Census Bureau publishes its EEO Tabulation for precisely this purpose: labor force availability by race, ethnicity, and sex, broken down by occupation and geography. Federal contractors are already required to use this data for the utilization and availability analyses in their affirmative action programs. BLS publishes occupational demographics on top of it.

The hiring industry benchmarks workforces against real community composition as a routine compliance exercise, right next door to the audit teams building candidate pools at a flat 50/50.

So when a vendor cannot use production data, a more realistic and more defensible test dataset is still available, and it looks nothing like perfect balance. It mirrors the actual availability of the occupations the tool screens: unbalanced groups, genuinely small minority populations, composition weighted to the labor markets where the tool is deployed.

That design puts the small-sample problem back into the exam, which is exactly the point, because that is where real audits find real trouble.

It still cannot replicate the correlation between demographics and lived circumstance, which is why production data must remain the standard. But it separates a good-faith simulation from a convenient one.

A vendor that reaches for utilization data built a test. A vendor that reaches for equal cells built a pass.

Ask the question that way and the perfectly balanced pool stops looking like a methodology choice and starts looking like what it is: the easiest version of the exam, selected for the party taking it.

“Our contracts won’t let us”

The standard explanation for skipping production data is that customer data agreements prohibit it. Sometimes that is partially true; enterprise contracts do carry purpose limitations. But the claim is unverifiable from the outside, so test it against the vendor’s own publicly available documentation.

Read their standard terms: in the case that started this series, the public terms restricted the clients, not the vendor. Read their privacy policy: if it discloses fairness testing on candidate data as the vendor’s own purpose, the data was available and the explanation collapses.

And watch for consistency: the same vendor told me insufficient data explained the missing disability testing and contractual prohibition explained the synthetic audit, while confirming it runs validation studies on production data for customers’ own purposes. The machinery exists. If contracts bar production data from audits, sample size was never the constraint, at any size.

A law that permits synthetic data gets synthetic audits. A test designed around the tested gets passed. A benchmark sitting in public Census data goes unused because nothing requires reaching for it.

FOR YOUR NEXT VENDOR CALL

Five questions that separate a certificate from an answer.

  1. Did your bias audit use real production data?
  2. If not, when will you begin using historical data, which we believe is more reflective of a candidate’s true experience with your technology?
  3. If test data was used, was it weighted to real labor-market availability, like the EEO utilization data, or was it artificially balanced?
  4. What are the real-time auditing data points available to me as the buyer? If real-time data is unavailable, what is the data refresh period?
  5. What are the options for contract termination if we find that your product results in biased outcomes?

At the end of the day, the industry needs vendors tested on real-world data, built from an appropriate mix of client data and enhanced to fill gaps, when needed, down to the job level.

What does it take to get there? It starts with a clear understanding: when a vendor pats itself on the back for passing an audit, ask what data earned the gold star.

The full exchange that started this, with the vendor’s answers verbatim, is here: The Full Maki People Q&A: Nine Questions on the Bias Audit, Disability Testing, and What Mochi Actually Measures

The tip line is always open.

Comments (0)

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Back To Top
Search