Refolk
StandardProcess, data, and compliance

The AEDT Bias Audit Standard: Publish, Remediate, or Hold

You can grade a vendor's AEDT bias audit report against named criteria and return a publish, remediate, or hold verdict two graders would agree on.

17 min readLast reviewed August 29, 2026Read as Markdown

This standard turns the bias-audit obligation for an AI sourcing or screening tool into a pass/fail grading rubric. It is for the compliance owner, recruiting operations lead, or revenue operations owner who has a vendor's audit report in hand and has to decide, this week, whether the tool is defensible enough to keep running. It gives you named criteria, an order of checks, and three verdicts - publish, remediate, or hold - stated so two graders reach the same call on the same document.

The pages that rank for this today describe the obligation. Law-firm summaries tell you an audit is required; auditor sales pages tell you they will provide one. Neither states what a complete report must contain, line by line, or how to grade the one you already have. That gap matters more than usual here, because the regulator's own detection is weak. Across the same 32 firms, the New York City agency responsible for enforcement flagged 1 issue while the State Comptroller found at least 17. The vendor's report, checked by you, is the real control.

What this standard grades, and what it does not

This standard grades a completed bias audit report and its published artifacts against fixed criteria, then returns publish, remediate, or hold. It does not perform the audit, certify the auditor, or give legal advice specific to your deployment.

The anchor here is New York City's Local Law 144 and the Department of Consumer and Worker Protection final rule that implements it, because that regime is the most detailed published definition of what a bias audit must contain. The federal Uniform Guidelines at 29 CFR Part 1607 supply the four-fifths rule that sits underneath it and that you should apply even for non-New-York use. Where a criterion comes from one and not the other, this guide says so.

Three facts set the stakes for why you grade the document yourself rather than trusting the regulator to catch problems.

17x
Detection gap between the Comptroller and DCWP across the same 32 firms
DCWP found 1 issue; the State Comptroller found at least 17 in the 2025 enforcement audit covering July 2023 to June 2025.

The law is disclosure-first, not prohibition-first. A sub-0.80 impact ratio triggers closer review; nothing in Local Law 144 bans the tool outright. Liability shifts from "did you fail the numbers" to "did you disclose and act." That is why this rubric scores documentation as heavily as the metrics, and why a report with clean numbers but missing counts still fails.

The grading thresholds you check against

Every numeric criterion in this standard comes from the regulation, not from judgment. Memorize these five, because most defects are a failure to meet one of them.

CriterionRequired value
Impact ratio concern thresholdbelow 0.80
Category exclusion floorunder 2% of audit data
Candidate notice lead time10 business days
Summary retention after last use6 months
Penalty range per day/violation$500 to $1,500

Read the impact ratio row precisely. The ratio is a category's selection rate divided by the selection rate of the most-selected category; for a scoring tool, it is the category's scoring rate divided by the highest scoring category's rate. A value below 0.80 flags potential adverse impact under the four-fifths rule. It does not, on its own, mean stop. It means the report must show closer review, and if it does not, that is a Hold.

The penalty row is why timing and completeness are not optional. Penalties run $500 to $1,500 per day, obtaining the audit and providing notice are separate requirements treated as separate violations, and each day a violation continues counts separately. Two missing obligations compound in parallel.

The metrics and categories a complete report must show

A complete report shows a selection rate and an impact ratio for every sex category, every race/ethnicity category, and every intersectional sex-by-race cell, computed from the EEO-1 component categories. Missing any one of those three axes is a defect.

The selection rate for a group is candidates selected divided by total candidates in that group. The impact ratio is that rate divided by the rate of the most-selected category. Do this for sex on its own, for race/ethnicity on its own, and for each combination of the two. The intersectional cells are where reports quietly fail, because a two-axis table can look clean while a specific sex-by-race combination sits well below 0.80.

How one impact ratio is built

  1. Count the group
    Selected divided by total in that sex, race/ethnicity, or intersectional cell
  2. Find the top rate
    Identify the most-selected category across the same axis
  3. Divide
    Category rate divided by top-category rate gives the impact ratio
  4. Compare
    Below 0.80 is a flag that must be paired with documented review
Every cell in the report is one selection rate divided by the top category's rate, and you recompute each one.

There is one more number the report must carry even when it is not a rate: the count of individuals the tool assessed who fall into an unknown category and are therefore not in the calculations. If that line is absent, you cannot tell whether the report understated adverse impact by dropping people, so its absence is a defect on its own.

The 2% exclusion and how it hides bias

The final rule lets an independent auditor exclude an EEO category that represents less than 2% of the audit data from the impact-ratio calculations. That allowance is the single most common silent failure point, because the easiest way to make a report look clean is to quietly drop the small groups most likely to show adverse impact.

The rule counters this, and your grading must enforce the counter. If a category is excluded, the summary must state the auditor's justification and must still report the number of applicants and the selection or scoring rate for the excluded category. A grader who accepts "excluded, too small" with no justification and no reported rate has passed a rigged report.

On statistical significance, the regulator set no fixed threshold. If the auditor determined there was insufficient historical data for a statistically significant audit, test data may be used, but the report must say so and explain why. This is also the answer to the four-fifths trap: passing 0.80 without any significance test is not full defensibility, because the federal guidelines say smaller differences can still be adverse impact where they are significant in both statistical and practical terms.

Auditor independence: the check that is not on the letter

Independence is a three-prong test, and the prong reports most often fake is the financial one. The auditor must not be employed by the employer or the vendor, must have no involvement in the tool's use, development, or distribution, and must have no direct or material indirect financial interest in either party.

The nuance to hold in mind: a vendor cannot audit its own tool itself, but the FAQ position is that a vendor can have an independent auditor audit the vendor's own tool, and the vendor can coordinate data collection for that third party. So "the vendor arranged it" is not disqualifying. "The vendor's data-science team ran it" is. And note that the regulator does not approve or list auditors, so a signed letter is the only artifact you get - which is exactly why you read it hard.

The pool of people who can actually do this work is thinner than the market implies. In Refolk's index, a search for explicit "AI Bias Auditor" or "Algorithmic Auditor" titles in the US returned zero matches - the named role barely exists. Only about 24 US professionals hold an explicit Responsible AI, AI Governance, or AI Ethics current title.

24
US professionals with an explicit Responsible AI / AI Governance / AI Ethics current title
In Refolk's index; a search for explicit "AI Bias Auditor" titles returned zero. The skill lives inside data-science and legal teams, not a titled job.

The practical consequence: you cannot assume a vendor's "auditor" is a specialist, and you cannot find one by searching for the title. You have to source by capability - adverse-impact analysis, four-fifths experience, disparate-impact metrics - which is a query about what someone has done, not what they are called.

Supply is also heavily concentrated in the US. In Refolk's index, the Responsible AI title pool runs about 3.4 times larger in the US than the UK, which means employers under copycat laws elsewhere face a thinner independent pool and a higher risk of the financial-entanglement failure above.

SegmentUS countUK countUS:UK ratio
Responsible AI / AI Governance / AI Ethics title2473.4x

The grading procedure

Run these eight steps in order. Steps 1 and 2 are gates: if the tool is not in scope, you stop; if independence fails, the verdict is Hold regardless of the numbers. Steps 4 and 5 are where you catch the quiet failures, and they are the reason this is a grading job and not a reading job.

From vendor report to signed verdict

  1. Scope the tool
    Confirm the tool meets the AEDT definition - output is the sole criterion, is weighted more heavily than any other, or is used to overrule or modify a human conclusion. Produce a written yes/no coverage determination.
  2. Verify auditor independence
    Check the three prongs against the signed statement - no employment by employer or vendor, no involvement in the tool, no direct or material indirect financial interest. File the attestation and confirm ties are disclosed.
  3. Confirm the data basis
    Confirm whether historical or test data was used and that the reason is documented, with source and date range stated. Test data is only acceptable when historical data was insufficient for significance.
  4. Recompute the metrics
    Independently recalculate selection or scoring rates and impact ratios for every sex, race/ethnicity, and intersectional cell using EEO-1 categories. Confirm your numbers match cell for cell.
  5. Check thresholds and exclusions
    Flag every impact ratio below 0.80, verify any excluded category is under 2% with justification plus its own count and rate, and confirm the unknown-category count is present.
  6. Grade the summary
    Confirm the published summary carries the audit date, distribution date, data source and explanation, unknown count, and all counts, rates, and ratios, posted clearly and conspicuously.
  7. Grade the candidate notice
    Confirm 10-business-day timing, the tool disclosure, the qualifications assessed, and alternative-process and accommodation instructions.
  8. Return the verdict
    Publish if all criteria are met, Remediate for fixable gaps, Hold for an independence failure, missing intersectional analysis, a stale audit, or sub-0.80 ratios with no documented review. Sign it with rationale.

One honest note on ordering. Sources agree on the criteria but not on where recomputation sits. Auditor-led guides fold recomputation into the engagement itself. This rubric pulls it out as a separate reviewer step, because the whole point of grading is that someone other than the person who produced the number checks it. If your second reviewer's recomputation does not match, that mismatch is a Hold until reconciled.

The three verdicts, stated so two graders agree

The verdict is a function of which criteria pass, not a feeling about the tool. Use this fixed mapping so two people grading the same report land in the same place.

Deciding publish, remediate, or hold

Numbers fail or missingNumbers pass thresholds
Missing count or stale artifact with passing numbers
Remediate - fix the artifact and re-post
Sub-0.80 with no review, or missing intersectional axis
Hold - the analysis itself is incomplete
All criteria met, artifacts complete
Publish - post and set the refresh date
Independence failure or fabricated data basis
Hold - no amount of editing cures it
Gap is fixable in placeGap is structural
The severity of the gap and whether it is fixable in place decide the verdict.

Publish means every criterion is met: the three axes are present, exclusions are documented with their rates, the unknown count is stated, independence holds, the summary is complete and posted clearly and conspicuously, and the notice cleared all four elements. Set the next refresh date before you close the file.

Remediate means the gaps are fixable without redoing the analysis: a missing exclusion justification, a missing unknown-category count, a summary that omits the distribution date, a notice missing accommodation instructions, or an audit approaching but not past its anniversary. List each defect with the criterion it fails and send it back.

Hold means the tool should not keep running on this report. Triggers: an independence failure of any prong, a missing intersectional analysis, an audit older than one year, or any impact ratio below 0.80 with no documented closer review. A Hold is not permanent; it is a stop until the underlying work exists.

Grade the document, not your confidence in the vendor - the verdict is which criteria passed, nothing else.

How this goes wrong: the failure modes

Most bad grades come from passing a report that looks complete but hides a defect. These are the false positives to hunt, each with the tell that gives it away.

  • Missing intersectional cells. The report shows sex and race/ethnicity but omits the sex-by-race combinations. The tell is a clean two-axis table that hides a failing cell. Check: count that every sex-by-race/ethnicity cell appears with its own rate and count.
  • Silent unknown-category drop. Unknowns are excluded from calculations without stating the number, which understates adverse impact for smaller groups. Check: confirm an explicit unknown-count line exists.
  • Undocumented 2% exclusion. A category is dropped as "too small" with no justification and no reported rate. Check: for each excluded category, confirm the justification plus applicant count and rate are all present.
  • Independence theater. The auditor is nominally external but paid on a recurring renewal relationship or embedded in deployment. The tell is a signed independence letter with no disclosure of financial ties. Check: verify no direct or material indirect financial interest, not just the employment prong.
  • Stale audit. The tool has been in use more than 12 months since the last audit. Check: compare audit date to today; older than one year is an automatic Hold.
  • Notice served but incomplete. The 10-day timing is met but the notice omits the qualifications assessed or the alternative-process and accommodation instructions. The tell is a job-posting blurb that just says "we use AI." Check: confirm all four notice elements.
  • Publish-then-forget retention. The summary is pulled the day the tool is retired. Check: confirm posting persists 6 months past last use.
  • Four-fifths as a green light. Passing 0.80 is treated as full defensibility. Check: look for a statistical-significance test alongside the ratio, because smaller differences can still be adverse impact when statistically and practically significant.

The verification checklist

Run this before you sign any verdict. Every item is a checkable statement, not a topic. If any item fails, the verdict cannot be Publish.

Bias audit report grading checklist

  • A written coverage determination confirms the tool meets the AEDT definition.
  • The auditor's signed statement clears all three independence prongs, including disclosure of financial ties.
  • The data source is stated as historical or test data, with a reason and a date range.
  • Selection or scoring rates and impact ratios appear for every sex, race/ethnicity, and intersectional cell.
  • Every impact ratio below 0.80 is paired with documented closer review, not just a number.
  • Any category excluded as under 2% carries a justification plus its applicant count and rate.
  • An explicit count of individuals in an unknown category is present.
  • A statistical-significance test accompanies ratios that pass 0.80.
  • The posted summary includes audit date, distribution date, data source and explanation, unknown count, and all rates and ratios.
  • The candidate notice met the 10-business-day lead time and includes tool disclosure, qualifications assessed, and accommodation instructions.
  • A retention plan keeps the summary posted 6 months past last use.
  • The audit date is within the last 12 months.

Keeping the verdict current

A Publish verdict has a shelf life. The audit anchors to a date, the tool keeps making decisions, and the four-fifths math drifts as the applicant pool changes. Treat the verdict as valid until the earlier of the audit's one-year anniversary or a material change to the tool, and set that date the day you sign.

Use this to record the outcome so the next grader inherits your reasoning rather than reconstructing it.

Bias audit verdict record
Tool / vendor:
Audit date / distribution date:
Data basis: historical / test  (reason if test: )
Independence: all three prongs cleared? Y/N  (ties disclosed: )
Axes present: sex / race-ethnicity / intersectional  (all Y/N: )
Ratios below 0.80: list cells + whether review documented
Excluded categories (<2%): list + justification + count + rate present? Y/N
Unknown-category count present? Y/N  (value: )
Significance test present for passing ratios? Y/N
Summary posted, clear and conspicuous, all fields? Y/N
Notice: 10 business days + 4 elements? Y/N
VERDICT: Publish / Remediate / Hold
Defect list (for Remediate/Hold):
Next refresh date:
Graded by / date:

One per graded report. Fill every field; a blank field is itself a defect worth noting.

Two habits keep the standard honest over time. First, re-source your auditor by capability before each refresh, because the titled role still barely exists and a vendor's "auditor" may have changed. A search for people who have actually published AEDT audits, or data scientists with adverse-impact and four-fifths experience, tells you more than any title. Refolk is built for exactly that kind of capability query when the job title you need does not exist as a title.

Second, watch the regulatory floor rather than a single deadline. For non-New-York use, the federal four-fifths rule already applies, and its own relevance floor - adverse-impact determinations at least annually for any group that is at least 2% of the relevant labor force - is a useful default cadence even where no city law forces one. Grade against the mechanism, refresh on the calendar, and the verdict stays defensible.

Questions practitioners ask

What must an AEDT bias audit report contain to be complete?

A complete report contains selection or scoring rates and impact ratios for every sex category, every race/ethnicity category, and every intersectional sex-by-race cell using EEO-1 categories; the count of individuals in an unknown category; a justification plus count and rate for any category excluded as under 2%; the data source and explanation; and the audit and distribution dates. Anything missing from that list is a defect, not a stylistic choice.

Does passing the four-fifths rule mean the tool is defensible?

No. An impact ratio at or above 0.80 clears the primary screen but the federal Uniform Guidelines say smaller differences may still constitute adverse impact where they are significant in both statistical and practical terms. A report that stops at 0.80 without a statistical-significance test leaves a defensibility hole. Treat 0.80 as a threshold that opens review, not one that closes it.

Can a vendor audit its own AI hiring tool?

No, a vendor cannot audit its own tool itself, but the DCWP FAQ position is that a vendor can have an independent auditor perform the audit and can coordinate data collection for that third party. The auditor must not be employed by the employer or vendor, must have no involvement in the tool's use or development, and must have no direct or material indirect financial interest. A recurring audit-renewal relationship is a financial-tie risk to check.

How is the adverse impact ratio calculated for a hiring tool?

Compute each category's selection rate as candidates selected divided by total candidates in that group, then divide that rate by the selection rate of the most-selected category. For scoring tools, use the scoring rate for a category divided by the scoring rate of the highest scoring category. A result below 0.80 flags potential adverse impact. Do this separately for sex, race/ethnicity, and each intersectional cell.

When is a bias audit too old to keep using the tool?

Treat an audit older than one year as an automatic Hold. The rule contemplates an annual cadence, and a tool in continuous use more than 12 months since its last audit no longer has a current basis for the posted summary. Compare the audit date to today during grading and refresh before the anniversary rather than after enforcement contact.

What are the penalties for skipping the audit or the notice?

DCWP can impose civil penalties from $500 to $1,500 per day. The first default and each additional same-day violation run up to $500 and rise to up to $1,500 for subsequent defaults. Obtaining the bias audit and providing candidate notice are separate requirements treated as separate violations, and each day a violation continues counts separately, so gaps compound quickly.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next