The AEDT Bias Audit Standard: Publish, Remediate, or Hold
You can grade a vendor's AEDT bias audit report against named criteria and return a publish, remediate, or hold verdict two graders would agree on.
This standard turns the bias-audit obligation for an AI sourcing or screening tool into a pass/fail grading rubric. It is for the compliance owner, recruiting operations lead, or revenue operations owner who has a vendor's audit report in hand and has to decide, this week, whether the tool is defensible enough to keep running. It gives you named criteria, an order of checks, and three verdicts - publish, remediate, or hold - stated so two graders reach the same call on the same document.
The pages that rank for this today describe the obligation. Law-firm summaries tell you an audit is required; auditor sales pages tell you they will provide one. Neither states what a complete report must contain, line by line, or how to grade the one you already have. That gap matters more than usual here, because the regulator's own detection is weak. Across the same 32 firms, the New York City agency responsible for enforcement flagged 1 issue while the State Comptroller found at least 17. The vendor's report, checked by you, is the real control.
What this standard grades, and what it does not
This standard grades a completed bias audit report and its published artifacts against fixed criteria, then returns publish, remediate, or hold. It does not perform the audit, certify the auditor, or give legal advice specific to your deployment.
The anchor here is New York City's Local Law 144 and the Department of Consumer and Worker Protection final rule that implements it, because that regime is the most detailed published definition of what a bias audit must contain. The federal Uniform Guidelines at 29 CFR Part 1607 supply the four-fifths rule that sits underneath it and that you should apply even for non-New-York use. Where a criterion comes from one and not the other, this guide says so.
Three facts set the stakes for why you grade the document yourself rather than trusting the regulator to catch problems.
The law is disclosure-first, not prohibition-first. A sub-0.80 impact ratio triggers closer review; nothing in Local Law 144 bans the tool outright. Liability shifts from "did you fail the numbers" to "did you disclose and act." That is why this rubric scores documentation as heavily as the metrics, and why a report with clean numbers but missing counts still fails.
The grading thresholds you check against
Every numeric criterion in this standard comes from the regulation, not from judgment. Memorize these five, because most defects are a failure to meet one of them.
| Criterion | Required value |
|---|---|
| Impact ratio concern threshold | below 0.80 |
| Category exclusion floor | under 2% of audit data |
| Candidate notice lead time | 10 business days |
| Summary retention after last use | 6 months |
| Penalty range per day/violation | $500 to $1,500 |
Read the impact ratio row precisely. The ratio is a category's selection rate divided by the selection rate of the most-selected category; for a scoring tool, it is the category's scoring rate divided by the highest scoring category's rate. A value below 0.80 flags potential adverse impact under the four-fifths rule. It does not, on its own, mean stop. It means the report must show closer review, and if it does not, that is a Hold.
The penalty row is why timing and completeness are not optional. Penalties run $500 to $1,500 per day, obtaining the audit and providing notice are separate requirements treated as separate violations, and each day a violation continues counts separately. Two missing obligations compound in parallel.
The metrics and categories a complete report must show
A complete report shows a selection rate and an impact ratio for every sex category, every race/ethnicity category, and every intersectional sex-by-race cell, computed from the EEO-1 component categories. Missing any one of those three axes is a defect.
The selection rate for a group is candidates selected divided by total candidates in that group. The impact ratio is that rate divided by the rate of the most-selected category. Do this for sex on its own, for race/ethnicity on its own, and for each combination of the two. The intersectional cells are where reports quietly fail, because a two-axis table can look clean while a specific sex-by-race combination sits well below 0.80.
How one impact ratio is built
- Count the groupSelected divided by total in that sex, race/ethnicity, or intersectional cell
- Find the top rateIdentify the most-selected category across the same axis
- DivideCategory rate divided by top-category rate gives the impact ratio
- CompareBelow 0.80 is a flag that must be paired with documented review
There is one more number the report must carry even when it is not a rate: the count of individuals the tool assessed who fall into an unknown category and are therefore not in the calculations. If that line is absent, you cannot tell whether the report understated adverse impact by dropping people, so its absence is a defect on its own.
The 2% exclusion and how it hides bias
The final rule lets an independent auditor exclude an EEO category that represents less than 2% of the audit data from the impact-ratio calculations. That allowance is the single most common silent failure point, because the easiest way to make a report look clean is to quietly drop the small groups most likely to show adverse impact.
The rule counters this, and your grading must enforce the counter. If a category is excluded, the summary must state the auditor's justification and must still report the number of applicants and the selection or scoring rate for the excluded category. A grader who accepts "excluded, too small" with no justification and no reported rate has passed a rigged report.
On statistical significance, the regulator set no fixed threshold. If the auditor determined there was insufficient historical data for a statistically significant audit, test data may be used, but the report must say so and explain why. This is also the answer to the four-fifths trap: passing 0.80 without any significance test is not full defensibility, because the federal guidelines say smaller differences can still be adverse impact where they are significant in both statistical and practical terms.
Auditor independence: the check that is not on the letter
Independence is a three-prong test, and the prong reports most often fake is the financial one. The auditor must not be employed by the employer or the vendor, must have no involvement in the tool's use, development, or distribution, and must have no direct or material indirect financial interest in either party.
The nuance to hold in mind: a vendor cannot audit its own tool itself, but the FAQ position is that a vendor can have an independent auditor audit the vendor's own tool, and the vendor can coordinate data collection for that third party. So "the vendor arranged it" is not disqualifying. "The vendor's data-science team ran it" is. And note that the regulator does not approve or list auditors, so a signed letter is the only artifact you get - which is exactly why you read it hard.
The pool of people who can actually do this work is thinner than the market implies. In Refolk's index, a search for explicit "AI Bias Auditor" or "Algorithmic Auditor" titles in the US returned zero matches - the named role barely exists. Only about 24 US professionals hold an explicit Responsible AI, AI Governance, or AI Ethics current title.
The practical consequence: you cannot assume a vendor's "auditor" is a specialist, and you cannot find one by searching for the title. You have to source by capability - adverse-impact analysis, four-fifths experience, disparate-impact metrics - which is a query about what someone has done, not what they are called.
Supply is also heavily concentrated in the US. In Refolk's index, the Responsible AI title pool runs about 3.4 times larger in the US than the UK, which means employers under copycat laws elsewhere face a thinner independent pool and a higher risk of the financial-entanglement failure above.
| Segment | US count | UK count | US:UK ratio |
|---|---|---|---|
| Responsible AI / AI Governance / AI Ethics title | 24 | 7 | 3.4x |
The grading procedure
Run these eight steps in order. Steps 1 and 2 are gates: if the tool is not in scope, you stop; if independence fails, the verdict is Hold regardless of the numbers. Steps 4 and 5 are where you catch the quiet failures, and they are the reason this is a grading job and not a reading job.
From vendor report to signed verdict
- Scope the toolConfirm the tool meets the AEDT definition - output is the sole criterion, is weighted more heavily than any other, or is used to overrule or modify a human conclusion. Produce a written yes/no coverage determination.
- Verify auditor independenceCheck the three prongs against the signed statement - no employment by employer or vendor, no involvement in the tool, no direct or material indirect financial interest. File the attestation and confirm ties are disclosed.
- Confirm the data basisConfirm whether historical or test data was used and that the reason is documented, with source and date range stated. Test data is only acceptable when historical data was insufficient for significance.
- Recompute the metricsIndependently recalculate selection or scoring rates and impact ratios for every sex, race/ethnicity, and intersectional cell using EEO-1 categories. Confirm your numbers match cell for cell.
- Check thresholds and exclusionsFlag every impact ratio below 0.80, verify any excluded category is under 2% with justification plus its own count and rate, and confirm the unknown-category count is present.
- Grade the summaryConfirm the published summary carries the audit date, distribution date, data source and explanation, unknown count, and all counts, rates, and ratios, posted clearly and conspicuously.
- Grade the candidate noticeConfirm 10-business-day timing, the tool disclosure, the qualifications assessed, and alternative-process and accommodation instructions.
- Return the verdictPublish if all criteria are met, Remediate for fixable gaps, Hold for an independence failure, missing intersectional analysis, a stale audit, or sub-0.80 ratios with no documented review. Sign it with rationale.
One honest note on ordering. Sources agree on the criteria but not on where recomputation sits. Auditor-led guides fold recomputation into the engagement itself. This rubric pulls it out as a separate reviewer step, because the whole point of grading is that someone other than the person who produced the number checks it. If your second reviewer's recomputation does not match, that mismatch is a Hold until reconciled.
The three verdicts, stated so two graders agree
The verdict is a function of which criteria pass, not a feeling about the tool. Use this fixed mapping so two people grading the same report land in the same place.
Deciding publish, remediate, or hold
Publish means every criterion is met: the three axes are present, exclusions are documented with their rates, the unknown count is stated, independence holds, the summary is complete and posted clearly and conspicuously, and the notice cleared all four elements. Set the next refresh date before you close the file.
Remediate means the gaps are fixable without redoing the analysis: a missing exclusion justification, a missing unknown-category count, a summary that omits the distribution date, a notice missing accommodation instructions, or an audit approaching but not past its anniversary. List each defect with the criterion it fails and send it back.
Hold means the tool should not keep running on this report. Triggers: an independence failure of any prong, a missing intersectional analysis, an audit older than one year, or any impact ratio below 0.80 with no documented closer review. A Hold is not permanent; it is a stop until the underlying work exists.
Grade the document, not your confidence in the vendor - the verdict is which criteria passed, nothing else.
How this goes wrong: the failure modes
Most bad grades come from passing a report that looks complete but hides a defect. These are the false positives to hunt, each with the tell that gives it away.
- Missing intersectional cells. The report shows sex and race/ethnicity but omits the sex-by-race combinations. The tell is a clean two-axis table that hides a failing cell. Check: count that every sex-by-race/ethnicity cell appears with its own rate and count.
- Silent unknown-category drop. Unknowns are excluded from calculations without stating the number, which understates adverse impact for smaller groups. Check: confirm an explicit unknown-count line exists.
- Undocumented 2% exclusion. A category is dropped as "too small" with no justification and no reported rate. Check: for each excluded category, confirm the justification plus applicant count and rate are all present.
- Independence theater. The auditor is nominally external but paid on a recurring renewal relationship or embedded in deployment. The tell is a signed independence letter with no disclosure of financial ties. Check: verify no direct or material indirect financial interest, not just the employment prong.
- Stale audit. The tool has been in use more than 12 months since the last audit. Check: compare audit date to today; older than one year is an automatic Hold.
- Notice served but incomplete. The 10-day timing is met but the notice omits the qualifications assessed or the alternative-process and accommodation instructions. The tell is a job-posting blurb that just says "we use AI." Check: confirm all four notice elements.
- Publish-then-forget retention. The summary is pulled the day the tool is retired. Check: confirm posting persists 6 months past last use.
- Four-fifths as a green light. Passing 0.80 is treated as full defensibility. Check: look for a statistical-significance test alongside the ratio, because smaller differences can still be adverse impact when statistically and practically significant.
The verification checklist
Run this before you sign any verdict. Every item is a checkable statement, not a topic. If any item fails, the verdict cannot be Publish.
Bias audit report grading checklist
- A written coverage determination confirms the tool meets the AEDT definition.
- The auditor's signed statement clears all three independence prongs, including disclosure of financial ties.
- The data source is stated as historical or test data, with a reason and a date range.
- Selection or scoring rates and impact ratios appear for every sex, race/ethnicity, and intersectional cell.
- Every impact ratio below 0.80 is paired with documented closer review, not just a number.
- Any category excluded as under 2% carries a justification plus its applicant count and rate.
- An explicit count of individuals in an unknown category is present.
- A statistical-significance test accompanies ratios that pass 0.80.
- The posted summary includes audit date, distribution date, data source and explanation, unknown count, and all rates and ratios.
- The candidate notice met the 10-business-day lead time and includes tool disclosure, qualifications assessed, and accommodation instructions.
- A retention plan keeps the summary posted 6 months past last use.
- The audit date is within the last 12 months.
Keeping the verdict current
A Publish verdict has a shelf life. The audit anchors to a date, the tool keeps making decisions, and the four-fifths math drifts as the applicant pool changes. Treat the verdict as valid until the earlier of the audit's one-year anniversary or a material change to the tool, and set that date the day you sign.
Use this to record the outcome so the next grader inherits your reasoning rather than reconstructing it.
Tool / vendor: Audit date / distribution date: Data basis: historical / test (reason if test: ) Independence: all three prongs cleared? Y/N (ties disclosed: ) Axes present: sex / race-ethnicity / intersectional (all Y/N: ) Ratios below 0.80: list cells + whether review documented Excluded categories (<2%): list + justification + count + rate present? Y/N Unknown-category count present? Y/N (value: ) Significance test present for passing ratios? Y/N Summary posted, clear and conspicuous, all fields? Y/N Notice: 10 business days + 4 elements? Y/N VERDICT: Publish / Remediate / Hold Defect list (for Remediate/Hold): Next refresh date: Graded by / date:
One per graded report. Fill every field; a blank field is itself a defect worth noting.
Two habits keep the standard honest over time. First, re-source your auditor by capability before each refresh, because the titled role still barely exists and a vendor's "auditor" may have changed. A search for people who have actually published AEDT audits, or data scientists with adverse-impact and four-fifths experience, tells you more than any title. Refolk is built for exactly that kind of capability query when the job title you need does not exist as a title.
Second, watch the regulatory floor rather than a single deadline. For non-New-York use, the federal four-fifths rule already applies, and its own relevance floor - adverse-impact determinations at least annually for any group that is at least 2% of the relevant labor force - is a useful default cadence even where no city law forces one. Grade against the mechanism, refresh on the calendar, and the verdict stays defensible.
Questions practitioners ask
What must an AEDT bias audit report contain to be complete?
A complete report contains selection or scoring rates and impact ratios for every sex category, every race/ethnicity category, and every intersectional sex-by-race cell using EEO-1 categories; the count of individuals in an unknown category; a justification plus count and rate for any category excluded as under 2%; the data source and explanation; and the audit and distribution dates. Anything missing from that list is a defect, not a stylistic choice.
Does passing the four-fifths rule mean the tool is defensible?
No. An impact ratio at or above 0.80 clears the primary screen but the federal Uniform Guidelines say smaller differences may still constitute adverse impact where they are significant in both statistical and practical terms. A report that stops at 0.80 without a statistical-significance test leaves a defensibility hole. Treat 0.80 as a threshold that opens review, not one that closes it.
Can a vendor audit its own AI hiring tool?
No, a vendor cannot audit its own tool itself, but the DCWP FAQ position is that a vendor can have an independent auditor perform the audit and can coordinate data collection for that third party. The auditor must not be employed by the employer or vendor, must have no involvement in the tool's use or development, and must have no direct or material indirect financial interest. A recurring audit-renewal relationship is a financial-tie risk to check.
How is the adverse impact ratio calculated for a hiring tool?
Compute each category's selection rate as candidates selected divided by total candidates in that group, then divide that rate by the selection rate of the most-selected category. For scoring tools, use the scoring rate for a category divided by the scoring rate of the highest scoring category. A result below 0.80 flags potential adverse impact. Do this separately for sex, race/ethnicity, and each intersectional cell.
When is a bias audit too old to keep using the tool?
Treat an audit older than one year as an automatic Hold. The rule contemplates an annual cadence, and a tool in continuous use more than 12 months since its last audit no longer has a current basis for the posted summary. Compare the audit date to today during grading and refresh before the anniversary rather than after enforcement contact.
What are the penalties for skipping the audit or the notice?
DCWP can impose civil penalties from $500 to $1,500 per day. The first default and each additional same-day violation run up to $500 and rise to up to $1,500 for subsequent defaults. Obtaining the bias audit and providing candidate notice are separate requirements treated as separate violations, and each day a violation continues counts separately, so gaps compound quickly.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.