# The Candidate-Scoring Tool Clearance Standard: Deploy, Restrict, or Hold

*You can grade any candidate-scoring tool against a deployment and return a defensible deploy, restrict, or hold verdict with the exact audit, notice, and review steps each jurisdiction requires.*

- Canonical URL: https://www.refolk.ai/guides/candidate-scoring-tool-clearance-standard
- Pillar: Process, data, and compliance
- Format: Standard
- Published: 2026-10-06
- Last reviewed: 2026-10-06
- Reading time: 15 min
- Keywords: is our sourcing tool an AEDT, bias audit required candidate ranking tool, AI hiring tool compliance checklist, local law 144 sourcing tool scope, human review AI hiring decision requirement

## Key takeaways

- NYC Local Law 144 puts a scoring tool in scope when its output is the highest-weighted criterion or overrules a human conclusion, so adding light-touch human review can increase exposure, not reduce it.
- Each day of unaudited AEDT use in New York City is a separate violation of up to $1,500, which makes holding a tool for a two-month audit far cheaper than running it non-compliant.
- A learned or data-tuned ranker is in scope even when a vendor brands it as a parser; scope turns on whether a computer identifies the inputs and their importance, not on the product name.
- The EU date split is a two-way trap: Article 50 transparency binds from August 2, 2026, while Annex III recruitment obligations moved to December 2, 2027 under the AI Omnibus.
- A Fundamental Rights Impact Assessment is not automatic for private recruiters; Article 27 reaches only public bodies, public-service providers, and credit or insurance deployers.
- In Refolk's index, 407 US profiles carry recruiting-operations titles versus 14 in the UK, a 29.1x gap, and none pair that title with explicit AI-bias-audit language.

Before a recruiting operation turns on an AI tool that scores or ranks sourced candidates for a role tied to a specific place, someone has to decide whether that tool is cleared to run. This standard is for recruiting operations, revenue operations, and anyone answerable for how the data was gathered. It gives you a repeatable pass/fail rubric that spans New York City, Illinois, Colorado, and the EU in one decision, so two reviewers grade the same tool the same way and return the same verdict: deploy, restrict, or hold.

The library already covers lawful sourcing by jurisdiction, legitimate-interest, retention clocks, and first-contact notice. None of those grade the automated ranking tool itself against the threshold that triggers bias-audit and notice duties. That is the gap this document fills.

## What this standard grades, and what a verdict means

This standard grades one thing: whether a specific candidate-scoring or ranking tool, used for a specific role and location, clears the legal duties that attach before first use. The output is one of three verdicts.

- **Deploy.** The tool is either out of scope, or in scope with every triggered duty already met: a current audit, matched notices, a named human reviewer, and configured logging.
- **Restrict.** The tool is in scope and usable, but only after you close named gaps. You can run it for roles or markets where the duty is already satisfied, and you hold it everywhere else until the gap is closed.
- **Hold.** The tool is in scope and a hard duty is unmet - no current bias audit, no candidate notice, no override authority. Running it now accrues liability each day.

A verdict is defensible when it names the prong that put the tool in scope, the jurisdiction whose duty is unmet, and the evidence on file. If two reviewers reach different verdicts, the disagreement is almost always about the AEDT threshold test, so that test comes first.

> **Rule:** Grade the deployment, not the tool
>
> A single tool can be deploy for a remote role, restrict for a New York City role, and hold for an EU candidate on the same day. Always scope a verdict to one role-and-location deployment.

## Is the tool in scope? The AEDT threshold test

A candidate-scoring tool is in scope as an Automated Employment Decision Tool when its machine-learned output substantially assists or replaces the hiring decision. The DCWP Final Rule draws that line with three prongs and one machine-learning clause.

The three prongs. A tool substantially assists or replaces discretionary decision-making when you do any one of these:

1. Rely solely on a simplified output - a score, tag, classification, or ranking - with no other factors considered.
2. Use a simplified output as one of a set of criteria where that output is weighted more than any other criterion in the set.
3. Use a simplified output to overrule conclusions derived from other factors, including human decision-making.

The machine-learning clause. The computational process must derive from machine learning, statistical modeling, data analytics, or AI for which a computer identifies the inputs, the relative importance placed on those inputs, and other parameters, in order to improve accuracy. This targets tools that learn or are tuned by data. A pure keyword resume parser with human-authored rules falls outside. A learned ranker whose score outweighs other criteria falls inside.

#### The AEDT threshold test

1. **Machine-learning clause** - Does a computer identify inputs and their importance to improve accuracy? If no, out of scope.
2. **Prong 1 - sole reliance** - Is the output used alone with no other factors?
3. **Prong 2 - highest weight** - Is the output weighted higher than any other criterion?
4. **Prong 3 - override** - Does the output overrule human conclusions?
5. **Verdict** - Clears the ML clause and any one prong, so it is in scope.

*A tool is in scope only when it clears the machine-learning clause and at least one of the three use prongs.*

The insight that trips people is prong three. The third prong puts a tool *in* scope precisely when its output overrules human conclusions, so adding light-touch human review can increase exposure if the reviewer rubber-stamps the score. Scope turns on weighting, not branding. A vendor re-labeling a learned ranker as a "parser" does not change the grade - check whether a computer identifies the parameters, not what the product is called.

**0.80 - Impact-ratio threshold under the four-fifths rule**

A selection or scoring rate below 0.80 relative to the most-selected group signals potential adverse impact in a Local Law 144 bias audit.

## The four-regime clearance matrix

Four regimes can attach to one deployment, each with its own trigger, date, notice, human-review posture, and retention rule. Grade the deployment against every regime whose candidates or jobs it touches.

| Regime | Scope trigger | Effective date | Pre-use notice | Human review | Retention |
| --- | --- | --- | --- | --- | --- |
| NYC LL144 | AEDT weighting/override test | July 5, 2023 | 10 business days | Not a safe harbor | Audit <12 months |
| Illinois HB 3773 | Discriminatory effect + ZIP proxy | Jan 1, 2026 | Notice required (rules pending) | No exemption | Limitation period |
| Colorado SB 26-189 | ADMT in consequential decision | Jan 1, 2027 | Clear/conspicuous pre-use | Meaningful, commercially reasonable | 3 years |
| EU AI Act Annex III | High-risk recruitment system | Dec 2, 2027 | Inform affected persons/workers | Competent oversight (Art 26(2)) | Logs ≥6 months |

Read the matrix as a set of independent gates, not a menu. New York City turns on the weighting and override test - a human being present does not exempt you. Illinois makes it a civil rights violation to use AI that discriminates on protected classes, to use ZIP codes as a proxy, or to fail to tell employees that AI is in use; it applies to employers with one or more employees in Illinois. Colorado's rewrite, SB 26-189, requires a clear and conspicuous pre-use notice, a 30-day adverse-outcome explanation, and meaningful human review to the extent commercially reasonable. The EU treats recruitment and ranking as an Annex III high-risk use case with deployer duties under Article 26.

> **Watch out:** The EU date split cuts both ways
>
> Article 50 transparency duties bind from August 2, 2026, but the Annex III high-risk recruitment obligations moved to December 2, 2027 under the AI Omnibus, Regulation 2026/1744. A single EU AI Act deadline line in your policy will be wrong for at least one duty.

## The NYC bias-audit math you must verify

The Local Law 144 bias audit is the single most common point of failure, so grade its math directly rather than accepting that an audit "exists." An independent auditor must evaluate potential disparate impact by sex, race, and ethnicity category, calculating selection and scoring rates and their corresponding impact ratios.

| Item | Value |
| --- | --- |
| Impact-ratio threshold | 0.80 |
| Small-category exclusion | <2% of data |
| Candidate notice | 10 business days |
| Max per-day penalty | $1,500 |

The impact ratio is the selection or scoring rate for each group divided by the rate for the most-selected group. It mirrors the EEOC four-fifths rule, so a ratio below 0.80 signals potential adverse impact. Categories representing less than 2% of the audit data may be excluded at the auditor's discretion. Candidates must be notified at least 10 business days in advance of use. Penalties run up to $500 for a first violation, up to $500 for same-day additional violations, and $500 to $1,500 for subsequent ones, with each day of unaudited use a separate violation.

That daily accrual changes the economics of the whole decision. Because each day of unaudited use is a separate violation of up to $1,500, a two-month delay to run the audit is far cheaper than two months of non-compliant running. Hold is often the financially conservative verdict, not the cautious one.

> Daily-accrual penalties make hold cheaper than deploy-and-fix, so holding a tool is the conservative financial choice.

## Grade a tool against a deployment: the procedure

Work the steps in order, with one caveat on evidence. Sources disagree on whether to procure vendor evidence before commissioning the audit or after: some checklists commission the audit first and treat vendor documentation as confirmation, others gate everything on vendor documentation. Either sequence works as long as both are complete before you grade.

#### From tool to signed verdict

1. **Inventory and scope the tool** - Record name, vendor, version, and whether it scores or ranks for a role tied to NYC, Illinois, Colorado, or an EU candidate. Done when a one-line scope statement exists per deployment.
2. **Run the AEDT threshold test** - Apply the three prongs and the ML clause with legal. Done when an in-scope or out-of-scope verdict names the triggering prong.
3. **Procure provider evidence** - Request conformity or CE evidence, instructions for use, the latest independent bias audit, and a training-data representativeness statement. Done when the evidence pack is on file.
4. **Commission or verify the bias audit** - Confirm selection and scoring rates and impact ratios by EEO-1 category, dated within 12 months, on data that reflects your pool. Done when a signed audit exists and the summary is posted.
5. **Build jurisdiction-matched notices** - Draft the NYC 10-business-day notice, the Illinois AI-use notice, and the Colorado pre-use and 30-day adverse-outcome notices. Done when matched copy sits in the ATS.
6. **Stand up human review and appeal** - Name a reviewer with authority to override and build a correction-and-reconsideration path. Done when an escalation workflow has a named owner.
7. **Configure logging and retention** - Set EU logs to at least six months and Colorado records to three years, per system and access-controlled. Done when retention is on and scoped to the tool.
8. **Grade and decide** - Score against the rubric and return deploy, restrict, or hold. Done when a signed verdict lists the gaps forcing restrict.

The person who should own this procedure barely exists as a title. In Refolk's index, 407 US profiles carry recruiting-operations titles against 14 in the UK, and a search for recruiting-ops or people-analytics titles paired with explicit AI-bias-audit compliance language returns zero profiles. The person answerable for this clearance is usually untitled, which is why clearances slip. Find them before the tool goes live.

Ask me this: `Recruiting operations managers at US tech companies who have run or commissioned a Local Law 144 bias audit.` - [run the search](https://www.refolk.ai/start?q=Recruiting%20operations%20managers%20at%20US%20tech%20companies%20who%20have%20run%20or%20commissioned%20a%20Local%20Law%20144%20bias%20audit.).

*Returns the named recruiting-ops owners who have actually carried a bias audit, so you can staff the reviewer role from people who have done it before.*

| Market | Titled recruiting-ops profiles | Top employers |
| --- | --- | --- |
| United States | 407 | Meta, Plaid, Box |
| United Kingdom | 14 | Anthropic, Monzo, Chainalysis |
| Derived: US-to-UK ratio | 29.1x | - |

When you need to find and brief the untitled owner fast, [Refolk](/) resolves a plain-English description of the role into named people across public LinkedIn and the open web, which is the friction this standard otherwise leaves you to solve by hand.

## How this goes wrong: failure modes and false positives

Most bad verdicts come from a short list of predictable errors. Each one is a false negative that lets a non-compliant tool run, or a false positive that wastes effort on a duty you do not owe. Check the specific thing named, not the reassuring summary.

| Failure mode | What it looks like | What to check instead |
| --- | --- | --- |
| Human-in-the-loop defense | "A person reviews every score, so we are out of scope" | The actual decision weighting - if the score is highest-weighted or overrules the human, it is in scope |
| "It is just a parser" | A learned ranker branded as resume parsing | Whether a computer identifies inputs and their importance |
| Stale audit | An audit on file but older than 12 months | The audit date, since each day of use past 12 months is a fresh violation |
| Vendor audit as yours | A generic vendor audit covering all customers | The data source and representativeness against your pool |
| FRIA confusion | Budgeting for a FRIA, or public-service deployer skipping it | The deployer category under Article 27 |

Two of these deserve weight. The **vendor-wide audit** passes on the vendor's data and fails on yours, because the impact ratio depends on who actually applied and how your tool is configured. A green summary built on someone else's applicant pool proves nothing about your deployment. The **stale audit** is the quiet one: the tool keeps running, nobody notices the audit expired, and each day accrues a separate violation up to $1,500 until someone checks the date.

Three more sit in the regime details. Assuming the EU deadline is August 2026 both over- and under-prepares, because the Omnibus moved Annex III to December 2027 while Article 50 transparency still binds in August 2026. Colorado pre-use notice alone is not compliance - the 30-day adverse-outcome explanation and the reconsideration path are separate duties, and all three must exist. And EU duties attach to a specific system, so a blanket policy without configured six-month logs per tool fails; check that retention is on and accessible for each system, not described in a document.

> **Watch out:** FRIA is narrower than vendors imply
>
> A Fundamental Rights Impact Assessment under Article 27 is required only for public bodies, private entities providing public services, and deployers of credit scoring or insurance pricing. A private recruiter deploying a hiring ranker generally does not owe one, so budgeting for a FRIA can waste effort while the real duties - logs, oversight, instructions - go unmet.

## The scoring matrix: deploy, restrict, or hold

Reduce the verdict to two axes: whether the tool is in scope, and whether the triggered duties are met. The four cells give you a defensible grade that two reviewers will reach the same way.

#### The clearance verdict

Horizontal axis runs from Out of scope to In scope. Vertical axis runs from Duties unmet to Duties met.

| Quadrant | What it means |
| --- | --- |
| Out of scope, duties unmet | Deploy - document the out-of-scope finding and the triggering absence of ML parameters |
| In scope, duties met | Deploy - current audit, matched notices, named reviewer, logging all on file |
| Out of scope, duties met | Deploy - over-compliant, keep the evidence and move on |
| In scope, duties unmet | Hold or restrict - restrict where a duty is met, hold where a hard duty is unmet |

*Place the deployment by scope and by whether every triggered duty is satisfied, then read the verdict.*

The restrict verdict is where judgment lives. A tool can be in scope with a current NYC audit and matched notice, so you deploy it for New York City roles, while the same tool has no Colorado pre-use notice configured, so you hold it for Colorado candidates until the notice ships. Restrict means you name the deployments that clear and the deployments that do not, rather than grading the tool as a single pass or fail.

**Clearance verdict record**

```
Deployment: <role + location, one line>
Tool: <name / vendor / version>
In scope? <yes/no> - triggering prong: <sole reliance | highest weight | override | ML clause excludes>
Regimes touched: <NYC | IL | CO | EU>
Bias audit: <date, auditor, representativeness note> - current within 12 months? <yes/no>
Notices live in ATS: <NYC 10-day | IL AI-use | CO pre-use | CO 30-day adverse | EU inform>
Human reviewer: <named owner with override authority>
Logging/retention: <per-system logs on? EU >=6mo / CO 3yr>
VERDICT: <deploy | restrict | hold>
Gaps forcing restrict/hold: <list, each with an owner and date>
Signed: <reviewer> <date>
```

*One record per deployment. Fill every field; a blank field is itself a gap that forces restrict or hold.*

## Records, retention, and keeping the verdict current

Retention duties differ by regime and attach to the specific tool, so set them per system rather than writing one blanket policy. EU deployers must retain automatically generated logs for at least six months, unless Union or national law requires otherwise. Colorado requires documentation and assessments retained for three years. For NYC, the operative rule is that the bias audit must be completed within one year before the tool is used and the summary must stay posted; a fixed multi-year document-retention figure is not established publicly, so hold the audit and summary for as long as the tool runs and keep each annual audit on file.

A verdict is a snapshot that decays. Re-grade on these triggers rather than on a calendar alone:

- **The audit clock.** Re-run the NYC verdict whenever the bias audit approaches twelve months old. Past that line the tool is non-compliant while running.
- **A configuration or model change.** Any change to inputs, weighting, or the model re-opens the AEDT threshold test and can invalidate a vendor audit.
- **A new jurisdiction.** Adding a role or candidate in a new covered market adds a regime and its duties.
- **A regulatory date.** Illinois rulemaking on notice, the Colorado January 2027 effective date, and the EU December 2027 Annex III deadline each shift what is owed. Watch the mechanism - effective dates and rulemaking status - rather than trusting a date you recorded once.

#### Before you sign a deploy verdict

- [ ] The AEDT threshold verdict names the triggering prong or the ML-clause exclusion.
- [ ] A bias audit exists, dated within 12 months, on data representative of your applicant pool, with its summary posted.
- [ ] Every jurisdiction-matched notice is live in the ATS, including Colorado's separate 30-day adverse-outcome path.
- [ ] A named human reviewer has real authority to override the score, not a rubber stamp.
- [ ] Per-system logging is on - EU logs at least six months, Colorado records three years.
- [ ] You confirmed whether a FRIA is actually owed, rather than assuming it is or is not.
- [ ] The verdict is scoped to a specific role-and-location deployment, not the tool in general.

Treat the verdict record as a living document tied to each deployment. The cost of getting this wrong is not a one-time fine - it is a daily accrual while a stale or unaudited tool keeps ranking candidates. Re-grade on the triggers above, keep the evidence where the next reviewer can find it, and name the owner before the tool goes live rather than after a candidate asks how the ranking was made.

## Frequently asked questions

### Is our sourcing tool an AEDT under Local Law 144?

It is an AEDT if its output is a machine-learned or data-tuned score, rank, tag, or classification, and that output is either relied on alone, weighted higher than any other criterion in the set, or used to overrule human conclusions. A human-authored, rule-based keyword filter with no learned parameters falls outside. The test is weighting and learning, not the product's branding or whether a person is present in the loop.

### Does having a human review the AI's output remove us from scope?

No. Local Law 144's third prong puts a tool in scope precisely when its output overrules human conclusions, so a reviewer who rubber-stamps the score can increase exposure. Neither Illinois nor Colorado offers a clean human-in-the-loop exemption either. A human reviewer only helps if they have genuine authority to override and the score is not the highest-weighted factor in the decision.

### Can we use the vendor's bias audit instead of commissioning our own?

Only if it reflects your applicant pool and your configuration. A vendor's generic audit can pass on their data and fail on yours, because the impact ratio depends on who actually applied. For NYC, confirm the audit was performed by an independent auditor, is dated within twelve months of use, and that its summary is posted. Treat vendor evidence as confirmation, not as a substitute for a representative audit.

### When does the EU AI Act apply to a recruitment ranking tool?

The date splits. Article 50 transparency duties bind from August 2, 2026, while the Annex III high-risk recruitment obligations moved from August 2026 to December 2, 2027 under the AI Omnibus Regulation 2026/1744. A single EU deadline line in your policy will be wrong for at least one duty. A private recruiter deploying a hiring ranker generally does not owe a Fundamental Rights Impact Assessment under Article 27.

### What happens if we keep running an AEDT with a stale bias audit?

In NYC, an audit older than twelve months voids compliance while the tool keeps running, and each day of unaudited use is a separate violation. Penalties run up to $500 for a first violation and $500 to $1,500 for subsequent ones. Because penalties accrue daily, holding the tool for the weeks it takes to re-audit is almost always cheaper than running it non-compliant.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/candidate-scoring-tool-clearance-standard*
