The AI Hiring Tool Register: Inventory, Risk Tier, and Evidence
You will be able to inventory every AI tool in your hiring stack, assign each a defensible EU AI Act risk tier, and hold the evidence file behind every classification.
You own the answer to a question an auditor, a works council, or a plaintiff's lawyer will eventually ask: which AI tools touched this candidate, and how do you know each one was allowed to. This guide is for the recruiting operations, revenue operations, and compliance leads answerable for that answer. It delivers a running deployer register - an inventory of every sourcing and screening tool, a defensible risk tier for each, and the evidence file behind it - plus the cadence that keeps it true as vendors and models change.
Most published material on this ranks tools against a shifting deadline and hands you a one-time checklist. That decays. A register is a durable object. It survives a deadline moving, a vendor swap, and a model update, because it stores the reasoning and the proof, not the date.
What makes a sourcing or screening tool high-risk
A sourcing or screening tool is high-risk when it is intended for recruitment or candidate selection, unless it clears the Article 6(3) filter and does no profiling. Employment sits in Annex III point 4 of the EU AI Act, which covers recruitment and candidate-selection systems and systems affecting the terms of a work relationship, promotion, termination, task allocation, or performance monitoring.
Annex III lists eight categories that are classified as high-risk under Article 6(2) when built for the regulated purposes: biometrics, critical infrastructure, education and vocational training, employment and worker management, essential private and public services, law enforcement, migration and border control, and administration of justice. Hiring tools land in point 4 the moment they screen, rank, or select people.
There is one escape route and one override. The escape route is Article 6(3): a system inside an Annex III use case may self-assess as not high-risk if it meets one of four conditions - it performs a narrow procedural task, improves a completed human activity, detects decision patterns without replacing human assessment, or performs a preparatory task. The override is decisive: a system that performs profiling of natural persons is always high-risk, regardless of whether it otherwise satisfies one of those four conditions.
That override is where most register mistakes hide. A tool marketed as a "preparatory" ranking helper that infers traits about candidates is profiling, and profiling forces high-risk. Your register needs a dedicated profiling flag that overrides the 6(3) column, not a footnote.
The Commission published draft high-risk classification guidelines on 19 May 2026 and opened a five-week consultation running to 23 June 2026. Guidance will keep moving; the classification logic in Article 6 will not. Anchor your register to the article, and re-read the guidance when it finalizes.
What the deployer owes, and what proves it
As a deployer you carry Article 26, which imposes 12 distinct obligations. The proof is a set of artifacts you can hand over on request. The core duties are: use the system according to the provider's instructions, with competent human oversight, and monitor its operation; manage input data, keep logs for at least six months, and inform the provider and authorities of any risk or incident; and notify workers' representatives and affected workers when the system is used in the workplace.
Each duty maps to a document. That mapping is the spine of the evidence file.
| Deployer duty (Article 26) | Evidence artifact you store |
|---|---|
| Use per instructions | Instructions-adherence procedure, referencing the vendor instructions for use |
| Human oversight | Named overseer, authority statement, training records, logged interventions |
| Input-data governance | Data-governance notes for inputs you control |
| Log retention (six-month floor) | Log exports with retention location and cutoff |
| Risk and incident reporting | Incident log and provider or authority notifications |
| Worker notification | Worker-notice records and representative notice |
Note two moving parts. First, the AI-literacy duty under Article 4 was softened in the Omnibus, so treat literacy as good practice rather than a hard artifact until the position settles. Second, the FRIA under Article 27 is not universal, covered later. Store what you have, and record plainly where a duty does not apply to you and why.
What the provider owes, and how to verify it
The provider owns the conformity documentation; your job is to collect it and check it is real. The provider must produce the technical documentation set out in Annex IV, the instructions for use under Article 13, the EU declaration of conformity, and the CE marking.
Verify each of these concretely rather than accepting a marketing PDF. The EU declaration of conformity must be drawn up as a machine-readable, physical or electronically signed document for each AI system, kept for 10 years, and made available to authorities on request; it confirms the system meets requirements and carries the information in Annex V. The CE marking must be affixed visibly, legibly, and indelibly for high-risk systems.
The tell for third-party assessment is a number. Where the conformity assessment under Article 43 involved a notified body, the CE marking is followed by that body's identification number. A CE mark with no notified-body number on a high-risk system either means self-assessment was used or that the marking is incomplete - both are things to record and question.
One reason to build the file now rather than later: the deferral bought time precisely because harmonized standards were not ready. Under the original timetable, organizations would have had to demonstrate compliance without finalized benchmarks. The deadline moved; the work did not shrink. A register assembled today is the hedge against the standards landing late.
The register schema that makes classification reproducible
No official standardized deployer-register schema is published, so treat any "canonical template" claim with suspicion. The defensible field set is derived from Articles 6, 11, 12, 26, and 27 and Annex VIII. The point of a schema is reproducibility: two people classifying the same tool a year apart should reach the same tier from the same fields.
System name Provider Deployer owner (internal) Intended purpose (one sentence) Annex III category (e.g. point 4, employment) Article 6(3) determination + written justification Profiling flag (Y/N - Y forces high-risk) Assigned risk tier DoC reference + CE / notified-body ID Instructions-for-use reference FRIA reference (or "N/A - regime recorded") Human-oversight design (owner + intervention points) Log-retention location Last bias-test date + impact ratios (per req) Monitoring cadence Incident log link
Copy as column headers. Fill one row per sourcing or screening tool that touches candidate data.
The two columns people skip are the two that matter most in a dispute: the written Article 6(3) justification and the profiling flag. A tier with no written reasoning is not defensible; it is an assertion. Write the sentence that says why this tool is or is not high-risk, and let the profiling flag override it when profiling is present.
Evidence file, outermost to innermost
- Register rowTier, owner, and pointers to everything below
- Classification reasoningAnnex III category, 6(3) justification, profiling flag
- Provider conformityDoC, CE / notified-body ID, instructions for use
- Deployer controlsOversight design, FRIA or regime note, worker notices
- Running evidenceSix-month logs, dated bias-test reports, incident log
Build the register: the procedure
Run these steps in order. The first pass takes a few weeks; after that the register lives on a quarterly rhythm. Roles and durations are drawn from a working deployment cadence, not from the Act, which sets no timetable for the build itself.
Standing up the register
- Discover the stackList every sourcing and screening tool that touches candidate data, from search sourcers to resume screeners and interview scorers. Done means one row per tool with named provider and internal owner.
- Classify each toolWalk Annex III point 4, then the Article 6(3) filter, and flag any profiling of natural persons, which forces high-risk. Done means a recorded tier with a written 6(3) justification per tool.
- Collect the provider conformity fileObtain the instructions for use, the EU declaration of conformity, the CE marking or notified-body ID, and any bias-test evidence from each vendor. Done means the file is stored and referenced from the register row.
- Design human oversightAssign a competent person with real authority to intervene, and document the intervention points in the workflow. Done means the oversight owner and procedure are recorded with training evidence.
- Run a FRIA where requiredFor deployers in scope of Article 27, complete the fundamental rights impact assessment elements before first use and notify the market-surveillance authority. Done means the FRIA is filed, or the register records that no strict FRIA duty applies.
- Stand up monitoringSet log retention to at least six months and schedule at-least-annual adverse-impact testing on your own candidate data using the four-fifths method. Done means the first test report is on file with its data source disclosed.
- Set the maintenance cadenceRe-review classifications on any vendor change or model update, and refresh evidence on a quarterly rhythm. Done means a dated review log showing the register was checked and what changed.
Discovery is where teams underestimate. A tool that "just ranks" candidates or "just writes outreach" from candidate data is in scope for the inventory even if it clears 6(3) later. The register lists everything that touches candidate data; the tier column decides how much evidence each row needs.
Finding the person to own this is its own problem. When you need an independent auditor or a governance lead to sanity-check a tier, Refolk turns a plain-English description of that person into a shortlist across the public GitHub graph, public LinkedIn records, and the open web, rather than a keyword hunt that misses people who never used your exact title.
Human oversight that survives scrutiny
Human oversight is a design, not a signature. Article 26(2) requires a competent person with the authority and the training to intervene, and the artifact that proves it is a record of interventions, not a name on a form. An overseer who cannot override the tool, or who has never once overridden it, is oversight in name only.
Design the intervention points explicitly. Where in the workflow can the overseer see the tool's output before it acts on a candidate? Where can they reverse it? Who trained them, and on what? A register row that answers those three questions is defensible. A row that names a person and stops is a false positive waiting to be exposed.
Governance staffing makes this harder in Europe. In Refolk's index, 19,607 US professionals list Responsible AI or AI Governance skills against 1,585 in Germany - a 12.4x gap. EU deployers will more often hand register ownership to a generalist, which is exactly why the schema must be self-documenting: the written justifications and intervention records carry the knowledge that a specialist would otherwise hold in their head.
| Market | People with Responsible AI / AI Governance skills | Share of US total |
|---|---|---|
| United States | 19,607 | 100% |
| Germany | 1,585 | 8.1% |
| US : Germany ratio | 12.4x | - |
That scarcity is sharper at the top: across the queried US and German set, only 24 people hold an explicit AI Governance, Compliance, or Responsible AI manager-level title. Do not design a register that assumes a dedicated manager will maintain it. Design one a competent generalist can run from the written record.
Oversight is a logged intervention, not a signature; a name with no override history is a gap, not a control.
Monitoring and adverse-impact testing
The Act sets no numeric cadence. Deployers must monitor the operation of the high-risk system on the basis of the instructions for use and inform the provider of any risk identified under Article 26(5), plus keep logs for at least six months. That is the floor. The useful practice comes from a more concrete regime.
Borrow the four-fifths method from NYC Local Law 144. The impact ratio is the selection rate of a group divided by the rate of the most-selected group, and a ratio below 0.80 generally indicates potential adverse impact under the EEOC's four-fifths rule. Use historical data from actual use; if that is insufficient, test data may be used, but the limitation must be disclosed. Audit at least annually, and re-audit whenever the tool is materially changed.
| Group | Selection rate | Impact ratio | Flag |
|---|---|---|---|
| Men (highest) | 60% | 1.00 | - |
| Women | 40% | 0.67 | Below 0.80 |
The worked example above is deliberately simple. The trap is that it looks fine when pooled. A 2026 Stanford study followed 3.4 million people submitting 4 million applications across 1,700 positions, and disparity surfaced only at the role level, because pooling averages out role-specific screens. A tool can pass an aggregate bias audit while systematically screening out candidates for specific roles. Run your impact ratios per requisition, and store the data source on every test so synthetic data never masquerades as live monitoring.
Where the evidence file gets built
- AllTools touching candidate data
Everything goes in the inventory
- Most sourcing and screening toolsInside Annex III point 4
Employment use case
- The high-risk subsetNot cleared by 6(3) or profiling
Full evidence file required
- The re-audit queueMaterially changed since last test
New bias test triggered
How this register goes wrong
The failure modes below are where a register that looks complete falls apart under scrutiny. Each one has a false positive - something that looks like compliance and is not - and a concrete check.
| Failure mode | What it looks like | The check |
|---|---|---|
| Aggregate audit hides role-level bias | A clean pooled pass on a tool that screens out groups per role | Impact ratios per requisition, never pooled |
| Treating the deferral as relief | Register work paused until December 2027 | Article 50 transparency artifacts dated on or before 2 August 2026 |
| Outsourcing liability to the vendor | A vendor audit report filed as full coverage | Your own notices, logs, and human review on file |
| Mis-firing the 6(3) filter | A profiling tool labeled a preparatory task | Mandatory profiling flag that overrides the 6(3) column |
| Oversight in name only | A signature with no intervention record | Logged interventions with authority to override |
| Test data as live monitoring | Fairness numbers from undisclosed synthetic data | Data-source field on every test report |
| DPIA reused as FRIA | A data-processing assessment filed under Article 27 | FRIA references the six prescribed rights elements |
Two of these deserve extra weight. The deferral is not relief. The Digital Omnibus moved Annex III standalone high-risk compliance to 2 December 2027, but Article 50 transparency has been enforceable since 2 August 2026, with only a legacy grace period to 2 December 2026. A register that stops until 2027 already carries a live transparency gap.
The FRIA trap runs the other way - people over-build. Article 27 requires FRIAs from public-law bodies, private entities providing public services, and specific private deployers using high-risk AI for credit scoring or insurance pricing. Many private recruiters carry Article 26 duties but not a strict FRIA obligation. Record which regime applies in the register so you neither skip a required FRIA nor waste weeks producing one you do not owe. And never reuse a completed DPIA as a FRIA; the rights scope is different.
Enforcement is thin but visibly tightening, which is the argument for a standing file over a one-time push. In a 2025 review, a state comptroller found at least 17 potential non-compliances where the enforcing agency had flagged just one among 32 firms. That gap invites proactive enforcement, and a standing evidence file is cheap insurance against it.
Before you call the register done
Run this checklist before you sign off the first full pass, and re-run the top half each quarter.
Register sign-off
- Every sourcing and screening tool touching candidate data has a row with named provider and internal owner.
- Each row records an Annex III category, an Article 6(3) determination with written justification, and a profiling flag.
- Each high-risk row references the vendor's declaration of conformity and CE or notified-body ID.
- Each high-risk row names a human overseer with documented authority and at least one logged intervention path.
- Article 27 applicability is recorded on every high-risk row, with a FRIA filed where required.
- Log retention is set to at least six months with a named storage location per tool.
- A first adverse-impact test is on file, computed per requisition, with its data source disclosed.
- Article 50 transparency artifacts are dated on or before 2 August 2026.
- A dated review log exists and shows the last classification review.
Keeping the register current
A register is only as good as its last review, so put the maintenance cadence in a calendar, not in someone's memory. The two triggers that force a re-classification are a vendor change and a model update; either can move a tool across the 6(3) line or turn a non-profiling tool into a profiling one. Between triggers, run a quarterly pass over the whole register and refresh the evidence that expires: logs age past six months, bias tests go stale after a year, and instructions for use get reissued.
The timeline table below is the one thing most likely to change under you, so track the mechanism rather than memorizing dates. Re-check the Official Journal and the Commission guidance when a deadline nears; the classification logic and the deployer duties are stable, but the dates and the finalized standards are not.
| Obligation | Original date | Post-Omnibus date |
|---|---|---|
| Annex III standalone high-risk | 2 Aug 2026 | 2 Dec 2027 |
| Annex I embedded high-risk | 2 Aug 2027 | 2 Aug 2028 |
| Article 50 transparency | 2 Aug 2026 | Unchanged (legacy grace to 2 Dec 2026) |
When the Commission's high-risk classification guidelines finalize after their consultation, re-read your Article 6(3) justifications against the final text and update any that no longer hold. When a vendor ships a new model, treat it as a new tool for classification purposes until you have confirmed the tier is unchanged. And when you swap a vendor entirely, the old row does not get deleted - it gets closed with a date, because the candidates it touched are still on record. A register that keeps its history is the one that answers the auditor's question years later, which is the only test that matters.
Questions practitioners ask
Is a candidate sourcing tool automatically high-risk under the EU AI Act?
Not automatically. Employment sits in Annex III point 4, so a sourcing tool falls into a regulated use case, but Article 6(3) lets a system self-assess out if it only performs a narrow procedural task, improves a completed human activity, detects patterns without replacing human assessment, or does a preparatory task. The override is profiling: any system that profiles natural persons is always high-risk regardless of that filter. Record the determination in writing per tool.
Does a vendor's bias audit satisfy my deployer obligations?
No. Most obligations sit on the deployer, not the provider. A vendor's audit report can be part of your evidence file, but you still need your own worker-notice practices, log retention of at least six months, human oversight with intervention records, and your own adverse-impact testing on your candidate data. Treating a vendor report as full coverage is one of the most common register failures.
When do I actually have to comply, given the Digital Omnibus deferral?
The Omnibus pushed Annex III standalone high-risk compliance from 2 August 2026 to 2 December 2027, and Annex I embedded high-risk to 2 August 2028. But Article 50 transparency has been enforceable since 2 August 2026, with only a legacy grace period to 2 December 2026. So transparency work is due now; the high-risk build has runway, but the register itself should exist today as the hedge.
Do I need a FRIA for AI recruitment tools?
Often not a strict one. Article 27 requires fundamental rights impact assessments from public-law bodies, private entities providing public services, and specific private deployers using high-risk AI for credit scoring or insurance pricing. Many private recruiters carry Article 26 duties without a strict FRIA obligation. Record which regime applies so you neither over-build nor under-build, and never reuse a DPIA as a FRIA since the rights scope differs.
How do I test a hiring tool for adverse impact?
Borrow the four-fifths rule. Compute the impact ratio as the selection rate of a group divided by the rate of the most-selected group; a ratio below 0.80 flags potential adverse impact. Use historical data from actual use where possible, and disclose the limitation if you must use test data. Run the ratios per role or requisition rather than pooled, and repeat at least annually or whenever the tool is materially changed.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.