The Sourcing Agent Oversight Playbook: Check Every Slate
You will run a repeatable per-batch oversight cycle that catches fabricated, mis-ranked, and unreachable profiles and leaves a defensible audit trail.
An autonomous sourcing agent runs your searches and hands you a ranked slate. This playbook is for the recruiting-operations and revenue-operations owner answerable for how that slate was built, and it gives you a recurring per-batch oversight cycle rather than a one-time tool clearance. Run it, and any slate an agent produces is defensible under NYC Local Law 144 and the EU AI Act, with the fabricated, mis-ranked, and unreachable profiles caught before you trust the list.
Most compliance guidance either clears a hiring tool once before it goes live or verifies a single flagged candidate. Neither fits an always-on agent that generates fresh slates on its own schedule. What follows specifies the sample rates, the exact audit-log fields, and the adverse-impact checks to run on every batch, not once a year.
Why an always-on agent needs a recurring oversight cycle
A sourcing agent that ranks or recommends people is a regulated decision tool, and because it produces a new slate every run, oversight has to be per batch rather than one-time. Under LL144 an automated employment decision tool (AEDT) is any algorithm, model, statistical tool, or AI system that scores, ranks, or recommends candidates to assist or replace human decision-making. A ranked slate is exactly that.
The one-time clearance model breaks here for a simple mechanical reason: the agent's behavior drifts between the moment you cleared it and the batch in front of you. Its sources change, its prompt or config changes, and the underlying population it draws from changes. A clearance you ran last quarter says nothing about whether today's slate contains a fabricated persona or a mis-ranked profile.
The exposure is real and cheap to trip. In a study of 391 recorded employers, only 18 posted the required audit reports and 13 posted transparency notices. LL144 penalties run from $500 for a first violation to $1,500 per day for continuing non-compliance, and candidates must get at least 10 business days' notice before the tool evaluates them. The load-bearing risk is not the audit statistics. It is failing to publish and notice, which is cheap to fix and expensive to miss.
The EU AI Act adds a second regime. AI systems used in recruitment, selection, targeted job advertising, candidate evaluation, and performance monitoring fall into the high-risk category. The Act took effect August 1, 2024, and multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk obligations, including the Article 26 deployer duties. As a deployer you must implement human oversight by natural persons who have the competence, training, and authority to do it, retain automated logs for at least six months, and conduct a Fundamental Rights Impact Assessment where required.
What you are checking for: the four defect classes
Every oversight cycle hunts four defect classes, and each has a different check because each lies in a different way. Naming them keeps the review from collapsing into a vague once-over.
- Fabricated or synthetic profiles. A plausible name, title, and employer with no corroborating live footprint. This is the critical defect. Gartner projects that by 2028 one in four candidate profiles globally will be fake, and Pindrop posted a single job listing and found roughly 12% of applicants used fake identities, about 100 of 827. The UN estimates North Korean IT worker schemes have funneled between $250 million and $600 million annually to the regime since 2018, and the DOJ announced coordinated actions in June 2025 including searches of 29 laptop farms across 16 states.
- Mis-ranked profiles. A real person placed above better-matched people, or placed with reasoning the profile does not support. The lie here is a confident evidence string that does not match the record.
- Unreachable or mismatched people. A real person who cannot be contacted, is duplicated across the slate, or whose location contradicts the brief.
- Suppressed records. Opted-out, do-not-contact, or prior-candidate people re-surfacing through a new source the agent reached into.
The fabrication defect deserves the strictest treatment because sampling alone will miss it. When the thing you are hunting is rare and hidden, a spot check finds it only by luck. That is why existence verification is a named step for every sampled record, not a discretionary spot check.
How large a sample catches a bad slate
There is no hiring-specific sample rate established publicly for sourcing slates, so borrow acceptance sampling from manufacturing. AQL sampling under ISO 2859-1 (also ANSI/ASQ Z1.4) tells you how many units to inspect from a batch and the maximum defects allowed before you reject the entire lot. It is the right transfer because a slate is a lot and a fabricated or mis-ranked profile is a defect.
At General Inspection Level II the sample size scales sub-linearly with the slate, so large batches get proportionally lighter review. Use this table as your starting plan and set fabrication as a critical defect with a low or zero accept number.
| Slate (lot) size | Sample size code | Units to review | Derived review rate |
|---|---|---|---|
| 1,500 | K | 125 | 8.3% |
| ~8,000 | L | 200 | 2.5% |
| 15,000 | M | 315 | 2.1% |
The review-rate column is derived as sample divided by lot; the lot, code, and sample columns come from the standard. The key move is the split accept number: for ordinary defects such as a weak ranking you tolerate a few and reject only past a threshold, but for fabrication you set the accept number at zero, so a single confirmed fake profile in the sample rejects the whole slate and sends it back to the agent.
The oversight cycle, step by step
Run these nine steps in order for every batch. Roles and rough timings are shown so you can staff the cycle; a 1,500-profile slate is a two-hour job for one reviewer plus short RecOps and compliance touches.
Per-batch oversight cycle
- Intake and scope the batchLog the agent run ID, model and version, config, prompt or rules, brief, and lot size. Done: an immutable header record exists before any human looks at a candidate. (RecOps lead, 10 min)
- Draw the sampleApply an AQL plan to the slate size, e.g. 125 for a lot near 1,500, with fabrication as a critical defect and a near-zero accept number. Done: a named sample list is frozen. (Reviewer, 5 min)
- Existence and single-person checkFor each sampled profile confirm a live, unique, reachable person matched to the brief; cross-check duplicates, mismatched locations, and unverifiable employers. Done: each sampled profile is real and reachable or flagged. (Reviewer, 15-30 min)
- Rank sanity and evidence checkVerify the agent's stated reason and evidence for each placement match the profile. Done: each ranking is supported or reclassified as a defect. (Reviewer, 20 min)
- Suppression and source checkConfirm opt-outs, do-not-contact, and prior-candidate suppression were honored and each source was permitted. Done: no suppressed record advances. (Reviewer, 10 min)
- Accept or reject the lotCount defects against the reject number; if exceeded, reject the whole slate and send it back. Done: a documented lot decision exists. (RecOps lead, 5 min)
- Record reason codes and audit trailWrite a mandatory reason code, evidence, reviewer identity, and action for every advanced or rejected profile. Done: the tamper-evident log is complete. (Reviewer, 15 min)
- Adverse-impact drift checkCompute selection rates by protected group at the advance step and flag any impact ratio under 0.80. Done: ratios logged and trended. (RecOps/compliance, 15 min)
- Escalate and noticeTrigger review on any four-fifths flag and confirm LL144 notice and published-audit obligations are current. Done: an escalation ticket or clean sign-off. (Compliance, as needed)
One batch through the oversight cycle
- IntakeFreeze run ID, model, config, and lot size before anyone views a candidate
- SampleDraw an AQL sample with fabrication set as a zero-accept critical defect
- VerifyConfirm existence, ranking evidence, and suppression on each sampled record
- DecideCount defects, accept or reject the lot, and write reason codes
- MonitorCompute impact ratios, trend them, and escalate any four-fifths flag
Verifying a real, reachable, single, correctly-matched person
The existence check is where fabricated personas are caught. Confirm government-issued identity through a reputable verification provider before an offer and match it to the name and work history on the application. For a live conversation, ask the person to change camera angle, hold up an ID, or respond to an unscripted prompt, since real-time deepfakes still struggle with sudden movement, occlusion, and spontaneity. Look for duplicate applicants, mismatched locations, and patterns across applications.
Identity verification should happen before trust is extended and before access is granted. For remote roles, add a real-world checkpoint: at least one in-person identity confirmation before onboarding is complete. It may be tempting to skip or delay this under time pressure. That would be a mistake, because verification is a gate, not a formality.
When the hard part is reachability rather than fraud, the friction is usually finding the right people to review or to fill an oversight bench in the first place. That is a plain-English retrieval problem.
I built Refolk to answer that kind of ask across public LinkedIn records, the public GitHub graph, and the open web, so you can staff the reviewer seat this playbook assumes rather than guessing who has done adverse-impact work before.
Why you cannot auto-reject on an AI-detector score
Never let a single AI-content-detector score drive a rejection, because detector false-positive rates are unreliable for individual decisions and skew against non-native English writers and neurodivergent people. This matters in sourcing specifically, because outreach replies and profile blurbs are short and often lightly edited, which sits in the detector's worst-calibrated region.
| Detector / study | Claimed or measured FPR | Source |
|---|---|---|
| Originality.ai Lite | 0.5% | Vendor self-report |
| Originality.ai Turbo | 1.5% | Vendor self-report |
| Turnitin (vendor claim) | <1% | San Diego law libguide |
| Turnitin (Washington Post test) | ~50% (small sample) | San Diego law libguide |
| GLTR on lightly-polished text | up to ~41-43% | arXiv 2502.15666 |
Two facts from that table decide the rule. First, the Washington Post figure is roughly fifty times the vendor's own claim, so accuracy claims do not travel to your data. Second, GLTR jumped from a 6.83% false-positive rate on pure human text to classifying 40.87% of extremely minor and 42.81% of minor-polished GPT-4o texts as AI. A candidate who ran their reply through a grammar checker lands in that region.
Base rates finish the argument. If 5% of submissions use AI and a detector has a 1% false-positive rate and a 95% true-positive rate, roughly 16% of flagged documents are false positives. A "99% accurate" tool is compatible with a majority-wrong flag list when the thing you are hunting is rare.
A 99 percent accurate detector still produces a majority-wrong flag list when the thing you hunt is rare.
The operational rule: use a detector flag only as one input, and require a second independent human signal before any rejection. A flag routes a record to the existence and single-person check; it never rejects on its own.
Recording the audit trail that survives an inquiry
The audit trail is the deliverable that makes a slate defensible, and it turns on structured reason codes plus system-state capture, stored so it cannot be silently edited. A generic "Rejected" provides zero legal context during an external inquiry. Build structured reason codes for common outcomes, especially rejections, and make at least one code mandatory.
run_id: <agent run ID from intake> model_version: <model or version identifier> config: <thresholds :: weighting :: enabled features> prompt_or_rules: <template or ruleset reference> profile_id: <candidate/profile identifier> reason_code: <MANDATORY: e.g. skills_mismatch | experience_level_mismatch | minimum_requirements_not_met | fabricated_identity | suppressed_record | assessment_below_threshold (threshold version)> evidence: <what was reviewed and where it was confirmed> reviewer_id: <human reviewer identity> action: <accepted | rejected | overridden | escalated> rationale: <free text, required on override or escalation> timestamp: <immutable, tamper-evident>
Store in the ATS/audit log; keep every field, and never overwrite the reviewer identity or timestamp.
Beyond the per-profile record, log the tool and component used, the model or version identifier, config parameters (thresholds, weighting, enabled features), prompt templates or rules, feature categories used, data exclusions, explanations the agent produced, confidence signals, and the human reviewer identity and action with a rationale note. These live in the ATS or audit log, and the log must be tamper-evident and access-controlled with controls that prevent silent edits. An audit trail that permits back-dating is worthless in an inquiry.
Catching adverse-impact drift on every batch
Compute selection rates by protected group at the advance step and flag any impact ratio under 0.80, on every batch, then trend the ratios. The four-fifths rule is codified at 29 CFR 1607.4(D): a selection rate for any race, sex, or ethnic group that is less than four-fifths of the rate for the highest group is generally regarded as evidence of adverse impact.
A four-fifths pass is not proof of fairness, and it is most dangerous on the large slates an always-on agent produces. Adverse impact can be present even when the ratio holds, particularly in high-volume decisions or when a statistical significance test reveals a pattern the ratio obscures. So on high-volume batches, back the ratio with Fisher's exact test, treating a p-value of .05 or less (about 1.96 standard deviations or greater) as statistically significant.
Reading the four-fifths ratio against statistical significance
The stakes are not abstract. The EEOC received 88,531 new discrimination charges in FY2024, a 9.2% jump, and recovered nearly $700 million for more than 21,000 victims. And the models themselves carry the risk you are policing: a University of Washington study of more than three million resume comparisons found LLMs preferred white-associated names 85% of the time and Black-associated names just 9%. An agent inherits that tendency unless you measure the output.
Where this goes wrong: failure modes and false positives
Most oversight failures are predictable, and each has a specific counter. Work this list as a fault catalog when a slate looks clean but you are not sure it is.
- Fabricated profile passes because the sample missed it. It looks like a plausible name, title, and employer with no corroborating live footprint. Counter: run independent existence and reachability confirmation on every sampled record, and set fabrication as a critical defect with accept number zero.
- Detector-driven rejection of a real person. A non-native or neurodivergent writer's outreach reply gets flagged as AI, and the base-rate math means many flags are false. Counter: never auto-reject on a detector score, and require a second human signal.
- Reason code is decorative. A generic "Rejected" gives zero legal context during an inquiry. Counter: mandatory specific code plus an evidence field.
- Four-fifths "pass" that still discriminates. The ratio can hold while a statistically significant pattern exists, especially in large-volume decisions. Counter: run Fisher's exact on high-volume batches.
- Log that can be silently edited. An audit trail that permits back-dating is worthless. Counter: tamper-evident, access-controlled storage.
- Suppression bypass. Opted-out or do-not-contact people re-surface through a new source. Counter: apply the suppression list at the advance step, not just at outreach.
- In-person or ID step skipped under time pressure. It is tempting to skip verification and onboard fast; that is a mistake. Counter: verification is a gate, not a formality.
A 1,500-profile slate through the cycle
- 1,500Full slate
Every profile the agent returned
- 125AQL sample
Code K under General Inspection Level II
- 125Existence-verified
Each sampled record checked for a live, single, reachable person
- 125Advanced with logged reason code
Sampled records carried to a documented lot decision
The pre-sign-off checklist
Before you call any batch done and let the slate advance, confirm every line below. This is the gate between "the agent produced a list" and "the list is defensible."
Before you sign off on a slate
- An immutable header record exists with run ID, model/version, config, prompt/rules, and lot size
- An AQL sample was drawn with fabrication set as a critical defect and a near-zero accept number
- Every sampled profile was confirmed a live, unique, reachable person or flagged as a defect
- Each ranking's stated evidence was checked against the profile and supported or reclassified
- Opt-out, do-not-contact, and prior-candidate suppression were applied at the advance step
- Defects were counted against the reject number and a lot decision was documented
- A mandatory reason code, evidence, reviewer identity, and action were logged for every advanced or rejected profile in tamper-evident storage
- Selection rates by protected group were computed, impact ratios below 0.80 flagged, and Fisher's exact run on high-volume batches
- LL144 notice and the published audit summary are current, and any four-fifths flag has an escalation ticket
Keeping the cycle current
The playbook is stable, but three inputs move under it, so re-check them on a schedule rather than trusting a fixed value. First, the EU AI Act timing: multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk deployer obligations, but one staffing-focused source cites a later date for some tools, so verify against the primary trackers before you rely on a deadline. Second, detector performance: false-positive rates shift with each model and each editing pattern, so the rule to require a second human signal survives even as specific numbers change. Third, the fraud base rate, which is trending up toward Gartner's one-in-four-by-2028 projection and pushes existence checks from spot check to named step.
There is a supply constraint worth planning around. In Refolk's index of professional profiles, the US sourcing and technical-recruiter pool numbers 21,686 people against 775 in the UK, roughly a 28-to-1 ratio.
| Market | Matching people | Derived share |
|---|---|---|
| United States | 21,686 | 96.6% |
| United Kingdom | 775 | 3.4% |
Counts come from Refolk's index; the share is derived. The mechanism to watch is that compliance demand and reviewer supply are decoupling by market: EU AI Act deployer duties land in 2026 while the UK oversight-reviewer bench is far thinner. If you run slates into UK-facing roles, staff the reviewer seat early, because the people who have run adverse-impact analyses or bias audits are scarce there. Finding them is a plain-English search - responsible-AI leads, I/O psychologists, or sourcers who name bias auditing - and it is the difference between a per-batch cycle you can actually sustain and one that exists only on paper.
Questions practitioners ask
Does an autonomous sourcing agent fall under NYC Local Law 144?
Yes, when its output ranks or scores people. An AEDT under LL144 is broadly defined as any algorithm, machine learning model, statistical tool, or AI system that scores, ranks, or recommends candidates to assist or replace human decision-making. If your agent hands you a ranked slate, it is in scope, which triggers the independent annual bias audit, a public summary of results, and at least 10 business days' notice to candidates before use.
How many profiles from a slate do I actually need to review?
There is no hiring-specific sample rate established publicly, so borrow AQL acceptance sampling from ISO 2859-1. At General Inspection Level II a lot near 1,500 units draws a sample of 125, roughly 8,000 draws 200, and 15,000 draws 315. Set fabrication as a critical defect with a near-zero accept number so a single fake profile can reject the lot.
Can I reject a candidate because an AI detector flagged their reply?
No. Detector false-positive rates are unreliable for individual decisions and skew against non-native English writers and neurodivergent people. At 5% prevalence a 1% false-positive rate still makes about 16% of flags wrong, and one measured rate reached roughly 50% against a vendor's sub-1% claim. Require a second independent human signal before any rejection, and never let a detector score stand alone.
What fields must the audit log actually contain?
Log a mandatory structured reason code plus an evidence field per action, and capture the tool and component, model or version identifier, config parameters, prompt templates or rules, feature categories, data exclusions, explanations produced, confidence signals, and the human reviewer identity and action with a rationale. Store it tamper-evident and access-controlled so silent edits and back-dating are impossible.
Is a four-fifths pass enough to prove no adverse impact?
No. Adverse impact can be present even when the four-fifths ratio holds, particularly in large-volume decisions or when a statistical significance test reveals a pattern the ratio obscures. Run Fisher's exact on high-volume batches and treat a p-value of .05 or less, equivalent to about 1.96 standard deviations, as significant. Always-on agents produce exactly the large batches where the rule of thumb under-warns.
When do EU AI Act obligations bite for a deploying employer?
Multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk obligations, including Article 26 deployer duties, though the Act itself took effect August 1, 2024. As a deployer you must implement human oversight by competent, trained, authorized people, retain automated logs for at least six months, and conduct a Fundamental Rights Impact Assessment where required. One staffing page cites a later date, so re-check against the primary trackers.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.