Refolk
PlaybookProcess, data, and compliance

The Sourcing Agent Oversight Playbook: Check Every Slate

You will run a repeatable per-batch oversight cycle that catches fabricated, mis-ranked, and unreachable profiles and leaves a defensible audit trail.

18 min readLast reviewed August 28, 2026Read as Markdown

An autonomous sourcing agent runs your searches and hands you a ranked slate. This playbook is for the recruiting-operations and revenue-operations owner answerable for how that slate was built, and it gives you a recurring per-batch oversight cycle rather than a one-time tool clearance. Run it, and any slate an agent produces is defensible under NYC Local Law 144 and the EU AI Act, with the fabricated, mis-ranked, and unreachable profiles caught before you trust the list.

Most compliance guidance either clears a hiring tool once before it goes live or verifies a single flagged candidate. Neither fits an always-on agent that generates fresh slates on its own schedule. What follows specifies the sample rates, the exact audit-log fields, and the adverse-impact checks to run on every batch, not once a year.

Why an always-on agent needs a recurring oversight cycle

A sourcing agent that ranks or recommends people is a regulated decision tool, and because it produces a new slate every run, oversight has to be per batch rather than one-time. Under LL144 an automated employment decision tool (AEDT) is any algorithm, model, statistical tool, or AI system that scores, ranks, or recommends candidates to assist or replace human decision-making. A ranked slate is exactly that.

The one-time clearance model breaks here for a simple mechanical reason: the agent's behavior drifts between the moment you cleared it and the batch in front of you. Its sources change, its prompt or config changes, and the underlying population it draws from changes. A clearance you ran last quarter says nothing about whether today's slate contains a fabricated persona or a mis-ranked profile.

The exposure is real and cheap to trip. In a study of 391 recorded employers, only 18 posted the required audit reports and 13 posted transparency notices. LL144 penalties run from $500 for a first violation to $1,500 per day for continuing non-compliance, and candidates must get at least 10 business days' notice before the tool evaluates them. The load-bearing risk is not the audit statistics. It is failing to publish and notice, which is cheap to fix and expensive to miss.

18 of 391
Employers that posted the LL144 bias-audit report they owed
In the same set only 13 posted a transparency notice, so the compliance gap is publishing and notice, not the math.

The EU AI Act adds a second regime. AI systems used in recruitment, selection, targeted job advertising, candidate evaluation, and performance monitoring fall into the high-risk category. The Act took effect August 1, 2024, and multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk obligations, including the Article 26 deployer duties. As a deployer you must implement human oversight by natural persons who have the competence, training, and authority to do it, retain automated logs for at least six months, and conduct a Fundamental Rights Impact Assessment where required.

What you are checking for: the four defect classes

Every oversight cycle hunts four defect classes, and each has a different check because each lies in a different way. Naming them keeps the review from collapsing into a vague once-over.

  • Fabricated or synthetic profiles. A plausible name, title, and employer with no corroborating live footprint. This is the critical defect. Gartner projects that by 2028 one in four candidate profiles globally will be fake, and Pindrop posted a single job listing and found roughly 12% of applicants used fake identities, about 100 of 827. The UN estimates North Korean IT worker schemes have funneled between $250 million and $600 million annually to the regime since 2018, and the DOJ announced coordinated actions in June 2025 including searches of 29 laptop farms across 16 states.
  • Mis-ranked profiles. A real person placed above better-matched people, or placed with reasoning the profile does not support. The lie here is a confident evidence string that does not match the record.
  • Unreachable or mismatched people. A real person who cannot be contacted, is duplicated across the slate, or whose location contradicts the brief.
  • Suppressed records. Opted-out, do-not-contact, or prior-candidate people re-surfacing through a new source the agent reached into.

The fabrication defect deserves the strictest treatment because sampling alone will miss it. When the thing you are hunting is rare and hidden, a spot check finds it only by luck. That is why existence verification is a named step for every sampled record, not a discretionary spot check.

How large a sample catches a bad slate

There is no hiring-specific sample rate established publicly for sourcing slates, so borrow acceptance sampling from manufacturing. AQL sampling under ISO 2859-1 (also ANSI/ASQ Z1.4) tells you how many units to inspect from a batch and the maximum defects allowed before you reject the entire lot. It is the right transfer because a slate is a lot and a fabricated or mis-ranked profile is a defect.

At General Inspection Level II the sample size scales sub-linearly with the slate, so large batches get proportionally lighter review. Use this table as your starting plan and set fabrication as a critical defect with a low or zero accept number.

Slate (lot) sizeSample size codeUnits to reviewDerived review rate
1,500K1258.3%
~8,000L2002.5%
15,000M3152.1%

The review-rate column is derived as sample divided by lot; the lot, code, and sample columns come from the standard. The key move is the split accept number: for ordinary defects such as a weak ranking you tolerate a few and reject only past a threshold, but for fabrication you set the accept number at zero, so a single confirmed fake profile in the sample rejects the whole slate and sends it back to the agent.

The oversight cycle, step by step

Run these nine steps in order for every batch. Roles and rough timings are shown so you can staff the cycle; a 1,500-profile slate is a two-hour job for one reviewer plus short RecOps and compliance touches.

Per-batch oversight cycle

  1. Intake and scope the batch
    Log the agent run ID, model and version, config, prompt or rules, brief, and lot size. Done: an immutable header record exists before any human looks at a candidate. (RecOps lead, 10 min)
  2. Draw the sample
    Apply an AQL plan to the slate size, e.g. 125 for a lot near 1,500, with fabrication as a critical defect and a near-zero accept number. Done: a named sample list is frozen. (Reviewer, 5 min)
  3. Existence and single-person check
    For each sampled profile confirm a live, unique, reachable person matched to the brief; cross-check duplicates, mismatched locations, and unverifiable employers. Done: each sampled profile is real and reachable or flagged. (Reviewer, 15-30 min)
  4. Rank sanity and evidence check
    Verify the agent's stated reason and evidence for each placement match the profile. Done: each ranking is supported or reclassified as a defect. (Reviewer, 20 min)
  5. Suppression and source check
    Confirm opt-outs, do-not-contact, and prior-candidate suppression were honored and each source was permitted. Done: no suppressed record advances. (Reviewer, 10 min)
  6. Accept or reject the lot
    Count defects against the reject number; if exceeded, reject the whole slate and send it back. Done: a documented lot decision exists. (RecOps lead, 5 min)
  7. Record reason codes and audit trail
    Write a mandatory reason code, evidence, reviewer identity, and action for every advanced or rejected profile. Done: the tamper-evident log is complete. (Reviewer, 15 min)
  8. Adverse-impact drift check
    Compute selection rates by protected group at the advance step and flag any impact ratio under 0.80. Done: ratios logged and trended. (RecOps/compliance, 15 min)
  9. Escalate and notice
    Trigger review on any four-fifths flag and confirm LL144 notice and published-audit obligations are current. Done: an escalation ticket or clean sign-off. (Compliance, as needed)

One batch through the oversight cycle

  1. Intake
    Freeze run ID, model, config, and lot size before anyone views a candidate
  2. Sample
    Draw an AQL sample with fabrication set as a zero-accept critical defect
  3. Verify
    Confirm existence, ranking evidence, and suppression on each sampled record
  4. Decide
    Count defects, accept or reject the lot, and write reason codes
  5. Monitor
    Compute impact ratios, trend them, and escalate any four-fifths flag
Every slate passes the same gates in the same order, from an immutable header to a logged lot decision.

Verifying a real, reachable, single, correctly-matched person

The existence check is where fabricated personas are caught. Confirm government-issued identity through a reputable verification provider before an offer and match it to the name and work history on the application. For a live conversation, ask the person to change camera angle, hold up an ID, or respond to an unscripted prompt, since real-time deepfakes still struggle with sudden movement, occlusion, and spontaneity. Look for duplicate applicants, mismatched locations, and patterns across applications.

Identity verification should happen before trust is extended and before access is granted. For remote roles, add a real-world checkpoint: at least one in-person identity confirmation before onboarding is complete. It may be tempting to skip or delay this under time pressure. That would be a mistake, because verification is a gate, not a formality.

When the hard part is reachability rather than fraud, the friction is usually finding the right people to review or to fill an oversight bench in the first place. That is a plain-English retrieval problem.

I built Refolk to answer that kind of ask across public LinkedIn records, the public GitHub graph, and the open web, so you can staff the reviewer seat this playbook assumes rather than guessing who has done adverse-impact work before.

Why you cannot auto-reject on an AI-detector score

Never let a single AI-content-detector score drive a rejection, because detector false-positive rates are unreliable for individual decisions and skew against non-native English writers and neurodivergent people. This matters in sourcing specifically, because outreach replies and profile blurbs are short and often lightly edited, which sits in the detector's worst-calibrated region.

Detector / studyClaimed or measured FPRSource
Originality.ai Lite0.5%Vendor self-report
Originality.ai Turbo1.5%Vendor self-report
Turnitin (vendor claim)<1%San Diego law libguide
Turnitin (Washington Post test)~50% (small sample)San Diego law libguide
GLTR on lightly-polished textup to ~41-43%arXiv 2502.15666

Two facts from that table decide the rule. First, the Washington Post figure is roughly fifty times the vendor's own claim, so accuracy claims do not travel to your data. Second, GLTR jumped from a 6.83% false-positive rate on pure human text to classifying 40.87% of extremely minor and 42.81% of minor-polished GPT-4o texts as AI. A candidate who ran their reply through a grammar checker lands in that region.

Base rates finish the argument. If 5% of submissions use AI and a detector has a 1% false-positive rate and a 95% true-positive rate, roughly 16% of flagged documents are false positives. A "99% accurate" tool is compatible with a majority-wrong flag list when the thing you are hunting is rare.

A 99 percent accurate detector still produces a majority-wrong flag list when the thing you hunt is rare.

The operational rule: use a detector flag only as one input, and require a second independent human signal before any rejection. A flag routes a record to the existence and single-person check; it never rejects on its own.

Recording the audit trail that survives an inquiry

The audit trail is the deliverable that makes a slate defensible, and it turns on structured reason codes plus system-state capture, stored so it cannot be silently edited. A generic "Rejected" provides zero legal context during an external inquiry. Build structured reason codes for common outcomes, especially rejections, and make at least one code mandatory.

Per-profile audit record (one row per advanced or rejected profile)
run_id: <agent run ID from intake>
model_version: <model or version identifier>
config: <thresholds :: weighting :: enabled features>
prompt_or_rules: <template or ruleset reference>
profile_id: <candidate/profile identifier>
reason_code: <MANDATORY: e.g. skills_mismatch | experience_level_mismatch | minimum_requirements_not_met | fabricated_identity | suppressed_record | assessment_below_threshold (threshold version)>
evidence: <what was reviewed and where it was confirmed>
reviewer_id: <human reviewer identity>
action: <accepted | rejected | overridden | escalated>
rationale: <free text, required on override or escalation>
timestamp: <immutable, tamper-evident>

Store in the ATS/audit log; keep every field, and never overwrite the reviewer identity or timestamp.

Beyond the per-profile record, log the tool and component used, the model or version identifier, config parameters (thresholds, weighting, enabled features), prompt templates or rules, feature categories used, data exclusions, explanations the agent produced, confidence signals, and the human reviewer identity and action with a rationale note. These live in the ATS or audit log, and the log must be tamper-evident and access-controlled with controls that prevent silent edits. An audit trail that permits back-dating is worthless in an inquiry.

Catching adverse-impact drift on every batch

Compute selection rates by protected group at the advance step and flag any impact ratio under 0.80, on every batch, then trend the ratios. The four-fifths rule is codified at 29 CFR 1607.4(D): a selection rate for any race, sex, or ethnic group that is less than four-fifths of the rate for the highest group is generally regarded as evidence of adverse impact.

A four-fifths pass is not proof of fairness, and it is most dangerous on the large slates an always-on agent produces. Adverse impact can be present even when the ratio holds, particularly in high-volume decisions or when a statistical significance test reveals a pattern the ratio obscures. So on high-volume batches, back the ratio with Fisher's exact test, treating a p-value of .05 or less (about 1.96 standard deviations or greater) as statistically significant.

Reading the four-fifths ratio against statistical significance

Statistically significantNot statistically significant
Ratio ok, not significant
Clear the batch; log the ratio and keep trending
Ratio below 0.80, not significant
Investigate; small samples can flag noise, so gather more batches before acting
Ratio ok but significant
The dangerous case on large slates; escalate despite the passing ratio
Ratio below 0.80 and significant
Reject and escalate; strongest evidence of adverse impact
Impact ratio at or above 0.80Impact ratio below 0.80
The ratio and the significance test disagree often enough that you need both before you clear a batch.

The stakes are not abstract. The EEOC received 88,531 new discrimination charges in FY2024, a 9.2% jump, and recovered nearly $700 million for more than 21,000 victims. And the models themselves carry the risk you are policing: a University of Washington study of more than three million resume comparisons found LLMs preferred white-associated names 85% of the time and Black-associated names just 9%. An agent inherits that tendency unless you measure the output.

Where this goes wrong: failure modes and false positives

Most oversight failures are predictable, and each has a specific counter. Work this list as a fault catalog when a slate looks clean but you are not sure it is.

  • Fabricated profile passes because the sample missed it. It looks like a plausible name, title, and employer with no corroborating live footprint. Counter: run independent existence and reachability confirmation on every sampled record, and set fabrication as a critical defect with accept number zero.
  • Detector-driven rejection of a real person. A non-native or neurodivergent writer's outreach reply gets flagged as AI, and the base-rate math means many flags are false. Counter: never auto-reject on a detector score, and require a second human signal.
  • Reason code is decorative. A generic "Rejected" gives zero legal context during an inquiry. Counter: mandatory specific code plus an evidence field.
  • Four-fifths "pass" that still discriminates. The ratio can hold while a statistically significant pattern exists, especially in large-volume decisions. Counter: run Fisher's exact on high-volume batches.
  • Log that can be silently edited. An audit trail that permits back-dating is worthless. Counter: tamper-evident, access-controlled storage.
  • Suppression bypass. Opted-out or do-not-contact people re-surface through a new source. Counter: apply the suppression list at the advance step, not just at outreach.
  • In-person or ID step skipped under time pressure. It is tempting to skip verification and onboard fast; that is a mistake. Counter: verification is a gate, not a formality.

A 1,500-profile slate through the cycle

  1. Full slate
    1,500

    Every profile the agent returned

  2. AQL sample
    125

    Code K under General Inspection Level II

  3. Existence-verified
    125

    Each sampled record checked for a live, single, reachable person

  4. Advanced with logged reason code
    125

    Sampled records carried to a documented lot decision

Volumes narrow from the full slate to the drawn sample to the records that reach a human decision.

The pre-sign-off checklist

Before you call any batch done and let the slate advance, confirm every line below. This is the gate between "the agent produced a list" and "the list is defensible."

Before you sign off on a slate

  • An immutable header record exists with run ID, model/version, config, prompt/rules, and lot size
  • An AQL sample was drawn with fabrication set as a critical defect and a near-zero accept number
  • Every sampled profile was confirmed a live, unique, reachable person or flagged as a defect
  • Each ranking's stated evidence was checked against the profile and supported or reclassified
  • Opt-out, do-not-contact, and prior-candidate suppression were applied at the advance step
  • Defects were counted against the reject number and a lot decision was documented
  • A mandatory reason code, evidence, reviewer identity, and action were logged for every advanced or rejected profile in tamper-evident storage
  • Selection rates by protected group were computed, impact ratios below 0.80 flagged, and Fisher's exact run on high-volume batches
  • LL144 notice and the published audit summary are current, and any four-fifths flag has an escalation ticket

Keeping the cycle current

The playbook is stable, but three inputs move under it, so re-check them on a schedule rather than trusting a fixed value. First, the EU AI Act timing: multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk deployer obligations, but one staffing-focused source cites a later date for some tools, so verify against the primary trackers before you rely on a deadline. Second, detector performance: false-positive rates shift with each model and each editing pattern, so the rule to require a second human signal survives even as specific numbers change. Third, the fraud base rate, which is trending up toward Gartner's one-in-four-by-2028 projection and pushes existence checks from spot check to named step.

There is a supply constraint worth planning around. In Refolk's index of professional profiles, the US sourcing and technical-recruiter pool numbers 21,686 people against 775 in the UK, roughly a 28-to-1 ratio.

MarketMatching peopleDerived share
United States21,68696.6%
United Kingdom7753.4%

Counts come from Refolk's index; the share is derived. The mechanism to watch is that compliance demand and reviewer supply are decoupling by market: EU AI Act deployer duties land in 2026 while the UK oversight-reviewer bench is far thinner. If you run slates into UK-facing roles, staff the reviewer seat early, because the people who have run adverse-impact analyses or bias audits are scarce there. Finding them is a plain-English search - responsible-AI leads, I/O psychologists, or sourcers who name bias auditing - and it is the difference between a per-batch cycle you can actually sustain and one that exists only on paper.

Questions practitioners ask

Does an autonomous sourcing agent fall under NYC Local Law 144?

Yes, when its output ranks or scores people. An AEDT under LL144 is broadly defined as any algorithm, machine learning model, statistical tool, or AI system that scores, ranks, or recommends candidates to assist or replace human decision-making. If your agent hands you a ranked slate, it is in scope, which triggers the independent annual bias audit, a public summary of results, and at least 10 business days' notice to candidates before use.

How many profiles from a slate do I actually need to review?

There is no hiring-specific sample rate established publicly, so borrow AQL acceptance sampling from ISO 2859-1. At General Inspection Level II a lot near 1,500 units draws a sample of 125, roughly 8,000 draws 200, and 15,000 draws 315. Set fabrication as a critical defect with a near-zero accept number so a single fake profile can reject the lot.

Can I reject a candidate because an AI detector flagged their reply?

No. Detector false-positive rates are unreliable for individual decisions and skew against non-native English writers and neurodivergent people. At 5% prevalence a 1% false-positive rate still makes about 16% of flags wrong, and one measured rate reached roughly 50% against a vendor's sub-1% claim. Require a second independent human signal before any rejection, and never let a detector score stand alone.

What fields must the audit log actually contain?

Log a mandatory structured reason code plus an evidence field per action, and capture the tool and component, model or version identifier, config parameters, prompt templates or rules, feature categories, data exclusions, explanations produced, confidence signals, and the human reviewer identity and action with a rationale. Store it tamper-evident and access-controlled so silent edits and back-dating are impossible.

Is a four-fifths pass enough to prove no adverse impact?

No. Adverse impact can be present even when the four-fifths ratio holds, particularly in large-volume decisions or when a statistical significance test reveals a pattern the ratio obscures. Run Fisher's exact on high-volume batches and treat a p-value of .05 or less, equivalent to about 1.96 standard deviations, as significant. Always-on agents produce exactly the large batches where the rule of thumb under-warns.

When do EU AI Act obligations bite for a deploying employer?

Multiple primary trackers treat August 2, 2026 as the binding enforcement date for high-risk obligations, including Article 26 deployer duties, though the Act itself took effect August 1, 2024. As a deployer you must implement human oversight by competent, trained, authorized people, retain automated logs for at least six months, and conduct a Fundamental Rights Impact Assessment where required. One staffing page cites a later date, so re-check against the primary trackers.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next