# The Inbound Triage Standard: From Applicant Flood to Ranked Human-Review Queue

*You will turn a raw inbound pile into a defensibly ranked review queue using public-footprint corroboration, without using AI-text detectors as a screen-out gate.*

- Canonical URL: https://www.refolk.ai/guides/inbound-applicant-triage-standard
- Pillar: Recruiting and sourcing
- Format: Playbook
- Published: 2026-08-26
- Last reviewed: 2026-08-26
- Reading time: 16 min

You have hundreds of inbound applications for one role, most of them touched by AI, and you need to decide which ones a human should actually screen. This guide is for in-house recruiters, sourcers, talent leaders, and founders doing their own hiring. It gives you a lawful, evidence-first method to take a raw batch, corroborate each applicant against their public footprint across GitHub, LinkedIn, and the open web, and produce a ranked queue of who earns a human screen, without leaning on AI-text detectors as a screen-out gate.

Most recruiting guides start from outbound sourcing or from verifying one already-flagged candidate. This one handles the flood: the resume itself has stopped being a usable signal, and public-footprint corroboration is the only cheap way to rank a batch on a Tuesday.

## Why the resume stopped being a usable signal

The resume stopped ranking applicants because the cost of tailoring one collapsed to seconds of compute. When producing a keyword-perfect resume takes no time and no skill, a clean resume is evidence of tool access, not of fit.

The volume numbers make this concrete. LinkedIn now processes about 11,000 job applications per minute, a 45 percent rise year over year, according to data reported by The New York Times. Overall applications rose 31 percent in the first half of 2024 compared with the same period a year earlier. On the supply side of the tooling, a Greenhouse report with more than 4,100 respondents found 74 percent of US job seekers use AI in their application process, and a Canva survey found 45 percent of applicants use AI to complete applications.

**11,000 - Job applications LinkedIn processes per minute**

A 45 percent year-over-year rise, per data reported by The New York Times.

The human cost of this is not hypothetical. One HR consultant received over 1,200 applications for a single remote role, was overwhelmed enough to pull the listing entirely, and spent three months sorting through the submissions. That is the job this guide exists to make survivable.

Corroboration is the only remaining separator. If two resumes read equally well and one applicant's claims check out against public records while the other's do not, the check is the signal. Everything below is built to run that check at batch scale.

## Why AI-text detectors cannot be your gate

An AI-text detector cannot be used to reject applicants, because it fails exactly where you would use it and the legal liability for that failure lands on you, not the vendor. The right move is to record the score as context if you must, and never let it advance or reject anyone by itself.

Detectors fail on short, polished, templated text - which is precisely what a resume bullet is. In the APT-Eval study, GLTR sat at a 6.83 percent false-positive rate on pure human writing but classified 41 to 43 percent of minor-polished GPT-4o text as AI-written. Turnitin's own whitepaper reports a 0.51 percent document-level false-positive rate on academic writing, yet a Washington Post study put it near 50 percent on a small sample. Detector vendors themselves warn against this use: Pangram recommends against screening short bullet lists, outlines, very short responses, and formulaic or template-based writing.

| Detector or study | False-positive rate on human text | Condition |
| --- | --- | --- |
| GLTR (APT-Eval) | 6.83% | pure human text |
| GLTR (APT-Eval) | 41 to 43% | minor-polished GPT-4o text |
| Turnitin (whitepaper) | 0.51% | academic documents |
| Turnitin (Washington Post) | ~50% | small sample |

There is a documented bias against non-native English writers, which turns a detector gate into a discrimination engine. And the liability does not transfer with the tool. EEOC guidance is explicit that employers may bear Title VII responsibility for adverse impact caused by third-party screening tools, and Title VII's disparate-impact provisions remain fully enforceable, so private plaintiffs can still sue. The EEOC's first AI settlement, involving a tool that auto-rejected more than 200 older applicants, shows this is enforced, not theoretical.

> **Watch out:** A detector score is not a disposition
>
> Rejecting on an AI-text score screens out non-native writers and polished-but-real humans, and GLTR flags over 40 percent of lightly-edited text. Require corroborating footprint evidence for every screen-out, and keep the detector score out of the reject decision.

## The order of proof: identity before depth

Verify who controls an account before you assess what that account did. If you assess the work first, you may be crediting a genuine repository to someone who does not own it.

The sequence matters because each check has a prerequisite. Ownership is the gate; depth is the measurement behind it. A resume can list any repository URL without proving anything, so a cited handle is a claim until corroborated. The only reliable confirmation of ownership is an OAuth challenge, which gives cryptographic proof rather than a self-report. Where OAuth is not available at triage, fall back to cross-source consistency: the same handle, photo, and name appearing together across GitHub, LinkedIn, and the open web.

#### The order of proof for one applicant

1. **Ownership** - Prove the applicant controls the cited accounts, via OAuth or cross-source consistency
2. **Depth** - Read pinned repos and per-author commit counts, not the contribution graph
3. **Consistency** - Match employer, title, and location claims across resume, LinkedIn, and the web
4. **Red flags** - Run a light DPRK and identity-fraud pass, escalate any match
5. **Disposition** - Assign a tier and record the one-line evidence basis

*Each stage gates the next, so no work is credited to an unproven account.*

Once ownership is settled, depth is what you measure, and the contribution graph is the wrong place to look. Git honors any GIT_AUTHOR_DATE, so a full two-year streak of green squares can be generated by a short script and back-dated to any year. Open-source tools exist specifically to fabricate this. GitHub only counts commits whose author email matches a verified account email and that land on the default branch, so real merged work is far harder to fake than a graph. Being listed as a contributor is also weak on its own: it could mean 2,000 commits or one typo fix, and the list alone will not tell you which. Practitioners who do this well look at pinned repositories with clean READMEs, live demos, real functionality, sensible commit history, and per-author commit counts. They do not care about the green graph.

> The graph measures willingness to run a script; per-author commit depth measures work that cannot be back-dated.

## Who runs this work, and why the playbook must be simple

This triage falls on generalist recruiters, not a dedicated OSINT team, so the method has to be runnable by one person with a browser and a spreadsheet. Any playbook that assumes specialist tooling will not get run on the pile it was written for.

In Refolk's index of professional profiles, there are 20,470 US profiles with a Technical Recruiter title against just 758 dedicated Sourcer profiles. That is roughly 27 technical recruiters for every dedicated sourcer, which tells you where this work actually lands.

| Role title | Matching profiles | Top employer | Recruiters per sourcer |
| --- | --- | --- | --- |
| Technical Recruiter | 20,470 | Experis | - |
| Sourcer | 758 | Deloitte | 27.0 (derived) |

The scarcity of the target pool changes the stakes of a wrong reject. In Refolk's index there are 603 US profiles with a Software Engineer title and Rust skill, against 87 in Germany - the US pool is about 6.9 times larger. A detector false positive in the German market could eliminate a meaningful fraction of the entire real pipeline in one careless batch. The narrower the pool, the more corroboration has to replace screen-out.

| Country | Matching profiles | Top employer in result | Share vs US |
| --- | --- | --- | --- |
| United States | 603 | Google | 100% (baseline) |
| Germany | 87 | Helsing | 14.4% (derived) |

The friction in generalist triage is finding and cross-matching the footprint at all: locating the real handle, confirming the same person appears across sources, and pulling the commit data. That evidence-gathering layer is exactly where a query-based tool earns its place, because it collapses the search that would otherwise eat your two-to-four-minute claim-extraction budget per applicant.

I ran this search: `Engineers who own the GitHub handle on their resume and have 200+ commits to a Rust project they didn't just fork.` - [see the full result list](https://www.refolk.ai/s/waybr605xn).

*Returns applicants whose cited handle is corroborated across sources and backed by real, non-fork commit depth, so you start triage from proven ownership rather than a claim.*

[Refolk](/) is built for this evidence-gathering layer: ask in plain English and get the corroborated footprint across GitHub, LinkedIn, and the open web, so the recruiter spends their minutes on judgment, not on hunting for a profile.

## The procedure: raw pile to ranked queue

Run these eight steps in order on the frozen batch. The first two build the working table, the middle four gather evidence, and the last two turn evidence into a defensible tier.

#### Inbound triage, start to finish

1. **Freeze the batch and set detectors aside** - Export the pile into one table and mark any AI-text score as non-dispositive context. Done: every applicant has a row with resume claims and a blank evidence column, and no one is rejected yet.
2. **Extract checkable claims per applicant** - Pull named employers, the GitHub or portfolio handle, the LinkedIn URL, and two or three specific technical claims. Done: each applicant has a locatable handle or a note that none exists.
3. **Prove identity ownership before assessing work** - Confirm the applicant controls the cited accounts, via OAuth where possible or cross-source handle, photo, and name consistency. Done: each account is marked ownership proven, claimed only, or no footprint.
4. **Assess depth, not decoration, on GitHub** - Ignore the graph, open pinned repos, read READMEs, and check commit counts per author against the claimed handle. Done: each technical applicant is scored on real contribution depth.
5. **Cross-check resume, LinkedIn, and web consistency** - Verify employer tenures, titles, and location claims match across sources, and note contradictions. Done: each applicant flagged consistent, minor-discrepancy, or material-contradiction.
6. **Run the fraud and nation-state red-flag pass** - Check for mismatched or shifting contact info, device-shipping requests, image-edited ID photos, and location-to-activity mismatches. Done: each applicant cleared or escalated to security and legal.
7. **Assign a disposition tier and record the reason** - Sort into advance-to-human-screen, unverifiable-hold, or do-not-advance, and write the one-line evidence basis. Done: every applicant has a tier and an auditable reason.
8. **Calibrate with a second reviewer** - Have a second person re-grade a random 10 percent sample to check tiers agree. Done: disagreement rate measured and rubric adjusted.

The batch narrows sharply as you move through it, and naming the stages helps you see where volume disappears and why.

#### How a batch narrows across triage

| Stage | Figure | Note |
| --- | --- | --- |
| Raw inbound | 1200 | The pile a single remote role can draw |
| Locatable footprint | 600 | Applicants with at least one findable public handle |
| Ownership proven | 300 | Accounts corroborated by OAuth or cross-source consistency |
| Advance to human screen | 60 | Consistent claims plus real contribution depth |

*Volumes are illustrative of one flooded remote role and narrow at each gate.*

## A disposition rubric you can defend

No regulator or industry body publishes named advance, hold, or reject cut-scores for footprint triage, so the rubric below is a proposed standard, not a cited authority. Its value is that every tier rests on a corroboration fact a second reviewer can audit, not on a black-box score.

The three tiers map to what the evidence actually established:

- **Advance to human screen.** Identity ownership proven, contribution depth real, and claims consistent across resume, LinkedIn, and the web. Minor discrepancies are allowed if the core claims hold.
- **Unverifiable hold.** No locatable footprint, or a handle that is claimed only and could not be corroborated. These are not rejects; the applicant may be private or early-career. They wait behind proven candidates and get a human screen only if the advance tier runs dry.
- **Do not advance.** A material contradiction between claims and public records, a fabricated contribution graph passed off as work, or a fraud red flag. Every entry here carries a written reason.

The load-bearing inputs are the ones you gathered: ownership proven versus claimed, per-author commit depth versus decoration, cross-source consistency, and the absence of fraud markers. Write the one-line reason for every applicant regardless of tier, because the reason is what survives an audit.

**Disposition row for one applicant**

```
Applicant: <name or ID>
Handle ownership: proven | claimed-only | none
Contribution depth: real (N commits, default branch) | decoration | n/a
Cross-source consistency: consistent | minor-discrepancy | material-contradiction
Red-flag pass: cleared | escalated (reason)
Tier: advance | hold | do-not-advance
Reason (one line): <the single fact that decided the tier>
```

*One row per applicant. Fill every field; the reason line is what makes the queue defensible.*

> **Rule:** The reject stays a documented human judgment
>
> Automate the evidence-gathering - locating profiles, pulling commit data, checking cross-source consistency. Never automate the screen-out decision. The employer retains Title VII liability for adverse impact even from a vendor's model, so every do-not-advance must carry a human-written reason.

## How this goes wrong

The failure modes below are where a triage queue quietly becomes indefensible or lets a fraud through. Each one has a false positive it produces and a specific check that catches it.

- **Detector-as-gate.** Rejecting on an AI-text score screens out non-native writers and polished-but-real humans, and GLTR flags over 40 percent of lightly-edited text. Check: never auto-reject on a score; require corroborating footprint evidence.
- **Green-graph trust.** A full contribution graph is decoration, because GIT_AUTHOR_DATE lets anyone back-date squares. The false positive is reading "committed daily for two years" as real work. Check: read commit content and per-author counts, not the graph.
- **Contributor-list inflation.** "Contributor to X" can mean a single typo fix. Check: confirm the commit count for the claimed handle against author data.
- **Handle spoofing.** A resume can list any repository URL without proving ownership. Check: OAuth or cross-source handle, photo, and name consistency before crediting any work.
- **Consistency theater.** AI-tailored resumes pass ATS keyword matching with plausible metrics, so a perfect keyword match reads as strong fit when it proves nothing. Check: verify tenure and title against LinkedIn and the open web, not against keywords.
- **Missing DPRK screen.** A strong-looking remote engineer may request device shipping to a mismatched address. Check: flag shipping-address and location-to-IP mismatches and route to security before offer, not after.
- **Over-automation.** Automating the screen-out decision, rather than only the evidence gathering, creates Title VII exposure the vendor will not absorb. Check: keep the reject a documented human judgment.
- **No calibration.** Two reviewers grading differently makes the queue indefensible. Check: a second-reviewer re-grade of a 10 percent sample, with the disagreement rate measured.

The fraud markers deserve their own emphasis because the cost of missing them is not a bad hire, it is legal exposure. Gartner predicts that by 2028, one in four candidate profiles worldwide will be fake, and in a Gartner survey of 3,000 candidates, 6 percent admitted to interview fraud. HireRight's 2025 benchmark found one in six employers has already experienced hiring identity fraud. The DPRK advisories give the canonical marker set: incorrect or frequently changing contact information, requests to ship company-issued devices to addresses not on identification documents, and stolen, altered, or image-edited identity documents. A common onboarding tell is a sudden family emergency and a request to send the laptop to an address that does not match HR records, usually a laptop farm. Paying such workers risks sanctions designation under US and UN authorities, which is why a match escalates to security and legal rather than to a rejection email.

> **Note:** The order of the fraud pass is contested
>
> Some practitioners place deep identity verification at the assessment stage rather than at triage. Run a light red-flag pass here to catch obvious tells, and reserve the deep verification for before any offer. Do not let the debate over ordering become a reason to skip the light pass entirely.

## Before you call the queue done

Verify the queue against this checklist before you hand it to hiring managers. Each item is a thing that has to be true, not a topic to think about.

#### Queue readiness

- [ ] Every applicant has a disposition tier and a one-line evidence reason a second reviewer could audit.
- [ ] No applicant was rejected on an AI-text detector score alone.
- [ ] Every credited GitHub contribution rests on a proven-ownership account, not a claimed handle.
- [ ] Contribution depth was judged on per-author commit counts on the default branch, not on the contribution graph.
- [ ] Employer, title, and location claims were checked against LinkedIn and the open web, not against keyword matches.
- [ ] Every applicant passed a light DPRK and identity-fraud red-flag pass, with any match escalated to security and legal.
- [ ] A second reviewer re-graded a 10 percent sample and the disagreement rate is low enough to trust the queue.
- [ ] The screen-out decisions are human judgments with written reasons, not automated outputs.

## Keeping the standard current

Re-check the parts of this method that move, because the ground shifts under it. Two things change fastest: what forgery is cheap, and what the law expects.

On forgery, assume any signal a script can generate will be generated. The contribution graph is already there; the next thing to fall will be whatever check you rely on that a tool can automate. Re-test your own checks periodically by asking whether a determined applicant with public tools could pass them without doing the work. If the answer is yes, move the check deeper - from the graph to commit depth, from commit depth to merged pull requests on a default branch under a verified email.

On the legal side, treat the automate-versus-human line as the durable boundary even as specific guidance evolves. The mechanism is stable: the employer owns adverse impact from screen-out, so the reject stays human and reasoned. When new EEOC guidance or a new settlement lands, re-read your own do-not-advance reasons and confirm each one still rests on a corroborated fact rather than a score.

Finally, re-run the calibration step on every new role, not just the first. A rubric that agreed at 90 percent on backend engineers can fall apart on a design role or a sales role where the public footprint means something different. The disagreement rate on a fresh 10 percent sample is your early warning that the standard needs adjusting for the market in front of you.

## Frequently asked questions

### Can I just reject every resume an AI detector flags?

No. AI-text detectors fail exactly where recruiters use them, on short and templated resume bullets. GLTR runs at 6.83 percent false positives on pure human text but flags 40 to 43 percent of lightly-polished text, and one study put Turnitin near 50 percent on a small sample. Rejecting on that score screens out non-native writers and polished-but-real humans, and the employer, not the vendor, owns the resulting adverse impact under Title VII.

### Is a full green GitHub contribution graph proof someone codes?

No. Git honors any GIT_AUTHOR_DATE, so a contribution graph can be back-dated to any year with a short script. A full streak measures willingness to run that script, not real work. Ignore the graph and instead read commit content and per-author commit counts on real repositories, and confirm the commits landed on the default branch under a verified account email.

### How do I confirm an applicant actually owns the GitHub account on their resume?

Prove ownership before you assess any work. The only reliable confirmation is an OAuth challenge, which gives cryptographic proof rather than a claim. Where OAuth is not available at triage, use cross-source consistency: does the same handle, photo, and name appear together across GitHub, LinkedIn, and the open web? A resume can list any repo URL, so treat an unproven handle as claimed only until corroborated.

### What are the red flags for a fake remote engineer applicant?

The DPRK IT-worker advisories give the canonical set: incorrect or frequently changing contact information, requests to ship company devices to an address not on identification documents, stolen or image-edited ID photos, and a stated location that does not match public activity or IP. A common onboarding tell is a sudden family emergency and a request to send the laptop to a mismatched address. Route any match to security and legal, because paying such workers carries OFAC sanctions exposure.

### How long should triaging one applicant against their public footprint take?

A benchmarked minutes-per-applicant figure is not established publicly, so treat these as working estimates. In practice, claim extraction runs two to four minutes and GitHub depth assessment three to six minutes per technical applicant, on top of a roughly 15-minute batch setup. Most applicants resolve faster once you find a proven handle or confirm no footprint exists, so budget time to the deeper checks and keep the light ones quick.

### Where is the line between automating triage and creating legal risk?

Automate the evidence-gathering layer: locating profiles, pulling commit data, checking cross-source consistency. Keep the screen-out decision a documented human judgment. EEOC guidance and the iTutorGroup settlement, where a tool auto-rejected more than 200 older applicants, show the employer retains Title VII liability for adverse impact even from a third-party model. The vendor will not absorb that, so the reject must stay human and reasoned.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/inbound-applicant-triage-standard*
