Refolk
FrameworkRecruiting and sourcing

The Candidate Authenticity Score: Real, AI Persona, or Fabricated

You will be able to score any candidate's public footprint on a fixed authenticity rubric and route them to book, verify-further, or reject with a written reason two reviewers would agree on.

16 min readLast reviewed August 27, 2026Read as Markdown

Before you book screening time, you have to answer one question about a promising sourced or inbound candidate: is this a genuine person, or an AI-generated or fabricated identity? This guide gives in-house recruiters, sourcers, and founders doing their own hiring a fixed rubric to score a candidate's public footprint across LinkedIn, GitHub, and the open web at once, then route them to book, verify-further, or reject with a written reason two reviewers would agree on. It is built to be used at the gate, before you spend interview time, and it carries explicit guards so real career-changers and privacy-conscious candidates are not wrongly cut.

This is a different job from two adjacent standards. It is not the operations dedup teardown that asks whether one record is a single human, and it is not the pre-offer authenticity standard for remote engineers at offer stage. This scores fraud risk earlier, at sourcing and screening, across all three surfaces together.

Why this gate exists now

Fake candidates are no longer an edge case, and a passed background check does not close the question. Gartner predicts that by 2028 one in four candidate profiles worldwide will be fake, which turns authenticity from a rare escalation into a standing step in the funnel.

The prevalence numbers disagree on magnitude, and the disagreement is the point. Show both ends rather than averaging them, because the true rate depends on role, channel, and geography.

MetricValueSource
Projected fake profiles by 202825%Gartner
Admitted interview fraud6%Gartner 2Q25
Hiring managers who met a deepfake candidate31%Greenhouse 2025
Hiring managers who met a deepfake candidate17%Resume Genius
Fake-identity applications in one real posting12.5%Pindrop

The Pindrop row is the most concrete: of 827 applications for one back-end developer role, roughly 100 came in on fake identities. Palo Alto Networks says a fake candidate can be assembled in as little as 70 minutes, which is why volume roles see clusters rather than one-offs. And the reason this gate is needed at all: SIA and Sterling research found 63% wrongly believed all background checks include identity verification. They do not. A passed-check note in your applicant tracking system is not evidence the person is who they claim.

25%
Share of candidate profiles Gartner projects will be fake by 2028
A one-in-four base rate means an authenticity read belongs in the standard funnel, not just in escalations.

Why no single check can carry the decision

Both the human eye and the automated detector sit near a coin flip on hard cases, so a defensible verdict only emerges from stacking independent signals. This is the load-bearing principle of the whole rubric: it is additive, never one gate.

MethodAccuracySource
Untrained human on deepfake video55.54%Wiley meta-analysis 2025
Open-source image detectors, out-of-distribution50-60%ScamAI
Enterprises trusting identity verification alone30% will not, by 2026Gartner

Read the table as a warning against confidence. A hiring manager who feels sure a video call is fake is right about 56% of the time. A detector that flags a headshot as AI-generated is right maybe 50 to 60% of the time on a generator it has not seen. Gartner expects that by 2026, 30% of enterprises will stop treating identity verification and authentication solutions as reliable on their own. The rubric answers this by requiring several weak signals to agree before any reject.

The five scoring dimensions and what each proves

Score every candidate on the same five dimensions. Each one states what a clean result proves and what it looks like when it lies, so the score means the same thing across reviewers.

  • Footprint age and depth. A multi-year, consistent presence is hard to fabricate at depth. It proves the identity predates the application. It lies when a farmed account fills a contribution graph to a goal, or when a real privacy-conscious candidate simply has a six-month-old LinkedIn.
  • Cross-platform match. The same person, same face, same work history across LinkedIn, GitHub, and the open web is a strong positive. It proves independent surfaces corroborate. It lies when a fabricator builds all three from one template, so check that the surfaces reference the same specific projects, not just the same name.
  • Contribution depth. For engineers, per-author commit counts and pre-creation-date work prove authored labor. Being listed as a contributor could mean 2,000 commits or one. The contributor list alone will not tell you which, and green squares are fabricable to an arbitrary goal.
  • Photo authenticity. Specific artifacts around ears, teeth, jewelry, hair boundary, and background raise suspicion. This proves little on its own given weak detectors; it lies in both directions, flagging real people who used AI headshot tools and clearing polished fakes.
  • Logistics and contact consistency. Contact details that recur across other identities, or a location that mismatches network and payment signals, are the most durable tell. This proves operational reuse, which the fraud economics require. It rarely produces a false positive because innocent candidates do not share a VoIP number with three other applicants.

The authenticity signal stack, weakest layer on top

  1. Photo and live-video artifacts
    Advisory only; detectors near coin-flip and gesture tests obsolete
  2. Footprint age and cross-platform match
    Corroborating, but thin footprints have innocent causes
  3. Contribution depth
    Per-author commits and pre-creation work resist fabrication
  4. Logistics and reused contact infrastructure
    The durable base; reuse is expensive for fraudsters to hide
Visual tells sit on top because they are the easiest to forge and the hardest to score; operational tells sit at the base because reuse is expensive to hide.

Notice the inversion. The signals people reach for first, the headshot and the video call, are the weakest and the most bias-prone. The signals that actually decide cases sit at the bottom of the stack and take longer to gather. Weight your rubric accordingly.

The eight-step scoring procedure

Run these in order. The first five gather and grade signals; the last three route, test, and escalate. Times assume a single candidate at the gate, not a batch.

Score a candidate from footprint to routing

  1. Intake and de-bias the record
    Pull name, headshot, LinkedIn, GitHub or portfolio, email, and phone into a signal sheet. Strip name-driven assumptions before scoring, because a foreign name or non-Western school is not a signal.
  2. Cross-platform corroboration
    Confirm the same person appears consistently across LinkedIn, GitHub, and the open web, reverse-image-search the headshot, and check whether the phone or email recurs on other identities.
  3. History-depth check
    For engineers, verify GitHub via OAuth where possible and inspect per-author commit counts, the account-creation-to-first-commit gap, and follower patterns rather than the contribution graph.
  4. Photo authenticity pass
    Zoom on ears, teeth, jewelry, hair boundary, and background, and run a detector for a probability score treated as advisory only. Note the specific artifact rather than the verdict.
  5. Score on the rubric
    Assign points across footprint age, cross-platform match, contribution depth, photo, and logistics or contact consistency. Produce a numeric score and a written one-line reason.
  6. Route to book, verify, or reject
    Map score bands to actions and never reject on a single signal. Log the routing decision with a reason two reviewers would agree on.
  7. Run a live authenticity test at interview
    For verify-further candidates, add friction in the call with unscripted questions, a camera-angle change, and ID on camera. Do not rely on the wave test alone.
  8. Escalate confirmed fraud
    For strong fraud signals in remote or IT roles, verify employment and education directly and report confirmed cases to IC3.

How the depth check actually works

Step three is where engineering candidates separate. GitHub join dates cannot be altered, but fixing commit history makes the contribution graph reflect any activity period you like, and public tools exist to fill a graph to an arbitrary goal. So the graph is theater. What resists fabrication is per-author commit counts, which GitHub does not summarize for you, and the gap between account creation and first commit. Farmed accounts tell on themselves: they host no original code, lots of tutorial-level code, an overcompensating README, and no commits before the creation date. The only reliable proof of ownership is GitHub OAuth, which gives cryptographic proof rather than a claim.

Market context changes what an empty GitHub means. Weigh it before you flag.

MarketEngineers listing GitHub skillShare of two-market total
United States11,47337%
India19,74463%

Both counts come from Refolk's index. The India pool is about 1.72x the US pool, which means a "no GitHub footprint" flag is a weaker signal in some markets than others. A flat global rule over-flags by geography, so let the depth check inform the score, not dominate it.

19,744
India-based software engineers listing GitHub as a skill, in Refolk's index
About 1.72x the 11,473 US-based pool, so an absent open-source footprint means different things by market.

Routing: turning a score into an action

Map the rubric score to one of three actions, and log the reason with every routing call. The point of the bands is that two reviewers reading the same signal sheet reach the same action.

  • Book. Consistent multi-year footprint, cross-platform match on specific projects, and clean contact consistency. Nothing suspicious in the depth check. Book screening time and move on.
  • Verify-further. A thin footprint with one innocent explanation, or a single artifact with no corroborating signal, or a market where absent GitHub is expected. Do not reject. Add friction at interview and confirm one hard artifact.
  • Reject. Multiple independent red flags stacking: reused contact details across identities, a fabricated-looking depth profile, and a location or logistics mismatch, with a written reason. Never a single-signal reject.

Routing a candidate on evidence and fraud signal

Multiple stacking fraud signalsWeak or no fraud signal
Verify-further
Add friction and confirm one hard artifact before booking
Reject
Strong signals plus rich footprint means active fabrication; log and escalate
Book
Corroborated identity with nothing suspicious; book screening time
Verify-further
Strong signal but thin data; corroborate before you act, do not reject on absence
Thin or unverifiable footprintRich, corroborated footprint
The horizontal axis is how much you can corroborate; the vertical is how strong the fraud signal is once bias is stripped.

The bottom-right and top-left quadrants are the traps. A rich footprint with stacking fraud signals is a fabricator who invested effort, so it escalates rather than passes. A strong-looking signal on a thin footprint is usually a data gap, not fraud, so it verifies rather than rejects.

Gathering the corroboration in step two is the slow part, because it means finding the same person across LinkedIn, GitHub, and the open web and checking that the surfaces reference the same specific work. This is exactly the retrieval Refolk is built for: you can ask for the corroborated footprint directly instead of stitching it together tab by tab.

How this goes wrong: failure modes and false positives

This is the most valuable part of the standard, because a rubric that over-flags is worse than no rubric. Every failure mode below has a false positive that cuts a real person, and a check that prevents it.

  • Photo detector over-trust. A confident AI-generated score is often wrong on new generators, which score 50-60% accuracy. The false positive is a real person who used an AI headshot tool, now common on LinkedIn. Check: require a second, differently-lit photo or a live camera moment, not the detector verdict.
  • Thin footprint read as fraud. A privacy-conscious candidate or recent immigrant looks identical to a fabricated one. The false positive is a genuine career-changer with a six-month-old LinkedIn. Check: corroborate one hard artifact such as verified employment or an OAuth-confirmed GitHub, rather than counting connections.
  • Contribution-graph theater. Green squares are fabricable to a goal, so a farmed account looks productive while a real developer on private or GitLab repos looks empty. Check: per-author commit counts and pre-creation-date commits, not the graph.
  • Name-based suspicion. Non-Western names and unfamiliar schools trigger a "doesn't line up" read that is bias, not signal. Resumes with Western names get 50% more callbacks at identical qualifications. Check: never let name or school contribute to the score, and audit rejections for name clustering.
  • The hand or turn test as proof. Modern face-swaps pass occlusion, so passing the test gives false confidence, and failing it can just be a bad webcam. Check: use unscripted knowledge probes plus logistics consistency, not one gesture.
  • Confusing AI-assisted with fraudulent. A polished, AI-written resume is not a fabricated identity. Nine in ten HR workers reported a surge in low-effort AI-generated applications, but a strong candidate who used a chatbot to write bullets is not a fraud. Check: score identity authenticity separately from application polish.
  • Reused-contact miss. A VoIP number or email recurring across different applicants is a strong tell that is easy to overlook at volume. Check: deduplicate contact fields across the whole pipeline, not per candidate.
  • Single-reviewer drift. Human-only judgment is inconsistent and bias-prone. Check: require a written reason and a second-reviewer agreement threshold before any reject.
The signals people reach for first, the headshot and the video call, are the weakest and the most bias-prone.

The live authenticity test, and where sources disagree

Only run the live test at interview, and only for verify-further candidates. It adds friction that a fabricated identity struggles with in real time: unscripted questions with no clean answer to prepare, a request to change camera angle, and ID shown on camera. Document anomalies without confronting the candidate mid-call, because a false accusation on a real person is a legal and reputational cost.

Here the evidence splits, and you should know it. Some vendors still promote the wave-your-hand and turn-your-head tests as reliable deepfake detection. Gem and others report the opposite: the simple tests are already obsolete because modern face-swap technology handles the gesture with minimal artifacts. Treat any gesture result as one weak input, never as proof. The durable substitutes are unscripted knowledge probes, watching for lip movement out of sync with speech, absent or unnatural blinking, artifacts at the edges of the face, and delays in expression reactions, always weighed against the possibility of a slow connection or an old webcam.

Verify-further note for the ATS
Candidate: [candidate ID, not full name if pipeline is anonymized]
Rubric score: [number] / band: verify-further
Footprint age/depth: [real / thin / suspicious] - [one specific observation]
Cross-platform match: [corroborated on projects X, Y / inconsistent because ...]
Contribution depth: [per-author commit finding, or N/A for non-engineering role]
Photo: [artifact noted, or clean; detector score advisory only]
Contact/logistics: [phone and email unique to pipeline: yes/no]
Reason to verify rather than book: [one line]
Live test to run at interview: [specific unscripted probe]
Second reviewer: [name] agrees: [yes/no]

Fill the bracketed fields with what you actually observed. Keep it factual and free of any name or school reference.

What to verify before you call the score done

Run this before logging any routing decision. It is the difference between a score two reviewers agree on and one reviewer's gut.

Authenticity score sign-off

  • The signal sheet contains the public footprint with no name-driven or school-driven assumptions baked in.
  • At least one hard artifact corroborates identity: verified employment, an OAuth-confirmed GitHub, or matching cross-platform project references.
  • Contribution depth was judged on per-author commit counts and pre-creation-date work, not on the contribution graph.
  • The photo detector result, if used, is recorded as advisory and is not the basis for any reject.
  • Phone and email were deduplicated across the pipeline to catch reused contact infrastructure.
  • No reject rests on a single signal, and every reject carries a written one-line reason.
  • A second reviewer read the reason and agreed with the routing band.
  • If routed reject, the decision was checked against recent rejections for name or geography clustering.

Keeping the rubric current

The rubric is stable but its inputs decay, so re-check the two moving parts on a schedule rather than trusting yesterday's calibration. First, detector reliability drifts as generators improve; the 50-60% figure on out-of-distribution images tells you to keep detectors advisory permanently, not to chase a better one. Second, prevalence shifts by role and channel, so if a single posting starts returning a Pindrop-style cluster of fake identities, tighten the verify-further band for that channel and log the change with its business reason.

Two controls keep the gate honest over time. Re-review a random sample of low-ranked and rejected candidates to see whether strong profiles are being missed, and document every criteria change and the reason for it. That audit is what protects real career-changers and privacy-conscious candidates from a rubric that quietly drifts toward over-flagging. It is also what lets you defend the gate: in Refolk's index there are 91,813 US-based recruiting and talent-acquisition professionals who apply judgment like this, and the ones whose decisions hold up are the ones who wrote down the reason and had a second reader agree.

The last thing to keep current is your own confidence. Humans read deepfake video at 55.54% and the best open detectors barely clear a coin flip, so the moment the rubric feels like it is giving you certainty from one glance, it has stopped working. Certainty comes from signals agreeing, from a reused number surfacing in a dedup, from a depth check that contradicts a claimed prominence. Stack the evidence, write the reason, and let a second reviewer break the tie.

Questions practitioners ask

How do I detect a fake candidate without wrongly cutting real people?

Score multiple independent signals additively rather than acting on any one. A thin footprint, an AI-looking headshot, or a non-Western name each has an innocent explanation, so a defensible reject needs several red flags stacking plus a written reason a second reviewer agrees with. The durable tells are operational, like a reused VoIP number or email across supposedly different applicants, not visual artifacts.

Are AI image detectors reliable enough to reject a candidate?

No. ScamAI research found leading open-source detectors score as low as 50-60% on out-of-distribution generated images, barely better than a coin flip, and modern generators have removed many obvious tells. Treat a detector verdict as advisory only. If a headshot scores as AI-generated, ask for a second differently-lit photo or a live camera moment rather than rejecting on the score alone, since real people now use AI headshot tools.

What is the strongest single tell of a fabricated candidate?

Reused contact infrastructure. The FBI reports that North Korean IT workers reuse VoIP phone numbers and email addresses across resumes purportedly belonging to different applicants, because the fraud economics require recycling infrastructure. Deduplicating phone and email fields across your pipeline catches identities that pass every visual and depth check. It is easy to overlook at volume, so make it a standing pipeline dedup rather than a per-candidate task.

Does a passed background check mean the candidate's identity is verified?

No, and assuming it does is common: SIA and Sterling research found 63% wrongly believed all background checks include identity verification. Many do not. A passed-check note in your ATS is not evidence the person is who they claim, which is exactly the gap this authenticity gate fills. Confirm identity through a hard artifact such as employment verified directly with the employer or an OAuth-confirmed GitHub account.

Is a candidate with no GitHub or a thin online presence a fraud risk?

Not on its own. A privacy-conscious candidate, a recent immigrant, or a developer who worked in private or GitLab repos looks identical to a fabricated one. Market context matters too: in Refolk's index the India pool of GitHub-visible engineers is about 1.72x the US pool, so an absent open-source footprint means different things by geography. Corroborate one hard artifact instead of counting connections.

Are the wave-your-hand and turn-your-head interview tests still useful?

They are largely obsolete as proof. Modern face-swap technology handles occlusion and head turns with minimal artifacts, so passing the test gives false confidence and failing it can just be a bad webcam. Sources disagree: some vendors still promote gesture tests. Use unscripted knowledge probes and logistics consistency instead, and treat any gesture result as one weak input among several.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next