1,100 Companies, 22 Personas: GitHub Tells That Beat the DPRK Funnel
A pre-interview checklist of GitHub, resume, and metadata signals that catch DPRK IT-worker personas before a hiring manager books the call.
If you run inbound for an engineering team, the DPRK IT-worker problem is no longer a SOC team's story. One PurpleDelta cluster alone hit 1,100+ companies, and the tells that would have stopped most of those hires live on the resume, the GitHub profile, and the application metadata, not on the Zoom call.
Recruiters are now the first line of defense, and the screen is cheaper than people think.
Why recruiters, not security teams, catch this first
The hiring funnel sees the artifacts before the laptop ever ships. By the time Huntress is finding a PiKVM on a corporate desk, three people (sourcer, recruiter, hiring manager) have already waved the application through on signals that were visible pre-interview.
The 2026 reporting makes the volume clear:
- Recorded Future's Insikt Group tracked a single PurpleDelta cluster that applied to 1,100+ companies between late 2024 and early 2025, concentrated in software, staffing, and healthcare/biotech.
- Insikt counted at least 22 fabricated personas in the clusters, with operators sending as many as 60 applications per day across 8+ job platforms, including LinkedIn and Upwork.
- Insikt assesses PurpleDelta operators are highly likely employed at 10+ organizations right now.
- Huntress's August 26, 2026 investigation confirmed five likely-DPRK workers hired at partner orgs in 2026, spanning IT, sales/marketing, and healthcare.
Do the math on 1,100 companies divided across roughly 22 personas and you get ~50 companies touched per persona, from one cluster. That is a volume game, and volume games are won in the top of funnel.
The niche-market math nobody is doing
Niche talent pools are structurally the most exposed, because the DPRK application rate is constant but the real haystack is small. If you hire Solidity engineers in the US, a single PurpleDelta persona could blanket your entire verifiable candidate pool in under two days.
Here is the shape of the US haystack a persona hides inside, from Refolk's index of professional profiles.
| Segment (US) | Count | Source |
|---|---|---|
| Senior/Entry SWE, Blockchain or Full-stack w/ React + Node.js | 41,824 | Refolk's index |
| Blockchain / Smart Contract / Web3 Engineer w/ Solidity | 83 | Refolk's index |
| Full Stack Developer, narrow "remote React Node" query | 3 | Refolk's index |
| Top employer inside the Solidity segment | Polygon Labs (2 of 83) | Refolk's index |
| PurpleDelta applications per persona per day | 60 | Recorded Future |
| PurpleDelta companies hit by one cluster | 1,100+ | Recorded Future |
The Solidity niche is roughly 0.2% of the broader React/Node senior-eng pool (83 of 41,824). Any Web3 recruiter screening inbound is statistically likely to see at least one PurpleDelta app per week, and the ratio of DPRK personas to real humans gets uglier the tighter the stack.
The practical move is to start every search from the verifiable pool and work toward inbound, not the other way around. That is the exact gap Refolk closes: describe the person in plain English ("US-based Solidity engineers with on-chain commits in the last 12 months, excluding anonymous handles") and get a ranked shortlist pulled from GitHub, LinkedIn, and the open web, so inbound noise is matched against a known universe instead of being the universe.
The GitHub tells, in order of signal strength
The strongest GitHub signal is not a quiet contribution graph. It is cross-persona co-authorship, where two "different" developers keep showing up on each other's commits.
Nisos's investigations give sourcers a concrete pattern library. Treat the following as a tiered checklist:
- Co-authored commits that cross-link personas.
nickdev0118andAnacondaDev0120kept appearing on each other's work. Click anyCo-authored-by:line and search the email. If the matching GitHub profile is also a stock-photo avatar with a Bootstrap portfolio site, you have a network, not two coincidences. - Template repos with identical scaffolding. Nisos flagged
imcode65-style repos: same README structure, samesrc/tree, same commit cadence, different "author." - Avatar and handle motifs. One documented network used lion avatars across the ring. Handles share tokens like
dev,century, or116. Weak alone, strong in combination. - Portfolio sites off the same Bootstrap template. View source on three "independent" candidate portfolios and you will often find the same grid component, same font stack, same hero copy with the name swapped.
- A GitHub link that doesn't resolve on inspection. In Nisos's June 2025 "Lead AI Architect" case, the resume claimed a portfolio that simply did not exist. Click the links.
- A healthy graph with no issue history, no PRs to other repos, no stars received. DPRK accounts tend to be monologues. Real engineers argue in issue threads.
Reference accounts documented publicly by Nisos, safe to search as pattern libraries: AnacondaDev0120, nickdev0118, superredstar, imcode65.
Why the "GitHub-first" bias backfires
Hiring managers love "has a GitHub." DPRK operators know this and build toward it. A tidy contribution graph is a lower-quality signal than a messy, argumentative one that includes PRs into other people's repos, issue comments going back years, and a star history that matches a real human's taste over time.
A tidy contribution graph is weaker evidence than a messy, argumentative one with years of PRs into other people's repos.
The resume and application-metadata tells
The highest-ROI pre-interview check is a diff between the candidate's resume skills section and your job description. If they match verbatim, you are looking at JD mirroring, and almost no ATS runs this check.
The Nisos June 2025 case cracked exactly this way: the "Lead AI Architect" resume had copy-pasted the skills list from the posting. Add these to your intake screen:
- JD mirroring: diff the skills section against the posting. 80%+ overlap in exact phrasing is a flag.
- Reused resume content across identities: paste a distinctive two-line sentence into Google in quotes. Nisos repeatedly found the same bullet copy across multiple personas.
- Address, license, and bank-state mismatch: the Ballena Azul first hire ("Angelo Espree") claimed Pasadena, TX with a California license and a New York bank. Any one of those is normal. All three together is not.
- EXIF stripped from "iPhone 15" photos: real phone pics carry GPS and camera metadata. ID photos that have been through an AI pipeline usually don't.
- SynthID watermark on ID or document images: Google's Gemini "Verify AI-generated images" tool will flag Gemini-made or Gemini-edited media. The Ballena Azul driver's-license image hit on this check.
SynthID is useful, and dangerously asymmetric
Gemini's SynthID detector only flags Google-AI-made or Google-AI-edited media. The moment an operator switches to Midjourney, Flux, or a self-hosted Stable Diffusion, the watermark check comes back clean. A negative is not a pass. Pair it with EXIF, reverse image search on the avatar, and the address/license/bank triangulation above.
The infrastructure signals, if the application makes it past the resume
If the candidate is already in interview, a handful of logistical questions surface the laptop-farm and KVM-over-IP layer that Huntress keeps finding.
The infrastructure tells documented by Huntress and Recorded Future:
- PiKVM, TinyPilot, or Guermok USB KVM devices on the shipping address request, or any pushback on where the laptop ships ("ship to my cousin in Pasadena").
- AstrillVPN or IPRoyal residential proxy egress on first login. Your IT team can look for these.
- SendGB URLs used to receive a profile photo. Huntress traced a stolen GitHub avatar back through SendGB.
- AnyDesk installed on day one, per Recorded Future, as the account-renting layer that lets the actual operator take over the desk.
- Monikers to recognize for the same actors across vendor reports: FAMOUS CHOLLIMA, Jasper Sleet, Nickel Tapestry, PurpleDelta (ex-TAG-121), UNC5267, Wagemole.
The scheme has outgrown "IT worker"
The GitHub-centric playbook does not help a recruiter screening a BDR. Huntress's 2026 cases include sales, marketing, and healthcare hires, so sourcing teams outside engineering now need a version of this screen too.
For non-engineering roles, the GitHub section of the checklist falls away but the application-metadata section gets more important:
- Resume reuse across LinkedIn personas (paste a distinctive bullet into Google in quotes).
- Address and ID/bank-state triangulation (the Ballena Azul mismatch pattern).
- SynthID + EXIF on any submitted ID or selfie, with the asymmetry caveat above.
- Shipping-address pushback for the laptop (the single biggest tell at onboarding).
- A LinkedIn profile with few connections, no mutuals with anyone at listed employers, and endorsements only from recently created accounts.
The DOJ sentenced two US-based laptop-farm facilitators in April 2026 tied to placements at 100+ companies, so a US shipping address does not clear a candidate. The facilitator layer is a US domestic problem now.
A 10-minute pre-interview screen
The whole pre-interview screen fits in about ten minutes and catches most personas before a hiring manager ever books a call. Give it to every inbound engineering application, regardless of referral source.
The screen, in order:
- Diff the resume skills section against the JD. >80% exact-phrase overlap is a flag.
- Google a distinctive two-line sentence from the resume in quotes. Hits on other personas end the process.
- Open the GitHub link. If 404, end the process. If live, open three random commits and search any
Co-authored-by:email. - Scan the contribution graph for PRs into other people's repos and issue comments pre-2024. Monologue graphs are a flag.
- Reverse-image-search the avatar and the LinkedIn photo. SendGB and stock-photo hits are flags.
- Run any submitted ID photo through Gemini's "Verify AI-generated images" tool. Positive = stop. Negative is not a pass.
- Check EXIF on submitted photos. Stripped metadata on claimed phone shots is a flag.
- Cross-check claimed city, ID-issuing state, and bank-of-record state. Three-way mismatch is a flag.
- Confirm LinkedIn connections overlap with claimed former employers. Zero mutuals across 3+ claimed jobs is a flag.
- Before scheduling, source the role from the verifiable pool. If your inbound candidate doesn't appear in a search of real humans with the claimed stack, that is information. Refolk is built for exactly this reverse check: describe the person you'd expect to see and compare the ranked results against the inbound name.
None of these steps require a security team, a new vendor contract, or a tool your ATS doesn't already have access to. The reason they work is that the DPRK funnel is a volume operation, and volume operations fail on specific checks a human can run in minutes. The two most expensive moments in a bad DPRK hire (the interview loop and onboarding) are both avoidable with a pre-interview screen that treats "has a GitHub" as the start of the investigation, not the end.
FAQ
What is the single highest-ROI check against DPRK candidate fraud?
Diff the candidate's resume skills section against your job description. The Nisos June 2025 "Lead AI Architect" case cracked because the skills list was lifted verbatim from the posting, and almost no ATS runs this check today. If you add one thing to intake this quarter, add a JD-mirroring diff with an 80% exact-phrase threshold.
Does Google's SynthID watermark detector reliably catch fake IDs?
No, and treating it as a pass/fail test is dangerous. Gemini's "Verify AI-generated images" tool only flags media made or edited by Google AI, so a candidate using Midjourney, Flux, or self-hosted Stable Diffusion will come back clean. Use it as a cheap positive signal (as in the Ballena Azul driver's-license hit) and pair it with EXIF checks, reverse image search, and address/ID/bank-state triangulation.
Are only engineering roles exposed?
Not anymore. Huntress's August 2026 report confirmed likely-DPRK hires across IT, sales/marketing, and healthcare at partner organizations in 2026, so BDR, CSM, and ops funnels now need a version of this screen too. The GitHub-specific checks fall away, but resume reuse, address mismatch, laptop-shipping pushback, and LinkedIn connection anomalies carry over directly.
How exposed is a small Web3 team compared to a generalist SaaS hirer?
Structurally much more exposed. The US Solidity pool is roughly 0.2% of the broader senior React/Node pool in Refolk's index (83 vs 41,824), while PurpleDelta operators fire up to 60 applications per persona per day. In a market that thin, a single persona can blanket the entire verifiable Web3 community in under two days, which means every Web3 recruiter should assume at least one DPRK-linked application per week and screen accordingly.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.