The Lookalike Shortlist Playbook: One Exemplar to 20-30 Ranked Comparables
You can decompose one exemplar into separable attributes, run them across GitHub, LinkedIn, and the open web, and produce a ranked shortlist of 20 to 30 that is not just one person cloned.
Your hiring manager points at one person and says "find me more like them." You now have to turn a single profile into a real shortlist. This playbook is for in-house recruiters, sourcers, and founders doing their own hiring, and it delivers the portable method underneath every vendor's "Similar Profiles" button: how to decompose one exemplar into separable attributes, run them across GitHub, LinkedIn, and the open web, rank the results, and stress-test the list so you are not just cloning one person's accidental biography.
The vendor button is seductive because it is one click. It is also a trap. LinkedIn Recruiter's Similar Profiles is restricted to Recruiter Corporate and Recruiter Professional Services seats, its Recommended Matches module shows up to 25 candidates from one seed, and it factors in your own save, message, and hide actions, which means it can faithfully reproduce whatever bias you already have. The method here out-scales that ceiling and builds in the step the vendor features skip entirely: testing the shortlist for homogeneity before you call it done.
What "similar profiles" features actually match on
Documented lookalike features match on a small, concrete set of attributes: job titles, skills, locations, industries, and companies. That is it. LinkedIn Recruiter's own feature page lists exactly those inputs, and its Similar Profiles help documentation describes a model that uses experience, job function, and social graph data. The marketing claims this cannot be replicated in a standard search, but that is a claim about convenience, not capability. Every input it uses is something you can query yourself.
Two properties of the vendor button matter for your strategy. First, it caps output. The legacy feature could generate up to 99 profiles from one seed, but the current Recommended Matches module shows up to 25. Second, it personalises to your behaviour. It considers all your hiring activities, including when you save, message, or hide a candidate. That personalisation is the exact mechanism by which a black-box tool launders your own preferences into "the algorithm's" recommendation.
The portable alternative is older than any button and more transparent. You take the exemplar apart into separable attribute groups, write one query per group, run those queries across more than one source, and rank what comes back. Nothing is hidden, nothing is capped at 25, and nothing is quietly learning from your hide button.
Decompose the exemplar into separable attributes
The first move is to stop treating the exemplar as one profile and start treating it as a bundle of independent attributes, each of which could be its own search. The load-bearing practice from Boolean sourcing is to split into separable attribute groups and OR the synonyms within each group. A title cluster becomes "Software Engineer" OR "Backend Engineer" OR "Developer." A skill cluster becomes "microservices" OR "distributed systems" OR "event-driven architecture." A domain becomes "fintech" OR "banking" OR "payments."
Write these down on an attribute sheet. The discipline is that each line must be answerable as a standalone query. If a line reads "works at the exemplar's current company," that is not a generalisable attribute; it is a pointer at one org chart. Here are the groups to extract from any exemplar:
- Title cluster. The job title and its reasonable synonyms, not the exact string on their badge.
- Skill cluster. The two or three skills that actually define the work, with synonyms OR'd together.
- Scope and seniority. Years of experience, team size led, whether they are an individual contributor or a manager.
- Domain. The industry or problem space, with adjacent domains included.
- Stack and repos (engineers only). The specific technologies they ship, and the repositories they contribute to.
The exemplar, decomposed
- Title clusterthe role and its synonyms, OR'd together
- Skill clusterthe two or three defining skills plus synonyms
- Scope and seniorityyears, team size, IC versus manager
- Domainthe industry and adjacent problem spaces
- Stack and reposthe technologies shipped and repos contributed to
Sources disagree on sequencing. Some practitioners open LinkedIn's Similar Profiles box first and reverse-engineer the attributes from what it returns; others decompose first and never touch the button. Either works, but if you start from the button, write the attribute sheet anyway. Without it you have no way to rank, and no way to run the homogeneity check at the end.
Tag each attribute: overfit versus generalise
Some attributes describe the person's skill and some describe the accidents of their career. The whole quality of your shortlist turns on telling them apart. Exact employer, exact title, and specific school are overfit signals: they describe where this one person happened to land, not what makes them good. Skill clusters, career trajectory, and scope are generalisable: they describe capability that another person could also have.
The clearest evidence that exact-employer and exact-title attributes overfit comes from GitHub. The contributors to a tool your team uses are not the people who listed that tool as a skill on LinkedIn. If you match on the exact title string, you miss them entirely. There is also a subtler trap in title-matching. In both of Refolk's Rust samples, Co-Founder, CTO, and CEO titles top the list, which means a naive title-match on a Rust exemplar drifts toward founders rather than individual contributors. That is a biography artifact of who writes Rust at small companies, not a skill signal about your exemplar.
A formal, published overfit-versus-generalise taxonomy for exemplar sourcing is not established publicly. The labels above are reasoned from documented practice, so treat them as a default you can override when you have reason to. If the exemplar's exact employer is genuinely rare and relevant, for example a single company that pioneered a technique, an employer filter may be worth keeping. The point is to make that a deliberate choice, not a silent default.
Run one query per group across GitHub, LinkedIn, and the open web
Run each generalisable attribute group as its own query, on the source best suited to it, then pool the results. Boolean search on LinkedIn and the open web handles function and domain. The GitHub Contributors tab plus language, location, and pushed qualifiers handles stack and proof-of-work. The goal of this stage is coverage, not precision: you want a raw pool of 60 to 150 candidates, with the source noted on every row so you can later check overlap.
For technical roles, GitHub is the highest-signal source and often the only one that works. Go directly to a repository your team uses or admires, click the Contributors tab, and you have an instant list of developers with hands-on experience with that exact technology. GitHub's user search supports qualifiers like language:javascript location:russia to find users whose repositories are majority a given language in a given place, and the pushed:>2025-01-01 qualifier filters for recently active committers. For a niche or emerging stack like Rust, Go, or Elixir, GitHub is often the only place where you can verify hands-on experience at all.
Proof-of-work has thresholds you can apply. Practitioner guides converge on a handful of community-validation numbers:
| Signal | Community-validation threshold | What it indicates |
|---|---|---|
| Stars on owned repo | 100+ | code valued by community |
| Forks | 20+ | others building on the work |
| Followers | 50+ | recognised domain expertise |
Use these as filters, not as gospel. High stars from bootcamp clones or tutorial forks are vanity metrics, not original work. Filter to original "Sources" rather than forks, and read the pinned repositories, since developers can pin up to six repos as their best work. A real contributor's pinned READMEs and recent commit dates tell you more than a raw star count.
The rarity of the anchor skill decides how aggressively you work each source. Refolk's index makes the point concrete.
| Anchor skill | Country | Matching profiles | Ratio vs smaller market |
|---|---|---|---|
| Rust | United States | 3,587 | 3.49x Germany |
| Rust | Germany | 1,028 | baseline |
| Kubernetes | United States | 143,714 | 40.1x US Rust |
Refolk runs that query across the public GitHub graph, public LinkedIn records, and the open web in one pass, which is the practical way to do the multi-source step without manually pivoting between a Contributors tab and a Boolean string. The index figures above are what let you calibrate before you start: a 40 to 1 difference between Kubernetes and Rust tells you, in advance, whether you are about to widen or narrow.
Widen or narrow to a workable sample
Adjust the pool by changing Boolean operators: add AND to narrow, add OR to widen, add NOT to exclude. This is the most mechanical stage and the one sourcers most often get stuck on, because the failure is silent. The target here is a deduped pool of roughly 40 to 60 that you can actually rank.
| Operator | Effect on result set | Use when |
|---|---|---|
| AND | narrows | sample too broad |
| OR | broadens | sample too thin |
| NOT / - | excludes | removing false positives |
If the sample is too thin, add synonyms with OR, use the asterisk wildcard so that financ* matches finance, financial, and financing, and remove a strict operator. For a rare anchor like Rust, you widen aggressively and lean on GitHub, because the LinkedIn pool simply is not large enough to narrow. If the sample is too broad, add AND terms and limiters: geography, company, seniority. For a common anchor like Kubernetes, with its 143,714 US profiles, you narrow hard or you drown.
Start broad, analyse the outcomes, and adjust filters to improve quality. Do not try to write the perfect query on the first attempt. The widening and narrowing is the search; the first query is just the opening bid.
The procedure, start to finish
This is the full method in order, with owner and rough duration per stage. It matches the attribute sheet you built and the sources you ran. Expect the whole pass to take a focused half-day for a first run on a new role.
One exemplar to a ranked, tested shortlist
- Extract the exemplar into attributesPull the one great profile apart into title cluster, skill cluster, scope and seniority, domain, and specific repos and stack. Done = a written attribute sheet where each line could be its own query. Recruiter or sourcer, 30 to 45 min.
- Tag each attribute overfit vs generaliseMark exact employer, exact title, and specific school as overfit; mark skill clusters, trajectory, and scope as generalisable. Done = each attribute labelled, overfit ones demoted to optional filters. Recruiter, 15 min.
- Run one query per generalisable group across sourcesBoolean on LinkedIn for function and domain; GitHub Contributors tab plus language, location, and pushed for stack. Done = a raw pool of 60 to 150 with source noted per row. Sourcer, 60 to 90 min.
- Widen or narrow to a workable sampleToo thin, add OR synonyms, wildcards, drop a strict filter. Too broad, add AND terms and limiters. Done = a deduped pool of roughly 40 to 60. Sourcer, 30 min.
- Rank against the attribute sheetScore each candidate on how many generalisable groups they hit, weighting proof-of-work over self-reported skills. Done = an ordered list of 20 to 30. Recruiter, 45 min.
- Run the homogeneity checkInspect spread of employer, school, region, and career path across the top 30. If any single value dominates, the list has collapsed toward the exemplar's biography. Done = a documented spread with no single attribute over-concentrated. Recruiter, 20 min.
- Remediate and re-seedIf homogeneous, re-run sourcing and ranking with a second exemplar or drop the overfit filters that caused clustering. Done = final ranked shortlist of 20 to 30 with diversity of path confirmed. Recruiter, 20 to 30 min.
For the ranking step, keep the scoring simple and auditable. A candidate who hits four of five generalisable groups and has verifiable GitHub contributions outranks one who hits five groups entirely by self-reported skills. Use this rubric and write the score next to each row so you can defend the order to the hiring manager.
Skill-cluster match (proof-of-work verified): 0-3 Skill-cluster match (self-reported only): 0-1 Scope and seniority match: 0-2 Domain match: 0-2 Career-trajectory shape resembles exemplar: 0-2 Overfit filters (employer, school): note only, do not add to score Total: ___ / 10
Score each candidate 0 to 10; adjust group weights to match the role's real anchor.
The homogeneity check: the step the vendor button skips
Before you call the list done, audit it for over-concentration. If any single employer, school, region, or career path dominates the top 30, the shortlist has collapsed toward the exemplar's biography rather than their capability. This is the step that separates a defensible shortlist from an accidental clone, and no vendor button does it for you.
The mechanism is well documented even if a single named test is not. If interviewers use themselves as a template and implicitly look for an exact fit, a vicious circle of homogeneity inevitably arises. Cloning is recruiting individuals who are similar to existing employees, which leads to a lack of diversity. On the algorithmic side, if every actor uses the same deterministic system, outcomes are necessarily homogeneous: an individual is accepted everywhere or rejected everywhere. That last point is why running at least one independent source is not optional polish. It is the mechanism that breaks systemic sameness.
Geography concentrates faster than you expect. In Refolk's Rust-Germany sample, Berlin alone accounted for 8 of 25 profiles. Talent for rare stacks clusters in one metro, so a lookalike list inherits that geographic monoculture unless you explicitly audit region spread. Run the same inspection across employer, school, and career path.
number: 8 of 25
label: Berlin's share of Refolk's Rust-Germany sample
note: A single-city cluster this large shows how quickly a rare-stack shortlist inherits geographic monoculture.
Questions practitioners ask
How many examples do I need before 'find more like them' is reliable?
No source publishes a threshold number. LinkedIn only notes that the count returned depends on the uniqueness of the profile and the size of related social graphs. Treat one seed as a starting point, not an answer: a single profile encodes one accidental biography, and the bias literature shows that cloning from a single template drives homogeneity. Re-seeding with a second exemplar is reasoned best practice, especially if your first shortlist clusters on one employer, school, or city.
Can't I just use LinkedIn Recruiter's Similar Profiles button?
You can, but it has two hard limits. It is restricted to Recruiter Corporate and Recruiter Professional Services seats, and its Recommended Matches module shows up to 25 similar candidates from one seed. It is also a black box that learns from your save, message, and hide actions, so it can reproduce your own bias. Use it as one input, then run at least one independent source like GitHub so you are not trusting a single model's judgement.
How do I avoid cloning my best employee's accidental biography?
Run the homogeneity check. After ranking, audit the spread of employer, school, region, and career path across your top 30. If any single value dominates, for example one metro supplying most of the list, the shortlist has collapsed toward the exemplar's biography rather than their skill. In Refolk's Rust-Germany sample, Berlin alone was 8 of 25 profiles, which shows how fast geographic concentration happens for a rare stack. Remediate by dropping the overfit filter or re-seeding.
What's the difference between 'People also viewed' and real similar-profile matching?
Co-viewing is browsing correlation, not similarity. Two people appear in 'People also viewed' because recruiters looked at both, often because both are interesting, not because they are alike. Documented similar-profile features and your own attribute queries match on job titles, skills, locations, industries, and companies. For lookalike sourcing, build from attributes and proof-of-work, and treat the co-view box as noise.
How does the rarity of the anchor skill change my approach?
Pool size decides your whole strategy. Refolk's index returns 143,714 US profiles with Kubernetes skills versus 3,587 with Rust, roughly 40 to 1. For a common skill like Kubernetes you narrow hard with AND terms and limiters. For a rare one like Rust you widen aggressively with OR synonyms and wildcards, and lean on the GitHub Contributors tab, which is often the only place to verify hands-on experience with niche stacks.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.