The Person Match-Key Reference: What Each Identifier Pins Down
You will be able to pick a person match key by its uniqueness, job-change survival, normalization, and named false-merge mode.
You need to merge records that describe the same human and never merge two who are not. This is a lookup reference for the people who own that decision in RevOps and RecOps: it grades each real sourcing identifier - verified email, corporate pattern email, phone, LinkedIn vanity slug, LinkedIn numeric ID, GitHub username, GitHub user ID, commit author email, name-plus-employer, and personal domain - on uniqueness, survival through a job change, how to normalize it, what a match proves, and the specific way it produces a false merge. Jump to the row you need. Each one is self-contained.
The abstract debate about deterministic versus probabilistic matching is covered everywhere. What is not written down anywhere public is a per-identifier grade of the keys you actually have in your source systems. That is what this reference is.
The one rule that governs every person key
The visible identifier is almost always the mutable one, and the stable identifier is hidden underneath it. This single pattern repeats across every platform you source from, and it is the reason most person-dedup rules are built on the wrong field.
LinkedIn shows people a vanity slug and keeps a numeric entity ID for its own systems. GitHub shows a username and keeps a numeric user ID. In both cases the thing that is easy to copy and paste is the thing that changes, and the thing that never changes is the thing you have to go dig out. Members can edit a LinkedIn slug up to five times in six months. A renamed GitHub username is released immediately for anyone else to claim. Both numeric IDs never change.
Every identifier that is easy to copy is the wrong one to store as your match key.
So the first move in any person match rule is the same: resolve the readable identifier to its hidden numeric one before you store it. If you skip that step, you have built your deduplication on a field whose owner can change it on a Tuesday.
The identifier grade table: strong versus weak
A key is strong if it is immutable and one-to-one, and weak if it is mutable or many-to-one. Strong keys can drive an automatic merge; weak keys can only form blocks or feed a score. Here is where each of the ten lands.
| Identifier | Strength | What a match proves |
|---|---|---|
| LinkedIn numeric member ID | Strong | Same LinkedIn account |
| GitHub numeric user ID | Strong | Same GitHub account |
| Verified unique email | Strong | Same mailbox, if truly one-to-one |
| ID-based commit noreply | Strong | Same GitHub account (carries user ID) |
| Corporate pattern email | Weak | Same mailbox now, breaks on job change |
| LinkedIn vanity slug | Weak | Current holder only; released slugs lie |
| GitHub username | Weak | Current holder only; renamed handles lie |
| Phone number | Weak | Current holder only; reassigned fast |
| Name-plus-employer | Weak | Nothing alone; collides at scale |
| Personal domain | Weak | Weak link unless uniquely owned |
Treat this as the top-level routing table. The rest of this reference explains why each weak key is weak and gives you the guard that stops it from producing a false merge.
Why a narrow query is a collision proxy
Before trusting name-plus-title as even a block key, look at how many people share one. In Refolk's index, volume is a direct proxy for collision risk.
| Query | People | Note |
|---|---|---|
| "Software Engineer", US | 352,540 | title is not a key |
| "Software Engineer", Germany | 22,661 | same title, 15.6x smaller |
| Rust skill, US | 3,593 | narrow skill |
| Go skill, US | 16,956 | 4.7x Rust |
A title shared by 352,540 people is not a key by any definition. A narrow skill like Rust, held by 3,593 people in the US, narrows a block but still does not uniquely identify anyone. Use these numbers to calibrate: the larger the population behind an attribute, the more a composite built on it will collide.
Email: three different keys wearing one coat
Email is not one identifier. A verified unique email is strong, a corporate pattern email is weak and expires on a job change, and a consumer Gmail address needs canonicalization that is dangerous if applied to the wrong domain. Handle each as its own key.
A verified unique email - one you have confirmed deliverable and tied to a single person - is as strong as a numeric ID while the mailbox exists. The weakness is reassignment: a corporate address like first.last@ gets reused for a new hire with the same name, which merges two employees. Guard by re-verifying on a cadence and never trusting an address you have not confirmed is still the same human's.
A corporate pattern email breaks by definition on a job change, because the domain leaves with the employer. It is fine as a block key for current-employer dedup but useless for tracking a person across moves. For that, see the job-change reference guides in this library.
A consumer Gmail address collapses many strings to one inbox. For gmail.com and googlemail.com, dots are ignored, capitalization is ignored, and anything after a plus sign delivers to the same mailbox: A.B.C@gmail.com, abc@gmail.com, and AbC+Shop@gmail.com all reach the same person. So you canonicalize by lowercasing, removing dots in the local part, and dropping everything from the first plus, then treating googlemail.com as gmail.com.
Here is the trap. On Google Workspace, dots change the address. Two coworkers on a company Google domain can differ only by a dot, and stripping it merges two distinct people.
Plus-tag sub-addressing is defined only in RFC 5233, an extension to mail filtering, not a universal delivery rule, so do not assume every provider routes +tag to the base inbox. Confirm per domain or treat plus-tags conservatively.
Phone: the least safe key that looks unique
A phone number looks like a one-to-one identifier and is not, because numbers are reassigned fast and often. It is the weakest "unique" key you can match on and should never drive an automatic merge without a recency check.
The FCC sets only a floor on how long a number ages before reuse: numbers previously assigned to residential customers may be aged no less than 45 days and no more than 90 days, and business numbers no less than 45 days and no more than 365 days. That is the minimum. In practice carriers move faster: one filing told the FCC that some large carriers reassign disconnected numbers in as few as two days.
And reassignment is not rare. A Princeton study found a sampled number was recycled 41.6% of the time, with a 95% confidence interval of 30.5% to 52.6%.
Normalize phones to E.164 before comparing so formatting differences do not split a true match. But normalization does not fix the reassignment problem. The guard is an aging and recency check: only merge on a phone when you have evidence the number still belongs to the same person, and re-verify before you rely on it.
Why phone fails as a merge key
- 100Numbers matched on digits
raw equality
- ~58Still held by same person
after 41.6% recycle rate
- fewerVerified recently
after aging and recency check
LinkedIn: slug lies, numeric ID holds
Match LinkedIn records on the numeric member ID, never on the vanity slug or the full URL. Every profile has a public vanity slug that members share and a numeric entity ID that LinkedIn's own systems use. The slug is editable up to five times per six months; the numeric ID never changes.
The false merge is specific. When a member edits their slug, the old one is released. If someone else later takes that slug, your record keyed on the slug now points to a different person. Resolve every profile to its numeric member ID before matching and the problem disappears.
One useful detail: default LinkedIn URLs append a hex hash for uniqueness when members share a name. That hash is a sign the person did not set a custom slug; it is still the slug, still mutable in principle, and still the wrong thing to key on.
GitHub: one field, three populations
GitHub identifiers need three separate rules because a commit email can be any of three things. The numeric user ID is immutable and is your strong key; the username is mutable and is not; and a commit noreply email is safe, unstable, or not even a person depending on its form.
GitHub usernames are explicitly not permanent. After a username change, the old username becomes available for anyone else to claim. GitHub retired its dormant-username release process, so support no longer reviews requests for inactive handles, and because not all activity is public, an empty-looking profile may be in active use. A deleted username is freed after 90 days; a renamed one is released immediately. Usernames max out at 39 characters, alphanumeric plus hyphens. None of that matters for matching except to tell you: do not key on the username.
The commit author email is the one field that splits into three populations:
- ID-based noreply - of the form
ID+USERNAME@users.noreply.github.com, used by accounts created after 18 July 2017 and by older accounts that enabled privacy after that date. It carries the immutable numeric user ID. This is safe. - Username-based noreply - of the form
USERNAME@users.noreply.github.com, used by older accounts that enabled privacy before that date. It detaches from the account on a rename, so old commits stop resolving. Treat it as unstable. - Bot addresses - of the form
USERID+APP-NAME[bot]@users.noreply.github.com, for example149130343+josh-issueops-bot[bot]. These belong to an app, not a human, and must be excluded.
Routing a GitHub commit email
- Read the addressinspect the local part before the @
- Contains [bot]?exclude - this is automation, not a person
- Starts with digits+username?safe - extract the numeric user ID
- Username only?unstable - detaches on rename, do not key on it alone
The payoff of getting this right: you can match the same human across GitHub and LinkedIn by resolving each side to its immutable numeric ID and ignoring every mutable handle in between. That is the hard part of cross-source person resolution, and it is exactly the kind of query that plain-English search removes the friction from.
When you want the matched records back without building the resolver yourself, Refolk keys on the stable identifiers across the public GitHub graph and public LinkedIn records so the people who come back are the same humans, not slug collisions.
The procedure: block, then match
The sequence is settled: block first to shrink the comparison space, then match within blocks, running strong keys deterministically and weak keys through a score. Sources agree on this order; they disagree only on whether a learned model should replace hand-tuned weights in the scoring step. Here is the full procedure.
From raw records to safe merges
- Inventory identifiers per recordList which of the ten person keys each source system actually stores, with coverage and native format. You end with a field map showing what identifier exists where.
- Normalize each key to canonical formLowercase emails; strip Gmail dots and plus only for gmail.com and googlemail.com; format phones to E.164; reduce LinkedIn to the numeric member ID; reduce GitHub to the numeric user ID. Keep a canonical column beside each raw column, raw preserved.
- Classify keys as strong versus weakMark a key strong if immutable and one-to-one, weak if mutable or many-to-one. You end with a graded key table.
- Block on a stable key to shrink comparisonsPartition records into blocks on an invariant key so comparison happens only within a block, avoiding the O(mn) cost of comparing every record to every other. Candidate pairs are generated only within blocks.
- Apply deterministic merges on strong keysMerge any two records that share an exact match on a single immutable ID. High-confidence clusters are formed and logged.
- Score remaining pairs probabilisticallyFor records sharing no unique ID, combine weak keys with weights or a model to produce a similarity score with a threshold.
- Hold a human-review band for mid-confidence pairsRoute pairs scoring between auto-merge and auto-reject to an analyst, and log every merge so it can be reversed.
Two framing notes on the matching step. Deterministic matching applies strict rules: two records match only if specific fields are identical. It is precise and transparent but misses valid matches when data is inconsistent, misspelled, or incomplete. Probabilistic matching uses statistical models to weigh similarity across multiple fields; it catches more true matches but introduces false positives. The standard answer is to use both - deterministic rules for high-confidence matches where a unique ID is shared, probabilistic scoring for the ambiguous remainder.
Blocking is not only a speed trick. It is a precision-and-recall lever. If the blocking key is not distinctive, many irrelevant records land in one block and matching slows down. If the blocking key is not invariant across records of the same entity, true matches get split across different blocks and are silently lost. That is why you block on a stable key, never on a mutable one.
How this goes wrong: the named false-merge modes
Every weak key has a specific failure with a specific guard. Here are the eight that recur, each one a false merge or a false miss you can prevent.
| Failure mode | What breaks | The guard |
|---|---|---|
| Gmail rule on Workspace | Two coworkers differing by a dot merged | Strip dots only for gmail.com / googlemail.com |
| Keying on LinkedIn slug | Released slug now points to a new person | Resolve to the numeric member ID |
| Keying on GitHub username | Renamed handle reclaimed by someone else | Use the numeric user ID |
| Username-only noreply after rename | Commits detach from the account (false miss) | Prefer the ID-based noreply form |
| Bot noreply treated as a person | Automation merged into a human cluster | Exclude any address containing [bot] |
| Phone reassignment | New holder merged to prior holder | Aging and recency check, re-verify |
| Name-plus-employer at scale | Two same-name people at one big employer collapsed | Require a second immutable key |
| Over-blocking on a mutable key | True matches split into different blocks (false miss) | Block on an invariant key |
The two worst in practice are name-plus-employer and phone, because both feel reliable and both fail most where you have the most data. Name-plus-employer fails hardest exactly where your records cluster: 352,540 US "Software Engineer" profiles with Google as the top employer means the composite collides most at your biggest targets. Phone fails because reassignment is fast and probable. Neither should ever trigger an automatic merge alone.
The two false misses - detached username-noreply commits and over-blocking on a mutable key - are quieter and more dangerous, because a false miss leaves two records for one person and you never see an error. Audit for them by sampling records that should have merged and checking why they did not.
What to verify before you call the merge rule done
Run this before you turn any match rule loose on live data. Each item is a specific check, not a topic.
Pre-merge verification
- Every LinkedIn record is keyed on the numeric member ID, not the vanity slug or URL.
- Every GitHub record is keyed on the numeric user ID, not the username.
- Commit emails are split into ID-based (keep), username-only (treat as unstable), and [bot] (excluded).
- Gmail dot and plus stripping is applied only to gmail.com and googlemail.com, never to Workspace domains.
- Phone numbers are in E.164 and carry a recency or re-verification flag before any merge.
- No automatic merge fires on name-plus-employer alone without a second immutable key.
- Blocking keys are invariant across records of the same person, confirmed by a sample of should-have-merged pairs.
- Every merge is logged with the key it fired on and is reversible.
Keeping the reference current
The grades in this document are stable because they describe mechanisms, not current values, but two things move and are worth a periodic re-check. First, platform rules: GitHub has already retired its dormant-username release process, and either platform could change how slugs, usernames, or noreply forms behave, so re-read the GitHub username and commit-email docs and confirm the LinkedIn slug edit limit when you revise your rules. Second, reassignment windows: the FCC floor of 45 days is a minimum, and carrier practice runs faster, so treat any phone match as time-bounded and set your recency threshold against current carrier behavior rather than the regulatory floor.
The core will not change. On every platform that matters, the easy-to-copy identifier is the mutable one and the stable identifier is hidden. Build on the hidden one, keep the readable one for display, and reserve probabilistic scoring for the records that share no immutable key at all.
LinkedIn: key = numeric member ID; slug = display only; auto-merge on exact ID. GitHub: key = numeric user ID; username = display only; auto-merge on exact ID. Commit email: keep ID-based noreply, flag username-only as unstable, drop any [bot]. Verified email: auto-merge on exact lowercased match; re-verify on a cadence. Gmail: canonicalize (lowercase, strip dots and +tag) only for gmail.com/googlemail.com. Phone: E.164; never auto-merge without a recency/re-verification flag. Name+employer: block key only; require a second immutable key before merge.
Paste into your dedup spec and adjust the thresholds to your data.
Questions practitioners ask
What is the single most reliable identifier to match two person records on?
A verified, immutable numeric ID is the most reliable: the LinkedIn numeric member ID and the GitHub numeric user ID never change and map one-to-one to an account. Both survive a job change and neither can be edited by the owner. The catch is that the owner-facing interface shows you the mutable slug or username instead, so you must resolve to the numeric ID before you store it as your match key.
Why should I not match people on their LinkedIn URL?
The readable part of a LinkedIn URL is the vanity slug, which members can change up to five times every six months. When a slug is released it can later point to a different person, so matching on it produces a false merge. Resolve the profile to its underlying numeric member ID, which never changes, and key on that instead.
Is a GitHub commit email a safe person key?
Only one of its three forms is safe. ID-based noreply addresses, used by accounts created after 18 July 2017, carry the immutable numeric user ID and are reliable. Username-based noreply addresses detach from the account on a rename, and any address containing [bot] belongs to an app, not a human. Key on the ID form and exclude [bot] addresses.
How do I normalize a Gmail address without merging different people?
For gmail.com and googlemail.com only, lowercase the address, remove dots in the local part, and drop everything from the first plus sign. Do not apply this to any other domain. On Google Workspace dots change the address, so two coworkers can differ by a single dot, and stripping it would merge two distinct people. Store the address exactly as typed and compare on the canonical key.
Can I safely deduplicate contacts on name plus employer?
No, not on its own. It is a weak, many-to-one key that collides worst where you have the most records. In Refolk's index, "Software Engineer" in the US returns 352,540 people with Google as the top current employer, so same-name collisions at large employers are common. Use name-plus-employer only to form blocks, then require a second immutable key before you merge.
When should I use probabilistic matching instead of deterministic rules?
Use deterministic rules first for any pair that shares an immutable ID or verified email: those merges are precise and transparent. Reserve probabilistic scoring for the pairs left over that share no unique identifier, where you combine weak keys with weights or a model. This two-tier approach catches more true matches than rules alone while keeping high-confidence merges auditable.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.