# From Patent Records to a Reachable Inventor Shortlist

*You can turn one technology area into a deduplicated shortlist of named inventors resolved to their current employer and a reachable contact.*

- Canonical URL: https://www.refolk.ai/guides/patent-records-to-inventor-shortlist
- Pillar: Recruiting and sourcing
- Format: Playbook
- Published: 2026-08-20
- Last reviewed: 2026-08-20
- Reading time: 15 min
- Keywords: sourcing engineers from patents, how to search USPTO for candidates, find inventors by patent classification, recruiting from google patents, how to source hardware engineers

## Key takeaways

- In Refolk's index there are 219 US profiles titled Semiconductor Engineer against 1,658 titled Robotics Engineer, so a patent-first strategy pays off roughly 7.6 times more where LinkedIn returns are thin.
- The employer named on any patent is 1.5 to 3 or more years stale by the time you read it, because publication trails filing by about 18 months and residence is frozen at filing under MPEP 719.02.
- German Semiconductor Engineer profiles collapse to 8 in Refolk's index, concentrated in Dresden, so the patent record plus one company map effectively enumerates that market.
- With 313,219 utility patents granted in 2023, even a single CPC subgroup yields hundreds of named engineers, so the constraint is disambiguation and contact resolution, not supply.
- PatentsView's disambiguated inventor IDs are probabilistic and fallible by design, so a common name must share a co-inventor or geography before you merge two records into one person.
- Never geo-target off a patent residence: it is city and state at a point in time and may be a corporate mailing city, not where the person lives now.

This is the procedure for finding deep-tech, hardware, and R&D engineers who never surface on LinkedIn or the public GitHub graph, using the public patent record, and turning them into a ranked shortlist you can actually reach. It is written for technical sourcers, in-house recruiters, and founders hiring in semiconductors, robotics, materials, RF, and biotech. By the end you can scope a technology to the right classification codes, extract and disambiguate inventors, filter out name-on-paper contributors, and resolve years-old records to a current employer and a reachable contact.

Existing guides on patent sourcing walk through one keyword search and stop at the inventor's name. They skip every hard part: the filing-to-publication lag that makes employer data stale, disambiguating a common name across assignees, filtering people who merely appear on a filing, and resolving an old inventor to a current person. This document handles all four.

## Why the patent record beats profile databases for these stacks

For narrow hardware and R&D fields, the patent record contains engineers who leave almost no trace in conventional sourcing channels. The public profile databases return thin, generic results for these titles, while a single patent classification code yields hundreds of named, technically qualified people.

The payoff is domain-specific, not "deep tech" generic. In Refolk's index of professional profiles, there are 219 US profiles titled Semiconductor Engineer against 1,658 titled Robotics Engineer. That is roughly a 7.6-to-1 gap for the exact titles, which tells you where the patent-first strategy earns its keep.

| Title | Profiles (US) | Top hub |
|---|---|---|
| Robotics Engineer | 1,658 | San Francisco Bay Area |
| Semiconductor Engineer | 219 | Albany / Cedar Park / St Paul |

Robotics engineers are visible; you do not need patents to find them. Semiconductor, RF, and materials engineers are not, and that is exactly where the record pays off. The volume is there to support it: the USPTO granted 313,219 utility patents in calendar 2023, so even a single CPC subgroup returns a deep pool.

**7.6x - US robotics-titled profiles per semiconductor-titled profile in Refolk's index**

1,658 Robotics Engineer profiles versus 219 Semiconductor Engineer profiles, which is why patents pay off for chips, not robots.

In the narrowest stacks, geography does most of the work for you. In Refolk's index there are only 8 profiles titled Semiconductor Engineer in Germany, concentrated in Dresden, with top employers Infineon, GlobalFoundries, Cypress, BMW, and DENSO. A market that small is effectively enumerated by the patent record plus one company map.

## What a public patent record actually exposes

A published US application or granted patent gives you the inventor names, the inventor residence, the assignee, the co-inventors, the CPC classification, and the filing and publication or grant dates. That is enough to build a shortlist, but only if you understand which fields are reliable and which lie.

Residence is the least trustworthy field for sourcing. It is limited to city and state or foreign country, with no street address, per MPEP 602.08(a), and it is a point-in-time value that is never maintained. USPTO regional statistics confirm this: patent origin is based on the residence of the first-named inventor, limited to the city and state at the time of grant.

| Field | Reliability for sourcing | Why |
|---|---|---|
| Inventor name | High, as recorded | Legally required; corrected under 37 CFR 1.48 |
| Co-inventors | High | Fixed on the document; anchors disambiguation |
| Assignee at filing | Historical | Frozen at filing; the company then, not now |
| CPC classification | High | Assigned by examiners; your search key |
| Inventor residence | Low | City and state only, point-in-time, never updated |

The name, co-inventors, assignee-at-filing, CPC, and dates are reliable as recorded. Residence and the employer relationship are historical the moment you read them. Treat the reliable fields as your extraction target and the historical fields as leads to be re-resolved, never as facts about the person today.

## The lag problem, and why it defines the whole method

The employer named on any patent is 1.5 to 3 or more years stale by the time you read it. Applications publish by default 18 months after the earliest filing date under 35 U.S.C. 122(b), grant runs later still, and the assignee and residence are frozen at filing and never maintained under MPEP 719.02. The sourcing edge is not extraction; it is resolution to current data.

#### Patent-record freshness against candidate freshness

| Stage | Figure | Note |
| --- | --- | --- |
| Application filed | 0 mo | employer is current here, but not public |
| Application publishes | ~18 mo | over half publish within 12 months |
| First office action | ~22 mo | record is now visible and aging |
| Patent grants | 18-36 mo | employer data is 1.5-3+ years old |

*Every stage of the patent lifecycle pushes the listed employer further out of date before you ever see it.*

The numbers behind that funnel matter because they set your recency window. Over half of US applications publish within a year of filing, first office action averages about 22 months, and grant typically spans 18 to 36 months. Two traps hide in these dates. A non-publication request means a US-only filing never publishes at all, and a provisional never publishes, so the record is never a complete map of a company's R&D. And a continuation of an old application can publish within weeks, making an old invention look recent; read the priority or earliest filing date, not the publication date.

| Stage | Typical lag from filing | Source |
|---|---|---|
| Application publication | ~18 months (over half under 12 mo) | ipwatchdog / MPEP 1120 |
| First office action | ~22 months | Baker Botts |
| Grant | 18-36 months | Baker Botts |

> **Rule:** Re-resolve every affiliation
>
> The assignee and residence on a patent are frozen at filing and never updated under MPEP 719.02. Treat them as historical leads, and confirm the current employer and location from a live source before any outreach.

## Scope the technology to the right classification codes

Finding candidates by patent classification means finding the right CPC codes first, then querying them as fields. CPC, the Cooperative Patent Classification, is co-managed by the USPTO and the European Patent Office. Do not reach for USPC: it was retired in June 2015, and applications filed after that date carry no USPC code.

The documented method is seed-and-harvest, and it beats guessing at codes cold. Collect 5 to 10 representative seed patents by keyword, extract their primary and secondary CPC and IPC codes, rank the most frequent codes across the seeds, and validate the scope in a CPC browser. Then, in Patent Public Search, query the code directly by entering it without spaces, adding `.cpc.`, and pressing Search.

Relying on the first-listed CPC is a common way to skew your pool thin. Harvest the secondary codes too and rank by frequency, or you will miss the adjacent art where half your candidates live.

**CPC seed-and-harvest worksheet**

```
Seed patent no.: __________
Primary CPC: __________
Secondary CPCs: __________ / __________ / __________
Assignee: __________
Earliest (priority) filing date: __________
--
After 5-10 seeds, list the 2-5 CPC groups that recur most and read each definition in a CPC browser before you commit.
Final query (Patent Public Search): CODEwithnospaces.cpc. AND @pd>=YYYYMMDD
```

*Fill one block per seed patent, then rank codes by how often they recur across all seeds.*

## The procedure, start to finish

This is the full method in order, with who does each step, roughly how long it takes, and what a good result looks like. The whole run for one technology area is a day of focused work, front-loaded on scoping and disambiguation.

#### Patent record to reachable shortlist

1. **Scope the technology to CPC codes** - Pull 5-10 seed patents by keyword, harvest their primary and secondary CPC codes, rank by frequency, and validate definitions in a CPC browser. Done when you have 2-5 CPC groups that return on-topic patents. (~30-60 min)
2. **Run the classification search** - In Patent Public Search, query CODE.cpc. optionally ANDed with a date range and an assignee. Done when you have a result set filtered to your technology and a recency window. (~30 min)
3. **Set the recency window for lag** - Search filings from the last 2-4 years to catch current-ish employers, because publication trails filing by ~18 months and grant by 18-36 months. Done when every record carries a documented as-of date and the listed employer is treated as historical. (~15 min)
4. **Extract inventors and co-inventor networks** - Export inventor names, residence, assignee, co-inventors, and dates against their source patent numbers. Done when you have a raw inventor table you can sort and group. (~1-2 hr)
5. **Disambiguate to real people** - Use PatentsView disambiguated inventor IDs, and for common names resolve manually via co-inventor overlap, assignee history, and geography, checking ORCID or The Lens. Done when each row is one resolved person with a confidence note. (~2-4 hr)
6. **Rank contributors** - Apply heuristic signals, position in the inventor list, filing count, and portfolio continuity. Flag the ranking as heuristic because no public benchmark validates the weightings. Done when you have a scored shortlist. (~1 hr)
7. **Resolve to current employer and contact** - Reconcile the stale patent affiliation against current professional-profile data to get present employer, location, and a reachable channel. Done when the shortlist is deduplicated with a current company and one contact path per person. (~2-3 hr)
8. **QA and dedupe** - Merge duplicate inventor IDs, drop mis-resolved names, and spot-check 10 percent against a second source. Done when the shortlist is clean and citable. (~30 min)

## Disambiguating a common inventor name

Disambiguation is the moat, and it is imperfect by design. The documented public method is PatentsView's probabilistic entity resolution, which uses similarity scores and clustering to decide whether two same-name records are the same person and assigns a disambiguated inventor identifier. Because the clustering is probabilistic, errors persist, so a sourcer who blindly trusts the IDs will both merge two people into one and split one person into two.

The manual method a sourcer mirrors leans on three anchors. In the hand-disambiguation study behind PatentsView's evaluation work, the predicted cluster was reviewed, wrongly assigned patents were removed, and additional mentions of similarly named inventors were found and added where appropriate. You do the same with:

- **Co-inventor overlap.** A shared co-inventor across filings is the strongest low-cost signal that two records are one person.
- **Assignee history.** A plausible employer path over time supports a merge; an implausible jump argues against it.
- **Geography.** Matching residence cities support a merge, but never on their own, because residence is point-in-time and unreliable.
- **ORCID and The Lens.** Where an inventor self-asserts an ORCID, The Lens auto-syncs their patent inventorship, giving you a self-declared anchor.

> **Watch out:** Do not merge on name alone
>
> "Wei Zhang" appearing under three assignees may be three people or one. Require at least one shared co-inventor, an overlapping assignee history, or corroborated geography before you merge records, or you will build a fictional super-inventor.

The scale explains why this is the bottleneck. PatentsView averaged more than 77,000 API queries per day in 2019, and its evaluation study sampled 100 inventors precisely to measure how often the algorithm is wrong. Trust the IDs as a first pass, then verify every merge that carries weight in your shortlist.

Once your rows are resolved to real people, the remaining work is turning a years-old affiliation into a present employer and a reachable channel. This is the exact step where the older public guides stop, noting only that the USPTO tells you a location and calling it a bonus. The location is stale and the employer is stale, so you need a current index to reconcile against.

I ran this search: `Semiconductor process-integration engineers in the Dresden area who are named inventors on power-device patents` - [see the full result list](https://www.refolk.ai/s/akpg041xhz).

*Returns resolved people with current employer, location, and a reachable channel, so a stale patent affiliation becomes an outreach-ready row.*

Reconciling a frozen assignee against live profile data by hand is slow and error-prone; [Refolk](/) does that resolution against a current index, which collapses the two-to-three-hour resolution step into a query. That is the difference between a list of names from 2021 and a shortlist you can send a first message to today.

## Ranking real contributors and filtering name-on-paper inventors

Ranking inventors is heuristic, and you should say so out loud. There is no published benchmark that quantifies how well any signal predicts that a named inventor is the technical contributor you want, so treat every weighting as a working assumption rather than a validated model.

US law helps a little. Each named person must have contributed to the conception of the invention, and practitioners note that despite the "inventor" label these are typically the engineers who did meaningful work. But conception is a weak filter against a manager or IP contact who appears for reasons you cannot see on the face of the document.

The practical signals, used with that caveat, are:

- **Position in the inventor list.** A first-named inventor is more often the lead technical contributor than a fifth-named one, though this is a tendency, not a rule.
- **Filing count.** Someone with many filings in the area is more likely a working engineer than a one-appearance name.
- **Portfolio continuity.** Repeated appearance across a coherent body of work over time signals sustained technical involvement.
- **Cross-assignee appearance.** The same person inventing under different employers over time is a mobility signal worth flagging, and one of the example searches worth running on its own.

#### Triage an inventor before you spend resolution time

Horizontal axis runs from Weak contributor signal to Strong contributor signal. Vertical axis runs from Low disambiguation confidence to High disambiguation confidence.

| Quadrant | What it means |
| --- | --- |
| Uncertain and thin | Drop or park; not worth resolution effort |
| Promising but unresolved | Verify identity first, then resolve |
| Confident but likely name-on-paper | Deprioritize; probably a manager or IP contact |
| Confident and clearly technical | Resolve to current contact now |

*Weigh how technical the contribution looks against how confident your disambiguation is.*

> The patent tells you who invented something years ago. Your job is to find who they are today.

## How this goes wrong

Most patent-sourcing failures are false positives that survive to your shortlist and embarrass you in outreach. Each has a specific check that catches it before it costs you.

- **Stale employer.** The assignee is the company at filing, and the inventor may have left years ago. The false positive is pitching someone as "at Intel" who left in 2021. Check: reconcile against current profile data, never the patent.
- **Common-name collision.** One name across three assignees may be three people or one. The false positive is a merged super-inventor. Check: require co-inventor or geography overlap before merging, and lean on disambiguated IDs.
- **Name-on-paper inventor.** A manager or IP contact listed for contribution reasons you cannot see. The false positive is ranking a non-technical person high. Check: use portfolio continuity and co-inventor role, and treat the ranking as heuristic.
- **Residence is not a home.** City and state is point-in-time and may be a corporate mailing city. The false positive is targeting the wrong metro. Check: never geo-target off patent residence alone.
- **Non-publication gap.** US-only filings with a non-publication request never publish, and provisionals never publish. The false positive is assuming full coverage of a company's R&D. Check: acknowledge the blind spot.
- **Wrong CPC scope.** Relying on the first-listed CPC misses secondary art. The false positive is a thin, skewed inventor pool. Check: harvest secondary codes and rank by frequency.
- **Continuation timing illusion.** A continuation of an old application can publish within weeks, making an old invention look recent. Check: read the priority or earliest filing date, not the publication date.

> **Tip:** Spot-check before you trust the batch
>
> After dedupe, pull 10 percent of the shortlist at random and confirm each person against a second source. If more than one in ten is mis-resolved, your disambiguation or contact resolution has a systematic error, and you should fix the method before you send anything.

## Before you call the shortlist done

A shortlist is ready when every row is one resolved person, tied to source patents, resolved to a current employer, and carrying one reachable contact path. Run this check before it leaves your hands.

#### Shortlist readiness

- [ ] Each row cites the source patent number(s) it came from.
- [ ] Every affiliation and location is re-resolved against current data, not taken from the patent.
- [ ] Each disambiguated identity carries a confidence note, and every weighty merge has a co-inventor, assignee, or geography anchor.
- [ ] Priority (earliest filing) dates were read, so no continuation is mistaken for recent work.
- [ ] Rankings are labeled heuristic, with the signals used stated.
- [ ] Duplicate inventor IDs are merged and mis-resolved names dropped.
- [ ] A 10 percent random sample passed verification against a second source.
- [ ] The non-publication and provisional blind spot is noted where coverage claims are made.

## Keeping the shortlist current

Patent-sourced shortlists decay from two directions at once. The record keeps aging because new filings publish on the 18-month lag, and the people keep moving because the affiliations were never maintained in the first place. Both mean a list is a snapshot, not an asset you can shelve.

Re-run the classification search on a schedule that matches your recency window. If you search filings from the last two to four years, refresh the pool at least twice a year so you catch newly published applications from filings you missed. Re-resolve the current employer for anyone you did not reach on the first pass, because the field data is now further out of date than when you built the list.

Two moves extend the life of the work. First, save your validated CPC groups and seed patents so the scoping step becomes a lookup, not a rediscovery, next quarter. Second, watch the cross-assignee movers, such as former Infineon power-semiconductor inventors who have since changed employer, because a mobility signal is often a hiring signal, and the patent record is one of the few public places it shows up early.

## Frequently asked questions

### Can I do this with Google Patents or do I need USPTO Patent Public Search?

Use USPTO Patent Public Search for the extraction step. Per practitioner accounts, USPTO exposes inventor location on the record while Google Patents historically did not, and Patent Public Search lets you query inventor, assignee, CPC, and dates as structured fields. Google Patents is fine for reading a single document, but the field-level classification and inventor queries this method depends on live in Patent Public Search at ppubs.uspto.gov/basic.

### How stale is the employer listed on a patent?

At minimum 1.5 to 3 or more years old by the time you read it. Applications publish by default 18 months after the earliest filing date under 35 U.S.C. 122(b), and grant typically runs 18 to 36 months from filing. Crucially, the assignee and residence are frozen at filing and not maintained under MPEP 719.02, so you must always re-resolve the person to a current source rather than trust the patent's affiliation.

### How do I tell two people with the same name apart?

Require a shared signal before you merge. Start with PatentsView's disambiguated inventor IDs, which use probabilistic clustering, but treat them as fallible. For a common name across multiple assignees, confirm a shared co-inventor, an overlapping assignee history, or matching geography before deciding two records are one person. Where the inventor self-asserts an ORCID, The Lens can link their patent inventorship, which gives you a self-declared anchor.

### Does this work for every deep-tech field equally?

No. Pool size is domain-specific. In Refolk's index, US Robotics Engineer profiles outnumber Semiconductor Engineer profiles about 7.6 to 1 (1,658 versus 219), so the patent-first approach pays off far more for semiconductors, RF, and materials, where profile databases return little, than for robotics, where they return plenty. Match the effort to how thin the conventional channels are for your stack.

### Which patent classification codes should I search, USPC or CPC?

Use CPC. The older USPC scheme was retired in June 2015 and applications filed after that date carry no USPC code. CPC, co-managed by the USPTO and the European Patent Office, is the live system. Find your codes by the seed-and-harvest method: collect 5 to 10 representative patents, extract their CPC codes, rank the most frequent, and validate the definitions in a CPC browser before you rely on them.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/patent-records-to-inventor-shortlist*
