Refolk
PlaybookRecruiting and sourcing

From Patent Records to a Reachable Inventor Shortlist

You can turn one technology area into a deduplicated shortlist of named inventors resolved to their current employer and a reachable contact.

15 min readLast reviewed August 20, 2026Read as Markdown

Key takeaways

  • In Refolk's index there are 219 US profiles titled Semiconductor Engineer against 1,658 titled Robotics Engineer, so a patent-first strategy pays off roughly 7.6 times more where LinkedIn returns are thin.
  • The employer named on any patent is 1.5 to 3 or more years stale by the time you read it, because publication trails filing by about 18 months and residence is frozen at filing under MPEP 719.02.
  • German Semiconductor Engineer profiles collapse to 8 in Refolk's index, concentrated in Dresden, so the patent record plus one company map effectively enumerates that market.
  • With 313,219 utility patents granted in 2023, even a single CPC subgroup yields hundreds of named engineers, so the constraint is disambiguation and contact resolution, not supply.
  • PatentsView's disambiguated inventor IDs are probabilistic and fallible by design, so a common name must share a co-inventor or geography before you merge two records into one person.
  • Never geo-target off a patent residence: it is city and state at a point in time and may be a corporate mailing city, not where the person lives now.

This is the procedure for finding deep-tech, hardware, and R&D engineers who never surface on LinkedIn or the public GitHub graph, using the public patent record, and turning them into a ranked shortlist you can actually reach. It is written for technical sourcers, in-house recruiters, and founders hiring in semiconductors, robotics, materials, RF, and biotech. By the end you can scope a technology to the right classification codes, extract and disambiguate inventors, filter out name-on-paper contributors, and resolve years-old records to a current employer and a reachable contact.

Existing guides on patent sourcing walk through one keyword search and stop at the inventor's name. They skip every hard part: the filing-to-publication lag that makes employer data stale, disambiguating a common name across assignees, filtering people who merely appear on a filing, and resolving an old inventor to a current person. This document handles all four.

Why the patent record beats profile databases for these stacks

For narrow hardware and R&D fields, the patent record contains engineers who leave almost no trace in conventional sourcing channels. The public profile databases return thin, generic results for these titles, while a single patent classification code yields hundreds of named, technically qualified people.

The payoff is domain-specific, not "deep tech" generic. In Refolk's index of professional profiles, there are 219 US profiles titled Semiconductor Engineer against 1,658 titled Robotics Engineer. That is roughly a 7.6-to-1 gap for the exact titles, which tells you where the patent-first strategy earns its keep.

TitleProfiles (US)Top hub
Robotics Engineer1,658San Francisco Bay Area
Semiconductor Engineer219Albany / Cedar Park / St Paul

Robotics engineers are visible; you do not need patents to find them. Semiconductor, RF, and materials engineers are not, and that is exactly where the record pays off. The volume is there to support it: the USPTO granted 313,219 utility patents in calendar 2023, so even a single CPC subgroup returns a deep pool.

7.6x
US robotics-titled profiles per semiconductor-titled profile in Refolk's index
1,658 Robotics Engineer profiles versus 219 Semiconductor Engineer profiles, which is why patents pay off for chips, not robots.

In the narrowest stacks, geography does most of the work for you. In Refolk's index there are only 8 profiles titled Semiconductor Engineer in Germany, concentrated in Dresden, with top employers Infineon, GlobalFoundries, Cypress, BMW, and DENSO. A market that small is effectively enumerated by the patent record plus one company map.

What a public patent record actually exposes

A published US application or granted patent gives you the inventor names, the inventor residence, the assignee, the co-inventors, the CPC classification, and the filing and publication or grant dates. That is enough to build a shortlist, but only if you understand which fields are reliable and which lie.

Residence is the least trustworthy field for sourcing. It is limited to city and state or foreign country, with no street address, per MPEP 602.08(a), and it is a point-in-time value that is never maintained. USPTO regional statistics confirm this: patent origin is based on the residence of the first-named inventor, limited to the city and state at the time of grant.

FieldReliability for sourcingWhy
Inventor nameHigh, as recordedLegally required; corrected under 37 CFR 1.48
Co-inventorsHighFixed on the document; anchors disambiguation
Assignee at filingHistoricalFrozen at filing; the company then, not now
CPC classificationHighAssigned by examiners; your search key
Inventor residenceLowCity and state only, point-in-time, never updated

The name, co-inventors, assignee-at-filing, CPC, and dates are reliable as recorded. Residence and the employer relationship are historical the moment you read them. Treat the reliable fields as your extraction target and the historical fields as leads to be re-resolved, never as facts about the person today.

The lag problem, and why it defines the whole method

The employer named on any patent is 1.5 to 3 or more years stale by the time you read it. Applications publish by default 18 months after the earliest filing date under 35 U.S.C. 122(b), grant runs later still, and the assignee and residence are frozen at filing and never maintained under MPEP 719.02. The sourcing edge is not extraction; it is resolution to current data.

Patent-record freshness against candidate freshness

  1. Application filed
    0 mo

    employer is current here, but not public

  2. Application publishes
    ~18 mo

    over half publish within 12 months

  3. First office action
    ~22 mo

    record is now visible and aging

  4. Patent grants
    18-36 mo

    employer data is 1.5-3+ years old

Every stage of the patent lifecycle pushes the listed employer further out of date before you ever see it.

The numbers behind that funnel matter because they set your recency window. Over half of US applications publish within a year of filing, first office action averages about 22 months, and grant typically spans 18 to 36 months. Two traps hide in these dates. A non-publication request means a US-only filing never publishes at all, and a provisional never publishes, so the record is never a complete map of a company's R&D. And a continuation of an old application can publish within weeks, making an old invention look recent; read the priority or earliest filing date, not the publication date.

StageTypical lag from filingSource
Application publication~18 months (over half under 12 mo)ipwatchdog / MPEP 1120
First office action~22 monthsBaker Botts
Grant18-36 monthsBaker Botts

Scope the technology to the right classification codes

Finding candidates by patent classification means finding the right CPC codes first, then querying them as fields. CPC, the Cooperative Patent Classification, is co-managed by the USPTO and the European Patent Office. Do not reach for USPC: it was retired in June 2015, and applications filed after that date carry no USPC code.

The documented method is seed-and-harvest, and it beats guessing at codes cold. Collect 5 to 10 representative seed patents by keyword, extract their primary and secondary CPC and IPC codes, rank the most frequent codes across the seeds, and validate the scope in a CPC browser. Then, in Patent Public Search, query the code directly by entering it without spaces, adding .cpc., and pressing Search.

Relying on the first-listed CPC is a common way to skew your pool thin. Harvest the secondary codes too and rank by frequency, or you will miss the adjacent art where half your candidates live.

CPC seed-and-harvest worksheet
Seed patent no.: __________
Primary CPC: __________
Secondary CPCs: __________ / __________ / __________
Assignee: __________
Earliest (priority) filing date: __________
--
After 5-10 seeds, list the 2-5 CPC groups that recur most and read each definition in a CPC browser before you commit.
Final query (Patent Public Search): CODEwithnospaces.cpc. AND @pd>=YYYYMMDD

Fill one block per seed patent, then rank codes by how often they recur across all seeds.

The procedure, start to finish

This is the full method in order, with who does each step, roughly how long it takes, and what a good result looks like. The whole run for one technology area is a day of focused work, front-loaded on scoping and disambiguation.

Patent record to reachable shortlist

  1. Scope the technology to CPC codes
    Pull 5-10 seed patents by keyword, harvest their primary and secondary CPC codes, rank by frequency, and validate definitions in a CPC browser. Done when you have 2-5 CPC groups that return on-topic patents. (~30-60 min)
  2. Run the classification search
    In Patent Public Search, query CODE.cpc. optionally ANDed with a date range and an assignee. Done when you have a result set filtered to your technology and a recency window. (~30 min)
  3. Set the recency window for lag
    Search filings from the last 2-4 years to catch current-ish employers, because publication trails filing by ~18 months and grant by 18-36 months. Done when every record carries a documented as-of date and the listed employer is treated as historical. (~15 min)
  4. Extract inventors and co-inventor networks
    Export inventor names, residence, assignee, co-inventors, and dates against their source patent numbers. Done when you have a raw inventor table you can sort and group. (~1-2 hr)
  5. Disambiguate to real people
    Use PatentsView disambiguated inventor IDs, and for common names resolve manually via co-inventor overlap, assignee history, and geography, checking ORCID or The Lens. Done when each row is one resolved person with a confidence note. (~2-4 hr)
  6. Rank contributors
    Apply heuristic signals, position in the inventor list, filing count, and portfolio continuity. Flag the ranking as heuristic because no public benchmark validates the weightings. Done when you have a scored shortlist. (~1 hr)
  7. Resolve to current employer and contact
    Reconcile the stale patent affiliation against current professional-profile data to get present employer, location, and a reachable channel. Done when the shortlist is deduplicated with a current company and one contact path per person. (~2-3 hr)
  8. QA and dedupe
    Merge duplicate inventor IDs, drop mis-resolved names, and spot-check 10 percent against a second source. Done when the shortlist is clean and citable. (~30 min)

Disambiguating a common inventor name

Disambiguation is the moat, and it is imperfect by design. The documented public method is PatentsView's probabilistic entity resolution, which uses similarity scores and clustering to decide whether two same-name records are the same person and assigns a disambiguated inventor identifier. Because the clustering is probabilistic, errors persist, so a sourcer who blindly trusts the IDs will both merge two people into one and split one person into two.

The manual method a sourcer mirrors leans on three anchors. In the hand-disambiguation study behind PatentsView's evaluation work, the predicted cluster was reviewed, wrongly assigned patents were removed, and additional mentions of similarly named inventors were found and added where appropriate. You do the same with:

  • Co-inventor overlap. A shared co-inventor across filings is the strongest low-cost signal that two records are one person.
  • Assignee history. A plausible employer path over time supports a merge; an implausible jump argues against it.
  • Geography. Matching residence cities support a merge, but never on their own, because residence is point-in-time and unreliable.
  • ORCID and The Lens. Where an inventor self-asserts an ORCID, The Lens auto-syncs their patent inventorship, giving you a self-declared anchor.

The scale explains why this is the bottleneck. PatentsView averaged more than 77,000 API queries per day in 2019, and its evaluation study sampled 100 inventors precisely to measure how often the algorithm is wrong. Trust the IDs as a first pass, then verify every merge that carries weight in your shortlist.

Once your rows are resolved to real people, the remaining work is turning a years-old affiliation into a present employer and a reachable channel. This is the exact step where the older public guides stop, noting only that the USPTO tells you a location and calling it a bonus. The location is stale and the employer is stale, so you need a current index to reconcile against.

Reconciling a frozen assignee against live profile data by hand is slow and error-prone; Refolk does that resolution against a current index, which collapses the two-to-three-hour resolution step into a query. That is the difference between a list of names from 2021 and a shortlist you can send a first message to today.

Ranking real contributors and filtering name-on-paper inventors

Ranking inventors is heuristic, and you should say so out loud. There is no published benchmark that quantifies how well any signal predicts that a named inventor is the technical contributor you want, so treat every weighting as a working assumption rather than a validated model.

US law helps a little. Each named person must have contributed to the conception of the invention, and practitioners note that despite the "inventor" label these are typically the engineers who did meaningful work. But conception is a weak filter against a manager or IP contact who appears for reasons you cannot see on the face of the document.

The practical signals, used with that caveat, are:

  • Position in the inventor list. A first-named inventor is more often the lead technical contributor than a fifth-named one, though this is a tendency, not a rule.
  • Filing count. Someone with many filings in the area is more likely a working engineer than a one-appearance name.
  • Portfolio continuity. Repeated appearance across a coherent body of work over time signals sustained technical involvement.
  • Cross-assignee appearance. The same person inventing under different employers over time is a mobility signal worth flagging, and one of the example searches worth running on its own.

Triage an inventor before you spend resolution time

High disambiguation confidenceLow disambiguation confidence
Uncertain and thin
Drop or park; not worth resolution effort
Promising but unresolved
Verify identity first, then resolve
Confident but likely name-on-paper
Deprioritize; probably a manager or IP contact
Confident and clearly technical
Resolve to current contact now
Weak contributor signalStrong contributor signal
Weigh how technical the contribution looks against how confident your disambiguation is.
The patent tells you who invented something years ago. Your job is to find who they are today.

How this goes wrong

Most patent-sourcing failures are false positives that survive to your shortlist and embarrass you in outreach. Each has a specific check that catches it before it costs you.

  • Stale employer. The assignee is the company at filing, and the inventor may have left years ago. The false positive is pitching someone as "at Intel" who left in 2021. Check: reconcile against current profile data, never the patent.
  • Common-name collision. One name across three assignees may be three people or one. The false positive is a merged super-inventor. Check: require co-inventor or geography overlap before merging, and lean on disambiguated IDs.
  • Name-on-paper inventor. A manager or IP contact listed for contribution reasons you cannot see. The false positive is ranking a non-technical person high. Check: use portfolio continuity and co-inventor role, and treat the ranking as heuristic.
  • Residence is not a home. City and state is point-in-time and may be a corporate mailing city. The false positive is targeting the wrong metro. Check: never geo-target off patent residence alone.
  • Non-publication gap. US-only filings with a non-publication request never publish, and provisionals never publish. The false positive is assuming full coverage of a company's R&D. Check: acknowledge the blind spot.
  • Wrong CPC scope. Relying on the first-listed CPC misses secondary art. The false positive is a thin, skewed inventor pool. Check: harvest secondary codes and rank by frequency.
  • Continuation timing illusion. A continuation of an old application can publish within weeks, making an old invention look recent. Check: read the priority or earliest filing date, not the publication date.

Before you call the shortlist done

A shortlist is ready when every row is one resolved person, tied to source patents, resolved to a current employer, and carrying one reachable contact path. Run this check before it leaves your hands.

Shortlist readiness

  • Each row cites the source patent number(s) it came from.
  • Every affiliation and location is re-resolved against current data, not taken from the patent.
  • Each disambiguated identity carries a confidence note, and every weighty merge has a co-inventor, assignee, or geography anchor.
  • Priority (earliest filing) dates were read, so no continuation is mistaken for recent work.
  • Rankings are labeled heuristic, with the signals used stated.
  • Duplicate inventor IDs are merged and mis-resolved names dropped.
  • A 10 percent random sample passed verification against a second source.
  • The non-publication and provisional blind spot is noted where coverage claims are made.

Keeping the shortlist current

Patent-sourced shortlists decay from two directions at once. The record keeps aging because new filings publish on the 18-month lag, and the people keep moving because the affiliations were never maintained in the first place. Both mean a list is a snapshot, not an asset you can shelve.

Re-run the classification search on a schedule that matches your recency window. If you search filings from the last two to four years, refresh the pool at least twice a year so you catch newly published applications from filings you missed. Re-resolve the current employer for anyone you did not reach on the first pass, because the field data is now further out of date than when you built the list.

Two moves extend the life of the work. First, save your validated CPC groups and seed patents so the scoping step becomes a lookup, not a rediscovery, next quarter. Second, watch the cross-assignee movers, such as former Infineon power-semiconductor inventors who have since changed employer, because a mobility signal is often a hiring signal, and the patent record is one of the few public places it shows up early.

Questions practitioners ask

Can I do this with Google Patents or do I need USPTO Patent Public Search?

Use USPTO Patent Public Search for the extraction step. Per practitioner accounts, USPTO exposes inventor location on the record while Google Patents historically did not, and Patent Public Search lets you query inventor, assignee, CPC, and dates as structured fields. Google Patents is fine for reading a single document, but the field-level classification and inventor queries this method depends on live in Patent Public Search at ppubs.uspto.gov/basic.

How stale is the employer listed on a patent?

At minimum 1.5 to 3 or more years old by the time you read it. Applications publish by default 18 months after the earliest filing date under 35 U.S.C. 122(b), and grant typically runs 18 to 36 months from filing. Crucially, the assignee and residence are frozen at filing and not maintained under MPEP 719.02, so you must always re-resolve the person to a current source rather than trust the patent's affiliation.

How do I tell two people with the same name apart?

Require a shared signal before you merge. Start with PatentsView's disambiguated inventor IDs, which use probabilistic clustering, but treat them as fallible. For a common name across multiple assignees, confirm a shared co-inventor, an overlapping assignee history, or matching geography before deciding two records are one person. Where the inventor self-asserts an ORCID, The Lens can link their patent inventorship, which gives you a self-declared anchor.

Does this work for every deep-tech field equally?

No. Pool size is domain-specific. In Refolk's index, US Robotics Engineer profiles outnumber Semiconductor Engineer profiles about 7.6 to 1 (1,658 versus 219), so the patent-first approach pays off far more for semiconductors, RF, and materials, where profile databases return little, than for robotics, where they return plenty. Match the effort to how thin the conventional channels are for your stack.

Which patent classification codes should I search, USPC or CPC?

Use CPC. The older USPC scheme was retired in June 2015 and applications filed after that date carry no USPC code. CPC, co-managed by the USPTO and the European Patent Office, is the live system. Find your codes by the seed-and-harvest method: collect 5 to 10 representative patents, extract their CPC codes, rank the most frequent, and validate the definitions in a CPC browser before you rely on them.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next