Refolk
PlaybookRecruiting and sourcing

From Patent Filings to a Ranked Inventor Shortlist

You can turn a technology area or competitor patent portfolio into a deduplicated, ranked list of inventors, each resolved to a current employer and a reachability signal.

17 min readLast reviewed October 8, 2026Read as Markdown

Key takeaways

  • In PatentsView's 2025-06-30 update, inventor disambiguation ran at 0.863 precision and 0.922 recall, so roughly 1 in 10 merged inventor records carries a split or lump error, and the errors skew toward false merges.
  • A patent shortlist is stale by construction: 18-month publication plus about 22.8 months average pendency means the assignee on a granted patent reflects where someone worked roughly two years ago, not today.
  • Inventor sequence order has no legal significance and forward citations are often examiner-added, so ranking on any single field systematically mis-orders primary versus padded inventors.
  • Specialism scarcity is steep inside AI: in Refolk's index, US computer-vision research scientists outnumber NLP ones about 3.9x (998 vs 256).
  • Geography multiplies scarcity: the same ML research-scientist profile is about 11x more common in the US than Germany (2,954 vs 263) in Refolk's index.
  • The PatentsView Search API allows 45 requests per minute per key and returns HTTP 429 over the limit, so a careless pull silently truncates into a date-skewed sample.

A technical sourcer needs the engineers and scientists credited on patents in a specific technology area, turned into a reachable, ranked hiring shortlist. This playbook is for in-house sourcers, talent leaders, and founders hiring deep-tech talent who want to work only from public patent data. It runs the job end to end: from a CPC class or a competitor's portfolio to a deduplicated, ranked list of inventors, each resolved to a current employer and a reachability signal.

Most published advice stops at "search the inventor's name." That advice is broken by three traps it never mentions: name disambiguation, the gap between the assignee printed on the patent and where the inventor actually works now, and the filing-to-publication lag that makes every patent a lagging indicator. This guide is built around those traps. Each one gets a step and a failure mode, so you finish with a shortlist you can defend rather than a pile of names.

What a patent shortlist can and cannot tell you

A patent shortlist is a list of people who provably contributed to an invention in a technology area, ranked by a composite of credible value signals, and resolved to where they work today. It is a strong source of hard-to-fake technical credibility and a poor source of current employment, and you must treat those two properties separately.

The credibility is real. To be named on a patent, an inventor must have contributed to the conception of at least one claim; if there are several claims, contributing to one is enough. That is a lower bar than "led the project," but it is a hard, legal bar that cannot be padded with a title. The employment, by contrast, is stale by construction, and the rest of this playbook exists to work around that.

22.8 months
Average total USPTO pendency to grant
Added to the 18-month publication lag, this is why the assignee on a patent reflects where an inventor worked roughly two years ago.

Three structural facts shape everything downstream. First, US non-provisional applications publish about 18 months after the earliest priority date under 35 U.S.C. 122(b), and the clock runs from the provisional date, not the later non-provisional filing. Second, average total pendency to grant runs about 22.8 months. Third, US-only filings with non-publication requests never publish at all. Together these mean the most recent 18-plus months of a target's R&D is invisible, and some inventors are structurally absent. Your shortlist is a map of where strong people were, not where they are.

The data source and its limits

The primary free source is the PatentsView Search API, at base URL https://search.patentsview.org/api/v1. It filters patents by CPC class and by disambiguated assignee organization, returns disambiguated inventors with persistent IDs, and exposes endpoints for patents, inventors, assignees, locations, and CPC classes. Every request needs an X-Api-Key header.

Two operational limits govern how you pull. The API allows 45 requests per minute per key; exceeding it returns HTTP 429 with a Retry-After header. Result windows are also capped, with some clients limiting to 1,000 records per query and others paging up to 10,000 per page against a 100,000-record ceiling. The database updates weekly, so a pull is reproducible within a week but will drift over time.

The reason to prefer PatentsView over a plain Google Patents search is disambiguation. PatentsView assigns a persistent inventor ID produced by a clustering method (Monath and McCallum, 2015) and a Jaro-Winkler model for assignees, and it publishes precision, recall, and F1 for each update. Those numbers are the single most important thing to internalize before you trust a count.

EntityUpdatePrecisionRecallF1
Inventor2024-09-300.8690.9220.895
Inventor2025-06-300.8630.9220.892
Assignee2025-06-300.9150.9850.949

Read this table as a warning, not reassurance. Inventor precision of 0.863 means false merges exist in roughly one of every eight to nine merged records, and because precision sits below recall, the errors skew toward lumping two people into one ID rather than splitting one person. In practice that means your most "prolific" inventor is also your most likely artifact. Assignee disambiguation is cleaner at 0.949 F1, which matters because you filter on assignee organization to scope a competitor's portfolio.

What each layer of the record proves

  1. Disambiguated inventor ID
    The real primary key; dedupe here, never on raw name, but QA the top ranks for merge errors.
  2. Patent + claims + CPC
    Proves a legal contribution to conception in a technology area; hard to fake.
  3. Forward citations
    A noisy value proxy, often examiner-added; only useful in a composite.
  4. Assignee organization
    A historical employer, roughly two years stale; never treat as current.
Trust descends as you move from the disambiguation ID toward the current employer.

Run the procedure end to end

The method has eight stages. Expect a focused day of work for a single technology area, split roughly as scoping and pulling in the morning, dedupe and ranking around midday, and employer and reachability resolution filling the afternoon. The time-consuming part is not the query; it is resolving current employers, which the API cannot do for you.

From CPC class to delivered shortlist

  1. Scope the target
    Define the technology as one or more CPC classes and/or a named competitor assignee, plus every spelling variant of the assignee name. Done is a written CPC list and a list of assignee name variants.
  2. Pull the raw set
    Query the PatentsView Search API filtered by CPC, assignee, and a date window, respecting 45 requests per minute and paging through everything. Done is full JSON of patents with inventor IDs, sequence, assignee, CPC, and citation counts, with rows reconciled against total_hits.
  3. Deduplicate on disambiguated inventor ID
    Collapse name mentions to the PatentsView inventor ID, never on the raw name. Done is one row per unique inventor ID with patent counts attached.
  4. Score and rank
    Rank by a composite of patent count in the CPC area, forward citations, claim counts, first or sole inventorship frequency, and recency. Done is a ranked list with every score column visible.
  5. Resolve current employer
    Map each top inventor to a present role using external profiles and the open web, and mark the patent assignee as historical. Done is each top candidate carrying a current-employer field with a confidence flag.
  6. Attach reachability signal
    Add a contact indicator such as a public profile, a GitHub account, or a published email on a recent paper. Done is each candidate with at least one reachable channel or flagged no signal.
  7. QA the merges
    Spot-check the highest-ranked inventors for split or lump disambiguation errors and assignee-name drift. Done is a reviewed shortlist with artifacts removed.
  8. Deliver the shortlist
    Export the ranked, deduplicated, employer-resolved list in a format the hiring manager can act on. Done is a handoff file with scores, current employer, and reachability per candidate.

Scoping notes that save a rerun

Scope on CPC class, not keyword. CPC is the classification examiners apply, so filtering by, for example, G06N for neural networks returns the right inventions regardless of how the applicant titled them. Collect assignee variants before you query: a single company often files under several legal entity names, subsidiaries, and acquired brands, and because assignee disambiguation is good but not perfect at 0.949 F1, a variant you miss is a block of inventors you never see. Pick a date window deliberately. Given the 18-month publication lag, a window ending "today" is mostly empty at the front; a window of the last five to seven priority years captures a settled body of work.

Pulling without truncating

The quiet killer in step two is rate-limit truncation. If you blow past 45 requests per minute, the pull stops with a 429 and you are left with a partial, often date-skewed sample that still looks complete. Always capture the API's reported total count and compare it to the rows you actually retrieved before you move on. If the two disagree, you are ranking a biased fragment.

Score inventors so the rank survives scrutiny

Rank on a composite, never on a single field, because every individual signal lies in a predictable way. The two tempting shortcuts, inventor sequence order and forward-citation count, are exactly the two that fail, so a defensible rank blends several weak signals into one that is harder to fool.

Sequence order carries no legal weight and no relationship to contribution. The convention of putting the primary contributor first is unreliable and firm-specific, so ranking by first-named inventor assumes a primacy that does not exist. Forward citations are the most widely used value proxy and correlate with patent value across many studies, but they are noisy: citations are frequently added by the examiner rather than by a later inventor building on the work, so a high count can reflect examiner behaviour or an incremental follow-on rather than quality.

Combine the following, and expose each as its own column so a hiring manager can see the reasoning:

  • Patent count in the target CPC - proves sustained focus in the area; lies when disambiguation has lumped two people, inflating one ID.
  • Forward citations - proves downstream influence; lies when citations are examiner-added or reward an incremental patent.
  • Claim counts - proves substance of contribution per filing; harder to pad than raw patent count.
  • First or sole inventorship frequency - proves repeated central involvement across filings, which is more credible than a single first-named position.
  • Repeat co-inventorship - proves a stable working relationship and a real research line, and helps confirm the disambiguation ID is one coherent person.
  • Recency - proves the work is still relevant and the person was recently active in the area.
No single patent field proves primacy, so rank on a composite or accept that your top names are noise.

The reason to combine rather than pick is that the errors are uncorrelated. A lumping error inflates patent count but scatters CPC classes incoherently; examiner citations inflate citations but not claim counts; sequence games the first-named field but not repeat co-inventorship. A candidate who scores high across all of them is far more likely to be a genuine primary contributor than one who tops any single column.

Resolve the assignee to a current employer

This is the step the old how-tos skip and the step that makes the shortlist usable. There is no published canonical method for mapping a disambiguated inventor to a current employer and reachable contact; no regulator or platform documents one. The practical workflow is to take the disambiguated inventor plus assignee plus location and cross-reference an external professional profile and the open web to find the current role, treating the patent assignee as historical throughout.

Why this cannot be skipped: publication lag plus pendency means the assignee data is roughly two years old. The person credited on a Siemens or Google patent may well have moved, and moving is exactly the signal a sourcer wants. The example queries that make this work are explicit about it, asking for inventors "now working somewhere else" or who "left Meta or Google in the last two years and are now at an early-stage startup." The move is the opportunity; the patent is only the proof of skill.

This is also the step where describing the person precisely in plain language removes the most friction. Instead of scraping and matching by hand, you can state the composite you just built as a single request.

Flag confidence on every resolution. A name, a location, and an employer that line up cleanly with a public profile is high confidence; a common name with no corroborating location is low, and low-confidence rows should not go to a hiring manager as facts. Refolk is where I close this gap, because asking for the current state of a person in English avoids the manual cross-referencing that eats the afternoon.

Attach a reachability signal

A ranked name with no channel is not actionable. For each top candidate, record at least one of a public professional profile, a GitHub account, or a published email on a recent paper, and flag anyone with none as "no signal" rather than quietly dropping them. Researchers in particular often publish contactable emails on papers, which is why the battery-materials example query asks for "a reachable email on a recent paper." A no-signal flag is honest information; a missing row is a hole the hiring manager will find later.

How this goes wrong

The failure modes below are the difference between a shortlist and a liability. Each has a tell and a check, and the most dangerous ones cluster at the top of your ranking, which is exactly where a hiring manager looks first.

Failure modeWhat it looks likeCheck
Lumping (false merge)A superstar with implausibly broad CPC spreadInspect CPC, location, and co-inventor coherence for top IDs
Splitting (false separation)A strong candidate ranks oddly lowSearch name variants, middle names, and location changes by hand
Sequence-order trapRanking leans on first-named positionUse claim count, citations, and repeat co-inventorship instead
Forward-citation inflationA thin patent scores surprisingly highSeparate applicant from examiner citations; weight by claims
Stale assigneeOutreach to a two-years-gone employerRe-resolve every top candidate to a live profile first
Non-publication blind spotA known expert never appearsTriangulate with conference papers and public GitHub

Lumping deserves the hardest scrutiny because the math predicts it. With precision at 0.863 and recall at 0.922, false merges are more common than false splits, and a merge inflates exactly the patent-count signal that lifts someone to the top of your list. When your number-one inventor spans wildly different CPC areas or co-invents with two disconnected clusters of people, suspect that the ID is two people, not one genius.

Splitting is the quieter loss. A strong inventor who changed name or moved cities can fracture into two IDs, each with half the patents, so neither clears your threshold and you never see the person. The only fix is manual: for names near your cutoff, search variants and location histories to see whether two low-ranked IDs are one high-ranked person.

The non-publication blind spot is structural and unfixable from patent data alone. US-only filings with non-publication requests and anything under a secrecy order never publish, so some of the best people in a sensitive area are simply absent. Treat the patent shortlist as one input and triangulate with conference papers and public GitHub for the same technology area, which surface people the patent record misses.

Scarcity: size the pool before you promise a shortlist

Before you commit to a patent search, sanity-check how many reachable people the technology area and geography can actually yield. A thin pool is not a sourcing failure; it is a fact to set expectations around, and the supply differences inside "AI" alone are larger than most hiring managers expect.

From filings to a reachable shortlist

  1. Patents in the CPC area
    full set

    lagging by 18-plus months

  2. Unique disambiguated inventors
    deduped

    ~1 in 10 records carries a merge or split error

  3. Composite-ranked top inventors
    shortlist

    sequence and citations alone mis-order this

  4. Resolved to current employer
    reachable

    scarcity multiplies here

Each stage narrows the set, and specialism plus geography can collapse the final pool to a fraction of the filings.

In Refolk's index, the specialism gap is steep. US computer-vision research scientists outnumber their NLP counterparts about 3.9 to 1, so an NLP-scoped patent search should expect a far thinner reachable pool and longer outreach cycles than a vision-scoped one, even if both start from a healthy set of filings.

SkillCountTop employer
Computer Vision998Google
Natural Language Processing256Meta
CV:NLP ratio (derived)3.9x-

Geography multiplies the same effect. The same ML research-scientist profile is about 11 times more common in the US than in Germany, so a Germany-scoped patent search converts a thin filing set into an even thinner hire-able set. If your target assignee is a European lab, plan for a smaller shortlist and a slower cycle from the start.

MarketCountTop employer
United States2,954Google DeepMind
Germany263Google
US:Germany ratio (derived)11.2x-
11.2x
US-to-Germany ratio of ML research scientists in Refolk's index
The same profile that is abundant in the US is scarce elsewhere, so geography-scoped patent searches need adjusted volume expectations.

A reusable score template and a final check

Lock your scoring into a written rubric so every run is comparable and every rank is defensible. Below is a skeleton you can paste into a sheet; the weights are a starting point, not a law, and you should tune them to the technology area.

Composite inventor score (starting weights)
score = 0.30 * patent_count_in_cpc
      + 0.20 * forward_citations
      + 0.20 * claim_count_total
      + 0.15 * first_or_sole_inventor_rate
      + 0.10 * repeat_co_inventor_rate
      + 0.05 * recency
columns to keep visible: inventor_id, name, patent_count, citations, claims, first_sole_rate, co_inventor_rate, last_filing_year, current_employer, employer_confidence, reachability_channel
flags: lumping_suspected (y/n), split_suspected (y/n), no_signal (y/n)

Normalize each raw column to 0-1 before weighting; adjust weights to the CPC area and document what you changed.

Before you hand anything over, run the shortlist against a closing checklist. The goal is that every claim in the file is one you could defend to a skeptical hiring manager.

Before you deliver the shortlist

  • Retrieved rows reconcile with the API's reported total_hits, so the pull is complete.
  • Dedupe was done on the disambiguated inventor ID, not the raw name string.
  • Every top candidate's CPC, location, and co-inventors were checked for a lumping artifact.
  • Names near the cutoff were checked for split IDs across variants and locations.
  • Rank uses the composite, not sequence order or citations alone.
  • Each top candidate has a current-employer field with a confidence flag, and the assignee is marked historical.
  • Each candidate has at least one reachability channel or an explicit no-signal flag.
  • Any sensitive area is triangulated with conference papers or GitHub to cover non-publication blind spots.

Keep the shortlist current

A patent shortlist decays, so treat it as a living artifact rather than a one-time deliverable. The PatentsView database updates weekly and new inventions keep publishing on the 18-month lag, so re-pull the CPC and assignee scope on a regular cadence and diff the new inventor IDs against your last run to catch people who have just crossed the publication threshold.

The faster-moving half is employment. Because the assignee is roughly two years stale the moment it publishes, the current-employer resolution is the field most likely to be wrong next quarter, especially for the people you most want: those in motion. Re-resolve current employers on any candidate you did not reach the first time, watch for the moves that turn a patent into an opening, and keep the reachability channels fresh so that when the hiring manager says go, the shortlist is ready rather than two years out of date.

Questions practitioners ask

Why can't I just search the inventor's name in Google Patents?

Because names are not unique and names change. Two different people with the same name get lumped into one record, and one person with a mid-career name or location change gets split across two. The PatentsView disambiguated inventor ID is the real primary key; it ran at 0.863 precision and 0.922 recall in the 2025-06-30 update, so even it is imperfect, but a raw name string is far worse. Dedupe on the ID, then QA the top of the list by hand.

How current is the employer on a patent?

It is not current. US applications publish about 18 months after the earliest priority date, and average total pendency to grant is around 22.8 months. The assignee on a granted patent therefore reflects where the inventor worked roughly two years ago. Treat the assignee as a historical employer and re-resolve every top candidate to a live profile before any outreach.

Is the first-named inventor the main contributor?

No. Inventor sequence order has no legal significance and no relationship to contribution. An inventor only needs to have contributed to the conception of one claim to be named at all. Rank on a composite of claim count, forward citations, repeat co-inventorship, and recency instead of position, because any single field, including sequence, will mis-order primary versus padded inventors.

What are the PatentsView API rate limits?

The Search API allows 45 requests per minute per key, and every request needs an X-Api-Key header. Exceeding the limit returns HTTP 429 with a Retry-After header. The real danger is silent truncation: if you hit the limit mid-pull, you get a biased, often date-skewed sample. Always compare retrieved rows against the reported total_hits and back off on 429.

Will I see a competitor's most recent research?

Not the last 18-plus months of it. Publication lag means the newest R&D is invisible, and US-only filings with non-publication requests never publish at all. For current signal in the same technology area, triangulate patents with conference papers and public GitHub activity, which surface people the patent record structurally misses.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next