The Domain-Authority Signal Reference: What Proves Someone Shapes a Field
You will assemble a ranked list of a niche's real shapers and, for every name, state which public signal earned the rank and how that signal could lie.
You need to name the handful of people who actually shape a niche field and defend the list to anyone who challenges it, using only public evidence. This reference is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It defines each authority signal precisely, states what that signal proves, and names exactly how it misleads, so the list you produce survives scrutiny instead of collapsing under the first sharp question.
Most expert-identification pages are either vendor demos or single-vertical checklists built for oncology or cardiology. This one is field-agnostic and treats each authority signal as evidence with a failure mode attached, because that is what separates a defensible list from a list of famous names.
What counts as a domain-authority signal, and where does it come from
A domain-authority signal is a public artefact that provides evidence someone shapes a field, drawn from a named, retrievable source. No single signal carries the weight alone; a rigorous map combines several families and looks for the same name across independent source types.
Practitioners in the most mature version of this work, pharma KOL mapping, combine evidence from publications, clinical trials, congress activity, digital signals, and field insights to see who holds authority, who is rising, and how information flows across the expert ecosystem. The signal families port cleanly to any field:
- Peer-reviewed publications and citation counts. The core scholarly-output signal.
- Clinical-trial and grant registries. Who is trusted to run or fund the work.
- Conference speaker and session-chair rosters. Who peers choose to hear from.
- Guideline and standards committees. Who writes the rules the field follows.
- Patents. Who files claims on the field's inventions.
- Share-of-voice and digital signals. Where the conversation concentrates.
The concept itself is not new. The key-opinion-leader idea originates in 1940s communications research, and the discipline has spent eight decades learning that the artefact is easy to count and hard to interpret. The rest of this reference is about the interpretation.
Authority versus reach: the distinction that decides the list
The separating evidence is the source of trust. A domain authority is trusted for credentials and proven knowledge; an influencer drives attention through audience connection and content reach. Confusing the two is the single most common way an expert list goes wrong.
A key opinion leader's credibility comes from qualifications and experience, not follower count. That means a follower-count ranking will systematically over-rank influencers, because reach and authority are built differently and measure different variables. The check is not to ignore digital presence but to read it correctly: focus on engagement quality, audience composition, and the ability to drive meaningful conversation, not the raw number next to the name.
A list built on reach alone is not a weak authority list. It is measuring the wrong variable entirely.
This is why the separation is its own step in the procedure below. If a name's rank rests only on visibility, and that visibility was bought or is broad but shallow, the name does not belong on an authority list no matter how large the audience.
The h-index and why it cannot cross fields
The h-index proves sustained, cited scholarly output within one field and career stage. It proves nothing across fields, and treating it as a universal yardstick is a category error rather than a calibration problem you can patch.
Hirsch derived his original figures from physics: an h-index of 12 as a typical value for tenure, 18 for membership in a prestigious national academy. Those numbers do not travel. A neurosurgery study found the h-index is not field-normalized and is inappropriate for comparison of authors from different fields. Hirsch himself acknowledged that citation cultures, field sizes, and publication formats differ so much that cross-disciplinary comparison is inherently misleading.
The reported field ranges make the gap concrete.
| Field | Reported "strong" range | Source |
|---|---|---|
| Medicine / life sciences | 30 to 60 | scirankglobal.org |
| Social sciences | 15 to 25 | scirankglobal.org |
| Humanities / history | 5 to 10 | hastewire.com |
These are illustrative, not normative, and both source blogs note that cross-field comparison is invalid. A medicine h-index of 40 and a mathematics h-index of 20 are not on the same axis. Ranking the medic above the mathematician for the same niche is not a close call resolved by the numbers; it is a mistake baked into the metric.
The disciplined rule is narrow: compare within field and career stage, hold the database constant, and never compare a Google Scholar figure against Scopus or Web of Science. State which database and which date every figure came from. Open bibliometric infrastructure now makes field-normalised comparison feasible: OpenAlex covers more than 250 million works and 90 million authors, and a Scopus data freeze captured roughly 230,333 career-long top-cited authors out of 10.9 million with five or more papers, so the peer set for normalisation exists.
The signal reference: what each one proves and how it lies
Every authority signal is evidence, and every piece of evidence has a failure mode. Read this table as the core of the reference: jump to the row for the signal you are holding, and read the "how it lies" column before you rank anyone on it.
| Signal | What it proves | How it lies |
|---|---|---|
| Citation count | Sustained cited output within a field | Lags by ~3.3 years plus 6-12 month indexing delay, so current leaders look weak |
| h-index | Depth of cited work at one career stage | Not comparable across fields or across databases |
| Authorship volume | Breadth of published contribution | Gift and honorary authorship inflate raw counts |
| Speaker slot | Peers chose to hear this person | Predatory and sponsored slots are pay-to-play |
| Patent count | Filed claims on field inventions | Many patents are defensive or unused |
| Follower count | Audience reach | Reach is not expertise; over-ranks influencers |
Two entries deserve expanding because they invert the naive reading.
Citations lag, so they penalise the present. A coauthorship study of statisticians found a mean citation delay of 3.30 years, with self-citation faster at 2.81 years versus 3.51 for distant citations. Add a 6 to 12 month indexing delay that creates invisible impact during evaluation windows. The consequence: a field's current shapers can look weaker than its past ones, and the faster the field moves, the worse the distortion. Check citation velocity, not just the stock.
Patent volume is the wrong count. Many patents are never used or are filed for defensive purposes, which undermines counts as an innovation signal. The authority lives in forward citations: each additional citation per patent was estimated to boost that patent's market value by about 3 percent. Patent-quality indicators also overlap only weakly and in technology-dependent ways, so a large portfolio is not leadership. Count forward citations and renewals, not filings.
How each signal gets gamed
Every signal in the table above can be manufactured, and the manufacturing methods are well documented. Knowing them is the difference between a list that reads as authoritative and one that is.
A bibliometric review catalogues the scholarly-output methods: authorship-based gaming through gift authorship and salami slicing of credit, citation-based gaming through massive self-citation and citation farms, editorial nepotism, journal impact-factor gaming, and paper mills. On the speaking side, predatory conferences run a pay-to-play model where researchers pay to speak and organizers show little concern for the rigour of abstracts or speakers. Even at legitimate events, thought leadership is for sale: pay to speak and you look like a thought leader.
The most useful pattern for a screen is that gaming clusters. It does not appear in isolation. The Research Integrity Risk Index combines delisted-journal publications, retractions, and self-citation shares across five risk tiers, and the study behind it flagged 21 of the world's fastest-growing research universities with bibliometric anomalies travelling together. So one integrity flag should trigger a full screen, not a single-point deduction.
Reading a candidate on authority versus reach
Halfway through a real map, the friction is rarely the ranking logic. It is assembling a candidate pool from public sources across publications, speaker rosters, committees, and digital signals without spending a week per field. A plain-language search that spans the public web, LinkedIn, and GitHub collapses that assembly step, which is where I built Refolk to save the time.
The procedure: from named niche to defended list
Run these steps in order. Each has a done-condition so you know when to move on, and each maps to a named source in the dossier. Total elapsed time runs roughly two to four weeks for a serious field.
Building and defending the list
- Scope the niche and the questionDefine the field boundary, geography, and the job the list serves, and write all three down. Done when the niche, geography, and use case are fixed on paper. Roughly one day.
- Assemble the candidate poolPull candidates from publications, trial and grant registries, congress speaker rosters, guideline committees, patents, and digital signals. Done when each candidate appears in at least two independent source types. Three to ten days.
- Retrieve per-signal evidenceFor each candidate, record the raw figure plus the database it came from and the date you pulled it. Done when every name carries source-stamped values. Two to five days.
- Normalise within field and career stageNever compare raw h-index across fields; benchmark against same-subfield peers of similar years since first publication. Done when scores are field-relative. One to two days.
- Separate authority from reachFlag anyone whose ranking rests only on follower count or paid visibility. Done when each name has an authority basis, not just a reach basis. One to three days.
- Run integrity screensCheck self-citation share, honorary-authorship patterns, predatory-venue speaking, and defensive patent volume. Done when each candidate is cleared or annotated. Two to four days.
- Weight, combine, and rankCombine signals into one ordered list and keep the weighting explicit and reproducible. Done when every rank traces to named evidence. One to three days.
- Peer review and set a refresh dateAdd qualitative expert judgement over the scores and fix a re-run date. Done when a named reviewer signs off and a cadence is set. Ongoing.
Two things vary across sources. First, order at step seven: some practitioners run network mapping before final ranking, others rank first and map networks after. Either works if the weighting stays explicit. Second, there is no documented public weighting formula; vendors keep theirs proprietary, one describing a Most Influential Index calculated between 0 and 1 per expert. So do not chase a formula. The defensible principle from research assessment is that a good process rests on qualitative expert judgement of actual contributions rather than a standalone bibliometric lookup table. The CoARA coalition and the Hong Kong Principles commit signatories to exactly that.
The evidence path for a single name
- Candidate foundAppears in at least two independent source types
- Evidence stampedEach figure carries its database and date
- NormalisedScored against same-field, same-stage peers
- Reach separatedCleared of follower-only or paid-visibility basis
- Integrity screenedSelf-citation, honorary authorship, predatory venues checked
- RankedPlaced with every input traceable to named evidence
Geography and base rates: why a global list skews without normalising
Authority supply is not evenly distributed, so a global list built without geographic normalisation over-represents whichever market has the larger base rate. Refolk's index makes the size of this effect visible.
In Refolk's index of professional profiles, the count of US machine-learning people whose headline includes "keynote speaker" is far larger than the UK equivalent, and not because US experts are individually more authoritative.
| Market | Skill | Count | Relative to US |
|---|---|---|---|
| United States | Machine Learning | 262 | 1.0x |
| United Kingdom | Machine Learning | 53 | 0.20x |
The US pool is about 4.9 times the UK pool for the same title and skill. If you build a global machine-learning authority list by volume, US names dominate on base rate alone, before any judgement of individual authority. Normalise by geography, or set market-specific quotas, so the list reflects who shapes the field rather than which country is larger.
The same base-rate caution applies across skills within one market. In Refolk's index, the US "keynote speaker" pools for two adjacent fields sit close together but are not identical.
| Skill | US count | Share of pair |
|---|---|---|
| Machine Learning | 262 | 53.5% |
| Cybersecurity | 228 | 46.5% |
Read this as calibration, not ranking. It tells you roughly how deep the speaker-signal pool runs in each field, so you know whether a shortlist of ten names is skimming the top or reaching into the middle.
The failure modes: how a defensible list still goes wrong
This is the most valuable section, because a standard that overclaims is worse than none. Each row below is a way a rank turns out false, the false positive it produces, and the specific check that catches it.
- h-index across fields. A medicine h=40 outranks a mathematics h=20 for the same niche. The medic looks dominant. Re-benchmark within subfield and career stage, and confirm the same database and date.
- Speaker slots. A paid or sponsored slot reads as peer selection. A vendor-funded keynote looks like recognition. Check whether the slot was tied to sponsorship and whether the venue peer-reviews abstracts.
- Citation counts. Lag makes recent work look weak. A stale name outranks a rising one. Check citation velocity and account for the roughly 3.3-year mean delay plus 6 to 12 month indexing lag.
- Self-citation. Insular citing inflates rank. A citation-cartel member looks widely cited. Check self-citation share and co-authorship network density.
- Authorship volume. Gift and honorary authorship pad counts. A hyper-prolific name looks central. Check first and corresponding-author share, not raw count.
- Patent count. Defensive and unused filings pad portfolios. A large portfolio looks like innovation leadership. Check forward citations and renewals, not volume.
- Follower count. Reach is mistaken for expertise. An influencer outranks a credentialed expert. Check credentials, engagement quality, and audience composition.
Use this template to annotate each name so the defence writes itself. When someone challenges a rank, you read the row.
Name: Field / subfield: Primary signal + value: Database + pull date: Second source type: Reach basis vs authority basis: Integrity flags checked: How this rank could be wrong:
Fill one per candidate. The "how it could lie" cell is the one that survives challenge.
Before you publish the list
Run this checklist before you hand the list to anyone who will act on it. Each item maps to a failure mode above, so passing all of them means the list can be defended row by row.
List-defence checklist
- Every ranked name appears in at least two independent source types.
- No h-index is compared across fields, and every figure carries its database and pull date.
- Citation-based ranks account for the ~3.3-year mean lag and 6 to 12 month indexing delay.
- Each name has an authority basis, not just a reach or paid-visibility basis.
- Self-citation share, honorary-authorship pattern, predatory venues, and defensive patent volume are screened per name.
- The list is normalised by geography, or market-specific quotas are set, so base rates do not skew it.
- The weighting is written down and reproducible, and every rank traces to named evidence.
- A named reviewer has signed off with qualitative judgement, and a refresh date is fixed.
Keeping the list current
A defended list decays, and the decay has a mechanism you can schedule against rather than a calendar rule to memorise. There is no established public refresh cadence, so set your interval from the drivers.
The main driver is citation lag: a mean of 3.30 years plus a 6 to 12 month indexing delay means a field's rising shapers keep getting under-counted until the citations catch up. The faster the field moves, the shorter your interval should be, because a machine-learning list ages faster than a history list. The second driver is integrity: flags emerge over time, and because they cluster, a name that was clean at first pass can surface a full pattern later. Re-run the integrity screen on the whole list at each refresh, not only on new names.
At sign-off, fix a re-run date and treat it as a floor. When a major congress publishes its next speaker roster, a new guideline committee forms, or a trial registry updates, refresh early rather than waiting. The point of the whole method is that the list stays true to who shapes the field now, and now keeps moving.
Questions practitioners ask
How do I identify key opinion leaders in a field I don't know well?
Start from public sources, not reputation. Pull publications and citation counts, clinical-trial or grant registries, conference speaker and committee rosters, patents, and digital signals, then require each candidate to appear in at least two independent source types before you rank them. Because you are new to the field, lean on structural signals like guideline-committee membership and first-author share rather than raw citation totals, which lag and reward volume.
What h-index counts as strong?
There is no universal threshold, and a single number is not established publicly. Reported strong full-professor ranges are 30 to 60 in medicine and life sciences, 15 to 25 in social sciences, and 5 to 10 in humanities. Hirsch's original physics figures were 12 for tenure and 18 for national-academy membership. Never compare an h-index across fields or across databases; state which database and date every figure came from.
How do I tell a real domain expert from a high-reach influencer?
Look at the source of trust. A domain expert's credibility comes from credentials, proven knowledge, and structural roles like committee seats and first-author publications, while an influencer's comes from audience connection and content reach. Check engagement quality and audience composition, not follower count, and flag any name whose rank rests only on reach or paid visibility.
How often should a KOL or expert map be refreshed?
A fixed cadence is not established publicly. Set your interval from the decay drivers rather than a rule of thumb: citations lag by a mean of about 3.3 years plus a 6 to 12 month indexing delay, and integrity flags emerge over time. Fast-moving fields need shorter intervals because citation-based ranks fall further behind current shapers. Fix a re-run date at sign-off and treat it as a floor, not a ceiling.
Can conference speaking slots be trusted as an authority signal?
Only after you check how the slot was earned. Predatory conferences run a pay-to-play model where speakers pay to present and organizers do little quality control, and even at legitimate events sponsored keynotes can be bought. Treat a speaker slot as authority only when the venue peer-reviews abstracts and the slot was not tied to sponsorship. Otherwise it measures spend, not recognition.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.