# Sourcing a Skill Before It Has a Job Title

*You can turn a fuzzy capability description into a documented, scored search vocabulary and proxy-signal set that reliably surfaces and qualifies the people doing that work.*

- Canonical URL: https://www.refolk.ai/guides/sourcing-skill-before-job-title
- Pillar: Market and talent intelligence
- Format: Playbook
- Published: 2026-08-31
- Last reviewed: 2026-08-31
- Reading time: 15 min
- Keywords: sourcing a role with no standard title, find candidates with emerging job title, build search vocabulary for a new skill, proxy signals to find unnamed skill, map talent pool for emerging AI role, how to source before a title exists

## Key takeaways

- Standard taxonomies lag the market by design: O*NET's occupational structure was last revised in 2019 and ESCO ships major versions roughly every two years, so a capability emerging now has no code to search against.
- In Refolk's index, US "AI Engineer" (4,391) is 8.4x "Prompt Engineer" (524), and the high-volume title is the least legible: high count plus a wide definition means title-count overstates the real qualified pool.
- US "AI Engineer" outnumbers the UK 5.0 to 1 in Refolk's index, so a US-tuned vocabulary structurally under-recalls abroad and must be re-harvested per market rather than translated.
- Recall is the binding constraint in sourcing because a missed candidate costs more than a filtered false positive, and retrieval degrades log-linearly with pool size, so every metric must be reported against N.
- Proxy signals over-include unless paired: a tool mention catches readers, a repository catches forkers, and a methodology term like "vibe coder" catches commentary, so require a second co-occurring signal before trusting any single one.
- Freeze the vocabulary as a versioned, dated artifact and re-score it quarterly, because terms crystallize in months while taxonomies revise in years.

Some capabilities get hired for before anyone agrees what to call them. A company needs someone who can build production applications on top of large language models, or design evaluation harnesses, or keep a GPU cluster serving inference, and the job requisition reaches for a title that does not yet mean one thing. This guide is for strategy, research, and talent-intelligence teams who have to build a defined, searchable talent pool for exactly that kind of work. It gives you a start-to-finish method that ends in a documented, scored, reusable set of search vocabulary and proxy signals.

Most skills-taxonomy and talent-pool guides assume the skill already has a settled title you can type into a search box. This one covers the harder, earlier job: finding people whose work the market has not yet named, by building vocabulary and proxy signals from scratch, then proving the result is good enough to trust.

## Why keyword search fails before a title exists

When a capability has no standard title, a search built on titles returns either near-zero results or a flood of the wrong people. The market names work last, not first, so the label you would search for is the one thing that does not yet exist.

The lag is structural, not accidental. Standard occupational taxonomies revise on multi-year cycles, while the vocabulary for a new capability crystallizes in months. O*NET updates its database quarterly, but its occupational structure was last revised in 2019 and its skill descriptions run on roughly a five-year survey cycle. ESCO, the European classification, ships major versions about every two years. Meanwhile Lightcast reports there is still no dedicated occupational code for AI even though 51% of AI-requiring postings now sit outside IT departments entirely.

| Source | Cadence or figure | Date |
|---|---|---|
| ESCO major versions | ~2 years (v1.1 to v1.2) | Jan 2022 to May 2024 |
| O*NET database | Quarterly; structure last revised 2019 | 2019 / ongoing |
| LinkedIn skill change | 70% of skills change by 2030 | 2025 |
| WEF skill instability | 39% by 2030 (was 44% / 57%) | 2025 |

The scale of change underneath makes the gap worse. LinkedIn's Work Change report estimates that by 2030, 70% of the skills used in most jobs will change, up from its earlier 65% figure. The World Economic Forum is more conservative at 39% of core skills by 2030, though it notes that instability has slowed from a 2020 peak of 57%. Both measure net change, not a clean "made obsolete by generative AI" figure, and you should not report either as one. The point for a sourcer is directional: the vocabulary you need is moving faster than any published taxonomy can follow.

**51% - AI-requiring job postings now outside IT**

Lightcast (2025); there is still no dedicated occupational code for AI, so title search has nothing to anchor on.

So proxy signals are not a shortcut you reach for when you are lazy. They are the only mechanism available in the window between a capability appearing and a title settling on it.

## What the same title actually contains

A title count overstates the real qualified pool, and the most-hunted role is usually the least legible. High volume plus a wide, contested definition means the number on the tin tells you very little about how many people can do the specific work you need.

In Refolk's index of professional profiles, US "AI Engineer" returns 4,391 people, which is 8.4 times the 524 who match "Prompt Engineer" and a large slice of the 10,336 who match "Machine Learning Engineer". But practitioners describe "AI Engineer" as the fastest-growing and most oversaturated title in tech, covering everything from API-based application building to inference optimisation. At the big labs the title does not even appear: engineering roles there are specific to a domain such as performance, tokenization, infrastructure, or inference, and "Research Engineer" acts as a unified label covering what others split into ML Research Engineer, Applied Scientist, or Alignment Engineer.

| Title | US count | Multiple vs Prompt Engineer |
|---|---|---|
| Machine Learning Engineer | 10,336 | 19.7 |
| AI Engineer | 4,391 | 8.4 |
| Prompt Engineer | 524 | 1.0 |

Read this table as a warning about false confidence. The 4,391 "AI Engineer" profiles are not 4,391 people who do the same job. If your capability is a sub-specialism inside that number - retrieval, evaluation, serving infrastructure - the title tells you nothing about who qualifies. That is precisely why you build vocabulary and proxy signals instead of trusting the count.

> The market names work last, not first, so the label you would search for is the one thing that does not yet exist.

## The proxy signals that stand in for a title

The most reliable substitutes for a missing title are proof-of-work and tool signals, in roughly this order: authored code, specific tools, conference talks, community membership, and publications. Each proves something different, and each lies in a predictable way, so you need to know both.

The best sourcers have already made this shift. Most still source on profiles, resumes, titles, and keyword matches, but the strong ones source on proof of work, because a profile tells you what a candidate claims to know, not what they can do. The public GitHub graph lets you evaluate people through their actual repositories and contributions. Tool names are the sharpest proxy for emerging AI work: infrastructure roles cluster around CUDA, Triton, Ray, vLLM, SLURM, Kubernetes, and Terraform, while model-layer work shows up as fine-tuning tools like LoRA and QLoRA, serving stacks like vLLM and TGI, and evaluation frameworks.

Here is what each signal proves and how it lies:

- **Authored repositories.** Proves someone built and shipped. Lies when the "contribution" is a fork or a star, not authored work. Filter to authored commits and original repos.
- **Specific tool names.** Proves familiarity with the exact stack. Lies by catching people who read about the tool but never built with it. Require a second co-occurring signal.
- **Conference talks and write-ups.** Proves depth enough to explain the work publicly. Lies by over-representing the loud minority and missing quiet builders.
- **Community membership.** Proves proximity to practitioners. Lies by catching students, hobbyists, and the merely curious.
- **Publications.** Proves research contribution. Lies by favouring academics over people who ship production systems.

> **Rule:** Never trust a single proxy
>
> Every proxy signal over-includes on its own. Require at least two co-occurring signals - an authored repository plus a named tool - before you treat a profile as a real match.

One category to exclude on sight: methodology terms masquerading as roles. "Vibe coder" and "AI-native engineer" describe a practice or a quality, not a job. Searching them returns commentary and think-pieces, not practitioners.

## The seed-to-vocabulary method

You expand a fuzzy capability into a searchable vocabulary by starting from real people, not from your own guesses about what the work should be called. The documented method is OR-chain expansion grounded in how practitioners actually self-describe.

The core sourcing skill here is thinking like your ideal candidate and how they would describe their role and skills on a profile. You harvest that language from a small seed set of people you already know do the work, then group the synonyms into chains. Keep the string focused: two or three core role terms, two or three must-have skills, and a few exclusions is usually enough.

#### From capability to validated pool

1. **Frame capability** - Define the work by tools and outputs, name 5 to 15 seeds
2. **Harvest vocabulary** - Read seed profiles and repos, record self-descriptions and tools
3. **Build proxy inventory** - List signals and their expected false positives
4. **Draft and run string** - OR-chain titles, AND skills, NOT known false positives
5. **Score and iterate** - Precision@K and Recall@K against a hand-labeled set

*The vocabulary crystallizes before the population, so harvest language first and validate against a labeled set.*

The order matters more than it looks. Boolean guides tend to expand terms first and validate later. The evaluation literature implies the opposite: fix a labeled known-good set before you expand, so you can measure whether each change to the string helps or hurts. Follow the evaluation order. Without a fixed reference set you are tuning blind, and you will convince yourself a broader string is better when it has quietly dropped real candidates.

This is also where a plain-language search tool removes the most friction, because it lets you express the capability by its behaviour before you have a settled string.

I ran this search: `Find people who build LLM applications with retrieval - vector databases, embeddings, reranking - who ship production code but do not train foundation models.` - [see the full result list](https://www.refolk.ai/s/cdea2xdfp1).

*Returns practitioners defined by what they build and ship, across the public GitHub graph and public profiles, without needing a title that does not exist yet.*

A tool like [Refolk](/) is useful precisely at the framing and harvesting stages, when you know the work but not the word for it, and you want a seed set to read before you commit to a Boolean string.

## Run the procedure

This is the end-to-end method. It runs in one to two working days for a first pass, plus one or two iteration cycles. Roles in parentheses assume an analyst and a sourcer, but one person can run all of it.

#### Build the pool, start to finish

1. **Frame the capability, not the title** - Write a one-paragraph definition of what the person does and ships, anchored on tools and outputs. Name 5 to 15 seed practitioners. (Analyst; 2 to 4 hours.)
2. **Harvest vocabulary from the seed set** - Read each seed's profile and repos, recording self-descriptions, tools, artifacts, and adjacent titles. (Analyst; half a day.)
3. **Build a proxy-signal inventory** - List repositories, tools, talks, communities, and publications, noting the false positive each one invites. (Analyst; half a day.)
4. **Draft the search vocabulary** - OR-chain two or three title terms, AND in must-have skills, add NOT only for false positives you have seen, quote-mark multi-word titles. (Sourcer; 2 to 3 hours.)
5. **Run and pull a candidate pool** - Execute across the public GitHub graph, public LinkedIn records, and the open web, then de-duplicate. Record the pool size. (Sourcer; 1 to 2 hours.)
6. **Hand-label a known-good set and score** - Label true positives by hand, then compute Precision@K and Recall@K against them. (Analyst; half a day.)
7. **Iterate until the operating point holds** - Fix the labeled set, then run one or two cycles until metrics are stable at your chosen precision, restated against pool size. (Analyst and sourcer; 1 to 2 cycles.)
8. **Freeze and document the vocabulary** - Save a versioned string with dated proxy notes and the pool size the metrics were measured against. (Analyst; 1 hour.)

A copy-pasteable string skeleton for step four, using the emerging-AI stack as an example:

**OR-chain search skeleton for an untitled capability**

```
("AI Engineer" OR "LLM Engineer" OR "Applied ML")
AND ("vLLM" OR "LoRA" OR "retrieval" OR "embeddings")
AND (repo OR "shipped" OR "production")
NOT ("student" OR "aspiring")
```

*Replace the role terms and tools with vocabulary harvested from your own seed set. Add NOT clauses only after you see a false positive.*

The NOT line is deliberately short. Every exclusion you add before seeing a matching false positive is a real candidate you may be dropping silently.

## How to measure whether the pool is good

Score the pool with Precision@K and Recall@K against a hand-labeled known-good set, and never report a metric without the pool size beside it. These two numbers are the difference between a vocabulary you can defend and one you merely hope works.

The formulas are standard. Recall@K is the fraction of ground-truth positives that appear in the top K results. Precision@K is the fraction of the top K that are actually positives. Sourcing weights recall higher, because the cost of a false negative - a qualified person you never surface - outweighs the cost of a false positive that a recruiter will filter out in seconds.

There is no published universal "ready" threshold, and you should be suspicious of anyone who claims one. Teams pick an operating point, such as recall measured at a fixed precision, and state it against a specific pool size. That last part is not optional: both retrieval methods degrade log-linearly as the pool grows. A string that recalls well against the 524-person "Prompt Engineer" pool will miss more against the 10,336-person "Machine Learning Engineer" pool, purely because of scale.

#### The qualified pool is far smaller than the title count

| Stage | Figure | Note |
| --- | --- | --- |
| MLE title match | 10,336 | US, Refolk's index |
| AI Engineer title match | 4,391 | US, Refolk's index |
| Prompt Engineer title match | 524 | US, Refolk's index |

*Refolk's index counts illustrate how title volume narrows once you require the specific capability.*

> **Watch out:** A precision figure with no N is noise
>
> Reporting "82% precision" without the pool size is meaningless, because recall at fixed precision degrades log-linearly with pool size. Always write the metric as precision and recall at K, against a stated N.

Practically: build a labeled set of at least 30 to 50 confirmed positives before you tune, then sample the top K of every run and count hits. When precision and recall stop moving between cycles at your chosen operating point, you have converged.

## How this goes wrong

The failure modes here are specific and they repeat. Each one produces a plausible-looking result that is quietly wrong, so build a check for each into your procedure rather than trusting the output.

- **Title-only search misses the pool.** When no title exists, a title OR-chain returns near-zero or wildly off-target results, or a large hit count that is mostly students and enthusiasts. Check: sample 20 hits and confirm shipped artifacts before trusting the count.
- **Tool-name proxies over-include.** A "vLLM" or "LoRA" mention catches people who read about the tool, not people who built with it. Check: require a second co-occurring signal, such as an authored repo alongside the tool.
- **Repository signal rewards forkers.** A GitHub presence can be forks and stars rather than authored work. Check: filter to authored commits and original repositories, never stars.
- **Methodology mistaken for a role.** "Vibe coder" and "AI-native engineer" describe practice, not a job, and searching them yields commentary rather than practitioners. Check: exclude these as role terms.
- **Vocabulary drift silently rots the string.** Terms crystallize in months while taxonomies lag years, so a frozen string decays. Check: re-score against a fresh labeled set quarterly.
- **Precision reported without pool size.** A precision figure is meaningless without N, since recall at fixed precision degrades log-linearly with pool size. Check: always report metrics with the pool size attached.
- **Cross-market label mismatch.** The same work carries different labels by country and by employer, so a US-tuned string under-recalls in the UK. Check: re-harvest vocabulary per market.
- **NOT clauses added preemptively.** Adding exclusions before seeing a false positive silently drops real candidates. Check: add NOT only after a labeled false positive appears.

Two of these deserve extra weight. The cross-market mismatch is large and directional: US "AI Engineer" outnumbers the UK 5.0 to 1 in Refolk's index, so a string tuned on US self-descriptions will structurally under-recall abroad. Do not translate the string. Re-harvest vocabulary from local seed practitioners, because the labels genuinely differ, not just the language.

The other is the tension between labs and startups. The same practitioner might be a "Research Engineer" at a frontier lab and an "AI Engineer" or "RAG Engineer" at a startup. If your seed set skews to one setting, your vocabulary inherits that skew and misses the other. Deliberately seed from both.

## Keeping the vocabulary current

A frozen vocabulary is a decaying asset, so treat the output of this procedure as a versioned artifact with an expiry, not a permanent answer. The whole reason proxy signals exist is that standards lag, and your own string lags too the moment you save it.

The mechanism is straightforward once you accept it. Vocabulary crystallizes before the population does: shared tools and phrasing appear first, and the title is a lagging label that arrives later. That means a string tuned this quarter is measuring a moving target. Re-run the scoring cycle every quarter against a freshly labeled set, and watch for two things: recall dropping at your fixed precision, which signals drift in how practitioners describe themselves, and new tools entering the seed set's vocabulary that your string does not yet contain.

#### Before you call the pool done

- [ ] The capability is defined by tools and outputs, not by a title.
- [ ] Vocabulary was harvested from named seed practitioners, not invented.
- [ ] Every proxy signal has a second co-occurring signal required alongside it.
- [ ] Repository signals are filtered to authored commits, not forks or stars.
- [ ] NOT clauses exist only for false positives actually observed.
- [ ] Precision@K and Recall@K are computed against a hand-labeled set.
- [ ] Every metric is reported with its pool size.
- [ ] The string is versioned and dated, with a quarterly re-score scheduled.
- [ ] Vocabulary was re-harvested per market if the pool spans countries.

When the market finally settles on a title, this work does not become wasted. The vocabulary and proxy signals you documented become the bridge that maps the new title back to the population you already found, and your labeled set becomes the ground truth for validating whoever's official definition arrives. You will have been sourcing the capability accurately for the entire period the rest of the market was waiting for a word.

## Frequently asked questions

### How do I source a role that has no standard title?

Stop searching titles and search for what the person does. Define the capability by its tools and outputs, name 5 to 15 people you already know do the work, then harvest their self-descriptions, tools, and repositories into a search vocabulary. Combine that vocabulary with proxy signals like authored repositories and specific tool usage, then score the resulting pool against a hand-labeled known-good set so you can prove it works.

### What are the best proxy signals when a skill has no keyword?

Authored code, specific tools, conference talks, community membership, and publications, in that rough order of reliability. Proof of work beats claims: a profile tells you what someone says they know, a repository shows what they built. Every proxy over-includes on its own, so require a second co-occurring signal - a tool mention plus an authored repo, not a tool mention alone.

### Why can't I just wait for the taxonomy to catch up?

Because the lag is structural. O*NET's occupational structure was last revised in 2019 and ESCO ships major versions roughly every two years, while a new capability's vocabulary crystallizes in months. Lightcast reports there is still no dedicated SOC code for AI even as 51% of AI-requiring postings sit outside IT. Proxy signals are not a stopgap; they are the only mechanism available before a code exists.

### How do I know my search vocabulary actually works?

Measure Precision@K and Recall@K against a hand-labeled set. Recall@K is the fraction of true positives returned in the top K; Precision@K is the fraction of the top K that are real. Sourcing weights recall higher because a missed candidate costs more than one a recruiter filters out. Always report the number alongside pool size, since retrieval degrades log-linearly as the pool grows.

### Will a US search vocabulary work in the UK?

Not reliably. In Refolk's index, US "AI Engineer" outnumbers the UK 5.0 to 1, and the same work carries different labels across countries and between labs and startups. A US-tuned string structurally under-recalls abroad. Re-harvest vocabulary per market from local seed practitioners rather than translating the string you already have.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/sourcing-skill-before-job-title*
