# The Repo Sourcing-Yield Score: Mine It, Skim It, or Skip It

*You will score any repository across six dimensions and reach a Mine, Skim, or Skip verdict that predicts its reachable, poachable, on-stack yield.*

- Canonical URL: https://www.refolk.ai/guides/repo-sourcing-yield-score
- Pillar: Recruiting and sourcing
- Format: Framework
- Published: 2026-09-16
- Last reviewed: 2026-09-16
- Reading time: 17 min

You have fifteen repositories open in an ecosystem and time to mine two of them properly. This guide is for in-house recruiters, sourcers, and founders doing their own hiring who need to decide which repositories are worth pulling contributor lists from before spending hours on the wrong ones. It gives you a six-dimension score that produces a Mine, Skim, or Skip verdict predicting how many reachable, poachable, on-stack engineers a repo will actually yield.

Most repo-sourcing guides start after the hard decision is already made. They tell you how to turn one chosen repo's contributor graph into a shortlist, which is a solved problem. What they never tell you is how to choose the repo, so sourcers burn tens of hours mining projects that turn out to be all bots, all locked at one employer, or all unreachable. This scores a repository as a sourcing pool, on the axis of yield per hour rather than code quality.

## Why score a repo before mining it

Because mining is expensive and most repos are bad pools. One recruiter source puts manually sourcing 50 qualified developers at 25 to 40 hours, which works out to roughly 30 to 48 minutes per qualified contributor. Mine fifteen repos blindly and you have committed tens of hours before you know whether any of them contain reachable people.

The score converts that blind spend into a triage decision. Three things routinely turn a rich-looking repo into a dead pool, and none of them are visible from the repo's front page:

- **Bots inflate the count.** Dependabot, CI, and org-service identities sit in the contributor graph and are indistinguishable from humans until you run a classifier.
- **Reachability is a fraction of the count.** GitHub defaults to private email, so a 200-name graph may surface around 30 directly contactable people.
- **Employer lock removes poachability.** A company-owned repo can look talent-rich while every contributor is locked at the same hard-to-poach employer.

**30-48 min - Time to source and qualify one developer manually**

Derived from a recruiter estimate of 25 to 40 hours for 50 qualified developers, so mining fifteen repos blindly is tens of hours you can avoid.

The framework's whole job is to spend a few minutes per repo up front so you spend the tens of hours only where they pay back. The output is not a code-quality rating. It is a bet on how many people you can actually reach, who are actually active, and who you could actually move.

## The six dimensions of sourcing yield

Score every repo on the same six dimensions, each answering one question about the pool rather than the code. Each dimension proves something specific, and each has a characteristic way of lying to you.

| Dimension | What it proves | What it looks like when it lies |
|---|---|---|
| Ecosystem population | The absolute candidate ceiling in your geography | A hot language with almost no local population |
| Distinct human contributors | The real headcount after bots are stripped | 400 contributors, half of them CI and dependabot |
| Reachability | The share you can actually email | Big active repo, every maintainer email masked |
| Activity recency | The share still committing, not disengaged | All-time count full of developers gone for a year |
| Employer concentration | Whether the pool is poachable or locked | Company-owned repo, one employer near 100% |
| Seniority fit | Whether the band matches your role | A large pool that is mostly the wrong level |

The order matters. Ecosystem population caps everything: if the language has no local population, no repo inside it can save you. Reachability and employer concentration are the two that most often turn a Mine into a Skip, so they carry the most weight in the final call. Seniority fit is a filter you apply against your specific brief rather than a property of the repo alone.

> **Rule:** Score the subset, not the badge
>
> Stars, forks, and headline contributor counts describe attention, not people you can hire. The only count that predicts yield is distinct-human, recently-active, reachable contributors with a poachable employer spread.

Weight the dimensions to match your constraint. If your bottleneck is reply rate, reachability dominates. If it is poaching from a locked competitor, employer concentration dominates. The dimensions are fixed; the weights are yours.

## Ecosystem population comes before any repo

Pick the ecosystem before you pick the repo, because ecosystem choice moves yield by multiples and repo choice moves it by fractions. Confirm the target language or framework has real population in your geography and get an absolute pool size, so you know the ceiling before you score a single project.

The gap between ecosystems is large enough to overwhelm any repo-level difference. In Refolk's index the US Go pool is 16,223 professionals against 3,526 for US Rust, so Go is 4.6x the pool before any single repo is scored. Geography compounds it: the US Rust pool is 3.6x the German Rust pool.

| Skill | Country | Matching professionals |
|---|---|---|
| Rust | United States | 3,526 |
| Rust | Germany | 977 |
| Go | United States | 16,223 |

*Counts from Refolk's index; multiples derived.*

Read this table as a ceiling. If your role needs Rust in Germany, the entire national population is 977 people, and no repo inside that ecosystem yields more than that pool allows. Knowing the ceiling changes how hard you should work each repo: in a 977-person pool you mine deeper and accept lower reachability, because there is nowhere else to go.

**16,223 - Go professionals in the United States, from Refolk's index**

Against 3,526 for US Rust, a 4.6x gap that no repo-level scoring can close.

Seniority is the second population cut, applied against the brief. In Refolk's index, 1,244 of the 3,526 US Rust professionals sit in the senior band, or 35.3% of the pool. If your role is senior-only, your real ceiling is that 1,244, and a repo full of junior committers scores low on fit no matter how reachable it is.

| Band | Count | Share of skill pool |
|---|---|---|
| All Rust (US) | 3,526 | 100% |
| Senior-band Rust (US) | 1,244 | 35.3% |
| Implied non-senior | 2,282 | 64.7% |

*Totals from Refolk's index; shares derived.*

## Reachability is the binding constraint

Score reachability early, because it is the dimension most likely to turn a large active repo into a dead pool. GitHub flipped to private-by-default email, so the share of contributors you can actually contact is often a small fraction of the graph.

The mechanism is worth understanding so you can measure it rather than assume it. Setting an email to private hides it from the profile but does not affect its visibility in commits to public repositories, so the commit metadata is your fallback. But the noreply.github.com address combined with the keep-email-private setting removes the real email from commit metadata entirely, which closes that fallback. Vendor-reported figures put profile-email exposure at roughly 15% of GitHub users, with direct discovery working for about 10 to 15% of developers, meaning for the other 85% you need enrichment.

#### From contributor graph to reachable people

| Stage | Figure | Note |
| --- | --- | --- |
| Headline contributors | 200 | What the repo page shows |
| Distinct humans | 100 | After de-botting the graph |
| Recently active | 55 | After filtering to last-year committers |
| Directly reachable | 30 | At roughly 15% profile-email exposure applied to the graph |

*A 200-name graph narrows sharply once bots, disengagement, and email masking are applied.*

The funnel numbers are illustrative of the shape, not a fixed law, but the shape is the point: score the reachable fraction on a 20-contributor sample before you commit to the full pull. If your sample surfaces two emails out of twenty, the repo is a Skip regardless of how many contributors it has.

Two cautions on reachability signals. First, do not treat a commit email as identity on its own: only around 10% of commits in one 60-repo analysis were verified, and unverified emails can be spoofed, so cross-reference the profile or account before trusting an address. Second, historical exposure is sometimes recoverable, because GH Archive has recorded GitHub's public event firehose hourly since 2011, but that is an enrichment path, not a reason to assume reachability.

## Employer concentration decides poachability

Tally current employers across the human contributors, because a pool locked at one employer is not a pool you can hire from. This is the dimension that most often makes a company's own repo a Skip despite looking talent-rich.

There is no published sourcing-specific employer concentration index, so the framework defines its own read using two contrasting shapes. The closest documented proxy from research is the truck factor: the minimum number of developers who need to leave for a project to stall. A repo with a truck factor of one or two has no poachable bench, just an owner. But truck factor measures fragility, not poachability, so for the hiring question the real signal is current-employer spread.

Refolk's index supplies the missing number. Across a 25-profile sample of a broad skill pool, no single current employer dominated:

| Pool | Top employer's share of 25-profile sample |
|---|---|
| Rust, US | 4% (1 of 25) |
| Rust, Germany | 8% (2 of 25) |
| Go, US | 4% (1 of 25) |
| Rust, US, Senior | 8% (2 of 25) |

*Sample composition from Refolk's index; shares derived.*

Read this against a single-org repo, where one employer would sit near 100%. That contrast is the entire poachability signal. A broad skill pool is employer-dispersed, so any one hire barely dents your relationship with a given company and no single hard-to-poach employer blocks the pool. A company-owned repo concentrates the risk and the lock in one place.

> The verdict hinges on which shape the repo is: a dispersed skill pool you can poach, or a single-org bench you cannot.

This is where the framework earns its keep, because employer spread is invisible from the repo's front page. You only learn it by resolving current employers across the contributor set, which is slow to do by hand across fifteen repos.

I ran this search: `Senior Rust engineers in the United States who have recently committed to public async-runtime or web-framework repositories.` - [see the full result list](https://www.refolk.ai/s/bzme00dv84).

*Returns a dispersed, recently-active, on-stack pool with current employers resolved, so you can read poachability before you mine a single repo.*

Asking [Refolk](/) for the people directly inverts the workflow: instead of mining repos to find engineers, you describe the engineers and get them back with employer and recency already attached. That collapses the reachability and employer-concentration scoring into one query.

## Activity recency separates the live bench from the graveyard

Filter to recent committers, because a name in the contributor graph is not a reachable engineer who still works in this stack. Ranking by cumulative contributors rewards stale projects and overstates the live bench.

The abandonment research is blunt about this. All open-source core developers take at least one break, 45% completely disengage for at least a year, and developers have a 35% to 55% chance of returning after abandoning a project. So nearly half of a repo's all-time core may be gone at any moment, and a returning committer is a coin flip.

The practical filter is commit recency. One recruiter source defines an active contributor as one making 3 or more weekly commits and links that to retention and response rates, though the methodology there is unverified, so treat the specific threshold as a starting point rather than a law. What is defensible is the direction: filter to committers active in the last 6 to 12 months and down-weight everyone else. An all-time count of 400 with 55 recent committers is a very different pool from an all-time count of 400 with 300 recent committers, and only the recent figure predicts a reachable, on-stack engineer.

> **Watch out:** All-time contributors are not the live bench
>
> Because 45% of core developers vanish for a year or more, ranking repos by cumulative contributors rewards dead projects. Rank by last-year committers or you will mine names that no longer code in the stack.

## The scoring procedure

Run these steps in order for each repo, then rank. The first three set up the field; the middle four apply the dimensions; the last combines them into a verdict.

#### Score a repository as a sourcing pool

1. **Pick the ecosystem before the repo** - Confirm the target language or framework has population in your geography and get an absolute pool size, so you know the ceiling before scoring any repo.
2. **Assemble the candidate repo shortlist** - Gather 10 to 20 repos in the ecosystem via search, topic tags, and dependency graphs, with stars, forks, and headline contributor counts pulled.
3. **Pull the contributor list via API** - Authenticate to get 5,000 requests per hour instead of 60, and produce a raw contributor and commit-email set per repo with request budget tracked.
4. **De-bot the list** - Run BoDeGiC or BoDeGHa, or a rules pass, to strip dependabot, CI, and org-service identities, leaving a distinct-human-contributor count.
5. **Score reachability** - Resolve profile and commit-metadata emails and flag noreply-masked accounts, producing a reachable fraction measured on a 20-contributor sample.
6. **Score activity** - Filter to committers active in the last 6 to 12 months and down-weight the disengaged, leaving an active-and-reachable count.
7. **Score employer concentration** - Tally current employers across the human contributors; a dispersed pool scores high, a single-org pool scores low, producing a concentration ratio.
8. **Compute the six-dimension verdict and rank** - Combine the dimensions into Mine, Skim, or Skip per repo and rank, leaving the two repos worth working out of fifteen.

Two operational notes. The de-bot pass uses a trained classifier: BoDeGiC scores commit messages at 0.80 precision across 6,922 contributors, and BoDeGHa scores issue and pull-request comments, both built on ground truth of 5,000 accounts with 527 labelled bots. The API pull has a hard ceiling: 60 requests per hour unauthenticated against 5,000 authenticated, so authenticate first and budget requests per repo or a plan to mine fifteen repos dies at request 60.

Sources disagree on order. Some recruiter guides start from a target company's org, skipping the ecosystem check, which is exactly how you end up with a single-employer pool you cannot poach. Do the ecosystem step first.

## Turning the score into a verdict

Combine the dimensions into one of three verdicts. The two axes that carry the most weight are reachable-and-active human contributors and employer dispersion, because those are the two that most often collapse a promising repo.

#### The Mine, Skim, Skip verdict

Horizontal axis runs from Few reachable-active humans to Many reachable-active humans. Vertical axis runs from Locked at one employer to Dispersed across employers.

| Quadrant | What it means |
| --- | --- |
| Skip | Small pool, all locked at one org; no yield and no poachability |
| Skim | Dispersed but thin; pull the handful of reachable names and move on |
| Skip | Rich but locked; a company repo where every contributor is unpoachable |
| Mine | Many reachable-active humans across many employers; work this one deep |

*Plot each repo on reachable-active headcount against employer dispersion to read its verdict.*

- **Mine** is the top-right: many reachable, recently-active human contributors spread across many employers. Work these deep, and expect only one or two of fifteen repos to land here.
- **Skim** is a dispersed but thin pool. Pull the handful of reachable names, do not build a whole workflow around it, and move on.
- **Skip** covers both a small locked pool and a rich locked pool. Employer lock is disqualifying regardless of size, and a small dispersed pool rarely repays the mining time.

The verdict is a bet on yield per hour. A Mine repo that yields 30 reachable, on-stack, poachable engineers at 30 to 48 minutes each is a real week of qualified pipeline. A Skip repo that looks identical from the front page yields nothing after you have spent the same hours discovering the lock.

## How this goes wrong

The failure modes below are the highest-value part of the framework, because each is a way a repo passes the eye test and fails the mine. Every one has a false positive and a cheap check.

- **Headline contributor count as the score.** A repo shows 400 contributors, but half are bots and most all-time committers have disengaged. Check by de-botting with BoDeGiC and filtering to last-12-month committers before you count anything.
- **Reachability assumed from the profile field.** A big active repo whose maintainers all masked their email looks minable; you pull 200 and reach 20. Check the reachable fraction on a 20-contributor sample first.
- **Stars or forks as a proxy for people.** A highly-starred repo can have a truck factor of one or two, one owner and no poachable bench. Check contributor concentration, not stars.
- **Verified-looking commits treated as identity.** Only around 10% of commits are verified and unverified emails can be spoofed, so do not treat a commit email as identity without cross-referencing the profile or account.
- **Single-employer org repos.** A company's own open-source repo looks talent-rich but every contributor is locked at that hard-to-poach employer. Check current-employer spread before mining.
- **Ignoring the API ceiling.** A plan to mine fifteen repos deep dies at request 60 unauthenticated. Authenticate for 5,000 per hour and budget requests per repo first.
- **Treating all-time contributors as active.** 45% of core developers vanish for a year, so a name in the graph is not a reachable engineer. Check last-commit recency.

> **Tip:** Test reachability on twenty before you pull two hundred
>
> The cheapest way to avoid the dead-pool trap is a 20-contributor reachability sample. If you surface two emails, the repo is a Skip; if you surface twelve, it earns the full pull. Twenty API calls beats two hundred wasted ones.

## Verify before you call a repo Mine

Before you commit the deep-mining hours to any repo, confirm every dimension has been actually measured rather than assumed. This is the gate between scoring and spending.

#### Before you mine a repo

- [ ] The ecosystem's local pool size is known, so the repo's ceiling is bounded.
- [ ] The contributor list was pulled authenticated, with request budget tracked against the 5,000-per-hour limit.
- [ ] A de-bot pass has run, and the distinct-human count is recorded separately from the headline.
- [ ] The reachable fraction was measured on a sample, not assumed from profile fields.
- [ ] Contributors are filtered to those active in the last 6 to 12 months.
- [ ] Current employers are tallied and the pool is confirmed dispersed, not locked at one org.
- [ ] The seniority band of the reachable-active set matches the role brief.
- [ ] The repo carries a written Mine, Skim, or Skip verdict, and the Skips are deliberately abandoned.

## Keeping the score current

Re-score rather than trust an old verdict, because the underlying signals decay. Activity recency is the fastest-moving dimension: because 45% of core developers disengage for a year, a repo you scored Mine six months ago may have lost its live bench. Re-pull the recent-committer set before you re-open a mined repo.

Reachability mechanics also shift under you. GitHub's private-by-default and noreply masking mean a contributor who was reachable last quarter may have enabled email privacy since, so re-resolve emails at pull time rather than caching them. Employer concentration is the slowest to move but the most decisive, so when a repo migrates from a dispersed community to a single corporate sponsor, its verdict should flip to Skip even if the contributor count grew.

The framework itself is stable. The dimensions do not change; only the weights and the measured values do. When your binding constraint changes, from reply rate to poaching a specific competitor, re-weight the dimensions and re-rank the same field. The template below is the artifact to keep open while you work.

**Repo sourcing-yield scoring row**

```
Repo: <name>
Ecosystem ceiling (local pool size): <n>
Headline contributors: <n>
Distinct humans (post de-bot): <n>
Reachable fraction (from 20-sample): <%>
Recently active (last 6-12mo): <n>
Reachable + active count: <n>
Employers in sample / top employer share: <n> / <%>
Seniority-band match to brief: <yes/partial/no>
Verdict (Mine / Skim / Skip): <verdict>
```

*One row per repo. Fill the measured values, not the front-page badges, then read the verdict off the last column.*

## Frequently asked questions

### How many contributors make a repo worth mining?

No published source establishes a fixed threshold, so treat any specific number as a judgement you set, not a fact you inherit. The count that matters is not the headline but the distinct-human, recently-active, reachable subset. A repo showing 400 contributors can collapse to 30 after de-botting, activity filtering, and reachability scoring, while a 60-contributor repo with a dispersed, reachable, active core can outperform it. Score the subset, not the badge.

### Which GitHub repo should I source from when I have fifteen candidates?

Score each across the six dimensions and mine only the one or two that clear Mine. Rank on active-and-reachable human contributors and employer dispersion, not stars or headline contributor counts. Stars measure attention, not people; a highly-starred repo can have a truck factor of one or two owners with no poachable bench. Pick the two worth working and deliberately skip the rest to protect your yield per hour.

### Why is reachability so low on GitHub repositories?

GitHub defaults to private-by-default email, and the noreply.github.com address plus the keep-email-private setting removes the real email from commit metadata. Vendor-reported figures put public profile-email exposure at roughly 15%, with direct discovery working for about 10 to 15% of developers. So a large active repo can still be a dead pool if its maintainers all masked their addresses. Test the reachable fraction on a 20-person sample first.

### How do I tell bots from human contributors?

Run a trained classifier. BoDeGiC scores commit messages and reached 0.80 precision across 6,922 contributors, while BoDeGHa scores issue and pull-request comments. The ground truth was 5,000 accounts of which 527 were manually labelled bots. Without this pass, dependabot, CI, and org-service identities silently inflate the count and push a repo toward a false Mine verdict. A rules pass on known bot patterns is a workable fallback.

### Why does a company's own open-source repo often score Skip?

Because employer concentration kills poachability. A company-owned repo concentrates one employer near 100%, so every contributor is locked at the same, often hard-to-poach, org. In Refolk's index, by contrast, no single employer exceeded 4% to 8% of a broad skill-pool sample. The verdict hinges on which shape the repo resembles: a dispersed skill pool you can poach, or a single-org bench you cannot.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/repo-sourcing-yield-score*
