# The Repo Traction Signal Reference: Proof, Gaming, Shelf Life

*You can grade any single repository traction metric in front of you for what adoption it proves, how it was faked, and when it goes stale, before wiring a check.*

- Canonical URL: https://www.refolk.ai/guides/repo-traction-signal-reference
- Pillar: Investing and deal sourcing
- Format: Reference
- Published: 2026-09-17
- Last reviewed: 2026-09-17
- Reading time: 16 min

You are diligencing an open-source company before a round, and someone has just put a repository metric in front of you as proof of adoption. This is the row-by-row reference for grading that one number: what genuine usage each repository signal proves, the exact way it gets inflated, and how long it stays valid before you have to look again. It is written for early-stage investors, platform and talent partners, and angels who need to make a call mid-diligence rather than read a whole-project score.

The published investing guides score a project overall - fork risk, commercial readiness, build versus wrapper. This one does not. It assumes you already have a snapshot in hand and you need to know whether the specific figure staring at you is proof, theatre, or already stale.

## Why one metric at a time, not one project score

Grade signals individually because they fail individually, and a good project score can average away a bought number sitting inside it. An investor who jumps to the single row for the metric on the table catches the manipulation a blended score hides.

The problem is now common enough to assume, not discover. A CMU, NC State, and Socket study presented at ICSE 2026 flagged around six million suspected fake stars across 18,617 repositories and roughly 301,000 accounts; a high-confidence filter narrowed that to 3.1 million fake stars across 15,835 repositories. As of July 2024, 16.66 percent of repositories with more than 50 stars had engaged in fake-star activity, and 78 repositories carrying fake-star activity reached GitHub's Trending page. Before 2022, fewer than 10 repositories a month were involved; by July 2024 that surged to 3,216 repositories and 30,779 accounts in a single month.

**16.66% - Repositories with 50+ stars that had engaged in fake-star activity by July 2024**

The base rate is high enough that inflation is the null hypothesis, not the exception.

The category matters. Among non-malicious cases, AI and LLM projects ranked first, accounting for 177,000 fake stars. Hype plus funding models that reward visible popularity concentrate the manipulation there, so diligence on an AI repo should start from a presumption of inflation rather than end with a spot check.

## Fork-to-star ratio: what it proves and when it lies

The fork-to-star ratio proves that people wanted to build on the code, not just bookmark it, because a fork is a heavier act than a star. Organic repositories carry forks at roughly 10 to 30 percent of their star count, which is 90 to 235 forks per 1,000 stars.

Below about 50 forks per 1,000 stars on a repository with more than 10,000 stars is worth a second look. The named baselines below are the band to score against.

| Repo | Forks per 1,000 stars | Zero-follower stargazers | Status |
|---|---|---|---|
| Flask | 235 | 10% | organic |
| LangChain | 155 | 12% | organic |
| AutoGPT | 90 | 6% | organic |
| Union Labs | ~52 (0.052 ratio) | 52% | 47.4% suspected fake |

Union Labs ranked #1 on Runa Capital's ROSS Index for Q2 2025 with 54.2x star growth and 74,300 stars, yet carried a fork-to-star ratio of 0.052, 52 percent zero-follower accounts, 32.7 percent zero-repo accounts, and a StarScout flag of 47.4 percent suspected fake stars. FreeDomain is starker: 157,000 stars against 168 watchers and 2,676 forks, a watcher-to-star ratio 26x lower than Flask, with 81.3 percent of sampled stargazers having zero followers.

The ratio lies in two directions. It throws false positives on template, tutorial, and course repositories, which are legitimately fork-heavy or fork-light by purpose; only 0.56 percent of systems have more forks than stars, and one such case just provides a forking tutorial. It throws false negatives because farms now buy forks too. Marc Bara found a flagged AI project with a near-healthy fork ratio but three-quarters zero-follower stargazers. So the ratio is a screen, never a verdict.

> **Rule:** Never clear a repo on the fork ratio alone
>
> A healthy fork-to-star ratio can be manufactured. Confirm it against a stargazer-profile sample before you treat it as evidence of real adoption.

## Stars: the price list you can read

A star count proves almost nothing about adoption on its own, because it is the cheapest signal to buy and the one screening funds watch most. That combination inverts its value: a public benchmark plus a near-zero manufacturing cost turns the metric into a target rather than a measurement.

Jordan Segall of Redpoint Ventures analysed 80 developer-tool companies and found a median GitHub star count of 2,850 at seed and 4,980 at Series A. Because those medians are public, and stars cost cents, the benchmark works as a shopping list.

| Stage | Median stars | Cost to buy | Typical round | Implied ROI |
|---|---|---|---|---|
| Seed | 2,850 | $85-$285 | $1-10M | 3,500x-117,000x |
| Series A | 4,980 | $990-$4,500 | (higher) | (derived, lower) |

Star farms charge $0.03 to $0.90 per star depending on account quality. At the low end of quality the seed median costs less than a team lunch. Higher-quality inflation costs more: aged accounts with a five-year commit history and the Arctic Code Vault Contributor badge sell for around $5,000 each, which is what makes a bought crowd look like real developers.

> A public star benchmark plus a near-zero manufacturing cost is a price list, not a measurement.

There is a genuine correlation underneath the noise, which is exactly why the metric is worth faking. An Organization Science paper found startups active on GitHub are 15 percentage points more likely to have raised a financing round, and 68 percent of projects on Runa Capital's ROSS list secured seed financing, cumulatively up to $169 million. The signal is real when organic, which is why manufacturing it pays.

## Stargazer profiles: the cheapest way to catch bought stars

Sampling stargazer profiles proves whether the crowd behind the count is made of real developers or empty accounts, and it is the fastest manual cross-check you have. Open 100 to 150 stargazers and count two things: zero-follower accounts and ghost accounts with no repositories, followers, or bio.

Score against the organic baseline of roughly 10 percent zero-follower accounts and 1 to 2 percent ghosts. Flask sits at 10 percent, LangChain at 12 percent, AutoGPT at 6 percent. Union Labs at 52 percent and FreeDomain at 81.3 percent are not close.

#### The three-signal fake-star read

1. **Fork ratio** - Screen for forks under ~90 per 1,000 stars
2. **Profile sample** - Count zero-follower and ghost accounts vs ~10% baseline
3. **Star history** - Classify growth as event-correlated or a vertical burst
4. **Converge** - Call manipulation only when two or more agree

*No single signal convicts; two or more agreeing is the standard for calling a count manufactured.*

The sampling error here is real. A sample of 100 to 150 has wide variance, and a new but real developer also has an empty profile: any single signal in isolation may characterize legitimate users, such as someone who starred a repo shortly after creating their account. That is why the standard is two signals agreeing, not one flagging.

## Star history: bursts versus adoption

A star-history curve proves whether growth arrived steadily from use or all at once from a purchase or a viral moment. Plot cumulative stars over time and look for vertical bursts against event-correlated slopes.

The misread cuts both ways. A Hacker News front page or an influencer post produces a legitimate spike that looks bought, and a project can accumulate thousands of stars from a single viral moment without becoming widely adopted. So a spike is not proof of fraud and not proof of adoption. The discipline is to correlate every spike to a datable external event: a launch, a conference talk, a news cycle. A burst you can tie to a real event is organic but shallow; a burst you cannot explain is the one to run through a detector.

I ran this search: `GitHub users with 3+ years of history and real followers who starred a major open-source AI agent framework during a recent growth spike` - [see the full result list](https://www.refolk.ai/s/3ksfz2ex1b).

*Returns the substantive accounts behind a star spike so you can see whether real, established developers - not ghost accounts - drove the momentum.*

Manually reconstructing who is behind a spike, across GitHub history and current employers, is slow. [Refolk](/) lets you ask for the established developers behind a growth event in plain English and get the people back, so you can judge a curve by the accounts that made it rather than by its shape.

## Contributor depth: the metric a farm cannot buy

Unique monthly active contributors proves sustained human work, because it counts anyone who opened an issue, submitted a PR, or committed code, and each of those is labour. Bessemer Venture Partners stopped relying on stars years ago and screens on this instead. Their benchmark of 250-plus unique contributors per month filters out under 5 percent of the top 10,000 repositories, which is why it is a moat rather than a vanity number.

Contributor depth is scarce by construction, and that scarcity is the point. In Refolk's index of professional profiles, US professionals publicly self-identifying with an open-source maintainer title are vanishingly few.

| Skill | US profiles with maintainer title | Multiple vs Rust (derived) |
|---|---|---|
| Rust | 3 | 1.0x |
| Go | 8 | 2.7x |

**3 - US professionals in Refolk's index publicly claiming a Rust open-source maintainer title**

Genuine multi-maintainer projects are rare, so real contributor depth cannot be conjured to order.

One caveat keeps this honest: drive-by typo-fix PRs pad contributor totals. Weight by substantive merged PRs and contributor retention, not raw headcount. As one commentator put it, a star count can be faked but a bug fix that saves someone's weekend cannot.

Dependents - packages that directly depend on the project - and named enterprise adopters sit in the same hard-to-fake tier, because both require a real third party to have chosen the code and lived with it.

## Downloads and the metrics that fail together

Package downloads prove nothing you can rely on, because the count is source-blind. npm has stated openly that the download statistic has no consideration for source, so a project with frequent CI runs or a repeated bot download inflates it. One developer got a package with little to no users to accumulate over one million downloads at no cost.

The deeper trap is that downloads and stars fail together. Both are source-blind and bot-inflatable, so stacking two gameable metrics gives no more assurance than one. Only dependents, substantive PRs, and named adopters break the correlation. When a founder walks you from stars to downloads, they have not corroborated anything - they have repeated the same weakness in a second color.

> **Watch out:** Two gameable metrics do not corroborate each other
>
> Six million fake stars and a package at one million bot downloads prove the same failure. Require a non-gameable signal - dependents, merged PRs, or named users - before you treat traction as verified.

## The metric reference: proof, gaming, shelf life

Use this as the lookup. Each row states what the metric proves when organic, how it gets inflated, and how long a reading stays valid before re-verification.

- **Stars** prove reach only. Inflated by farms at $0.03 to $0.90 each; the seed median buys for $85 to $285. Shelf life: short, re-check at start of diligence and again before signing, since a burst can land inside a quote period.
- **Fork-to-star ratio** proves intent to build on the code. Inflated by buying forks alongside stars. Shelf life: medium, but re-score if the star count moves.
- **Stargazer profile mix** proves the crowd is real developers. Inflated by aged accounts at ~$5,000 each. Shelf life: medium; re-sample after any growth event.
- **Star history shape** proves growth was gradual or event-driven. Confounded by legitimate viral spikes. Shelf life: re-plot before signing.
- **Unique monthly contributors** prove sustained human work; 250-plus filters under 5 percent of top repos. Hardest to fake. Shelf life: long, holds a quarter or more.
- **Dependents and named adopters** prove third-party production use. Very hard to fake. Shelf life: long.
- **Package downloads** prove nothing reliable; source-blind and bot-inflatable. Shelf life: never treat as adoption at any age.

#### Grading a repository's traction, front to back

1. **Pull the raw counts** - Record stars, forks, watchers, open and closed issues, PRs, contributor count, and latest release date into one dated row.
2. **Compute the fork-to-star ratio** - Convert to forks per 1,000 stars and flag below ~0.05 with 10,000+ stars, or under 90 per 1,000 stars, against the 90-235 band.
3. **Sample stargazer profiles** - Open 100-150 stargazers; count zero-follower and ghost accounts against the ~10% and ~1-2% baselines.
4. **Plot star history** - Chart cumulative stars and tie any vertical burst to a datable external event.
5. **Run a detection tool** - Execute StarScout or an equivalent lockstep and low-activity heuristic and record the suspected-fake percentage.
6. **Verify work-intensive signals** - Count unique monthly contributors against 250+, substantive merged PRs, dependents, and named enterprise adopters.
7. **Cross-check registry data** - Reconcile downloads and dependents with repo activity, treating downloads as gameable and dependents as stronger.
8. **Run the representation review** - Have deal counsel confirm the data room does not repeat manufacturable metrics as representations.

## Detection tools and the validator you can lean on

Detection tools prove coordination and low activity that a human sample would miss, but their output is probabilistic, not proof. StarScout is the peer-reviewed option, built on two heuristics: a low-activity heuristic and a lockstep heuristic that catches account groups starring the same repositories within a short window, an approach adapted from CopyCatch, an algorithm for detecting fraudulent patterns in social networks. Dagster Labs maintains an open tool, star-gazer, that automates star-pattern analysis and visualization, and RealStars offers open detection heuristics.

The strongest external validator is not a tool score at all - it is platform enforcement. 90.42 percent of repositories flagged by StarScout were later deleted, and 57.07 percent of flagged accounts were removed. That means GitHub's own enforcement independently agreed on roughly nine of ten flagged repos, which is corroboration you can lean on harder than any single ratio.

#### From detected to confirmed fake stars

| Stage | Figure | Note |
| --- | --- | --- |
| Suspected fake stars | 6,000,000 | initial StarScout flag |
| High-confidence fake stars | 3,100,000 | after tighter filter |
| Flagged repos later deleted | 90.42% | GitHub enforcement agreed |

*A high-confidence filter and then GitHub's own deletions narrow the flagged set to a corroborated core.*

Treat the tool output as a lead, not a conviction. A known limitation is the use of ad-hoc heuristics and parameters, notably a 50-star threshold in much of the detection, so a small or unusual repo can slip past or get miscounted.

## How this goes wrong: the failure modes

Every signal in this reference has a way to mislead you, and reading them wrong is how a clean-looking deal turns out to be inflated. Work through these before you score.

- **Fork-ratio false positive.** Tutorial, template, and course repos are legitimately fork-heavy or fork-light. Only 0.56 percent of systems have more forks than stars, and one example just provides a forking tutorial. Check the repo's purpose before scoring the ratio.
- **Fork-ratio false negative.** Farms now buy forks, so a healthy ratio can hide manipulation. Cross-check stargazer profiles rather than clearing on the ratio.
- **Star-history burst misread.** A Hacker News or influencer spike looks bought. A project can gain thousands of stars from one viral moment without adoption. Correlate spikes to a datable event.
- **Download-count trust trap.** Downloads are trivially inflated by CI and bots. Never treat weekly downloads as adoption; require dependents plus named users.
- **Ghost-account sampling error.** A sample of 100-150 has wide variance, and a new but real developer also has an empty profile. Require two or more signals agreeing.
- **StarScout limitation.** Ad-hoc heuristics and the 50-star threshold make the output probabilistic. Treat it as a lead, not proof.
- **Contributor-count inflation.** Drive-by typo-fix PRs pad totals. Weight by substantive merged PRs and retention, not raw contributor count.

> **Tip:** Where to point diligence first
>
> On an AI or LLM repo, start from a presumption of inflation. That category alone accounted for 177,000 fake stars, so the profile sample and detector are worth running before you read the pitch.

## The regulatory and representation angle

Buying fake traction has moved from vanity to liability, and the exposure now reaches the buyer, not just the seller. The FTC's final rule, effective October 21, 2024, prohibits selling or buying fake indicators of social media influence such as followers or views generated by a bot or hijacked account, and authorizes civil penalties up to $51,744 per violation for knowing violators. Whether GitHub stars specifically fall under the rule's social-media-indicator definition is untested and not established publicly, so treat it as a live risk to check with counsel rather than a settled fact.

The fundraising precedent is settled, though. Skael co-founder Baba Nadimpalli was indicted on securities and wire fraud charges for inflating revenues after the company raised more than $40 million across three rounds, and the SEC charged former HeadSpin CEO Manish Lachwani with defrauding investors out of $80 million by falsely claiming strong and consistent growth. Inflated fundraising metrics are prosecutable. Your representation review should trace every traction claim in the data room to a non-gameable source, so a bought star count never becomes a warranty someone later has to defend.

## Keeping the read current

A traction read has a shelf life, so re-verify rather than trusting a snapshot from the start of a deal. Stars and star-history are the shortest-lived: re-pull them just before signing, because a burst can land inside a quote period. Contributor depth and dependents move slowly and hold for a quarter or more, so a single reading carries further. Platform deletion is worth re-checking near close, since GitHub's enforcement is retrospective and a flagged repo may only be removed weeks after you first look.

#### Before you wire the check

- [ ] Every headline metric has been scored against its organic band, not accepted at face value.
- [ ] The fork-to-star ratio is corroborated by a stargazer-profile sample, not read alone.
- [ ] At least two signals agree before any count is called manufactured.
- [ ] Every star-history spike is tied to a datable external event or run through a detector.
- [ ] A work-intensive signal - contributors, dependents, or named adopters - confirms real adoption.
- [ ] Package downloads were treated as gameable and not counted as adoption evidence.
- [ ] Deal counsel has confirmed no manufacturable metric appears as a representation.
- [ ] Short-lived metrics were re-pulled just before signing.

The verification that survives is the human work behind the repo: who merged real code, who runs it in production, who left another infrastructure company to build this. That is the layer a farm cannot reach, and it is the layer worth asking for directly. When you can name the outside engineers contributing to a repo and the companies depending on it, you are grading adoption instead of counting stars.

## Frequently asked questions

### What fork-to-star ratio should make me suspicious?

Organic repositories carry forks at roughly 10 to 30 percent of their star count, or 90 to 235 forks per 1,000 stars. Treat anything below about 0.05 on a repo with more than 10,000 stars as a flag worth chasing. Union Labs sat at 0.052 and had 47.4 percent of its stars flagged as fake. But a healthy ratio is not a clean bill of health, because sophisticated farms now buy forks too, so always cross-check stargazer profiles.

### How much does it cost to fake the GitHub stars a fund screens on?

Star farms charge $0.03 to $0.90 per star depending on account quality. The published median at seed is 2,850 stars, which manufactures for $85 to $285, and Series A territory at 4,980 stars costs $990 to $4,500. Against a typical seed round of $1 to 10 million, that is an ROI between 3,500x and 117,000x, which is exactly why a public star benchmark functions as a price list.

### Which repository metric is hardest for a founder to fake?

Unique monthly active contributors is the hardest headline metric to fake, because it counts people who opened an issue, submitted a PR, or committed code. Bessemer's benchmark of 250-plus per month filters under 5 percent of the top 10,000 repos. Substantive merged PRs, dependents, and named production adopters are similarly work-intensive. A bug fix that saves someone's weekend cannot be bought the way a star can.

### Can I trust weekly package downloads as an adoption signal?

No. npm has stated the download statistic has no consideration for source, so CI runs and repeated bot downloads inflate it freely. One developer pushed a package with little to no real users past one million downloads at no cost. Downloads and stars fail together because both are source-blind. Require dependents, the packages that directly depend on yours, plus named users before treating registry data as adoption.

### Is buying GitHub stars actually illegal?

The direct application to stars is untested and not established publicly. But the FTC rule effective October 21, 2024 bans buying and selling fake indicators of social media influence and authorizes penalties up to $51,744 per violation for knowing violators, which now reaches the buyer, not just the seller. Separately, inflated fundraising metrics have drawn securities and wire fraud charges, as in the Skael and HeadSpin cases. Have counsel confirm the data room does not repeat gameable metrics as representations.

### How often should I re-verify these metrics during a live deal?

Re-check star and fork counts and the star-history curve at the start of diligence and again just before signing, because a burst can appear inside a quote period. Contributor depth and dependents move slowly and hold their value for a quarter or more. The strongest external validator, platform deletion of flagged repos and accounts, is worth re-running near close because GitHub removed 90.42 percent of StarScout-flagged repos over time.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/repo-traction-signal-reference*
