The Repo Traction Signal Reference: Proof, Gaming, Shelf Life
You can grade any single repository traction metric in front of you for what adoption it proves, how it was faked, and when it goes stale, before wiring a check.
You are diligencing an open-source company before a round, and someone has just put a repository metric in front of you as proof of adoption. This is the row-by-row reference for grading that one number: what genuine usage each repository signal proves, the exact way it gets inflated, and how long it stays valid before you have to look again. It is written for early-stage investors, platform and talent partners, and angels who need to make a call mid-diligence rather than read a whole-project score.
The published investing guides score a project overall - fork risk, commercial readiness, build versus wrapper. This one does not. It assumes you already have a snapshot in hand and you need to know whether the specific figure staring at you is proof, theatre, or already stale.
Why one metric at a time, not one project score
Grade signals individually because they fail individually, and a good project score can average away a bought number sitting inside it. An investor who jumps to the single row for the metric on the table catches the manipulation a blended score hides.
The problem is now common enough to assume, not discover. A CMU, NC State, and Socket study presented at ICSE 2026 flagged around six million suspected fake stars across 18,617 repositories and roughly 301,000 accounts; a high-confidence filter narrowed that to 3.1 million fake stars across 15,835 repositories. As of July 2024, 16.66 percent of repositories with more than 50 stars had engaged in fake-star activity, and 78 repositories carrying fake-star activity reached GitHub's Trending page. Before 2022, fewer than 10 repositories a month were involved; by July 2024 that surged to 3,216 repositories and 30,779 accounts in a single month.
The category matters. Among non-malicious cases, AI and LLM projects ranked first, accounting for 177,000 fake stars. Hype plus funding models that reward visible popularity concentrate the manipulation there, so diligence on an AI repo should start from a presumption of inflation rather than end with a spot check.
Fork-to-star ratio: what it proves and when it lies
The fork-to-star ratio proves that people wanted to build on the code, not just bookmark it, because a fork is a heavier act than a star. Organic repositories carry forks at roughly 10 to 30 percent of their star count, which is 90 to 235 forks per 1,000 stars.
Below about 50 forks per 1,000 stars on a repository with more than 10,000 stars is worth a second look. The named baselines below are the band to score against.
| Repo | Forks per 1,000 stars | Zero-follower stargazers | Status |
|---|---|---|---|
| Flask | 235 | 10% | organic |
| LangChain | 155 | 12% | organic |
| AutoGPT | 90 | 6% | organic |
| Union Labs | ~52 (0.052 ratio) | 52% | 47.4% suspected fake |
Union Labs ranked #1 on Runa Capital's ROSS Index for Q2 2025 with 54.2x star growth and 74,300 stars, yet carried a fork-to-star ratio of 0.052, 52 percent zero-follower accounts, 32.7 percent zero-repo accounts, and a StarScout flag of 47.4 percent suspected fake stars. FreeDomain is starker: 157,000 stars against 168 watchers and 2,676 forks, a watcher-to-star ratio 26x lower than Flask, with 81.3 percent of sampled stargazers having zero followers.
The ratio lies in two directions. It throws false positives on template, tutorial, and course repositories, which are legitimately fork-heavy or fork-light by purpose; only 0.56 percent of systems have more forks than stars, and one such case just provides a forking tutorial. It throws false negatives because farms now buy forks too. Marc Bara found a flagged AI project with a near-healthy fork ratio but three-quarters zero-follower stargazers. So the ratio is a screen, never a verdict.
Stars: the price list you can read
A star count proves almost nothing about adoption on its own, because it is the cheapest signal to buy and the one screening funds watch most. That combination inverts its value: a public benchmark plus a near-zero manufacturing cost turns the metric into a target rather than a measurement.
Jordan Segall of Redpoint Ventures analysed 80 developer-tool companies and found a median GitHub star count of 2,850 at seed and 4,980 at Series A. Because those medians are public, and stars cost cents, the benchmark works as a shopping list.
| Stage | Median stars | Cost to buy | Typical round | Implied ROI |
|---|---|---|---|---|
| Seed | 2,850 | $85-$285 | $1-10M | 3,500x-117,000x |
| Series A | 4,980 | $990-$4,500 | (higher) | (derived, lower) |
Star farms charge $0.03 to $0.90 per star depending on account quality. At the low end of quality the seed median costs less than a team lunch. Higher-quality inflation costs more: aged accounts with a five-year commit history and the Arctic Code Vault Contributor badge sell for around $5,000 each, which is what makes a bought crowd look like real developers.
A public star benchmark plus a near-zero manufacturing cost is a price list, not a measurement.
There is a genuine correlation underneath the noise, which is exactly why the metric is worth faking. An Organization Science paper found startups active on GitHub are 15 percentage points more likely to have raised a financing round, and 68 percent of projects on Runa Capital's ROSS list secured seed financing, cumulatively up to $169 million. The signal is real when organic, which is why manufacturing it pays.
Stargazer profiles: the cheapest way to catch bought stars
Sampling stargazer profiles proves whether the crowd behind the count is made of real developers or empty accounts, and it is the fastest manual cross-check you have. Open 100 to 150 stargazers and count two things: zero-follower accounts and ghost accounts with no repositories, followers, or bio.
Score against the organic baseline of roughly 10 percent zero-follower accounts and 1 to 2 percent ghosts. Flask sits at 10 percent, LangChain at 12 percent, AutoGPT at 6 percent. Union Labs at 52 percent and FreeDomain at 81.3 percent are not close.
The three-signal fake-star read
- Fork ratioScreen for forks under ~90 per 1,000 stars
- Profile sampleCount zero-follower and ghost accounts vs ~10% baseline
- Star historyClassify growth as event-correlated or a vertical burst
- ConvergeCall manipulation only when two or more agree
The sampling error here is real. A sample of 100 to 150 has wide variance, and a new but real developer also has an empty profile: any single signal in isolation may characterize legitimate users, such as someone who starred a repo shortly after creating their account. That is why the standard is two signals agreeing, not one flagging.
Star history: bursts versus adoption
A star-history curve proves whether growth arrived steadily from use or all at once from a purchase or a viral moment. Plot cumulative stars over time and look for vertical bursts against event-correlated slopes.
The misread cuts both ways. A Hacker News front page or an influencer post produces a legitimate spike that looks bought, and a project can accumulate thousands of stars from a single viral moment without becoming widely adopted. So a spike is not proof of fraud and not proof of adoption. The discipline is to correlate every spike to a datable external event: a launch, a conference talk, a news cycle. A burst you can tie to a real event is organic but shallow; a burst you cannot explain is the one to run through a detector.
Manually reconstructing who is behind a spike, across GitHub history and current employers, is slow. Refolk lets you ask for the established developers behind a growth event in plain English and get the people back, so you can judge a curve by the accounts that made it rather than by its shape.
Contributor depth: the metric a farm cannot buy
Unique monthly active contributors proves sustained human work, because it counts anyone who opened an issue, submitted a PR, or committed code, and each of those is labour. Bessemer Venture Partners stopped relying on stars years ago and screens on this instead. Their benchmark of 250-plus unique contributors per month filters out under 5 percent of the top 10,000 repositories, which is why it is a moat rather than a vanity number.
Contributor depth is scarce by construction, and that scarcity is the point. In Refolk's index of professional profiles, US professionals publicly self-identifying with an open-source maintainer title are vanishingly few.
| Skill | US profiles with maintainer title | Multiple vs Rust (derived) |
|---|---|---|
| Rust | 3 | 1.0x |
| Go | 8 | 2.7x |
One caveat keeps this honest: drive-by typo-fix PRs pad contributor totals. Weight by substantive merged PRs and contributor retention, not raw headcount. As one commentator put it, a star count can be faked but a bug fix that saves someone's weekend cannot.
Dependents - packages that directly depend on the project - and named enterprise adopters sit in the same hard-to-fake tier, because both require a real third party to have chosen the code and lived with it.
Downloads and the metrics that fail together
Package downloads prove nothing you can rely on, because the count is source-blind. npm has stated openly that the download statistic has no consideration for source, so a project with frequent CI runs or a repeated bot download inflates it. One developer got a package with little to no users to accumulate over one million downloads at no cost.
The deeper trap is that downloads and stars fail together. Both are source-blind and bot-inflatable, so stacking two gameable metrics gives no more assurance than one. Only dependents, substantive PRs, and named adopters break the correlation. When a founder walks you from stars to downloads, they have not corroborated anything - they have repeated the same weakness in a second color.
The metric reference: proof, gaming, shelf life
Use this as the lookup. Each row states what the metric proves when organic, how it gets inflated, and how long a reading stays valid before re-verification.
- Stars prove reach only. Inflated by farms at $0.03 to $0.90 each; the seed median buys for $85 to $285. Shelf life: short, re-check at start of diligence and again before signing, since a burst can land inside a quote period.
- Fork-to-star ratio proves intent to build on the code. Inflated by buying forks alongside stars. Shelf life: medium, but re-score if the star count moves.
- Stargazer profile mix proves the crowd is real developers. Inflated by aged accounts at ~$5,000 each. Shelf life: medium; re-sample after any growth event.
- Star history shape proves growth was gradual or event-driven. Confounded by legitimate viral spikes. Shelf life: re-plot before signing.
- Unique monthly contributors prove sustained human work; 250-plus filters under 5 percent of top repos. Hardest to fake. Shelf life: long, holds a quarter or more.
- Dependents and named adopters prove third-party production use. Very hard to fake. Shelf life: long.
- Package downloads prove nothing reliable; source-blind and bot-inflatable. Shelf life: never treat as adoption at any age.
Grading a repository's traction, front to back
- Pull the raw countsRecord stars, forks, watchers, open and closed issues, PRs, contributor count, and latest release date into one dated row.
- Compute the fork-to-star ratioConvert to forks per 1,000 stars and flag below ~0.05 with 10,000+ stars, or under 90 per 1,000 stars, against the 90-235 band.
- Sample stargazer profilesOpen 100-150 stargazers; count zero-follower and ghost accounts against the ~10% and ~1-2% baselines.
- Plot star historyChart cumulative stars and tie any vertical burst to a datable external event.
- Run a detection toolExecute StarScout or an equivalent lockstep and low-activity heuristic and record the suspected-fake percentage.
- Verify work-intensive signalsCount unique monthly contributors against 250+, substantive merged PRs, dependents, and named enterprise adopters.
- Cross-check registry dataReconcile downloads and dependents with repo activity, treating downloads as gameable and dependents as stronger.
- Run the representation reviewHave deal counsel confirm the data room does not repeat manufacturable metrics as representations.
Detection tools and the validator you can lean on
Detection tools prove coordination and low activity that a human sample would miss, but their output is probabilistic, not proof. StarScout is the peer-reviewed option, built on two heuristics: a low-activity heuristic and a lockstep heuristic that catches account groups starring the same repositories within a short window, an approach adapted from CopyCatch, an algorithm for detecting fraudulent patterns in social networks. Dagster Labs maintains an open tool, star-gazer, that automates star-pattern analysis and visualization, and RealStars offers open detection heuristics.
The strongest external validator is not a tool score at all - it is platform enforcement. 90.42 percent of repositories flagged by StarScout were later deleted, and 57.07 percent of flagged accounts were removed. That means GitHub's own enforcement independently agreed on roughly nine of ten flagged repos, which is corroboration you can lean on harder than any single ratio.
From detected to confirmed fake stars
- 6,000,000Suspected fake stars
initial StarScout flag
- 3,100,000High-confidence fake stars
after tighter filter
- 90.42%Flagged repos later deleted
GitHub enforcement agreed
Treat the tool output as a lead, not a conviction. A known limitation is the use of ad-hoc heuristics and parameters, notably a 50-star threshold in much of the detection, so a small or unusual repo can slip past or get miscounted.
How this goes wrong: the failure modes
Every signal in this reference has a way to mislead you, and reading them wrong is how a clean-looking deal turns out to be inflated. Work through these before you score.
- Fork-ratio false positive. Tutorial, template, and course repos are legitimately fork-heavy or fork-light. Only 0.56 percent of systems have more forks than stars, and one example just provides a forking tutorial. Check the repo's purpose before scoring the ratio.
- Fork-ratio false negative. Farms now buy forks, so a healthy ratio can hide manipulation. Cross-check stargazer profiles rather than clearing on the ratio.
- Star-history burst misread. A Hacker News or influencer spike looks bought. A project can gain thousands of stars from one viral moment without adoption. Correlate spikes to a datable event.
- Download-count trust trap. Downloads are trivially inflated by CI and bots. Never treat weekly downloads as adoption; require dependents plus named users.
- Ghost-account sampling error. A sample of 100-150 has wide variance, and a new but real developer also has an empty profile. Require two or more signals agreeing.
- StarScout limitation. Ad-hoc heuristics and the 50-star threshold make the output probabilistic. Treat it as a lead, not proof.
- Contributor-count inflation. Drive-by typo-fix PRs pad totals. Weight by substantive merged PRs and retention, not raw contributor count.
The regulatory and representation angle
Buying fake traction has moved from vanity to liability, and the exposure now reaches the buyer, not just the seller. The FTC's final rule, effective October 21, 2024, prohibits selling or buying fake indicators of social media influence such as followers or views generated by a bot or hijacked account, and authorizes civil penalties up to $51,744 per violation for knowing violators. Whether GitHub stars specifically fall under the rule's social-media-indicator definition is untested and not established publicly, so treat it as a live risk to check with counsel rather than a settled fact.
The fundraising precedent is settled, though. Skael co-founder Baba Nadimpalli was indicted on securities and wire fraud charges for inflating revenues after the company raised more than $40 million across three rounds, and the SEC charged former HeadSpin CEO Manish Lachwani with defrauding investors out of $80 million by falsely claiming strong and consistent growth. Inflated fundraising metrics are prosecutable. Your representation review should trace every traction claim in the data room to a non-gameable source, so a bought star count never becomes a warranty someone later has to defend.
Keeping the read current
A traction read has a shelf life, so re-verify rather than trusting a snapshot from the start of a deal. Stars and star-history are the shortest-lived: re-pull them just before signing, because a burst can land inside a quote period. Contributor depth and dependents move slowly and hold for a quarter or more, so a single reading carries further. Platform deletion is worth re-checking near close, since GitHub's enforcement is retrospective and a flagged repo may only be removed weeks after you first look.
Before you wire the check
- Every headline metric has been scored against its organic band, not accepted at face value.
- The fork-to-star ratio is corroborated by a stargazer-profile sample, not read alone.
- At least two signals agree before any count is called manufactured.
- Every star-history spike is tied to a datable external event or run through a detector.
- A work-intensive signal - contributors, dependents, or named adopters - confirms real adoption.
- Package downloads were treated as gameable and not counted as adoption evidence.
- Deal counsel has confirmed no manufacturable metric appears as a representation.
- Short-lived metrics were re-pulled just before signing.
The verification that survives is the human work behind the repo: who merged real code, who runs it in production, who left another infrastructure company to build this. That is the layer a farm cannot reach, and it is the layer worth asking for directly. When you can name the outside engineers contributing to a repo and the companies depending on it, you are grading adoption instead of counting stars.
Questions practitioners ask
What fork-to-star ratio should make me suspicious?
Organic repositories carry forks at roughly 10 to 30 percent of their star count, or 90 to 235 forks per 1,000 stars. Treat anything below about 0.05 on a repo with more than 10,000 stars as a flag worth chasing. Union Labs sat at 0.052 and had 47.4 percent of its stars flagged as fake. But a healthy ratio is not a clean bill of health, because sophisticated farms now buy forks too, so always cross-check stargazer profiles.
How much does it cost to fake the GitHub stars a fund screens on?
Star farms charge $0.03 to $0.90 per star depending on account quality. The published median at seed is 2,850 stars, which manufactures for $85 to $285, and Series A territory at 4,980 stars costs $990 to $4,500. Against a typical seed round of $1 to 10 million, that is an ROI between 3,500x and 117,000x, which is exactly why a public star benchmark functions as a price list.
Which repository metric is hardest for a founder to fake?
Unique monthly active contributors is the hardest headline metric to fake, because it counts people who opened an issue, submitted a PR, or committed code. Bessemer's benchmark of 250-plus per month filters under 5 percent of the top 10,000 repos. Substantive merged PRs, dependents, and named production adopters are similarly work-intensive. A bug fix that saves someone's weekend cannot be bought the way a star can.
Can I trust weekly package downloads as an adoption signal?
No. npm has stated the download statistic has no consideration for source, so CI runs and repeated bot downloads inflate it freely. One developer pushed a package with little to no real users past one million downloads at no cost. Downloads and stars fail together because both are source-blind. Require dependents, the packages that directly depend on yours, plus named users before treating registry data as adoption.
Is buying GitHub stars actually illegal?
The direct application to stars is untested and not established publicly. But the FTC rule effective October 21, 2024 bans buying and selling fake indicators of social media influence and authorizes penalties up to $51,744 per violation for knowing violators, which now reaches the buyer, not just the seller. Separately, inflated fundraising metrics have drawn securities and wire fraud charges, as in the Skael and HeadSpin cases. Have counsel confirm the data room does not repeat gameable metrics as representations.
How often should I re-verify these metrics during a live deal?
Re-check star and fork counts and the star-history curve at the start of diligence and again just before signing, because a burst can appear inside a quote period. Contributor depth and dependents move slowly and hold their value for a quarter or more. The strongest external validator, platform deletion of flagged repos and accounts, is worth re-running near close because GitHub removed 90.42 percent of StarScout-flagged repos over time.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.