# Telling Real Category Adoption From Manufactured Momentum

*You can take one technology category, run manipulation-resistant checks on its repos, and land on a growth read you can defend line by line in a review.*

- Canonical URL: https://www.refolk.ai/guides/real-adoption-vs-manufactured-momentum
- Pillar: Market and talent intelligence
- Format: Teardown
- Published: 2026-08-17
- Last reviewed: 2026-08-17
- Reading time: 15 min

## Key takeaways

- Fake stars sell for as little as $0.10 each, and their promotion effect lasts under two months before reversing, so any adoption read anchored on star velocity captures the pump, not the trend.
- Contributor density is the cheapest star-independent correction: AutoGPT converted under 9 contributors per 1,000 stars against LangChain's 41, a roughly 4.5x depth gap invisible in raw star totals.
- Package-registry presence is a high-specificity filter because only 1.23% of repos in confirmed fake-star campaigns are published in package registries.
- In Refolk's index of professional profiles keyworded to AI agent frameworks, LangChain-skilled practitioners in the US outnumber AutoGen's 9 to 0, showing practitioner mass concentrates in one or two projects despite category-wide hype.
- Flagged repositories showed deletion ratios up to 90%, sixteen times the baseline, a rare after-the-fact ground truth that bought stars cannot fake.
- When star velocity and manipulation-resistant signals disagree, the resistant signals win, and every line of the final read should trace back to a specific number.

A hyped technology category lands on your desk with a chart of surging GitHub stars, and someone wants to know whether to put the trend in a market or talent-intelligence report. This guide is for strategy analysts, talent-intelligence researchers, and operators sizing a category. It carries one worked example - the AI agent framework category - all the way through, with real counts and the wrong turns, so you can run the same checks on your own case and land on a growth read you can defend line by line.

The short version: stars are the cheapest signal to fake and the first to reverse. Everything below is about replacing star velocity with signals that cost real effort to manufacture.

## Why star velocity lies, and what it lies about

Star velocity lies because attention is purchasable and adoption is not. A CMU-led study, "Six Million (Suspected) Fake Stars on GitHub," scanned over 6 billion GitHub events from July 2019 to October 2024 and, after its filtering step, identified 18,617 repositories with fake star campaigns, 301,000 participating accounts, and 3.81 million fake stars. Fake stars sell for as little as $0.10 each. That price is the whole problem: a signal you can buy for pennies cannot separate a real trend from a marketing budget.

The lie has a specific shape and a specific expiry. Bought stars produce a lone spike, not a rising tide. And the lift does not last - the research found that fake stars only have a promotion effect in the short term, less than two months, and become a liability in the long term. So the danger is not just that a spike is fake; it is that a spike is *temporary*, which means a four-week read can capture the pump precisely when it looks most convincing.

**$0.10 - Cost per fake GitHub star at the disposable-account tier**

A signal this cheap to buy cannot distinguish a real trend from a marketing spend.

There is a media-version trap worth naming before you cite anything. The first arXiv version of the paper reported 4.53 million suspected stars narrowing to 3.1 million; the final ICSE 2026 version reports 6.0 million narrowing to 3.81 million. Both sets circulate. Cite the final numbers.

| Metric | arXiv v1 (reported) | Final ICSE 2026 |
|---|---|---|
| Suspected fake stars (pre-filter) | 4.53M | 6.0M |
| Repos (pre-filter) | 22,915 | 26,254 |
| Fake stars (post-filter) | 3.1M | 3.81M |
| Repos with campaigns | 15,835 | 18,617 |

The category detail matters for you specifically. The majority of fake stars promote short-lived phishing malware; the remaining ones mostly promote AI/LLM, blockchain, tool, and tutorial repositories, and a separate analysis names AI and LLM repositories as the largest non-malicious category of recipients. If your hyped category is AI-adjacent, assume the star signal is contaminated until proven otherwise.

## The worked example: is the AI agent framework category actually being adopted?

The job is to decide whether the AI agent framework category's open-source momentum is real before it goes in a report. I frame it around a fixed set of repos - AutoGPT, LangChain, MetaGPT, LangFlow, AutoGen, and their peers - and I record star, fork, contributor, and dependent counts with a date stamp so the read is reproducible.

The naive read is seductive. AutoGPT gained 111,967 stars in a single month. On a star chart that is the biggest thing in the category, and a report anchored on it would call AutoGPT the runaway leader. That read is wrong, and the correction is the whole point of this teardown.

The Cisco agent-framework study covers 15 major AI agent framework repositories from late 2022 to early 2026, using 808,042 stars, 73,997 pull requests, 86,241 commits, and 987,330 user profiles. Its headline correction: AutoGPT converted fewer than 9 contributors per 1,000 stars, while LangChain converted 41. That is a roughly 4.5x depth gap the star chart cannot see. LangChain also functions as shared infrastructure, attracting 82.5% of cross-ecosystem contributors, which is what real adoption looks like from the inside.

| Framework | One-month star gain | Contributors per 1,000 stars |
|---|---|---|
| AutoGPT | 111,967 | under 9 |
| LangChain | (not stated) | 41 |
| MetaGPT / LangFlow | (not stated) | under 5 |

The empirical agent-developer study adds the downstream check. MetaGPT holds 48.7K stars and 5.8K forks, yet it appears in only 2 of the 1,575 agent-related repos in that dataset. Big star total, near-zero footprint in what other people actually build. That divergence - loud stars, quiet dependents - is the exact pattern manufactured momentum produces, and it is the pattern you are hunting.

> Loud stars and quiet dependents is the signature of momentum someone bought rather than earned.
> </pull>
>
> ## The signal ladder: what each signal proves and how it lies
>
> Rank signals by how hard they are to fake, and read from the resistant end down. The rule is simple: never let a gamed signal outvote a resistant one.

figure
kind: stack
title: The signal ladder, most resistant on top
caption: Read adoption from the resistant layers first and treat stars as the least trustworthy input.
layer: Dependents and downloads :: Other projects depend on it; near-impossible to fake at scale
layer: Contributor density and retention :: Real humans file PRs and keep contributing past 90 days
layer: Forks, watchers, issues :: Costlier to game than stars, and rise together in real growth
layer: Trending placement :: Reachable but filtered; weak proof on its own
layer: Star count and velocity :: Cheapest to buy, first to reverse; anchor nothing here
```

Here is what each signal proves and what it looks like when it lies:

- **Dependents ("used by") and registry downloads.** Proves other software relies on the code. It lies rarely: only 1.23% of campaign repos are published in package registries, so a real dependents footprint is a high-specificity filter. A snapshot lies by youth - a genuinely new project has no dependents yet, so read the trend, not the number.
- **Contributor density.** Proves humans do work, not just click star. It lies when a big institutional name front-loads early contributors; check retention alongside it. Retention drops most steeply in the first 30 days and stabilizes near 90 days, so measure past 90.
- **Forks, watchers, issues.** Proves engaged use. In real growth these rise together with stars; a star-only spike with flat forks and issues is the tell. Practitioner ratio checks help - a healthy repo like Flask shows roughly 29 watchers per 1,000 stars, against a padded repo's 1.
- **Trending placement.** Proves visibility, not adoption. It lies by association: 78 flagged repos reached GitHub Trending, so never treat Trending as validation on its own.
- **Star count and velocity.** Proves almost nothing by itself. It lies for $0.10 a star and reverses within two months.

> **Rule:** Resistant signals win ties
>
> When star velocity and a manipulation-resistant signal disagree, the resistant signal decides the read. Stars break ties for nothing.

## The detection procedure, run on the category

Run these seven steps in order. Some practitioners run the resistant-signal cross-check (step 4) before any fake-star tooling, treating stars as untrustworthy from the outset; the sources agree on the signals, not the sequence, so either order is defensible as long as resistant signals win the final call.

#### From a hyped category to a defensible growth read

1. **Frame the category and fix the repo set** - List the 8 to 15 repos that define the category and record current star, fork, contributor, and dependent counts with a date stamp. Done means a fixed, dated repo set you can rerun later.
2. **Plot star history and read the shape** - Chart each repo's star history and classify it as gradual or spike. Flag any single month that contributes an outsized share of total stars.
3. **Run a burst and low-activity check on flagged repos** - Run Dagster's simple REST-API model or StarGuard for a fast read, then escalate to StarScout over BigQuery for clustering depth. Done means a percent-suspected-fake figure per flagged repo plus a stargazer sample.
4. **Cross-check manipulation-resistant signals** - Compute forks-to-stars, watchers-to-stars, distinct contributors, contributor density per 1,000 stars, and release cadence. Done means a table where each repo has star-independent signals beside its star count.
5. **Measure downstream adoption** - Pull dependent-package counts, registry download trends, and dependents growth in the 60 to 90 days after major releases. Done means an adoption trend that corroborates or contradicts the star trend.
6. **Size the talent pool** - Query Refolk's index for practitioners with the category's skills, split by geography and seniority. Done means a headcount with named employers per segment.
7. **Reconcile and write the defensible read** - Where star velocity and resistant signals disagree, let the resistant signals win. Done means a one-line growth verdict per repo, each traceable to a number.

The tooling for step 3 is documented and mostly free. StarScout applies two heuristics plus a postprocessing step over GHArchive in BigQuery. The two heuristics are worth knowing because every tool reuses them: the **low-activity signature** finds accounts with only one WatchEvent, meaning they have starred exactly one repo, plus at most one additional event in that repo the same day; the **lockstep signature** finds clusters of N users and M repos where each repo received stars from at least P percent of the N users in a short window. Dagster's detector ships a simple heuristic using only the GitHub REST API and a complex clustering detector using GH Archive in BigQuery. StarGuard adds a BurstDetector (a sliding-window MAD algorithm that catches inorganic spikes) and a User Profiler that samples stargazers and checks account age, avatar, follower count, and repo history. Socket productized the same two heuristics into a "Suspicious Stars on GitHub" alert.

#### How the CMU-led scan narrowed 6 billion events to confirmed campaigns

| Stage | Figure | Note |
| --- | --- | --- |
| Suspected fake stars (pre-filter) | 6.0M | across 26,254 repos |
| Fake stars after postprocessing | 3.81M | across 18,617 repos |
| Participating accounts | 301K | the human (and bot) side of the campaigns |
| Campaign repos also in package registries | 229 | only 1.23%, the high-specificity tell |

*The filtering steps show how much raw volume must be discarded before a fake-star claim is safe.*

## Sizing the talent pool without inheriting the star hype

Once you trust the adoption read, size the people. This is where category-level hype most often smuggles a false conclusion into a talent report: the category looks broadly adopted, but real practitioners cluster in one or two projects and one or two markets.

In Refolk's index of professional profiles keyworded to "AI agent framework," LangChain-skilled practitioners in the United States number 9 and in Germany number 2, while AutoGen-skilled practitioners in the United States number 0. These are keyword-constrained, small counts - treat them as directional, not absolute market size - but the shape is the point. The US-to-Germany ratio for the same skill is about 4.5x, and one framework's practitioner footprint dwarfs another's even though both sit inside the same hyped category.

| Segment (skill / geography) | Practitioner count |
|---|---|
| LangChain, United States | 9 |
| LangChain, Germany | 2 |
| AutoGen, United States | 0 |

That table mirrors the adoption read exactly. LangChain's 41 contributors per 1,000 stars and its role as shared infrastructure predict a real practitioner base; AutoGen's zero US count in the index is consistent with a smaller lived footprint than category buzz implies. When talent depth and star hype diverge this sharply, the talent read is the more honest description of where the category actually lives.

I ran this search: `Engineers who have contributed to LangChain, AutoGen, or CrewAI repositories in the last 12 months, based in the US or Germany.` - [see the full result list](https://www.refolk.ai/s/b2f7knfc73).

*Returns named practitioners with public contribution history, so you can size the real talent pool per framework instead of inferring it from star counts.*

Sizing this by hand means scraping contributor graphs, resolving identities, and geolocating people one profile at a time. Describing the pool in plain English and getting named practitioners back is the friction [Refolk](/) removes - you ask for contributors to specific repos in specific markets and get people, not star totals. For the talent read specifically, prefer queries that demand contribution evidence over popularity, such as "senior ML engineers who shipped multi-agent orchestration systems and have public GitHub commit history, not just stars." Refolk's index lets you split those results by seniority and employer so each segment lands with named companies attached.

## How this read goes wrong

This is the section to reread before you file. Every failure mode below has a specific check, and most of them are ways a careful analyst still reaches the wrong verdict.

> **Watch out:** A four-week window can capture the pump
>
> Fake-star lift lasts under two months and then reverses. A one-month adoption snapshot can measure a manipulation campaign at its peak and read it as growth. Always measure the trend at 60 to 90 days post-release.

- **Star spike misread as a launch.** A real coordinated launch - a conference or a Hacker News front page - also spikes. The check: do forks, issues, and downloads rise with the stars? A star-only spike is the tell.
- **The "empty profile" heuristic over-trusted.** Modern campaigns build months of authentic-looking fake history, so naive profile checks both miss real campaigns and falsely flag legitimate new developers. The check: require two or more independent signals before flagging anything.
- **Trending treated as validation.** 78 flagged repos reached GitHub Trending. The check: pair any Trending appearance with dependents growth.
- **Raw cross-category star comparison.** Different languages carry different star norms - Python overtook JavaScript driven by AI, and TypeScript later became the most-used language on GitHub - so comparing a Python AI repo to a JS tool by raw stars misleads. The check: compare within a category and normalize to contributor density or dependents.
- **Registry adoption assumed either way.** Only 1.23% of campaign repos appear in registries, so absence of a published package can mean youth *or* hollowness. The check: read the dependents trend over time, not a single snapshot.
- **Deletion ratio taken alone.** Flagged repos showed deletion ratios up to 90%, sixteen times the baseline, because manipulation infrastructure is disposable by design. But high deletion can also reflect churny experimentation. The check: combine deletion with lockstep clustering before concluding manipulation.
- **Institutional backing mistaken for depth.** A big name behind a repo does not guarantee community depth; contributor bases stay thin even under strong sponsors. The check: contributor density and retention, not the logo.

The AutoGPT read is the live example of failure mode one and eight together. Its 111,967-star month is real activity, but its under-9 contributors per 1,000 stars and thin presence in downstream projects mean the star chart overstates its adoption depth. A report that ranked frameworks by star velocity would have led with the wrong project.

## Reconcile the signals into one defensible read

The final read is a one-line growth verdict per repo, and every line traces to a number. When signals disagree, the resistant ones win, and you write down why. For the worked category, the reconciliation reads roughly like this: LangChain shows real depth (41 contributors per 1,000 stars, shared-infrastructure role, the largest verified practitioner base in the index), so its momentum is genuine; AutoGPT's star velocity overstates its adoption (under 9 contributors per 1,000 stars, thin downstream presence); MetaGPT's star total is not matched by dependents (2 of 1,575 repos). Each verdict cites a figure, so a reviewer can challenge the number rather than argue with your judgment.

**One-line growth verdict per repo**

```
Repo: <name>
Star trend: <gradual | spike> (<share of stars in peak month>)
% suspected fake: <StarScout/Dagster/StarGuard figure> from <n> sampled stargazers
Contributors per 1,000 stars: <figure>
Dependents trend (60-90d post-release): <up | flat | down>
Registry footprint: <published package? | dependents count>
Talent pool: <practitioner count by geography, from Refolk's index>
Verdict: <genuine adoption | inflated | inconclusive, recheck at 90d> because <the one signal that decided it>
```

*Fill one row per repo; every cell must be a number or a direct classification, never an adjective.*

Then verify before you call it done.

#### Before the trend goes in the report

- [ ] The repo set is fixed and date-stamped so the read is reproducible.
- [ ] Every flagged spike was checked against forks, issues, and downloads rising together.
- [ ] At least two independent signals support any fake-star claim, never profile checks alone.
- [ ] Contributor density per 1,000 stars is recorded beside every raw star count.
- [ ] Dependents and download trends were read at 60 to 90 days post-release, not a single snapshot.
- [ ] Cross-category comparisons were normalized to density or dependents, never raw stars.
- [ ] The talent-pool count is split by geography and carries named employers.
- [ ] Each repo has a one-line verdict that traces to a specific number.

## Keeping the read current

Adoption reads decay, so treat this as a cadence, not a one-off. The mechanism to re-check is straightforward: rerun the fixed repo set from step one on a schedule and watch whether the resistant signals move in the same direction they did last time. Contributor density and dependents change slowly and honestly; a sudden divergence between them and stars is your re-audit trigger.

Two moving parts deserve a standing note. First, the detection tools evolve because campaigns evolve - bot accounts increasingly build authentic-looking history, so the "empty profile" tell weakens over time and the two-signal threshold matters more each cycle. Re-check which heuristics your tool of choice currently ships. Second, language and hype baselines shift; because Python, JavaScript, and TypeScript have traded the top spot on GitHub, the norm for a "normal" star count in your category is not fixed. Compare within the category every time rather than trusting a baseline you cached last quarter. When both the tooling and the baseline are current, and every verdict still traces to a number, the read is one you can defend in review.

## Frequently asked questions

### How do I tell if GitHub stars are fake?

Look for the shape and the corroboration, not the count. Genuine growth is a rising tide where forks, issues, and downloads climb alongside stars, while inflation is a lone star spike. Then check star-independent signals: contributor density, package-registry dependents, and stargazer account quality. Tools like StarScout, Dagster's detector, and StarGuard automate the low-activity and clustering heuristics that flagged 18,617 repositories in the CMU-led study.

### What is the single best predictor of genuine open source adoption?

Dependent-package growth, often shown as "used by," paired with contributor density. Real adoption requires humans who file pull requests and other projects that depend on the code, neither of which can be bought like stars. MetaGPT held 48.7K stars and 5.8K forks yet appeared in only 2 of 1,575 dataset repositories, which is exactly the divergence you are watching for.

### Can a repo reach GitHub Trending on fake stars alone?

It can appear there, but Trending is not proof of organic adoption. The CMU-led study found 78 flagged repositories also reached GitHub Trending, though the authors note Trending seems to filter out most superficially popular repos. Treat a Trending appearance as one weak signal and always pair it with dependents growth before drawing a conclusion.

### How long does fake-star momentum last before it reverses?

Under two months. The fake-star research shows the promotion effect lasts less than two months and then becomes a liability. That is why a four-week adoption snapshot is dangerous: it can capture the pump rather than the truth. Measure the trend at 60 to 90 days after a major release to see whether real usage followed the attention.

### Why can't I compare star counts across programming languages?

Star norms shift with language mix and hype cycle, so a raw cross-ecosystem comparison misleads. Python overtook JavaScript on GitHub driven by AI, and TypeScript later became the most-used language, each reshaping what a normal star count looks like. Compare within a category and normalize to contributor density or dependents rather than trusting raw stars across ecosystems.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/real-adoption-vs-manufactured-momentum*
