# Verifying a Consumer App's Traction Claims From Public Store Signals

*You will turn one app's public store signals into modeled download and active-user bands with a confidence level, and reach a defensible holds/inflated/decaying verdict before you wire.*

- Canonical URL: https://www.refolk.ai/guides/verify-app-traction-store-signals
- Pillar: Investing and deal sourcing
- Format: Teardown
- Published: 2026-09-18
- Last reviewed: 2026-09-18
- Reading time: 17 min

You are about to wire into a consumer app, and the deck leads with a download count and an active-user number. This guide is for early-stage investors, platform and talent partners, and angels who need to know whether those figures are real before the money moves. It carries one app end to end from raw public store signals to a traction verdict, with the actual queries, the intermediate counts, the wrong turns, and honest confidence bands attached instead of a bought number.

The worked example is deliberately generic so you can substitute your own app: call it a mid-tier consumer photo app, iOS and Android, deck claiming 2,000,000 lifetime downloads and 400,000 monthly active users. Follow along on your own case and you will end with a defensible one-paragraph verdict.

## What store signals can and cannot prove

Public store signals can corroborate a download claim to an order of magnitude and can expose rank decay. They cannot produce MAU, DAU, or retention at all. Knowing which claim you are testing before you pull anything is the difference between a defensible check and a plausible-looking guess.

The store gives you three public fields per app: chart rank, star rating, and rating count. Apple exposes these per country storefront and publishes no download number. Google Play publishes an install count, but as a lower-bound threshold like "10,000+", showing only the bottom of the range. Neither store shows active users, sessions, or retention. Every active-user claim therefore needs a fallback outside the store.

| Deck claim | Store signal proves it? | Fallback |
|---|---|---|
| Monthly downloads | Partially, order-of-magnitude | Google Play bucket plus modeled iOS |
| MAU / DAU | No | Paid panel estimate plus cohort export |
| Retention D1/D7/D30 | No | Founder analytics export vs benchmarks |
| Team plausibility | No | LinkedIn headcount cross-check |

The load-bearing point: a download measures one thing, your ability to get someone to click in a store, and traction is read in the D1/D7/D30 retention curves. With average 30-day retention across all apps around 6%, a big download number implies almost nothing about a live user base. So the download claim is the cheapest to test and the least informative, and the active-user claim is the one that matters and the one store signals cannot touch.

> **Rule:** Define active before you test it
>
> You still have to define active explicitly, because a vague definition is as good as no metric. Write the deck's MAU claim as a testable sentence with its denominator and window before pulling a single number.

## Fix the exact claim

Start by turning the deck's numbers into testable sentences. This is 15 minutes of copying, and skipping it is how investors end up arguing about the wrong figure two weeks later.

For the example app, the deck says "2M downloads" and "400K MAU". Neither is testable yet. Rewrite them:

- "Cumulative lifetime installs across iOS and Android exceed 2,000,000 as of the deck date."
- "In the calendar month before the deck date, 400,000 distinct devices had one or more foreground sessions."

Now the claims have denominators and windows. The MAU rewrite forces the definition question immediately: is a monthly active user a device with one session, or a device that completed a core action? If the founder cannot answer, the number is undefined and you note that as your first finding. A vague definition is as good as no metric.

Watch for a subtler trap here that store signals will never surface. Holloway's diligence chapter documents a company claiming 85,000 customers that turned out to be acquired when the product was an add-on to a large ecosystem. If a cohort was bundled or inherited rather than won, no download model will tell you. Add "how was each cohort acquired?" to your founder list now.

## Pull the current store snapshot

Record both listings today, dated, one row per relevant country storefront. This is your baseline; without it, tomorrow's daily pull has nothing to compare against.

For iOS, capture the star rating and total rating count per country you care about (usually the app's home market plus its top two). For Google Play, capture the install threshold and the rating count. For the example app, assume the US listing shows a 4.5 rating on 41,000 ratings, and Google Play shows "1,000,000+" installs.

That Google threshold is your first hard fact. It is a cumulative lifetime floor, not a current or active figure. "1,000,000+" means true lifetime Android installs are at or above one million and could be 4,900,000. It is a floor only, and reading it as precise or as live users is the most common false positive in this entire process.

**200 - Ranked apps per country in Apple's public feed**

Apple caps its public feeds at 200 ranked apps per storefront and keeps no chart history, so anything you want to track over time you must snapshot yourself.

## Start a daily rank and ratings time series

Schedule a daily pull of chart position and rating count at a fixed hour, because Apple keeps no chart history and you cannot reconstruct it later. Done means your second run produces deltas.

This is the step most investors skip and the one that pays off most. Apple publishes where an app stands today and keeps no history of it. That amnesia is your edge: an investor who starts a daily snapshot early owns rank-decay evidence the founder cannot retroactively dispute. A DIY daily snapshot needs no login, and vendors that scan 35 metadata points hourly across 50-plus countries exist if you want density. Daily at a fixed hour is enough for comparability; hourly is overkill for a single diligence.

#### The signal-to-verdict pipeline

1. **Snapshot** - Record rank, rating, rating count per country, dated
2. **Delta** - Daily pulls produce rating-count and rank movement
3. **Model** - Back out two independent download bands
4. **Floor** - Check bands against the Google Play threshold
5. **Verdict** - Attach confidence and classify holds, inflated, decaying

*Each stage narrows a claim into a band, and the daily snapshot is the only stage a founder cannot reconstruct after the fact.*

Run the pipeline for at least a week before you need the answer. Two weeks of daily rank data separates a genuine growth curve from a paid spike better than any single estimate.

## Back out a download band from the signals

Model downloads two independent ways and keep the wider band where they disagree. The band, not the number, is the deliverable, because download modeling compounds error instead of averaging it.

**Method one: rating-count delta.** Take the change in total rating count over your observation window and divide by a raters-per-install range. There is no clean benchmark. The legacy rule is about one rating per 30 paid-app buyers, but commenters found apps far off that, suggesting one per 100 might be closer, and one analytics vendor found a minority of apps at a ratio of 1 to 3 while outliers ran into the thousands, with no definitive correlation. So apply the range: if the example app gained 700 ratings in seven days, that back-outs to roughly 21,000 to 70,000 installs that week, or very roughly 3,000 to 10,000 daily.

**Method two: rank-to-downloads mapping.** Independently map the app's chart rank to a daily download figure using published thresholds.

| Chart | Free (daily downloads) | Paid (daily downloads) |
|---|---|---|
| Overall top 25 | 38,400 | 3,530 |
| Games top 25 | 25,300 | 2,280 |
| Photography top 25 | not given | ~270 |

These are a dated snapshot from a small sample, so treat them as order-of-magnitude anchors, not a lookup table. If the example photo app sits around the photography top 25 on the free chart, the free thresholds tell you it needs meaningful daily volume to hold that rank, consistent with the rating-delta band. Where the two methods disagree, keep the wider band.

> **Watch out:** Drop the rank method for paid and niche apps
>
> The rank-to-downloads mapping was built on free-app charts and overstates a paid app. Modern estimators do not estimate paid downloads at all. Confirm whether the app is on the Top Free or Top Paid chart first; for paid, the rating-delta method and the Google floor are all you have.

Now reconcile against the Google Play floor. The example app's rating-delta band plus its iOS rank suggest a total lifetime figure in the low millions across both platforms, and Google's "1,000,000+" Android floor is consistent with that. A deck claiming 2,000,000 lifetime downloads against a "500,000+" bucket, by contrast, is a red flag you can state plainly.

**10x - Revenue range from a compounding model**

When each input in a download times ARPU model shifts by 30 to 40%, the output range spans more than 10x, because the errors multiply rather than average.

## Attach a confidence level

State a documented error band on every figure and widen it for the hard cases. A point estimate that hides a 10x compounded range is worse than an honest band, and it is a documented failure mode.

The published anchor is a median absolute percent error running from 5% in the best case to 25% in the worst. That best case applies to popular, top-100, free apps in dense markets. Widen it for three conditions:

- **Paid apps:** downloads are not estimated at all, so you have only the rating-delta method and the Google floor.
- **Small categories or markets:** limited ranking signals produce less stable estimates.
- **Mid-tier apps ranked #300 to #1000:** panel overlap is thin and the models lean harder on assumptions.

That third condition is the trap. Panel density collapses below the top 100, and mid-tier is exactly where seed-stage consumer deals live. The apps investors most often diligence are the least reliably estimated, so your default confidence word for a seed-stage app should be "low to moderate", not "high".

> The band, not the number, is the deliverable, because download modeling compounds error rather than averaging it away.

## Corroborate active users and retention through the fallback

Store signals cannot yield MAU or DAU, so pull a panel estimate and demand the founder's cohort export, then reconcile the two. This is where the real traction claim lives, and there is no store shortcut.

A paid panel product returns estimated MAU, WAU, or DAU for an app with historical data back 37 months, defining an active user as a device with one or more foreground sessions in the period. That gives you an independent read on the 400,000 MAU claim. Pull two independent estimators if you can, because they use different panels and often disagree, especially in the mid-tier; treat a single tool's number as truth and you have cited a failure mode as a fact.

Then sanity-check the ratio. If lifetime installs are around 2,000,000 and D30 retention across all apps averages 6% to 8%, most of those installs are dead. A 400,000 MAU claim implies a downloads-to-MAU relationship far healthier than benchmark, which is possible for a genuinely sticky app but demands the founder's cohort export to defend.

| Metric | Value | Source |
|---|---|---|
| D1 retention (all apps) | ~26-28% | Getstream / growth-onomics |
| D7 retention (all apps) | ~13-18% | Getstream / growth-onomics |
| D30 retention (all apps) | ~6-8% | Adjust / Getstream |
| DAU/MAU stickiness (NA gaming) | ~31-32% | Mixpanel |

Install-cohort retention and active-user retention are not directly comparable, so use this table to spot claims that are physically implausible, not to grade a company against a single line. A deck claiming 30% D30 for a general consumer app when the all-app average is 6% to 8% is not disqualifying on its own, but it is a claim the cohort export must prove.

## The headcount cross-check

Compare the claimed scale against employee count and the presence of growth staffing, because store signals and analytics both originate with the founder and the professional-profile count does not. This is the one un-gameable cross-check in the process.

If a consumer app attributes its growth to paid acquisition, it should have User Acquisition and App Store Optimization people on or near the team. The talent pool exists and is countable.

| Segment | Count | Derived |
|---|---|---|
| ASO skill, US | 1,166 | baseline |
| ASO skill, UK | 440 | US = 2.65x UK |
| User Acquisition skill, US | 672 | US ASO = 1.74x US UA |

The logic is simple: a claimed paid-growth engine with no UA or ASO hires nearby is internally inconsistent, and investors will notice a headcount-to-traction discrepancy. If the LinkedIn page shows five employees but the deck claims a large paid engine and high scale, that gap is a founder question, not a footnote. The counts also give you a benchmark for what "staffed" looks like: with only 672 US UA professionals total, a real UA function is a small, identifiable set of people you can name.

I ran this search: `Growth and user-acquisition leads who worked at consumer mobile apps in San Francisco in the last two years` - [see the full result list](https://www.refolk.ai/s/3sh74gfbbg).

*Returns named growth and UA people you can use as a comparator cohort and to check whether the paid-acquisition story has real staffing behind it.*

Refolk resolves the headcount question in plain English without scraping profiles by hand. You can also ask [Refolk](/) for "Product analytics or data engineers at consumer social apps" to confirm the target can even produce the cohort data it claims, and for "Former employees now at other companies" to line up reference calls with people who saw the internal metrics.

## The procedure, end to end

Run these nine steps in order. Steps 1 to 6 are a single sitting plus a week of daily pulls; steps 7 to 9 need the fallback data and the founder's cooperation.

#### From store signals to a traction verdict

1. **Fix the exact claim** - Copy the deck's literal numbers and definitions with dates and windows, and rewrite each as a testable sentence with its denominator. Define any vague word like active explicitly.
2. **Pull the current store snapshot** - Record iOS star rating and rating count per country and the Google Play install threshold and rating count, dated, one row per country.
3. **Start a daily rank and ratings time series** - Schedule a daily pull of chart position and rating count at a fixed hour, because Apple keeps no chart history; the second run produces deltas.
4. **Back out a download band** - Apply a one-rating-per-30-to-100-installs range to the review delta and independently apply a rank-to-downloads mapping; keep the wider band and drop the rank method for paid apps.
5. **Cross-check against Google Play buckets** - Treat the install threshold as a hard lifetime floor and judge whether the deck's lifetime claim sits inside or above the bucket.
6. **Attach a confidence level** - State a 5 to 25% MAPE band and widen it for paid, mid-tier, or small-market apps so every number carries a band and a confidence word.
7. **Corroborate active users and retention** - Pull a panel MAU estimate and demand the founder's cohort export, then sanity-check the downloads-to-MAU ratio against benchmark stickiness.
8. **Run the team and headcount cross-check** - Compare claimed scale against employee count and UA/ASO staffing; flag any gap as a founder question.
9. **Write the verdict** - Classify the app as holds, inflated, or decaying, with declining rating velocity and rank indicating decay, in one paragraph carrying bands and a confidence word.

## How this check goes wrong

Every failure mode below produces a plausible number that is false. Knowing what each signal looks like when it lies is the most valuable part of this standard, because a confident wrong verdict costs more than no verdict.

- **Rating-count back-out on a review-prompted app.** An app that aggressively prompts ratings has a raters-per-install rate far below 30, so the back-out inflates downloads. Check: compare rating velocity before and after a known "rate us" prompt update in the app's release notes.
- **Rank-to-downloads on a paid or niche app.** The free-chart mapping overstates a paid app, and estimators do not estimate paid downloads. Check: confirm the Top Free versus Top Paid chart and drop the rank method for paid.
- **Google Play bucket read as precise.** "1,000,000+" read as exactly 1M when the truth could be 4.9M, or read as current users when it is lifetime installs. Check: treat it as a cumulative lifetime floor only.
- **MAU inferred from downloads.** Treating cumulative downloads as active users; with 6% to 8% D30, most installs are dead. Check: require a panel MAU and a cohort export and reconcile them.
- **Single-tool estimate treated as truth.** One vendor's number cited as fact when panels disagree, especially at #300 to #1000. Check: pull two independent estimators and report the range.
- **Confidence band omitted.** A point estimate that hides a 10x compounded range. Check: every figure carries a stated band, 5 to 25% MAPE, wider for paid and small markets.
- **Customers that were bundled.** A large user count acquired as an add-on to a bigger ecosystem, as in the documented 85,000-customer case. Check: ask how each cohort was acquired.
- **Decay masked by a launch spike.** A favourable short window flattering the trend. Check: pull the full rank history; short paid campaigns distort projections.

#### Reading the download claim against retention evidence

Horizontal axis runs from Download band far below deck claim to Download band supports deck claim. Vertical axis runs from Weak or absent retention evidence to Strong cohort retention evidence.

| Quadrant | What it means |
| --- | --- |
| Inflated downloads, no live base | Verdict inflated; the number is both wrong and irrelevant |
| Downloads plausible, users dead | Downloads real but traction is decaying; probe D30 |
| Underclaimed but unproven | Ask why the deck undersells; verify the cohort export |
| Claim holds | Verdict holds; downloads and a live base both corroborated |

*Where an app lands here decides whether the download number is worth trusting at all.*

## Writing the verdict and keeping it current

Close with one paragraph that classifies the app as holds, inflated, or decaying, and carries a band and a confidence word for every figure. A verdict is defensible only as a band, never as a borrowed point estimate.

For the example app, a defensible verdict reads: "Modeled lifetime downloads of 1.5M to 3M, moderate confidence, consistent with the Google Play 1,000,000+ Android floor and the iOS rating-count trajectory; the 2M deck claim holds within band. MAU claim of 400,000 unverifiable from store signals and awaiting the founder's cohort export; the implied stickiness is above the 6% to 8% all-app D30 benchmark and requires that export to stand. Rank has drifted down over the two-week snapshot window, so watch for decay. UA and ASO staffing consistent with the paid-growth story."

Run the check against the list below before you call it done.

#### Before you sign off on the traction verdict

- [ ] Each deck claim is written as a testable sentence with its denominator and window.
- [ ] A dated store snapshot exists per relevant country, iOS and Google Play.
- [ ] A daily rank and ratings series has run long enough to show deltas.
- [ ] Two independent download bands were computed and the wider one kept.
- [ ] The lifetime claim was checked against the Google Play threshold as a floor.
- [ ] Every figure carries a stated confidence band, widened for paid or mid-tier.
- [ ] A panel MAU estimate and a founder cohort export were reconciled.
- [ ] Headcount and UA/ASO staffing were checked against the claimed scale.
- [ ] The verdict paragraph names bands and a confidence word, not point estimates.

To keep the read current, remember that store signals age fast and the download thresholds here are a dated snapshot from a small sample; re-anchor them against a fresh public source rather than reusing the figures as constants. Because Apple keeps no history, the only durable evidence is the daily snapshot you started, so keep it running through the raise. And re-check the headcount cross-check the week you decide: growth hires and departures move the internal-consistency read faster than any store metric, and a UA team that quit between the deck and the wire is the loudest signal of all.

## Frequently asked questions

### Can I estimate app store downloads from the ratings count alone?

Only into a wide band, never a point estimate. The raters-per-install rate is not established publicly as a stable number: the legacy rule is about one rating per 30 buyers, but commenters found figures closer to one per 100, and one analytics vendor found outliers with ratios in the thousands. Apply a 30x to 100x range to the review delta over a fixed window, treat the result as order-of-magnitude, and cross-check it against a rank-based method and the Google Play floor.

### How do I check an app's MAU from public data?

You cannot derive MAU or DAU from public store signals at all. Store feeds expose rank, star rating, and rating count, but no active-user data. The honest route is a paid panel product that models active users from a device panel, defined as a device with one or more foreground sessions in the period, combined with the founder's own analytics or MMP cohort export. Reconcile the two and sanity-check the downloads-to-MAU ratio against benchmark stickiness.

### Why not just buy a download estimate from a vendor?

Buying a number hides the band. Published estimator error runs from about 5% in the best case to 25% in the worst, and it degrades sharply for paid apps, small categories, and mid-tier apps ranked #300 to #1000 where panel overlap is thin. A single vendor figure cited as fact is a documented failure mode. Pull two independent estimators, expect them to disagree, and report the range with a confidence word rather than one purchased point.

### What does a Google Play install count like 1,000,000+ actually mean?

It is a cumulative lifetime lower bound, not a current or active figure. Google used to show ranges and now shows only the threshold, so 1,000,000+ means the true lifetime installs are somewhere at or above one million and could be several times that. It says nothing about how many of those users are still active. Read it as a floor for the deck's lifetime download claim and nothing more.

### How do I tell traction decay from a healthy launch curve?

Pull the full rank history rather than a favourable recent window. A short paid campaign or launch produces a spike that flatters projections, while genuine decay shows as declining rating velocity and a falling chart position over weeks. Because Apple keeps no chart history, you must snapshot rank daily yourself to build the evidence; starting early is the whole advantage.

### What headcount would I expect behind a real paid-growth story?

Named growth staffing that matches the claim. A consumer app attributing scale to paid acquisition should have User Acquisition and App Store Optimization people on or near the team. In Refolk's index there are 672 US professionals with a User Acquisition skill and 1,166 with an ASO skill, so the pool exists; an app claiming a large paid engine with none of these hires is internally inconsistent and worth a direct founder question.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/verify-app-traction-store-signals*
