Refolk
TeardownInvesting and deal sourcing

Pressure-Testing a B2B Startup's ARR Claim Before the Term Sheet

You can take a claimed ARR figure for a B2B startup and corroborate it to a defensible range from public signals, or quantify the gap, before you wire.

16 min readLast reviewed October 9, 2026Read as Markdown

You have a founder's claimed ARR number and a term sheet in front of you. This guide is for early-stage investors, platform and talent partners, and angels who need to decide whether that number is real using public signals, before the 30 days of confirmatory diligence that follow a signature. It carries one B2B SaaS deal all the way through: the real queries, the intermediate counts, the arithmetic that turns counts into an ARR range, and the wrong turns that make a claimed number look bigger than it is.

The worked example is a mid-market workflow SaaS company. The founder claims $4M ARR at what presents as a late-seed stage. There is no app store to check, no download counter, no public MAU. Everything below is about bounding paying customers from hiring, named logos, review velocity, and pricing math instead.

Why a B2B ARR claim needs its own method

A B2B SaaS ARR claim cannot be checked the way a consumer app's traction can, because there is no public install base; you have to reconstruct paying-account count from indirect signals and multiply by a defensible contract value. That reconstruction is the whole job.

The claim here is a single revenue number that stands in for the health of the business. Unlike a consumer product, where a download count is public and roughly honest, a B2B number is assembled internally from contracts, and the founder controls both the arithmetic and what counts as recurring. Your task is not to reproduce their internal number. It is to build an independent range from signals they do not control, and then see whether their number sits inside it.

Three things make this tractable. Named customers leak onto logo walls and case studies. Paying customers leave reviews, and each review marks a real account. And companies hire customer-facing staff in proportion to the accounts they must serve, which is expensive and therefore hard to fake. None of these is sufficient alone. Agreement across them is what makes a verdict defensible.

28,475
Customer success managers at US software companies
From Refolk's index. This is the base against which a claimed account count must be plausible.

The three public signals and what each one proves

The three signals worth building on are named logos, published review counts, and customer-facing hiring; each bounds paying customers differently, and each lies in a specific way. Know the failure direction before you trust the number.

Named logos and case studies are a hard floor but a severe undercount. Most vendors showcase only marquee accounts, so a logo wall of twelve names tells you there are at least twelve customers and nothing about the ceiling. It is a selection, not a census.

Review-platform counts are a defensible floor because each review maps to at least one real customer account. They still undercount, because only a fraction of customers ever review. And they can be inflated: a 50,000-review scrape found 25.5% were incentivized, ranging from 3.5% for one vendor to 78% for another, with a standard $25 gift card per review. Velocity is where the lie lives. Learnerbly went from 3 to 467 G2 reviews in two weeks, a 155% jump. Use the cumulative count, never the slope.

Hiring for customer success, onboarding, and implementation is the mid-range proxy and the hardest to fake, because headcount costs money and scales with the active account base. It lags growth and varies by contract size, but a claimed large customer base with near-zero onboarding staff is internally inconsistent.

SignalWhat it provesHow it lies
Named logos / case studiesHard floor of named accountsMarketing selection, severe undercount
Review countFloor of one account per reviewIncentivized campaigns, velocity spikes
CS / onboarding hiringActive base roughly tracks headcountLags growth, varies by contract size

From customer count to an ARR range: the arithmetic

ARR is roughly paying customers times blended annual contract value, and you bound it by applying a low and a high ACV from the company's segment band rather than a single figure. The error that dominates this step is not arithmetic. It is choosing the wrong segment.

Annual contract value, or ACV, is the average annualized revenue per customer contract. Public benchmarks exist by segment, and the spread between segments is enormous. SMB and micro-business land at a $4K-$15K range with a median near $8K. Mid-market runs $20K-$75K with a median of $42K. Enterprise runs $100K-$350K with a median of $185K. A cross-segment anchor from a survey of over 1,000 private B2B SaaS companies puts the median at $26,265.

SegmentMedian ACVTypical range
SMB / micro$8K$4K-$15K
Mid-market$42K$20K-$75K
Enterprise$185K$100K-$350K
Cross-segment$26,265-

Applying the enterprise median to an SMB self-serve product, or the reverse, swings the resulting ARR range by more than 20x. That is why step two of the procedure is reading the pricing page, not multiplying. Ground your low and high ACV in the vendor's own list prices first, and only then reach for the segment benchmark to fill gaps or set the band width.

For the worked example, the pricing page lists three published tiers topping out around $45,000 a year, and the sales language targets "revenue teams at growing companies." That places it squarely mid-market. I take a low ACV of $25,000 and a high of $50,000.

Now the customer floor. The logo wall shows 9 named accounts. G2 and Capterra together show 210 cumulative reviews, with no visible velocity spike - the dates spread evenly over three years, so I treat the full count as organic. Reviews are a floor of accounts, and only a fraction of customers review, so I estimate an active base somewhere between the 210-review floor and a plausible multiple of it. I carry two account estimates: a conservative 210 and a working 320.

The range: 210 to 320 accounts times $25,000 to $50,000 ACV gives a corroborated ARR band of roughly $5.25M at the low-account, low-ACV corner up to $16M at the high corner, with a central estimate around $9.6M. The claimed $4M sits below even my conservative corner. That is the first surprise, and it points the other way from the usual failure: the public signals suggest the founder may be understating, or that the reviews are inflated. Hold that tension and move to the cross-check.

Narrowing a mid-market SaaS to an ARR range

  1. Named logos
    9

    Hard floor only

  2. Cumulative reviews
    210

    Floor of real accounts

  3. Working active base
    320

    Review floor plus unreviewed accounts

  4. ARR at $25K-$50K ACV
    $5.25M-$16M

    Claimed $4M sits below the floor

Each stage converts a public signal into a tighter bound on paying accounts, then into revenue.

The efficiency cross-check that catches what arithmetic misses

Dividing claimed ARR by observed headcount and comparing to the stage benchmark is the fastest way to flag a number that the customer-count math alone would pass. It is a different lens, so it catches a different lie: annualization and services padding.

A claim can survive the customer-count arithmetic yet fail here. The benchmark is ARR per employee by revenue band. Companies under $5M run a median of about $126,499 per employee. The $5M-$20M band sits near $178,000. The $20M-$50M band reaches $278,848, from a sample of 17 companies. The $50M-$100M band is around $240,000. The cross-segment median rose 29% to $193K, so treat these as a slowly moving target and re-pull them rather than hard-coding.

ARR bandMedian ARR per employee
Under $5M$126,499
$5M-$20M$178,000
$20M-$50M$278,848
$50M-$100M$240,000

The red flag is a small team claiming scale-stage efficiency: a sub-$5M-looking company reporting figures near the $20M-$50M median of $278,848 is almost certainly annualizing a spike or counting one-time services. In the worked example, the company shows roughly 55 employees on LinkedIn. The founder's $4M claim divided by 55 is about $73,000 per employee, well under the $126,499 under-$5M median. That is the opposite of a padding signal. A team of 55 running at $4M ARR is inefficient, not inflated. Combined with the customer-count range pointing higher, the efficiency check reinforces the read that $4M is conservative or that the review count overstates the real base.

Here is the fork in the worked example, and the wrong turn I nearly took. My first instinct was to trust the $5.25M-$16M arithmetic range and write the founder up as sandbagging. That would have been an error. The efficiency check at $73,000 per employee is low, which is also consistent with a heavily incentivized review base inflating the account floor. I cannot resolve which story is true from the desk alone, so I flag both and carry the tension into the failure-mode tests rather than picking a winner.

A claim that survives the arithmetic can still fail the efficiency check, and the two lies look nothing alike.

The hiring signal: sizing the CS team against a real benchmark

Customer-facing hiring is the signal a founder cannot fake cheaply, because customer success, onboarding, and implementation headcount scales with the number of live accounts and shows up in public professional data. Benchmark the observed team against the index to turn a raw count into a judgement.

Refolk's index gives the population to calibrate against. There are 28,475 customer success managers at US software companies and 5,983 in the UK, a 4.76x ratio that matters when the target operates across both. There are 13,113 implementation and onboarding specialists in the US, which works out to roughly one implementation specialist for every 2.17 customer success managers across the US software population.

Role groupCountryHeadcount
Customer Success ManagerUnited States28,475
Customer Success ManagerUnited Kingdom5,983
Implementation / Onboarding SpecialistUnited States13,113

For the worked example, I count the target's current customer success and onboarding staff and its open roles for the same functions. The company shows 6 CS and onboarding people and 2 open reqs. For a mid-market product at roughly 250-320 accounts, 6 CS people means each manager carries 40-50 accounts, which is plausible for mid-market and would be implausible for enterprise. That consistency is a point in favor of the account estimate. If I had found 2 CS staff against a claimed 300 enterprise accounts, the story would collapse: nobody onboards 150 enterprise accounts per person. The two open reqs also tell you the active base is growing, not shrinking, which bears on churn.

Refolk removes the manual scraping this step would otherwise require. Instead of hand-counting titles on a company page, you ask Refolk for the current customer-facing staff and get a named list back, which is also the starting point for reaching references in a later step.

The step-by-step procedure

The full procedure runs eight steps and about four to six analyst hours, which fits comfortably inside the desk-diligence window before a term sheet. Each step has a clear done condition so you know when to move on.

Pressure-test the ARR claim

  1. Fix the claim and the ARR definition
    Record the exact figure, stage, date, and whether the founder said ARR or run rate. Done when you have a single number and a date to test.
  2. Segment the company by ACV tier
    Read the pricing page and target-customer language to place it in SMB, mid-market, or enterprise. Done when you have a low-high ACV band.
  3. Build the customer-count floor
    Count named logos, case studies, and current review totals on G2, Capterra, and FeaturedCustomers. Done when you have a defensible minimum account count.
  4. Pull the hiring signal
    Count current CS, onboarding, and implementation staff plus open roles, and benchmark the ratio against the index. Done when headcount is logged.
  5. Compute the ARR range
    Multiply the customer-count floor and a plausible active-account estimate by low and high ACV. Done when you have a low-high ARR band.
  6. Run the plausibility cross-check
    Divide claimed ARR by total headcount and compare to the stage ARR-per-employee benchmark. Done when the claim is inside or outside the band.
  7. Test for the failure patterns
    Probe run-rate-versus-ARR, LOIs and pilots, churn, and services revenue. Done when each pattern is marked present, absent, or unknown.
  8. Write the gap
    State the claim against the corroborated range and the size of any gap. Done when the memo gives a range and a confidence note.

Sources disagree on sequencing. Some teardowns start with the arithmetic and only then run the efficiency cross-check; others use the efficiency check as a first filter to decide whether the arithmetic is even worth building. Either order works. The only rule is that you never skip the cross-check, because it catches the failure the arithmetic is blind to.

Two lenses on one claim

  1. Bottom-up
    Customer floor times ACV band gives an ARR range
  2. Top-down
    Claimed ARR divided by headcount versus stage median
  3. Reconcile
    Agreement raises confidence; disagreement names the lie to probe
The arithmetic bounds the number from the bottom up; the efficiency check tests it top down.

How this goes wrong: failure modes and false positives

Most bad ARR claims survive a quick read because they exploit a specific definitional gap, not an outright lie. The eight patterns below are where a number inflates, and each has a concrete tell you can check from the desk.

Logo wall read as full customer list. The false positive is counting 12 logos and inferring 12 customers. Logos are a marketing selection. Treat them only as a floor and reconcile against reviews and headcount.

Review count inflated by a campaign. A velocity spike, dozens of reviews in days as in the Learnerbly case, reads as rapid growth. Inspect review dates for clustering and look for incentivization language. Discount the burst and use the pre-spike cumulative count.

Run rate dressed as ARR. Taking the best month's revenue and multiplying by twelve is annual run rate, not ARR, and conflating the two will get you caught in the first ten minutes of real diligence. Demand three consecutive months at the claimed rate; absent that, label it run rate.

LOIs and pilots counted as revenue. Letters of intent are non-binding and can be withdrawn at any time, and startups with impressive LOI pipelines sometimes fail to convert a single one. Separate contracted-and-live from contracted-future.

Services revenue bundled into ARR. Implementation and onboarding fees are one-time and should not sit in a recurring number. Back them out before comparing.

ARR per employee implausibly high. A seed team claiming scale-stage efficiency, for instance figures above the $278,848 median while clearly sub-$5M, is a flag for annualization or services padding.

Gross versus net of churn. A claim can ignore departed accounts. Dormant named logos and stale case studies are churn tells; look for accounts that have disappeared from the wall.

ACV mis-segmentation. Applying enterprise ACV to an SMB self-serve product inflates the range by more than an order of magnitude. Match the ACV band to the pricing page and sales motion before you multiply.

The deepest trap is that contamination is uneven across vendors. With incentivized-review rates ranging from 3.5% to 78%, raw review counts are not comparable between companies. Use them as a floor for one company at a time, never as a ranking across your pipeline.

Writing the gap and keeping the read current

The deliverable is a one-page memo that states the claim against your corroborated range and names the size and direction of any gap, with a confidence note. A range plus a confidence note is more useful to an investment committee than a false point estimate.

For the worked example, the memo reads: founder claims $4M ARR at late seed. Public signals corroborate a range of $5.25M to $16M from 210-320 accounts at $25K-$50K ACV, with an efficiency check of $73,000 per employee that sits below the under-$5M median and is consistent with either a conservative claim or an inflated review base. Hiring at 6 CS staff for roughly 300 mid-market accounts is internally consistent. Verdict: the claim is not inflated; the open question is whether the real figure is higher than stated or the review floor overstates the base. Next action is to confirm contracted-and-live versus pipeline and to reach two references.

ARR corroboration memo skeleton
Claim: $____ ARR, stage ____, as of ____, said "ARR" / "run rate"
Segment and ACV band: ____ ($____ low, $____ high)
Customer floor: ____ logos, ____ reviews, working base ____
Corroborated ARR range: $____ to $____ (central $____)
Efficiency check: claimed ARR / ____ employees = $____/employee vs $____ band median
Hiring consistency: ____ CS/onboarding staff, ____ open reqs - consistent / inconsistent
Failure patterns: run-rate [ ] LOIs [ ] services [ ] churn [ ] (present/absent/unknown)
Gap: claim is ____% above / below / inside the corroborated range
Confidence: high / medium / low because ____
Next action: ____

Fill each line from your own counts. Keep it to one page.

Before you call the ARR read done

  • You recorded whether the founder said ARR or run rate, with a date
  • The ACV band is anchored to the pricing page, not just the segment benchmark
  • Review dates were inspected for velocity spikes and discounted if clustered
  • The customer floor was reconciled against logos, reviews, and hiring
  • Claimed ARR divided by headcount was compared to the stage band median
  • Each failure pattern is marked present, absent, or unknown
  • The memo states a range and a confidence note, not a single number

Keep the read current by re-pulling the benchmarks rather than trusting these figures forever. ARR-per-employee medians move; the cross-segment figure rose 29% to $193K in one cycle, so treat the bands as a slowly drifting reference. ACV benchmarks shift with the market too. The mechanism stays fixed: build a floor from signals the founder does not control, multiply by a segment-anchored ACV band, cross-check against efficiency, and name the gap. Do this at the desk, inside the two-to-four-week seed window or the four-to-eight-week Series A window, because once the term sheet is signed you have only about 30 days of confirmatory diligence and far less leverage to walk.

Questions practitioners ask

How do you verify startup ARR from only public signals?

Build a customer-count floor from named logos and published review totals, where each review maps to at least one paying account. Place the product in an ACV segment using its pricing page, then multiply the floor and a plausible active-account estimate by low and high ACV to get a range. Cross-check by dividing claimed ARR by observed headcount against the stage benchmark. No single signal is enough; the corroboration comes from agreement across three or four.

What is the difference between ARR and run rate, and why does it matter in diligence?

ARR should reflect confirmed, recurring, cash-in-bank revenue from fully executed contracts. Run rate takes a strong recent month and multiplies by twelve, which flatters a seasonal or one-time spike. Conflating the two is one of the fastest ways a claim collapses in diligence. Demand three consecutive months at the claimed rate; absent that, label the number run rate and discount it.

Can I trust G2 or Capterra review counts as a measure of customers?

Treat review count as a conservative floor, not a census. Each review is at least one real account, but only a fraction of customers review, so the true base is larger. Reviews can also be inflated: a 50,000-review scrape found 25.5% were incentivized, ranging from 3.5% to 78% by vendor. Because contamination is uneven, use the cumulative count for a single company and never rank companies on raw review totals.

How does hiring data help confirm a revenue claim?

Customer success, onboarding, and implementation headcount scales with the active account base and is expensive to fake, which makes it the hardest signal to manufacture. A large claimed customer base with near-zero onboarding staff is internally inconsistent. Benchmark the observed CS team against peers; Refolk's index shows 28,475 customer success managers at US software companies and 13,113 implementation specialists to calibrate against.

How much time do I have to check an ARR claim before committing?

The screen is fast and the verification window is short. Investors spend about four minutes reviewing a pre-seed deck and get roughly 30 days of confirmatory diligence after the term sheet. Most of the public-signal corroboration described here takes four to six analyst hours, so it belongs at the desk stage before you sign, when a bad claim is cheapest to catch.

What is a realistic ARR-per-employee figure to sanity-check against?

Use the stage band, not a single number. Companies under $5M ARR run a median of about $126,499 per employee; $5M to $20M about $178,000; $20M to $50M about $278,848. A sub-$5M-looking team claiming figures near the $20M-$50M median is a red flag for annualization or services padding. The cross-segment median rose 29% to $193K, so re-check the benchmark periodically rather than hard-coding it.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next