Refolk
StandardMarket and talent intelligence

The Market Sizing Standard: When a TAM Estimate Is Defensible

You will be able to grade any TAM/SAM/SOM estimate pass or fail against explicit criteria and know exactly what to fix before it reaches a board packet.

15 min readLast reviewed August 15, 2026Read as Markdown

You built a TAM, SAM, and SOM for a category and you need to know whether it will survive a leadership or investor review before you present it. This guide is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It gives you a gradeable definition of done: a pass/fail rubric plus a verification checklist that two analysts can apply to the same estimate and reach the same verdict, built around company-count-derived sizing rather than analyst-report top-down figures.

Every public guide explains the formulas. Almost none states when the output is defensible, so estimates get waved through on confidence or torn apart on taste. This standard fixes the grading, not the arithmetic.

What makes a TAM estimate defensible

A TAM estimate is defensible when the base count is real, the methods are independent, the assumptions are sourced, the output is a range, and a second reviewer can reproduce the verdict. Anything short of that is a narrative wearing a number.

The three layers are not interchangeable, and confusing them is the most common reason a review goes badly. Define them tightly before you defend anything.

  • TAM is the total annual revenue available if every possible buyer in the category used your product.
  • SAM is the slice you can actually reach given your current product, segment, and geography.
  • SOM is the realistic share you can capture in a defined period, given competition, sales capacity, and budget.

The spine of a defensible model is bottom-up: count the accounts that match your ICP, estimate the annual revenue each would generate, and multiply. This produces a smaller number than a top-down analyst headline, but a more defensible one, because every input traces to something you can inspect. Consulting practice and venture investors both favor bottom-up as the number you present, with top-down kept only as a sanity ceiling.

A single number signals you either grabbed a two-year-old headline figure or made it up.

The base count is where models quietly break

The base count is the number of ICP-matching accounts you multiply by revenue, and it is the input most likely to be wrong by an order of magnitude while everything downstream looks clean. Two decisions set it: which geography you scope, and which buyer title you name.

Geography is the single biggest lever on a bottom-up TAM, before any pricing assumption. In Refolk's index of professional profiles, the same VP-of-Sales persona counts 46,257 in the United States and 635 in the United Kingdom. That is a 72.8x gap. A model that leaves geography vague can be off by nearly two orders of magnitude while every other input looks defensible.

72.8x
US-to-UK gap for the same VP-of-Sales buyer persona
In Refolk's index, VP of Sales counts 46,257 in the US against 635 in the UK. A vague geography scope can move TAM by that much on its own.
Buyer personaCountryCountIndex vs UK
VP of SalesUnited States46,25772.8x
VP of SalesUnited Kingdom6351.0x

The second silent lever is the seniority band you call "the buyer." In the United States, Directors of Sales outnumber VPs by 1.58x in Refolk's index. Define the buyer as "VP and above" versus "Director and above" and your base count changes by 2.58x before you have touched pricing.

Buyer personaCountMultiple vs VP
VP of Sales46,2571.0x
Director of Sales72,9901.58x
VP + Director (combined universe)119,2472.58x

Column source: counts from Refolk's index; multiples derived.

This is why the rubric forces an explicit buyer-title definition in the first step. The same category, sized honestly, produces very different numbers depending on choices that never appear on the summary slide. To see how much the choice matters, look at the illustrative TAM it produces at two placeholder price points.

Buyer universe (US)Count× ACV $30k× ACV $50k
VP of Sales only46,257$1.39B$2.31B
VP + Director119,247$3.58B$5.96B

Column source: counts from Refolk's index; TAM figures derived. ACV values are illustrative placeholders, not sourced pricing.

The point of that table is not the dollar figures, which are illustrative. It is the spread: two reasonable definitions of the same buyer, at the same price, produce TAMs that differ by more than 4x. Your base count decision has to be written down and defended, because a reviewer who picks a different band will get a different answer and think you got it wrong.

The evidence bar for each input

Every input in a defensible model clears a stated evidence bar, and an input traceable only to an analyst headline is treated as unsourced. This is the difference between a model a reviewer can interrogate and one they have to take on faith.

The base count comes from counting real ICP-matching accounts, not assuming them, using census data, an industry database, or a sourcing tool that returns actual people and companies. Revenue per customer uses actual pricing: annual contract value for SaaS, or average annual spend for usage and transaction models. Penetration uses primary research plus competitor benchmarks, never a flat 1%.

InputMinimum evidenceWhat it looks like when it lies
Base countCounted ICP accounts from a named sourceA round number with no query behind it
Revenue per customerReal ACV or average annual spendA single price applied to every segment
Penetration ratePrimary research plus competitor benchmarkA flat 1% with no derivation
Any headline figureSourced or dated primary-research noteTraceable only to an analyst headline

Counting the base is where a sourcing tool earns its place. Rather than assume a buyer universe, you can pull the actual count of people who match the persona, geography, and company profile you defined, which is exactly the number your bottom-up model needs.

Refolk turns the base-count step from a guess into something reproducible: two analysts running the same query get the same universe, which is the precondition for two people grading the same estimate the same way.

The pass/fail rubric

An estimate passes when it clears every criterion below, and fails if it misses any one of them. The criteria are written so two reviewers grading the same model reach the same verdict.

CriterionPassFail
Layer definitionsTAM, SAM, SOM each one sentence with geography, ICP, buyer title, time windowAny layer vague or geography unstated
Bottom-up spineTAM = counted accounts × real ACVHeadline figure with no count behind it
TriangulationTwo independent methods within 15%One method, or agreement from back-solving
Primary validation30 to 50 ICP-matched strangers per segmentValidated on friendly or beta users
Output formRange with sensitivity-set boundsSingle point estimate
SOM basisBuilt bottom-up, 5 to 15% of SAMFlat percentage of a big TAM
SourcingEvery input cited or datedAny input traceable only to a headline
Stage fitClears the low end of the stage bandBelow the published floor

Two of these deserve a note on how the numbers were set.

On triangulation, the widely copied rule is that top-down and bottom-up within 15% means your assumptions are likely solid, and a gap above 15% is a signal to revisit your ratios. That rule is necessary but not sufficient, because bottom-up can be back-solved to match a top-down headline. Agreement within 15% proves nothing if the inputs were not derived independently first. The rubric checks derivation order, not just the two outputs.

On SOM, a realistic near-term figure is typically 5 to 15% of SAM. A SOM above 20% of SAM generally requires competitive data showing low incumbent concentration and high switching activity to be credible. For context, when tech companies go public they usually capture only 0.1% to 2% of their TAM, so a SOM that implies double-digit TAM capture in a few years should draw hard questions.

The stage threshold is a band, not a number

An estimate is graded against the low end of the published TAM band for its funding stage, because sources disagree on the exact floor. Grading against one number would be false precision that two reviewers could never agree on.

The rationale is consistent even where the floors differ: venture-scale businesses need a credible path to $100 million in annual revenue to create the exit potential fund economics require, which is why seed investors look for billion-dollar-range markets. The documented empirical basis for the thresholds is an analysis of 200-plus early-stage decks, consistent across US and European markets.

StageCited TAM floor (range across sources)How to grade
Seed$100M to $1B+Clears the low end of the band
Series A$1B+Meets or exceeds $1B
Series B+$5B to $10B+Meets the lower bound with room

Grade "does it clear the low end of the published band for this stage," not "does it hit one number." A worked case shows the honest output form: triangulation produced a defensible range of $100M to $300M rather than a single figure, which is what a range looks like when the methods are doing their job.

From category to capturable revenue

  1. TAM
    every possible buyer

    whole category spend

  2. SAM
    your geography and ICP

    filters actually subtracted

  3. SOM
    5 to 15% of SAM

    built from sales capacity

Each layer subtracts a filter, and the SOM should be built bottom-up, not taken as a percentage of the top.

How to grade an estimate step by step

Grade in this order, because each step depends on the one before it, and a failure early makes later checks moot. Assign an owner and a rough duration to each so the review is a scheduled task, not a hallway conversation.

The grading procedure

  1. Scope and define the three layers
    Write TAM, SAM, and SOM as one sentence each with explicit geography, ICP filters, buyer title, and time window. Done when all three are unambiguous.
  2. Build the bottom-up model
    Count ICP-matching accounts from a named source, estimate ACV per account, and multiply. Done when TAM = sourced count × real pricing.
  3. Build the top-down check
    Filter a published industry figure to the same ICP so it stands as the outer bound. Done when it sits above the bottom-up number on identical filters.
  4. Triangulate the methods
    Confirm the independently derived methods land within 15%. If they diverge, log which ratio is responsible and re-derive it. Done when they agree or the gap is documented.
  5. Validate assumptions with primary buyers
    Interview 30 to 50 ICP-matched strangers per segment on adoption and willingness to pay. Done when adoption and WTP are backed by transcripts.
  6. Run sensitivity and set the range
    Flex each critical input plus and minus 20% and record low, base, and high. Done when the bounds come from sensitivity, not comfort.
  7. Document sources per assumption
    Attach a citation or dated primary-research note to every number. Done when each input has a source line a reviewer can follow.
  8. Review against threshold and grade
    Grade pass or fail on the rubric, check the stage band, and list every gap. Done when a second reviewer reaches the same verdict.

The whole grade takes roughly two to three days of analyst work plus the validation window. Primary validation is the long pole at one to two weeks, but it also has the sharpest payoff. Moving from friendly-sample guesswork to 30 to 50 ICP-matched interviews is documented to drop no-market-need failure below 15%, and 8 to 12 interviews already surface most failure points. Given that poor product-market fit was a factor in 43% of failures in a study of 431 shut-down venture-backed companies, the marginal cost of clearing the validation bar is low relative to the risk it removes.

Triangulation quality

Estimates agree within 15%Estimates diverge >15%
Fake divergence
rework, likely a broken shared ratio
Real divergence
log which ratio diverges and re-derive
Fake agreement
reject, bottom-up was back-solved
Real agreement
pass the triangulation criterion
Methods derived from each otherMethods derived independently
Independent derivation matters more than raw closeness; agreement from back-solving is the trap.

How this goes wrong

Most failing estimates fail in one of a handful of predictable ways, and each has a false positive that looks fine on the slide. Learn to spot the false positive, because that is what a rushed review misses.

  • The 1% fallacy. The estimate is presented as "capture 1% of a huge TAM." The false positive is a big round number with no bottom-up SOM behind it. Check that the TAM is paired with a bottom-up SOM showing how customers are actually acquired.
  • TAM mislabeled as SAM or SOM. The false positive is a headline that is really the whole category spend, cited as your addressable market. Citing "the $200 billion cloud infrastructure market" without filtering is a narrative, not a strategy. Check that geography and ICP filters were actually subtracted.
  • Firmographics-only SAM. Filtering by industry code and headcount alone. Two companies look identical on paper but have opposite budgets. Check for a spend or technology-adoption filter.
  • Friendly-sample validation. WTP "validated" by beta users or the founder's network. Interviews with five friendly users who already like you tell you nothing reliable about a broader market. Check that respondents are ICP-matched strangers with budget authority.
  • Stale data. A clean-looking number sourced to a two-year-old report. Check the source date against the re-size trigger.
  • Fake triangulation. Two methods "agree" because bottom-up was back-solved to the top-down figure. Check that inputs were derived independently before comparison.
  • Point estimate. A single dollar figure with no bounds. The moment you state a point estimate you lose credibility; the reviewer assumes you grabbed a headline or made it up. Check for a sensitivity-set range.
  • Wrong geography scoped. Sizing global while selling in one country. This is the most frequent case-interview error. Check that the SAM geography matches the actual go-to-market footprint.

The verification checklist

Run this before the estimate goes into a deck or board packet. Every item is a checkable statement, not a topic, so two reviewers ticking the same list should agree on the verdict.

Before it goes in the deck

  • TAM, SAM, and SOM are each written as one sentence with explicit geography, ICP, buyer title, and time window.
  • The base count comes from a named source and the buyer title and seniority band are stated, not implied.
  • TAM equals counted ICP accounts times real ACV, not a top-down headline.
  • A top-down figure filtered to the same ICP stands above the bottom-up number as the outer bound.
  • Bottom-up and top-down were derived independently and land within 15%, or the divergence is logged.
  • 30 to 50 ICP-matched strangers per segment were interviewed on adoption and willingness to pay.
  • The output is a range with low, base, and high set by sensitivity analysis, not a single point.
  • SOM is built bottom-up and sits within 5 to 15% of SAM, or has competitive data justifying more.
  • Every input traces to a citation or a dated primary-research note.
  • The estimate clears the low end of the published TAM band for its funding stage.
  • TAM and SAM dates are within a year and SOM was rebuilt this quarter.
  • A second reviewer independently reaches the same pass or fail verdict.

Keeping a passing estimate current

A passing grade has a shelf life, so re-size on a cadence rather than treating the estimate as done. TAM and SAM drift slowly and should be reviewed at least annually or when a major market shift occurs, such as a regulatory change, a new competitor, or a pivot. SOM moves far faster.

The re-size cadence is also where the base count pays off a second time. Because the count came from a query rather than a static spreadsheet cell, you can rerun it when you re-size and see whether the buyer universe grew, shrank, or shifted geography. Comparing buyer-universe size across two markets before you set a new SAM is a query, not a research project. That is what keeps the standard alive: the estimate is not a one-time artifact but a model whose inputs you can refresh and regrade on schedule, so the number in front of leadership is always one a second reviewer would pass today.

Try it on your own search

Stop building boolean strings. Just describe the person.

Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.

  • One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
  • Read live at search time, not from a database that went stale last quarter.
  • Watch every step as it runs, and see why each name made the list.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next