The Market Sizing Standard: When a TAM Estimate Is Defensible
You will be able to grade any TAM/SAM/SOM estimate pass or fail against explicit criteria and know exactly what to fix before it reaches a board packet.
You built a TAM, SAM, and SOM for a category and you need to know whether it will survive a leadership or investor review before you present it. This guide is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It gives you a gradeable definition of done: a pass/fail rubric plus a verification checklist that two analysts can apply to the same estimate and reach the same verdict, built around company-count-derived sizing rather than analyst-report top-down figures.
Every public guide explains the formulas. Almost none states when the output is defensible, so estimates get waved through on confidence or torn apart on taste. This standard fixes the grading, not the arithmetic.
What makes a TAM estimate defensible
A TAM estimate is defensible when the base count is real, the methods are independent, the assumptions are sourced, the output is a range, and a second reviewer can reproduce the verdict. Anything short of that is a narrative wearing a number.
The three layers are not interchangeable, and confusing them is the most common reason a review goes badly. Define them tightly before you defend anything.
- TAM is the total annual revenue available if every possible buyer in the category used your product.
- SAM is the slice you can actually reach given your current product, segment, and geography.
- SOM is the realistic share you can capture in a defined period, given competition, sales capacity, and budget.
The spine of a defensible model is bottom-up: count the accounts that match your ICP, estimate the annual revenue each would generate, and multiply. This produces a smaller number than a top-down analyst headline, but a more defensible one, because every input traces to something you can inspect. Consulting practice and venture investors both favor bottom-up as the number you present, with top-down kept only as a sanity ceiling.
A single number signals you either grabbed a two-year-old headline figure or made it up.
The base count is where models quietly break
The base count is the number of ICP-matching accounts you multiply by revenue, and it is the input most likely to be wrong by an order of magnitude while everything downstream looks clean. Two decisions set it: which geography you scope, and which buyer title you name.
Geography is the single biggest lever on a bottom-up TAM, before any pricing assumption. In Refolk's index of professional profiles, the same VP-of-Sales persona counts 46,257 in the United States and 635 in the United Kingdom. That is a 72.8x gap. A model that leaves geography vague can be off by nearly two orders of magnitude while every other input looks defensible.
| Buyer persona | Country | Count | Index vs UK |
|---|---|---|---|
| VP of Sales | United States | 46,257 | 72.8x |
| VP of Sales | United Kingdom | 635 | 1.0x |
The second silent lever is the seniority band you call "the buyer." In the United States, Directors of Sales outnumber VPs by 1.58x in Refolk's index. Define the buyer as "VP and above" versus "Director and above" and your base count changes by 2.58x before you have touched pricing.
| Buyer persona | Count | Multiple vs VP |
|---|---|---|
| VP of Sales | 46,257 | 1.0x |
| Director of Sales | 72,990 | 1.58x |
| VP + Director (combined universe) | 119,247 | 2.58x |
Column source: counts from Refolk's index; multiples derived.
This is why the rubric forces an explicit buyer-title definition in the first step. The same category, sized honestly, produces very different numbers depending on choices that never appear on the summary slide. To see how much the choice matters, look at the illustrative TAM it produces at two placeholder price points.
| Buyer universe (US) | Count | × ACV $30k | × ACV $50k |
|---|---|---|---|
| VP of Sales only | 46,257 | $1.39B | $2.31B |
| VP + Director | 119,247 | $3.58B | $5.96B |
Column source: counts from Refolk's index; TAM figures derived. ACV values are illustrative placeholders, not sourced pricing.
The point of that table is not the dollar figures, which are illustrative. It is the spread: two reasonable definitions of the same buyer, at the same price, produce TAMs that differ by more than 4x. Your base count decision has to be written down and defended, because a reviewer who picks a different band will get a different answer and think you got it wrong.
The evidence bar for each input
Every input in a defensible model clears a stated evidence bar, and an input traceable only to an analyst headline is treated as unsourced. This is the difference between a model a reviewer can interrogate and one they have to take on faith.
The base count comes from counting real ICP-matching accounts, not assuming them, using census data, an industry database, or a sourcing tool that returns actual people and companies. Revenue per customer uses actual pricing: annual contract value for SaaS, or average annual spend for usage and transaction models. Penetration uses primary research plus competitor benchmarks, never a flat 1%.
| Input | Minimum evidence | What it looks like when it lies |
|---|---|---|
| Base count | Counted ICP accounts from a named source | A round number with no query behind it |
| Revenue per customer | Real ACV or average annual spend | A single price applied to every segment |
| Penetration rate | Primary research plus competitor benchmark | A flat 1% with no derivation |
| Any headline figure | Sourced or dated primary-research note | Traceable only to an analyst headline |
Counting the base is where a sourcing tool earns its place. Rather than assume a buyer universe, you can pull the actual count of people who match the persona, geography, and company profile you defined, which is exactly the number your bottom-up model needs.
Refolk turns the base-count step from a guess into something reproducible: two analysts running the same query get the same universe, which is the precondition for two people grading the same estimate the same way.
The pass/fail rubric
An estimate passes when it clears every criterion below, and fails if it misses any one of them. The criteria are written so two reviewers grading the same model reach the same verdict.
| Criterion | Pass | Fail |
|---|---|---|
| Layer definitions | TAM, SAM, SOM each one sentence with geography, ICP, buyer title, time window | Any layer vague or geography unstated |
| Bottom-up spine | TAM = counted accounts × real ACV | Headline figure with no count behind it |
| Triangulation | Two independent methods within 15% | One method, or agreement from back-solving |
| Primary validation | 30 to 50 ICP-matched strangers per segment | Validated on friendly or beta users |
| Output form | Range with sensitivity-set bounds | Single point estimate |
| SOM basis | Built bottom-up, 5 to 15% of SAM | Flat percentage of a big TAM |
| Sourcing | Every input cited or dated | Any input traceable only to a headline |
| Stage fit | Clears the low end of the stage band | Below the published floor |
Two of these deserve a note on how the numbers were set.
On triangulation, the widely copied rule is that top-down and bottom-up within 15% means your assumptions are likely solid, and a gap above 15% is a signal to revisit your ratios. That rule is necessary but not sufficient, because bottom-up can be back-solved to match a top-down headline. Agreement within 15% proves nothing if the inputs were not derived independently first. The rubric checks derivation order, not just the two outputs.
On SOM, a realistic near-term figure is typically 5 to 15% of SAM. A SOM above 20% of SAM generally requires competitive data showing low incumbent concentration and high switching activity to be credible. For context, when tech companies go public they usually capture only 0.1% to 2% of their TAM, so a SOM that implies double-digit TAM capture in a few years should draw hard questions.
The stage threshold is a band, not a number
An estimate is graded against the low end of the published TAM band for its funding stage, because sources disagree on the exact floor. Grading against one number would be false precision that two reviewers could never agree on.
The rationale is consistent even where the floors differ: venture-scale businesses need a credible path to $100 million in annual revenue to create the exit potential fund economics require, which is why seed investors look for billion-dollar-range markets. The documented empirical basis for the thresholds is an analysis of 200-plus early-stage decks, consistent across US and European markets.
| Stage | Cited TAM floor (range across sources) | How to grade |
|---|---|---|
| Seed | $100M to $1B+ | Clears the low end of the band |
| Series A | $1B+ | Meets or exceeds $1B |
| Series B+ | $5B to $10B+ | Meets the lower bound with room |
Grade "does it clear the low end of the published band for this stage," not "does it hit one number." A worked case shows the honest output form: triangulation produced a defensible range of $100M to $300M rather than a single figure, which is what a range looks like when the methods are doing their job.
From category to capturable revenue
- every possible buyerTAM
whole category spend
- your geography and ICPSAM
filters actually subtracted
- 5 to 15% of SAMSOM
built from sales capacity
How to grade an estimate step by step
Grade in this order, because each step depends on the one before it, and a failure early makes later checks moot. Assign an owner and a rough duration to each so the review is a scheduled task, not a hallway conversation.
The grading procedure
- Scope and define the three layersWrite TAM, SAM, and SOM as one sentence each with explicit geography, ICP filters, buyer title, and time window. Done when all three are unambiguous.
- Build the bottom-up modelCount ICP-matching accounts from a named source, estimate ACV per account, and multiply. Done when TAM = sourced count × real pricing.
- Build the top-down checkFilter a published industry figure to the same ICP so it stands as the outer bound. Done when it sits above the bottom-up number on identical filters.
- Triangulate the methodsConfirm the independently derived methods land within 15%. If they diverge, log which ratio is responsible and re-derive it. Done when they agree or the gap is documented.
- Validate assumptions with primary buyersInterview 30 to 50 ICP-matched strangers per segment on adoption and willingness to pay. Done when adoption and WTP are backed by transcripts.
- Run sensitivity and set the rangeFlex each critical input plus and minus 20% and record low, base, and high. Done when the bounds come from sensitivity, not comfort.
- Document sources per assumptionAttach a citation or dated primary-research note to every number. Done when each input has a source line a reviewer can follow.
- Review against threshold and gradeGrade pass or fail on the rubric, check the stage band, and list every gap. Done when a second reviewer reaches the same verdict.
The whole grade takes roughly two to three days of analyst work plus the validation window. Primary validation is the long pole at one to two weeks, but it also has the sharpest payoff. Moving from friendly-sample guesswork to 30 to 50 ICP-matched interviews is documented to drop no-market-need failure below 15%, and 8 to 12 interviews already surface most failure points. Given that poor product-market fit was a factor in 43% of failures in a study of 431 shut-down venture-backed companies, the marginal cost of clearing the validation bar is low relative to the risk it removes.
Triangulation quality
How this goes wrong
Most failing estimates fail in one of a handful of predictable ways, and each has a false positive that looks fine on the slide. Learn to spot the false positive, because that is what a rushed review misses.
- The 1% fallacy. The estimate is presented as "capture 1% of a huge TAM." The false positive is a big round number with no bottom-up SOM behind it. Check that the TAM is paired with a bottom-up SOM showing how customers are actually acquired.
- TAM mislabeled as SAM or SOM. The false positive is a headline that is really the whole category spend, cited as your addressable market. Citing "the $200 billion cloud infrastructure market" without filtering is a narrative, not a strategy. Check that geography and ICP filters were actually subtracted.
- Firmographics-only SAM. Filtering by industry code and headcount alone. Two companies look identical on paper but have opposite budgets. Check for a spend or technology-adoption filter.
- Friendly-sample validation. WTP "validated" by beta users or the founder's network. Interviews with five friendly users who already like you tell you nothing reliable about a broader market. Check that respondents are ICP-matched strangers with budget authority.
- Stale data. A clean-looking number sourced to a two-year-old report. Check the source date against the re-size trigger.
- Fake triangulation. Two methods "agree" because bottom-up was back-solved to the top-down figure. Check that inputs were derived independently before comparison.
- Point estimate. A single dollar figure with no bounds. The moment you state a point estimate you lose credibility; the reviewer assumes you grabbed a headline or made it up. Check for a sensitivity-set range.
- Wrong geography scoped. Sizing global while selling in one country. This is the most frequent case-interview error. Check that the SAM geography matches the actual go-to-market footprint.
The verification checklist
Run this before the estimate goes into a deck or board packet. Every item is a checkable statement, not a topic, so two reviewers ticking the same list should agree on the verdict.
Before it goes in the deck
- TAM, SAM, and SOM are each written as one sentence with explicit geography, ICP, buyer title, and time window.
- The base count comes from a named source and the buyer title and seniority band are stated, not implied.
- TAM equals counted ICP accounts times real ACV, not a top-down headline.
- A top-down figure filtered to the same ICP stands above the bottom-up number as the outer bound.
- Bottom-up and top-down were derived independently and land within 15%, or the divergence is logged.
- 30 to 50 ICP-matched strangers per segment were interviewed on adoption and willingness to pay.
- The output is a range with low, base, and high set by sensitivity analysis, not a single point.
- SOM is built bottom-up and sits within 5 to 15% of SAM, or has competitive data justifying more.
- Every input traces to a citation or a dated primary-research note.
- The estimate clears the low end of the published TAM band for its funding stage.
- TAM and SAM dates are within a year and SOM was rebuilt this quarter.
- A second reviewer independently reaches the same pass or fail verdict.
Keeping a passing estimate current
A passing grade has a shelf life, so re-size on a cadence rather than treating the estimate as done. TAM and SAM drift slowly and should be reviewed at least annually or when a major market shift occurs, such as a regulatory change, a new competitor, or a pivot. SOM moves far faster.
The re-size cadence is also where the base count pays off a second time. Because the count came from a query rather than a static spreadsheet cell, you can rerun it when you re-size and see whether the buyer universe grew, shrank, or shifted geography. Comparing buyer-universe size across two markets before you set a new SAM is a query, not a research project. That is what keeps the standard alive: the estimate is not a one-time artifact but a model whose inputs you can refresh and regrade on schedule, so the number in front of leadership is always one a second reviewer would pass today.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.