# The Why-Now Score: Grading a Deal's Market Timing From Public Signals

*You will be able to grade any deal's timing across five catalyst dimensions from public evidence and reach one verdict: real inflection, too early, or no catalyst.*

- Canonical URL: https://www.refolk.ai/guides/why-now-score-market-timing
- Pillar: Investing and deal sourcing
- Format: Framework
- Published: 2026-09-05
- Last reviewed: 2026-09-05
- Reading time: 17 min
- Keywords: how to evaluate why now startup, assess startup market timing, market inflection point signals, why now slide investor evaluation, too early startup investment

## Key takeaways

- Bill Gross's Idealab analysis put timing at 42% of the difference between startup success and failure, ahead of team and idea, which is why the why-now claim deserves its own scored verdict.
- A live inflection is one you can date: the catalyst must have changed within roughly the last 18 to 24 months, and with AI projected to hit 50% adoption in about 3 years versus the smartphone's 5, that window is tightening.
- General Magic sold only 3,000 to 4,000 devices and lost over $74M before 1996, proving that too-early is measurable in real time, not just hindsight.
- Instacart won the identical idea Webvan lost by waiting for a specific enabler: smartphone ubiquity and an asset-light model, so score the enabler, not the market's attractiveness.
- In Refolk's index, 3,316 US founders list Generative AI as a skill, 4.37x the UK's 759, and that density can confirm an inflection or warn of consolidation depending on S-curve position.
- DocSend found investors spent 36% more time on the why-now section in funded decks, meaning weak timing claims are penalized in the room before diligence even starts.

You have a pitch in front of you and the why-now claim is doing a lot of work. Before you spend a meeting on it, you need to decide whether it is a real market inflection or manufactured urgency. This guide is for early-stage investors, platform and talent partners, and angels who want a repeatable rubric that turns a founder's timing story into one of three verdicts against public evidence: real inflection, too early, or no catalyst.

Every public why-now resource is written for the founder building the slide, not the investor stress-testing it. This one runs the other direction. It gives you the specific signals that separate a genuine inflection from a narrative dressed up as urgency, and the ones that tell "too early" apart from "right time now."

## Why grade timing at all, and why score it

Timing is the single most predictive dimension of startup outcomes, and it is also the one founders manufacture most easily, which is exactly why it needs a scored verdict rather than a gut read. Bill Gross's Idealab analysis of roughly 200 companies found that timing accounted for 42% of the difference between success and failure, ahead of both team and idea. That is the strongest primary evidence that this judgement is worth building a rubric around.

There is a second number worth keeping in view and worth keeping separate. CB Insights' post-mortem work ranks "no market need" as the leading cause of failure at 42%, ahead of running out of cash at 29% and being outcompeted at 19%. Those two 42% figures are a coincidence, not a confirmation. Gross measures share of success variance; CB Insights measures share of failure cause. Stack them and you overstate the case. Keep them apart and each earns its place.

**42% - Share of the difference between startup success and failure attributable to timing**

Bill Gross's Idealab analysis of ~200 companies ranked timing first, ahead of team and idea.

Investors already price timing behaviorally, before any rubric. DocSend found the why-now slide appeared in nearly 54% of successful decks versus 38% of unsuccessful ones, and in decks that included it, investors spent 36% more time on that section when the deck ultimately got funded. That extra dwell time is the market doing informally what this guide does formally: allocating scrutiny to the timing claim. Scoring it just makes the allocation deliberate and repeatable.

## The five catalyst dimensions

A real why-now rests on one of five external forces, and naming which one is the first act of grading. Practitioner sources converge on the same short list: technology inflection, cost-curve decline, regulatory change, behavior or adoption shift, and incumbent gap. The point of the taxonomy is not tidiness. It tells you what evidence to demand, because each category maps to a different public source.

| Dimension | What it claims | Public evidence to demand | What it looks like when it lies |
|---|---|---|---|
| Technology inflection | A new capability just became usable | Benchmark reports, adoption-curve data | "Advances in AI" with no dated capability jump |
| Cost-curve decline | A price crossed a threshold | Component or unit-price benchmarks | Spend front-loaded on a curve that hasn't moved |
| Regulatory change | A rule created or removed a market | Government reports, effective dates | A rule cited with no effective date |
| Behavior / adoption shift | Users started doing something new | Penetration data, forum chatter | "Growing awareness" true for a decade |
| Incumbent gap | A leader left an opening | Public retreats, discontinued lines, pricing | An asserted gap with no observable retreat |

The taxonomy also gives you a red flag for free. A why-now that claims all five at once is not comprehensive, it is hedging. Practitioner consensus is blunt about this: pick the one or two most compelling catalysts and communicate them clearly, because a single sharp proof point outperforms five vague ones. When a founder spreads urgency across four categories, they are covering the fact that no single one is load-bearing.

> **Rule:** One catalyst, one sentence, one date
>
> A real why-now can be stated as a single external force with a date attached. If the founder needs four forces or cannot name a date, treat the timing claim as unproven until an external source supplies both.

## The three tests that separate real from manufactured

A catalyst is real only if it survives three questions in order: has this always been true, did something dateable change, and does the change directly enable this product. Fail any one and the timing verdict drops regardless of how attractive the market looks.

The first is the "hasn't this always been true?" test. If the catalyst the founder names has been true for a decade, it is not a catalyst. "Growing consumer awareness of sustainability" is the canonical failure: real, important, and useless as timing because it has no edge in time. The remedy the founder should have applied is the before-and-after framing: eighteen months ago this company was impossible because X, and today it is inevitable because Y. If they cannot fill in both halves with dated facts, the claim is a mood, not an inflection.

The second is the timestamp test. Reject language with no timestamp, no percentage, and no named event. A live inflection is a set of conditions that exist right now in a way they did not two years ago and may not two years from now. That gives you a concrete window: roughly the last 18 to 24 months. Anything older needs a fresh dated event to still count as live.

The third is the connection test, and it is the one that voids scores. If the founder says "AI costs dropped 90%" but the product does not use AI, something is wrong. The timing trend has to directly enable or accelerate what they are building. A borrowed macro narrative is a favorite because it sounds current and requires no proof specific to the company. Draw the causal line from catalyst to product on paper. If you cannot draw it, the score is zero no matter how real the trend is in the abstract.

#### The three-test gate a catalyst must pass

1. **Hasn't this always been true?** - Reject timeless trends; demand what specifically changed
2. **Did it change in 18-24 months?** - Demand a dated event, percentage, or named change
3. **Does it enable this product?** - Draw the causal line from catalyst to product or void the score

*A why-now claim only reaches scoring after it survives all three tests in order.*

The window matters more than it used to because adoption is compressing. The telegraph took 56 years to reach roughly 50% adoption, radio 22, the internet 7, the smartphone 5, and AI tools are projected to get there in about 3. A catalyst goes stale faster now, so the "did this change in the last 18 to 24 months?" bar is stricter today than it was a decade ago.

| Technology | Years to ~50% adoption | Intro year |
|---|---|---|
| Telegraph | 56 | 1844 |
| Radio | 22 | 1922 |
| Internet | 7 | 1991 |
| Smartphone | 5 | 2007 |
| AI tools | ~3 (projected) | 2022 |

## The too-early signal, and how to read it in real time

Too-early is not a hindsight verdict; it is measurable while the deal is live, through the gap between demand pull and cost commitment. The instructive pattern is a single idea attempted at two different times. General Magic tried to ship a handheld connected device in 1990 and sold only 3,000 to 4,000 Magic Link units, mostly to family and friends, while losing more than $74 million by mid-1996. Webvan tried asset-heavy grocery delivery in 1996, raised roughly $396 million, and went bankrupt. Instacart took the identical grocery-delivery idea in 2012 and now processes tens of billions in GMV.

| Company | Founded | Outcome / demand signal | Enabler absent then, present later |
|---|---|---|---|
| General Magic | 1990 | 3,000-4,000 devices sold; -$74M by 1996 | Consumer smartphone readiness |
| Webvan | 1996 | ~$396M raised; bankrupt ~2001 | Smartphone ubiquity, asset-light model |
| Instacart | 2012 | ~$30B GMV (2025) | Arrived after both enablers existed |

The lesson is not that grocery delivery was a bad idea in 1996. It is that the enabler, not the idea, is the unit of judgement. Instacart won by waiting for smartphone ubiquity and by using existing stores instead of building warehouses. When you check the too-early signal, you are not asking whether the market is attractive. You are asking what specific enabling condition is present now that was absent in the prior attempt. If you cannot name it, the deal is too early even if the demand is genuine.

> Score the presence of the enabler, not the attractiveness of the market.

Two observable signals told the story at the time in both failed cases. General Magic's near-zero unit sales were market pull that never arrived. Webvan's billion-dollar infrastructure build against unproven demand was cost front-loaded on hope. Both were visible in real time. The failure was ignoring the data, not lacking it. That is why the too-early check is a step you can actually perform on a live deal: prior failed attempts are public, and so is the enabler that was missing.

## Founder supply as a live-adoption reading

Founder-supply density is a timing signal that cuts both ways: a crowded field can confirm that an inflection is real, or warn that the market is already consolidating. You read it against S-curve position, not in isolation.

In Refolk's index of professional profiles, 3,316 founders and co-founders in the US list Generative AI as a skill, concentrated in Seattle, Los Angeles, and the SF Bay Area. The UK count is 759, heavily concentrated in London. That is a 4.37x gap. Separately, 1,923 US founders list Robotics, with the SF Bay Area as the top hub and a notable share of "Stealth Startup" entries.

| Category / Market | Founders (count) | Reading |
|---|---|---|
| Generative AI - US | 3,316 | Dense field, mid-curve crowding risk |
| Generative AI - UK | 759 | Thinner field, earlier on the curve |
| Robotics - US | 1,923 | Concentrated, many still in stealth |
| US GenAI ÷ UK GenAI | 4.37x | Geographic supply gap |
| US GenAI ÷ US Robotics | 1.72x | Cross-category supply gap |

The same 3,316 supports opposite verdicts. Early on the S-curve, where adoption is between 1% and 5%, a dense founder field confirms a real inflection: serious people are moving. Once the market climbs into the exponential middle, that same density signals consolidation, and a new entrant is late, not early. So the number is only meaningful paired with an adoption reading. This is the sixth failure mode made concrete: a real catalyst plus thousands of founders already in-category is a too-late problem, not a too-early one.

Reading founder supply by market, skill, and geography by hand is slow, and public profiles are scattered across sources. This is where a plain-English query beats manual scraping.

I ran this search: `UK-based generative AI founders in London who previously held senior ML research roles` - [see the full result list](https://www.refolk.ai/s/k340505v4j).

*Returns named founders with prior research pedigree, letting you gauge how dense and how credentialed a geographic field already is before you decide the window is still open.*

**3,316 - US founders and co-founders listing Generative AI as a skill**

From Refolk's index; 4.37x the UK's 759, a supply gap that reads as opportunity or crowding depending on S-curve position.

## The scoring procedure

Run these seven steps in order for any deal in front of you. Each has a clear done-state so you know when to move on. Budget about 75 to 105 minutes total; the evidence-gathering step is the one that stretches.

#### Grading a why-now from public signals

1. **Extract the claimed catalyst** - Read the why-now claim and name the single external force. Done when you can write it in one sentence with a date. Founders often cite internal progress instead; that does not count.
2. **Classify the catalyst** - Assign it to one of five categories: technology inflection, cost-curve, regulatory, behavior/adoption, or incumbent gap. Done when one category is assigned. Spanning four at once is a kitchen-sink red flag.
3. **Apply the hasn't-this-always-been-true test** - Identify what specifically changed and when. Done when you have a dated event within the last 18-24 months, or the claim fails.
4. **Find independent public evidence** - Pull a benchmark, a regulation with an effective date, or penetration data from a source that is not the deck. Done when you have at least one primary source per catalyst. Allow 20-40 minutes.
5. **Run the connection test** - Confirm the catalyst directly enables this product. Done when you can draw a causal line from trend to product. A disconnect voids the score.
6. **Check the too-early signal** - Find prior failed attempts at the same idea and what has since changed. Done when you can name the specific enabler now present that was absent before.
7. **Score each dimension and reach a verdict** - Rate the catalysts and land on real inflection, too early, or no catalyst. Done when you have one verdict with the load-bearing dimension named.

Sources disagree on order at the margins. Some gather evidence before classifying, and some run the before-and-after framing first as the organizing device. Both work. What does not vary is that the connection test and the too-early check are gates, not weightings: either can void an otherwise strong score, so run them before you commit to a verdict.

### Turning dimension scores into a verdict

No source publishes a numeric per-dimension rubric, so this mapping is my own construction rather than a cited standard. Treat it as a default you calibrate to your own thesis, not a law. Score each of the five dimensions the founder invokes on evidence strength: 2 if backed by a dated primary source, 1 if plausible but thinly evidenced, 0 if it fails a gate or has no dated anchor.

**Why-now verdict rubric**

```
Dimension: __________  Category: [tech / cost / reg / behavior / incumbent]
  Dated event in last 18-24 months? [Y / N]        (N = 0 on this dimension)
  Independent public source found?  [Y / N]        (N = max 1 on this dimension)
  Directly enables the product?     [Y / N]        (N = voids the score)
  Score: [0 / 1 / 2]

Verdict:
  REAL INFLECTION  = at least one dimension at 2, all gates passed
  TOO EARLY        = catalyst real but enabler still absent (too-early check fails)
  NO CATALYST      = no dimension reaches a dated anchor, or connection test fails

Load-bearing dimension: __________
One-line rationale: __________
```

*Score only the dimensions the founder actually claims. Any gate failure (connection or too-early) caps the verdict at "no catalyst" or "too early" regardless of points.*

The verdict is a routing decision, not a score for its own sake. Real inflection means the timing claim earns the meeting. Too early means the enabler is watch-list material: track the missing condition and re-check when it moves. No catalyst is a timing decline, though the deal may still merit attention on other grounds.

#### Verdict from catalyst strength and enabler presence

Horizontal axis runs from Catalyst weak or timeless to Catalyst dated and real. Vertical axis runs from Enabler absent to Enabler present now.

| Quadrant | What it means |
| --- | --- |
| No catalyst | Decline on timing; revisit only if a real change appears |
| Too early | Watch-list; track the missing enabler and re-check when it moves |
| No catalyst | Decline; a present enabler with no real change is not a why-now |
| Real inflection | Timing earns the meeting; name the load-bearing dimension |

*Two variables decide the timing verdict: whether the catalyst is dated and real, and whether the enabling condition is present now.*

## How this goes wrong

The scoring fails in predictable ways, and most failures are false positives that look comprehensive or current. Learn the seven and you will catch the manufactured urgency the rubric is built to expose.

- **Kitchen-sink why-now.** Five equal-weight forces read as thorough but signal hedging. Test: can the founder name one catalyst in a sentence? If not, no single force is load-bearing.
- **Timeless trend dressed as urgency.** "Growing awareness of X" that has been true for a decade. Test: apply "hasn't this always been true?" and demand a dated event.
- **Borrowed macro narrative.** "AI costs dropped 90%" on a product that does not use AI. Test: draw the causal line from catalyst to product, and reject on any disconnect.
- **Confusing why-now with why-us.** Product launches and team milestones presented as timing. Test: is the driver external? Internal progress is disqualified.
- **Strong demand, wrong economics (the Webvan trap).** Customers love it but unit economics fail. The false positive is high satisfaction scores. Test: is the enabling cost-curve actually present, or is spend front-loaded on hope?
- **Real inflection, over-crowded (the index trap).** A genuine catalyst plus thousands of founders already in-category is too-late, not too-early. Test: read founder-supply density against how far up the S-curve the market already is.
- **Number laundering.** Citing Gross's 42% and CB Insights' 42% as one finding. Test: they are separate studies, success variance versus failure cause. Keep them apart.

> **Watch out:** Satisfaction is not timing
>
> A pilot with delighted users can still be too early. General Magic's few thousand devices went mostly to family and friends, and Webvan's customers liked the service before the economics sank it. High enthusiasm from a tiny base is a demand signal, not proof the enabler has arrived.

The most dangerous of these is the index trap, because it inverts your instinct. A crowded founder field feels like validation. But paired with a mature adoption curve, it means the window you are grading has already closed. The founder-supply counts are only useful when you also know where on the S-curve the market sits, which is why the too-early check and the supply reading belong together.

> **Note:** When the evidence is genuinely thin
>
> Some catalysts are too new to have public benchmarks yet, which is honest rather than disqualifying. When you cannot find an independent source, say so in the verdict and name the single datapoint that would confirm the inflection, then set a date to re-check it.

## Before you call the score done

Run this checklist before you record a verdict. It catches the shortcuts that produce a confident-looking score built on the founder's own narrative.

#### Verdict readiness check

- [ ] The catalyst is written as one external force with a date, in one sentence.
- [ ] The catalyst is assigned to exactly one of the five categories, not spread across four.
- [ ] A specific change is dated within the last 18-24 months, or a fresh event justifies an older one.
- [ ] At least one independent public source per catalyst, none of them the founder's deck.
- [ ] A causal line runs from catalyst to product; the connection test passed.
- [ ] A prior failed attempt was checked, and the now-present enabler is named.
- [ ] Founder-supply density was read against S-curve position, not in isolation.
- [ ] The verdict names its load-bearing dimension in one line.

## Keeping the rubric current

Timing standards decay faster than most diligence tools, so the maintenance job is to keep two things fresh: the adoption window and the crowding read. Because adoption is compressing, revisit your definition of "live" periodically. The 18-to-24-month window is a reasonable default now, but as more categories follow the AI curve toward 50% in roughly 3 years, you may need to tighten it for fast-moving sectors and relax it for slow physical ones like robotics.

The crowding read is the part that ages by the week. Founder supply in a category moves as an inflection plays out, and the same number that confirmed an opening last quarter can signal consolidation this one. Re-pull the supply density for any category you are actively grading rather than trusting a figure you cached. [Refolk](/) lets you re-run a supply query by market, skill, and geography in plain English, so refreshing the S-curve read is a query, not a research project. When a category's founder count climbs sharply while adoption is still early, the window is open; when the count plateaus and adoption has crossed into the exponential middle, start scoring new entrants as too-late by default.

Keep the two 42% figures labeled wherever you cite them, keep the enabler as your unit of judgement rather than the market, and the rubric will keep separating real inflections from urgency for as long as timing keeps deciding outcomes.

## Frequently asked questions

### How do I evaluate a why-now slide as an investor?

Treat it as a claim to falsify, not a story to enjoy. Extract the single external force, classify it into one of five catalyst categories, then apply three tests: hasn't this always been true, does the catalyst directly enable this product, and did something dateable change in the last 18 to 24 months. Verify each with a public source that is not the deck. A claim that survives all three earns a real-inflection verdict.

### What is the difference between a too-early deal and a no-catalyst deal?

Too-early means the catalyst is real but the enabling condition is still missing, so demand or unit economics fail in practice. General Magic and Webvan are the archetypes: the idea was right, the enabler absent. No-catalyst means there is no dateable external change at all, just a timeless trend dressed as urgency. The fix differs: too-early deals become watch-list items, no-catalyst deals are declines.

### How recent does a market inflection have to be to count?

Roughly the last 18 to 24 months. Practitioner framing prescribes a before-and-after: eighteen months ago this was impossible, today it is inevitable. Because adoption is compressing, with AI projected to reach 50% in about 3 years versus the smartphone's 5, catalysts go stale faster now, so treat anything older than two years as needing a fresh dated event to still qualify as live.

### Can a crowded founder field prove the timing is right?

It cuts both ways. In Refolk's index, 3,316 US founders list Generative AI, 4.37x the UK's 759, which can confirm an inflection or warn that the market is already consolidating. Read the density against S-curve position: high founder supply early in the curve confirms the window, but high supply plus a mature curve signals too-late rather than right-time.

### Is the 42% timing statistic reliable?

Bill Gross's Idealab analysis of about 200 companies found timing accounted for 42% of the difference between success and failure. It is directional, not a scoring input. Do not conflate it with CB Insights' separate finding that no market need causes 42% of failures. Those are two different studies that share a number by coincidence, one about success variance and one about failure cause.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/why-now-score-market-timing*
