Scoring an AI Startup's Defensibility From Public Evidence
You will score any AI startup's moat across six named dimensions from public evidence and reach invest, dig deeper, or pass with a defensible per-dimension reason.
You are looking at an AI startup and have to decide whether it holds a durable moat or is a feature a foundation-model provider will absorb on its next release. This guide is for early-stage investors, platform and talent partners, and angels who need to make that call from public evidence, fast, and defend it. It gives you six named moat dimensions, a proving signal for each, financial thresholds, and a weighting that resolves into invest, dig deeper, or pass with a reason attached to every dimension.
Two things this framework does that a code teardown cannot. It scores the non-code moats that actually decide survival, so it can rate even a thin-code product as defensible when the evidence supports it. And it treats "will a lab copy this" as a checkable question, not a gut feeling.
Why code diligence is the wrong test
The decisive asset in an AI startup is no longer the code. Because model inference prices fell more than 280-fold between November 2022 and October 2024, the model itself is a commodity input, and a repos teardown that confirms proprietary code has answered the least important question. The moat has moved from code to contract.
The load-bearing evidence sits in the customer relationship, not the repository. The single provision that determines whether a company keeps the asset its product generates is the derived-data and model-improvement clause, and enterprise buyers now strike training rights as a default redline. A scorer who reads only the repo misses the exact place the moat lives or dies.
This is why a "wrapper" is not automatically doomed. A wrapper with only AI access plus a basic UI is vulnerable. A wrapper with real workflow integration, a data flywheel, or owned distribution is fine. The pejorative use of the word assumes no other moats are present, and that assumption is where most quick dismissals go wrong. The job of this framework is to check the assumption rather than inherit it.
The six moat dimensions and their proving signals
Independent investor sources converge on five durable moats: a proprietary data flywheel, deep workflow integration, owned distribution, brand and trust, and genuine network effects. None of these can be replicated by a lab shipping a feature update. Wrapper-focused sources add a sixth, regulatory and compliance depth, which matters most in high-stakes verticals. A dimension is only worth scoring if it has a public proving signal, so each one below carries the signal that proves it and what it looks like when it lies.
| Dimension | Proving signal (public) | When it lies |
|---|---|---|
| Data flywheel | Changelog or usage page showing usage measurably improves output | A static dataset that was merely expensive to collect |
| Workflow integration | Deep embed in customer workflow, marketplace or integration listings | A demo that a competitor's positioning copies in a weekend |
| Distribution / vertical | Named design partners, customer logos, niche ownership | Broad logos with no vertical concentration |
| Brand and trust | High-stakes use cases, references, trust center | Trust badge with no underlying report |
| Network effects | Each user makes the product more valuable to others | Multi-tenant hosting mistaken for a network |
| Regulatory depth | SOC 2 Type II, HIPAA, ISO in a neutral registry | Claimed compliance, vendor-controlled badge only |
The data flywheel deserves the tightest definition because founders overclaim it most. A real data moat is living and compounding: usage generates data that measurably improves the product for the next user, which attracts more usage. The flywheel is the moat, not the dataset. Score inference quality and compounding, not raw volume, because incremental data has diminishing returns and a well-funded competitor can replicate most datasets.
Order of attack matters less than coverage. Flywheel-led investors start with data; wrapper-led sources start with the substitution test. Both agree distribution and vertical ownership are the durable tiebreaker, so that is where the weighting lands regardless of where you begin.
The substitution test: wrapper or real
The fastest disqualifier is a ten-minute test. Paste the product's apparent core prompt into a frontier model and estimate what fraction of its output you reproduce. If a technical user gets 80 percent of the product's output that way, the code classifies as a wrapper and the burden of proof shifts entirely onto the non-code moats.
The test has a well-known false negative. A demo can look unique when the moat is only prompt engineering: you spend six months building prompt optimization and a competitor copies your positioning in a weekend. That is the commodity trap. So the substitution test never stands alone. A high reproduction score does not end the analysis; it tells you the product must earn its defensibility from distribution, workflow lock-in, data compounding, or regulatory depth instead of from the model.
A wrapper with real distribution outscores a deep-tech product with no channel.
Run it early because it is cheap and it reframes everything after it. If reproduction is low, you still verify the other moats, because a defensible-looking model is not the same as a defensible business.
The unit economics floor
Unit economics are the second fast filter because they are public-ish, comparable, and hard to spin once you adjust them correctly. A defensible AI startup clears the thresholds below; a fragile one sits on the wrong side of most of them. Treat these as a floor, not a target.
| Metric | Healthy / defensible | Fragile |
|---|---|---|
| Software gross margin | 80%+ (86%+ top-tier) | 55-65% (AI-company median) |
| LTV:CAC | 3:1 floor, 4:1+ elite | below 3:1 |
| Burn multiple | under 1x to 1.2x | over 3x |
| NRR (AI apps) | above 108% | below 108% |
Read the gross-margin row carefully. The 80 percent target comes from full-year data across 342 software companies, but AI-native reality is structurally lower: CRV puts private SaaS median gross margin at 77 percent against 55 to 65 percent for AI companies. Margin compression here is structural, not transient. Mature features add retrieval and self-critique calls faster than token prices fall, so unit economics will not grow into the threshold on their own. Treat sub-70 percent AI margins as the baseline.
Two moat signals hide inside these numbers. The burn multiple, net burn divided by net new ARR on David Sacks' scale, runs from under 1x amazing to over 3x bad, and AI-native companies are hitting 0.8x to 1.2x when they are efficient. And NRR above 108 percent is where a retention moat becomes demonstrated rather than asserted; below that line the moat is a claim.
Absorption risk: will a lab absorb this
The question "will a horizontal provider copy this feature" has a 20-year-old name. "Sherlocking" describes a platform absorbing a third party's core feature into its own free product, after Apple's Sherlock killed the startup Watson in 2002. The modern version is a foundation-model provider's product cadence walking into an app category.
The precedents below are drawn from blog and council sources, so treat them as directional claims and verify before you lean on any single figure. What they establish is a pattern and its shape, not a citable market statistic.
| Precedent | Signal | Status |
|---|---|---|
| Apple Sherlock kills Watson, 2002 | Origin of "Sherlocking" | Historical |
| Jasper valuation cut 20% after ChatGPT features | Incumbent feature launch | Secondary claim |
| 200+ GPT-wrapper startups cannibalized in 2024 | Provider product cadence | Secondary claim |
| AI-app down-rounds rose to 11.4% in 2024 | Investor reassessment | Secondary claim |
The test that resolves absorption risk is category ownership. A horizontal provider will never build a product specifically for dental-office scheduling. If a startup owns that vertical with dedicated automation, it has distribution into a niche the horizontal players cannot justify targeting. So the absorption check is really two questions: is this category on a provider's roadmap, and if it were copied, does at least one non-code moat survive? Jasper is the cautionary counterweight to over-dismissal here. It survived not because it had better AI but because it built distribution first.
Absorption risk versus stacked moats
Team replicability and the talent scarcity read
A founding team with genuine LLM depth is materially harder to clone than a general ML team, which makes team replicability a live moat input rather than a soft factor. The scarcity is measurable. In Refolk's index of professional profiles, the pool of US ML/AI engineers listing both machine learning and deep learning skills is 7,128, but engineers who also list large language model skills number just 75.
| Segment | Count | Derived ratio |
|---|---|---|
| US, ML/AI engineers with ML + deep learning | 7,128 | Baseline |
| Germany, same profile | 973 | US is about 7.3x Germany |
| US, ML/AI engineers with LLM skills | 75 | About 1 in 95 of the US ML pool |
That 1-in-95 figure is the point. LLM-specialist talent is roughly 1 percent of the US ML pool, so a team with real LLM depth sits in a scarce band that a well-funded competitor cannot staff overnight. Geography compounds it: US ML talent is about 7.3 times as deep as Germany's, which means a startup defending a European vertical faces a thinner local hiring pool. That raises its own switching cost and raises a US entrant's cost to localize.
To turn this into a score, you need to size the founding team's skill against the specific market they compete in. That is a sourcing question: who else could actually build this, and how many of them exist in the relevant geography.
When the replicable pool is thin, team is a moat input. When it is abundant, team is neutral and the decision has to be carried by the other five dimensions. Refolk collapses that pool-sizing from a manual search into a single plain-English query, which is what lets team replicability become a scored dimension rather than a hunch.
The scoring procedure
Run these seven steps in order. The first two are cheap and reframe everything after them; the last one is the partner's call. No published source gives a numeric weighted scorecard with fixed cutoffs, so the aggregation math below is the framework's own construction. What is sourced is the inputs and the weighting priority; what is constructed is the arithmetic that turns them into a verdict.
From cold look to invest, dig deeper, or pass
- Run the substitution testPaste the core prompt into a frontier model and estimate output reproduced. Above 80 percent flags wrapper risk and shifts the burden to the non-code moats.
- Score each moat dimensionFor all six dimensions, record the single public artifact that proves or fails it. Every dimension gets evidence or an explicit "none found."
- Pull unit economicsGather gross margin, LTV to CAC, burn multiple, NRR, and whether AI is priced separately. Each metric sits clearly above or below its threshold.
- Check absorption riskMap the core feature against provider changelogs and cadence. End with a yes or no on whether the category is on a horizontal provider's roadmap.
- Verify compliance depthFind the trust center, confirm SOC 2 Type II versus Type I and HIPAA or ISO scope, and request the actual report. Confirm type and audit window, not just a badge.
- Assess team replicabilitySize the founding team's skill against the local talent pool. Land on a scarcity read of abundant or scarce.
- Aggregate the verdictWeight distribution and vertical ownership heaviest, require at least two stacked moats, and resolve into invest, dig deeper, or pass with a per-dimension reason.
For the aggregation, a workable rule that respects the sourced priorities: score each dimension 0, 1, or 2. Distribution and vertical ownership count double. Invest requires at least two dimensions at 2 with distribution or vertical among them, unit economics clearing three of four thresholds, and no fatal absorption exposure. Pass is any product that fails the substitution test with zero surviving non-code moats, or one that has a copyable feature squarely on a provider's roadmap with nothing stacked behind it. Everything in between is dig deeper, which is a real verdict, not a dodge: it names the two or three artifacts you would need to see to move it either way.
How candidates fall out of the pipeline
- 100All AI startups reviewed
Full inbound
- 60Survive substitution + at least one non-code moat
Not a bare wrapper
- 35Clear three of four unit-economics thresholds
Economics support a moat
- 20No fatal absorption exposure
Category not on a lab's roadmap
- 10Two stacked moats incl. distribution/vertical
Reach an invest verdict
How this scoring goes wrong
This is the most valuable part of the framework, because a confident wrong score costs more than no score. Seven failure modes recur, each with a check.
- Substitution test false negative. A demo looks unique but the moat is prompt engineering that a competitor copies in a weekend. Check whether positioning is copyable over a weekend before you credit uniqueness.
- Static dataset mistaken for a flywheel. Founders call expensive-to-collect data a "moat." Check whether usage measurably improves output for the next user, which is the only signal that proves a living flywheel.
- Healthy gross margin hiding a money-losing account. A strong company-wide margin can hide an enterprise customer that loses money every month. Check contribution margin per account, not the blended number.
- LTV/CAC inflated by measurement error. Failing to margin-adjust overstates LTV by about 30 percent. Check that LTV is gross-margin-adjusted and CAC is fully loaded before you trust the ratio.
- Compliance claimed but unverifiable. A trust-center badge is vendor-controlled, and there is no public registry of issued SOC 2 reports because the AICPA maintains no verification database. Check Type II versus Type I and the audit window; a rushed short window signals thin evidence.
- "Wrapper" dismissed too fast. Thin code can still be defensible via distribution, the way Jasper survived. Check the non-code moats before you write pass.
- Data-moat overweighted. a16z's own point is that incremental data has diminishing returns and a well-funded competitor can replicate most datasets. Score inference quality and compounding, not raw volume.
What to verify before you file the verdict
Before you write invest, dig deeper, or pass, confirm the evidence is real and not inherited. This checklist is the last gate.
Before you commit the score
- Substitution test run, with a rough percent of output reproduced recorded.
- Every one of the six dimensions has a named public artifact or an explicit "none found."
- Gross margin, LTV:CAC, burn multiple, and NRR each placed above or below threshold, with LTV margin-adjusted and CAC fully loaded.
- Contribution margin checked at the account level, not just blended.
- Absorption verdict written: is the category on a horizontal provider's roadmap, yes or no.
- SOC 2 type and audit window confirmed from the actual report, not a badge.
- Team scarcity read landed as abundant or scarce for the specific geography.
- At least two moats stacked, including distribution or vertical, before any invest verdict.
Keeping the score current
Two inputs in this framework move and will invalidate an old score. Model inference pricing keeps falling, which keeps shifting value away from the model and toward the contract, so re-read the derived-data clause every time you revisit a company rather than trusting a prior read. And provider product cadence changes the absorption map continuously; a category that was safe last quarter can appear on a changelog the next.
Re-run the substitution test against the current frontier model, not the one you used last time, because a feature that survived paste-in a year ago may not survive today. Re-check the trust center and request a fresh report, since a certification with a lapsed audit window is a claim, not a moat. The framework is stable; the evidence under it is not, and the discipline is to re-pull the evidence rather than re-use the verdict.
title: Per-startup defensibility scorecard
note: Fill one per company. Distribution and vertical count double in the total.
Startup: __________ Reviewer: __________ Date: __________
Substitution test: __% output reproduced (>80% = wrapper risk)
Data flywheel (0/1/2): ____ Evidence: __________
Workflow integration (0/1/2): ____ Evidence: __________
Distribution / vertical (0/1/2, x2): ____ Evidence: __________
Brand and trust (0/1/2): ____ Evidence: __________
Network effects (0/1/2): ____ Evidence: __________
Regulatory depth (0/1/2): ____ SOC 2 type + window: __________
Unit economics: GM __% LTV:CAC __ Burn __x NRR __% AI priced separately? Y/N
Absorption: on a provider roadmap? Y/N Survives if copied? Y/N
Team scarcity: abundant / scarce Pool sized against: __________
Verdict: INVEST / DIG DEEPER / PASS
Per-dimension reason: __________
Questions practitioners ask
Will OpenAI kill this startup?
Ask whether the startup's core feature sits on a horizontal provider's roadmap and whether it owns a niche the provider will not prioritize. OpenAI will not build dental-office scheduling; a startup that owns that vertical with dedicated automation has distribution into a niche horizontal players cannot justify targeting. Map the core feature against provider changelogs and product cadence, then check whether at least one non-code moat, usually distribution, survives if the feature is copied.
How do I tell a wrapper from a real AI startup?
Run the substitution test: can a technical user reproduce 80 percent of the output by pasting the core prompt into a frontier model? If yes, the code is a wrapper. But a wrapper with real workflow integration, a data flywheel, or owned distribution is still defensible. Jasper survived not on better AI but on distribution built first. Score the non-code moats before you dismiss thin code.
What gross margin should an AI startup have?
Software gross margin of 80 percent or higher is the target and 86 percent is top-tier, but AI-native companies run structurally lower at 55 to 65 percent against a 77 percent SaaS median. Treat sub-70 percent AI margins as the baseline, not a fixable anomaly, because mature features add retrieval and self-critique calls faster than token prices drop. Check contribution margin per account so one money-losing enterprise customer does not hide inside a healthy blended number.
Is a data moat real or overrated?
A real data moat is living and compounding: usage generates data that measurably improves the product for the next user, which attracts more usage. The flywheel is the moat, not the dataset. Incremental data has diminishing returns and a well-funded competitor can replicate most static datasets, so score inference quality and compounding rather than raw volume. A dataset that was merely expensive to collect is not a flywheel.
Can I verify a startup's SOC 2 from public sources?
Only partially. A trust-center badge is vendor-controlled, and there is no public registry of issued SOC 2 reports because the AICPA maintains no verification database. Confirm the type publicly where you can, then request the actual report and note whether it is Type II or Type I and the audit window. A rushed short window signals thin evidence. Neutral registries list some HIPAA-compliant companies but do not replace the report itself.
How many moats does a startup need to be defensible?
At least two stacked. One moat is fragile; two becomes a real barrier. Weight distribution and vertical ownership most heavily because they have emerged as the most durable moats for early-stage companies and are the ones horizontal providers will not prioritize. A thin-code product that owns a niche can outscore a deep-tech product with no channel.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.