The AI Defensibility Read: Durable Moat, Wrapper, or Feature
You can score an AI target on five public moat dimensions and assign one of three verdicts, each backed by evidence you gathered from the outside.
Key takeaways
- A disclosed gross margin below 50% with a flat inference-to-revenue ratio is the cheapest wrapper tell there is, because ICONIQ's 23% inference figure does not meaningfully decline as AI companies scale.
- In Refolk's index the US holds 3,625 PyTorch ML and research engineers against 441 in Germany, so a European startup claiming custom models with no ML individual contributors is fighting a pool that barely exists.
- US strategist-class titles number 1,353 against a ~2.7:1 ratio of core ML engineers, and the single most common one is Head of AI, a leadership title that can masquerade as build depth.
- A static dataset is not a moat; apply the 12-month replication test, and if a well-funded rival could acquire or generate the data in a year it fails regardless of how expensive it was to collect.
- Patents and AI-hiring velocity separate washing from building better than any demo: the documented contrast was 14 patents plus a 340% rise in AI postings versus zero and zero.
- Fail two or more of the substitution, API-shutdown, feature-announcement, and data-accrual tests and the target is a wrapper, not a product.
Before the investment committee, you have to decide one thing about an AI startup: does it have a defensible edge, or is it a thin layer over someone else's model. This guide is for early-stage investors, platform and talent partners, and angels who need that verdict from the outside, before the founder hands over a single slide. It gives you five externally-observable dimensions, a way to score each one from public evidence, and three verdicts you can defend in the room.
The founder-facing tests you will find elsewhere assume you can paste the product's prompt into ChatGPT, interrogate the CTO, and open the data room. You usually cannot, not yet. So this read is built entirely on signals you can gather yourself: the team's ML depth from profiles, the model-stack clues in job posts and model cards, the repo and Hugging Face footprint, the margin economics, and the distribution evidence. A verdict built this way survives before diligence formally opens.
What the three verdicts mean
The read ends in one of three places: durable moat, wrapper with a path, or feature to pass. Each is a decision about defensibility, not about the quality of the product today.
A durable moat company owns something a well-funded rival cannot replicate inside a year: a compounding data flywheel, deep workflow lock-in, or distribution control, backed by in-house ML individual contributors. A wrapper with a path calls a foundation model through an API today but is accruing a real moat - law-firm relationships, ingested domain data, switching costs - that will outlast the model dependency. Harvey is the archetype: raised at a $3 billion valuation in 2025 on law-firm relationships and ingested case data while technically calling GPT-4 class models through an API. A feature to pass is a thin layer that a foundation lab could ship as a feature, that fails the substitution test, and that shows no compounding mechanism.
The five dimensions you can score from the outside
Five dimensions are observable without the data room: stack architecture, team ML depth, public footprint, margin economics, and the moat trio of data, workflow, and distribution. Each produces a score and a piece of cited public evidence.
The order matters. Architecture and team depth are the cheapest to read and the hardest to fake. Margin is the single cheapest wrapper tell when it is disclosed. The moat trio is the partner-level judgement that the earlier signals feed. Score them in that order so a clear fail on an early, high-confidence dimension can save you the later work.
The five dimensions, outermost first
- Stack architectureProprietary, fine-tuned, or pure API wrapper, read from model cards and dependencies
- Team ML depthRatio of ML individual contributors to strategist-class titles
- Public footprintModels, repos, patents, and AI-hiring velocity across eight quarters
- Margin economicsGross-margin band and inference-to-revenue ratio against published bands
- Moat trioProprietary data, workflow lock-in, distribution, scored two-of-three
Dimension one: stack architecture
Place the target in a three-tier taxonomy. Proprietary models, custom-trained on proprietary data, carry low AI-washing risk. Fine-tuned foundation models carry medium risk, where value depends on fine-tuning depth. API wrappers, which call third-party APIs with minimal customisation, carry high risk. Model cards and GitHub are your primary footprint sources, and dependency manifests and SDK references in public repos often reveal the underlying provider directly.
Dimension two: team ML depth
This is countable. Diligence frameworks explicitly ask for ML engineers versus "AI strategists" on the org chart, with a kill criterion of no in-house ML talent and outsourced model development. Count the individual contributors - ML engineers, research engineers, MLOps engineers - and divide by the leadership and advisory titles. The absence of foundational roles like data engineers, MLOps engineers, and AI governance leads is a documented red flag.
The talent-depth ratio and why scarcity makes it load-bearing
A strategist-heavy team signals narrative over capability. The scarcity of real ML builders is what makes this ratio decisive: in Refolk's index of professional profiles, the US holds far more ML individual contributors than any other market, so an absence that looks ordinary is actually structurally suspicious.
The table below is the index-wide pool, not any single company. Use it two ways: as scarcity context, and as a template for the per-target count you build in step three.
| Segment | Count | Derived |
|---|---|---|
| PyTorch ML/Research Eng, US | 3,625 | baseline |
| PyTorch ML/Research Eng, Germany | 441 | 8.2x smaller than US |
| "AI Strategist"-class titles, US | 1,353 | US ML:strategist ≈ 2.7:1 |
The read on geography is direct. With only 441 PyTorch ML engineers in Germany against 3,625 in the US, a European startup claiming custom-trained models while showing no ML individual contributors is fighting a pool that barely exists. The absence of builders is more suspicious outside US hubs, not less, because there is almost nowhere for those builders to have come from. And the strategist-class pool is led by "Head of AI" - a leadership title, not a builder's title. A company whose "AI team" is strategist-heavy is telling you where its weight sits.
Building this per-target count by hand means paging through profiles and job posts and classifying each title. Refolk runs the same query in plain English and returns the named people, so the numerator and denominator in step three are a search rather than an afternoon.
Reading the margin economics
The gross-margin band is the single cheapest external wrapper tell, because inference sits in cost of goods sold and does not fall meaningfully with scale. A disclosed sub-50% margin with a flat inference-to-revenue ratio implies a model dependency the founder cannot explain away.
The published bands are consistent. a16z's 2020 "The New Business of AI" by Casado and Bornstein pegged AI gross margins at 50 to 60%, well below the 60 to 80% or higher SaaS benchmark. ICONIQ's January 2026 reporting puts AI gross margins averaging 52%, up from 41% in 2024, with inference averaging 23% of revenue at scaling-stage AI B2B firms.
| Business type | Gross margin band |
|---|---|
| Legacy SaaS (pre-AI) | 75-90% |
| Mature AI-first SaaS | ~50-60% (ICONIQ avg 52%, 2026) |
| Pure application-layer AI (2026 proj.) | ~45% |
| Fast-scaling early-stage AI | ~25% |
The second table turns the band into a fragility test. The diligence threshold is a gross margin below 50% for software or below 30% for AI-infra with no path to improvement.
| Metric | Value |
|---|---|
| Traditional SaaS COGS | 10-25% |
| AI inference at scaling stage | 23% |
| AI-first total COGS | 40-50%+ |
| Diligence fragility threshold (margin) | <50% software / <30% infra |
The trajectory point is why "margin" is a signal and not a verdict. What makes it damning is a sub-50% margin and a flat inference ratio and no routing strategy to bring cost down. One without the others is noise.
The four public tests
Four tests are runnable from the outside with no cooperation from the founder. They convert "is this a wrapper" into four yes-or-no answers.
The substitution test: can a technical user get 80% of the output by pasting the core prompt into ChatGPT or Claude directly? If yes, you are looking at a wrapper. The API shutdown test: if the primary model provider revoked the API key tomorrow, does the product still deliver value? If no, it is a reskinned API. The feature announcement test: if OpenAI, Anthropic, or Google shipped the same feature, how many customers would stay? If most would leave, switching costs are near zero. The data test: does the product get better the more a specific customer uses it, delivering more value on day 90 than day 1 because of data only the company holds?
The AI defensibility read, start to finish
- Scope the claimRead the company's own messaging and separate architecture claims from marketing. State whether they claim proprietary models, fine-tuning, or just "AI-powered", and whether AI appears in product docs or only in press.
- Classify the stackPlace the target in the proprietary / fine-tuned / wrapper taxonomy using model cards, GitHub, and SDK or dependency evidence. Record a tier with supporting URLs.
- Read team ML depthCount ML and research engineers against "AI strategist" and "Head of AI" titles from profiles and job posts. Record a numerator, denominator, and ratio for the specific company.
- Check the footprintSearch Hugging Face for the org and models and GitHub for repos, then check patents and AI-hiring velocity across eight quarters. Record counts of models, repos, patents, and AI job posts.
- Run the four testsApply the substitution, API-shutdown, feature-announcement, and data-accrual tests against the public product. Record a clean pass or fail on each.
- Score the margin economicsBenchmark disclosed or inferred gross margin and inference-to-revenue against the bands. Record a band placement and raise or clear the fragility flag.
- Weigh the three moatsScore proprietary data, workflow lock-in, and distribution. Record a two-of-three verdict, noting that sources disagree on which ranks first.
- Assign the verdictCombine into durable moat, wrapper with a path, or feature to pass, each carrying its cited public evidence.
A documented pass/fail threshold exists for this battery: fail two or more of the tests and the target is a wrapper. Treat one fail as a flag to investigate and two as a verdict.
Weighing the moat trio: data, workflow, distribution
The partner-level call scores three moats and asks whether at least two hold. Proprietary data that compounds, workflow lock-in customers would miss if it disappeared, and distribution control are the three dimensions cited most consistently across frameworks.
Be honest about where the sources disagree. There is no single named "two of three" rule established publicly; it is a synthesis. Equidam lists five persisting moats - workflow capture, proprietary context and feedback loops, multi-model orchestration, distribution control, and compliance. Other frameworks count three or four, and they disagree on ordering: some weight distribution most heavily now, while others still weight proprietary data first. Score all three, name which one carries the company, and flag the disagreement rather than pretending it away.
The data moat is where most false positives hide, so hold it to a strict standard. Simply possessing a dataset that was expensive to collect is not durable, because it is static and can eventually be matched. A real data moat is living and compounding, where usage generates data that measurably improves the product for the next user. The flywheel is the moat, not the dataset.
The flywheel is the moat, not the dataset, and the 12-month clock is how you tell them apart.
Scale AI is the clean positive case: valued at $29 billion after Meta's 2025 investment, it built defensibility on proprietary data-labeling infrastructure rather than a model. The data accrues with use, which is what a static dataset never does.
How this read goes wrong
The most valuable part of a standard is the list of ways it lies to you. Each failure mode below has a specific check that defuses it.
Reading the signal against the risk of a false read
The specific traps, and the check for each:
- Patent and job-post signal lag. Absence of patents can be timing, not absence of AI. Filings and postings are incomplete, skew toward larger firms, and carry known lags. Check across eight quarters, not one, before you read an absence as a fail.
- Model-card theater. A polished card proves documentation, not ownership. Over 90% of cards include architecture, evaluation metrics, and compute requirements, yet those evaluation scores are often created by the author rather than the community, and open-source weights get re-uploaded as "ours". Check commit history and base-model lineage.
- Substitution-test false negative. A target can pass because the UI hides a thin prompt. Reproduce the core output in a raw model yourself rather than trusting the demo.
- Scripted demos. Prepared demos are easy to stage. Request a live demonstration on data not previously processed, and treat resistance as a signal in its own right.
- Static "data moat". A large one-time dataset reads as a moat and is not one. Check for compounding usage loops, not volume.
- Margin mirage. A 50% margin can be fine if inference prices are falling. Condemning a company whose costs drop 80% next year is a false positive; read the routing strategy and the trajectory.
- Headcount ratio gaming. "Head of AI" is the single most common strategist-class title in the US index - a senior non-builder title can masquerade as ML depth. Check for individual contributors: ML engineers, research engineers, MLOps.
- Credential inflation. AI roles or partnerships get announced without the infrastructure or follow-through to support them. Cross-check job posts for the actual required skills, not the headline.
Google's Darren Mowry framed the underlying tell well: wrappers putting "very thin intellectual property around Gemini or GPT-5" signal a startup not distinguishing itself. The failure modes above are the ways that thin IP disguises itself as something more.
The pre-IC checklist
Run this before you write the verdict. Every item is something you can verify from public evidence.
Before you call the read done
- You have classified the stack as proprietary, fine-tuned, or wrapper with supporting URLs
- You have a company-specific ML-IC to strategist-class ratio, not just the index baseline
- You have counts of Hugging Face models, GitHub repos, AI patents, and AI job posts across eight quarters
- You have a pass or fail recorded on all four tests: substitution, API shutdown, feature announcement, data accrual
- You have placed gross margin in a band and checked the inference-to-revenue trajectory, not just the snapshot
- You have applied the 12-month replication test to any data-moat claim
- You have scored all three moats and named which one carries the company
- Each dimension score links to the specific public evidence that justifies it
Verdict: [DURABLE MOAT | WRAPPER WITH A PATH | FEATURE TO PASS] Stack tier: [proprietary / fine-tuned / wrapper] - evidence: [model card / repo / dependency URL] ML depth: [N] ML ICs : [M] strategist titles = [ratio] - evidence: [profile/job-post source] Footprint: [X] HF models, [Y] repos, [Z] patents, [hiring velocity] over 8 quarters Four tests: substitution [pass/fail], shutdown [pass/fail], feature [pass/fail], data [pass/fail] Margin: [band] with inference ratio [trend]; fragility flag [raised/cleared] Carrying moat: [data / workflow / distribution] - 12-month replication: [yes = not a moat / no = holds]
Fill each bracket from your own evidence; delete the two verdicts you did not choose.
Keeping the read current
Re-run the footprint and margin dimensions on a cadence, because both move. Hiring velocity, new patent filings, and new Hugging Face models change the footprint read quarter to quarter, and the whole point of the eight-quarter window is that a single snapshot misleads. Margin economics move faster still: inference prices fall, routing strategies change, and a company you flagged on a 50% margin can cross into durable territory as its cost base drops.
The dimensions that move slowly are team ML depth and the data flywheel, which is exactly why they anchor the verdict. A compounding data moat and a credible bench of ML individual contributors take years to build and years to replicate, so a positive read on those two ages well. When you revisit a deal, re-read the fast-moving dimensions in full and spot-check the slow ones for a step change - a founding ML hire leaving, or an exclusive data relationship ending. The read is a living document, not a one-time verdict.
Questions practitioners ask
How can I tell if a startup is an AI wrapper without access to the data room?
Run the four public tests. Ask whether a technical user gets 80% of the output by pasting the core prompt into ChatGPT or Claude, whether the product still works if the model provider revoked the API key tomorrow, how many customers would stay if a foundation lab shipped the same feature, and whether the product measurably improves per customer with use. Fail two or more and it is a wrapper. Cross-check with the gross-margin band and the ML-engineer-to-strategist ratio.
What gross margin signals an AI wrapper versus a real product?
The diligence fragility threshold is a gross margin below 50% for software or below 30% for AI-infra with no path to improvement. Mature AI-first SaaS runs roughly 50 to 60% (ICONIQ averaged 52% in 2026), against 75 to 90% for legacy SaaS. Margin alone does not condemn a company, because falling inference prices can lift a 50% margin toward 90%, so always read the trajectory and model-routing strategy, not the single number.
Is proprietary data always a defensible moat?
No. Simply possessing a dataset that was expensive to collect is not durable because it is static and can eventually be matched. A real data moat is living and compounding, where usage generates data that improves the product for the next user. Apply the 12-month replication test: if a well-funded competitor could acquire or generate the dataset within a year, it is not a moat. Scale AI built real defensibility on compounding labeling infrastructure, not a one-time dataset.
Does a Hugging Face model card prove a startup owns its model?
No. A polished model card proves documentation, not ownership. Over 90% of cards include architecture, evaluation metrics, and compute requirements, but the evaluation scores are often created by the author rather than the community. Open-source weights can be re-uploaded as 'ours'. Check commit history and base-model lineage before you trust a card, and treat it as a footprint signal rather than proof of in-house training.
How do I read in-house ML depth from public profiles?
Count individual-contributor ML roles - ML engineers, research engineers, MLOps - against leadership and advisory titles like Head of AI or AI Strategist, then compute the ratio for the specific company. A strategist-heavy team signals narrative over capability. The kill criterion is no in-house ML talent with outsourced model development. In Refolk's index the US baseline is roughly 2.7 core ML engineers per strategist-class title, which gives you a reference for whether a target looks top-heavy.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.