# The Competitor AI-Maturity Read: Shipping, Piloting, or Marketing-Only

*You will score any named competitor as Shipping, Piloting, or Marketing-only, citing the specific public artifacts that separate real production use from AI washing.*

- Canonical URL: https://www.refolk.ai/guides/competitor-ai-maturity-read
- Pillar: Market and talent intelligence
- Format: Framework
- Published: 2026-10-11
- Last reviewed: 2026-10-11
- Reading time: 17 min

Deciding how far a competitor has actually gone with AI is a judgement call most teams get wrong because the public record is a sea of tool round-ups and vendor landing pages that treat a press release as a shipped capability. This guide is for strategy and research teams, talent-intelligence analysts, and operators who need to score a named competitor on a defensible AI-maturity verdict. It gives you a weighted rubric across four signal classes that produces the same verdict no matter who runs it, and it tells you which artifacts separate real production use from marketing claims.

The verdict has three values: **Marketing-only**, **Piloting**, or **Shipping**. The rest of this document is how to assign one and cite the artifact that earned it.

## What the verdict means and why three bands

The three bands describe how far a capability has traveled from claim to production. Marketing-only means the claim exists but no operational substance does. Piloting means the capability runs for some users, often behind a beta label, with the plumbing half-built. Shipping means a generally available feature serves real users on real infrastructure, corroborated by hiring and engineering traces.

The reason to score bands rather than yes or no is that most of the gap between competitors lives in the middle. Adoption is broad but shallow by base rate. In Federal Reserve analysis, the AI share of job postings is only 1.6 percent across all firms but 8.6 percent among firms that have ever posted an AI role. The useful comparison is a competitor against AI-active peers, not against the whole economy.

> **Rule:** Cite an artifact for every score
>
> A verdict without a named public artifact is an opinion, not a read. Each band assignment must point to a specific filing, release label, model card, merged PR, or hire count that a second analyst could pull independently.

The single most diagnostic test underneath all three bands is the simplest one. If you turned off the AI component of a given capability, what would be left? Everything that survives is cosmetic. A chatbot bolted onto unchanged infrastructure as a presentation layer survives the test with its real product intact, which tells you the AI is a veneer. A fraud model whose removal breaks the product does not survive it, which tells you the AI is load-bearing.

## The four signal classes and what each one proves

Four classes of public evidence carry the verdict, and no single one is sufficient. Practitioners triangulate across all four because each proves something different and each lies in a characteristic way.

| Signal class | What it proves | How it lies |
|---|---|---|
| Hiring mix | Intent and depth of the AI team | Open roles with generic requirements and no shipped surface |
| Engineering output | Real engineering capability | Contribution-graph volume inflated by AI-generated commits |
| Product surface | A capability users can reach | A polished demo on curated data, or an invite-only beta |
| Infrastructure and filings | Governance and investor-grade honesty | Dense AI language in investor decks with no operational detail |

Hiring data proves intent and depth but not shipped product. Open repositories prove engineering capability but are now noisy. Case-study pages prove a named customer exists but not scale. Regulatory filings are the most falsifiable class because the SEC has treated false AI claims as securities fraud under existing anti-fraud rules, so investor-facing claims are the ones a company has the most incentive to keep honest.

#### Signal classes from softest to hardest

1. **Marketing copy** - Homepage and press releases, zero liability, discard on conflict
2. **Hiring mix** - Open roles and title composition, proves intent and depth
3. **Product surface** - Release labels and changelogs, proves reachable capability
4. **Engineering traces** - Merged PRs and model cards, proves real output
5. **Filings and governance** - Investor claims under fraud liability, hardest to fake

*Weight the harder, more falsifiable layers over the softer ones when they conflict.*

When two classes disagree, the harder layer wins. A homepage that promises an AI platform while filings describe only a pilot tells you to score Piloting and note the gap. The gap between what a firm tells investors and what it tells visitors is itself a finding worth recording.

## Reading the hiring mix for depth, not headcount

Depth hides in title composition, not in headcount. A firm claiming AI transformation with a dozen generically titled AI Engineers but no research scientists and no MLOps engineers is almost certainly running a single bolt-on team, not an AI operation.

Refolk's index gives you the bands to measure this against. The applied-engineer base is large, the research band is a meaningful minority, and the MLOps band is scarce enough to be the sharpest tell of all.

| Role band | Count | Share of applied engineers |
|---|---|---|
| Applied ML/AI/MLOps engineers | 15,269 | 100% |
| ML research / applied scientists | 3,535 | 23% |
| MLOps / ML-platform / ML-infra | 188 | 1.2% |

Counts are from Refolk's index of professional profiles; shares are derived against the applied-engineer base. The interpretation lever is direct. A higher research share signals depth of modeling capability. A near-zero MLOps share signals a bolt-on with no production plumbing, because models do not reach users without deployment infrastructure that someone builds and runs.

**81x - Applied ML engineers per MLOps or ML-infra engineer in Refolk's index**

MLOps is the scarcest band, which makes its presence the most load-bearing proof that AI is actually in production.

That scarcity is the core insight of the whole read. In Refolk's index there are 4.3 applied engineers per research scientist but 81 per MLOps engineer. The firm that has hired into the 81x band has paid for the hard, rare skill that production demands. The firm that has not is telling you, through its own hiring, that its AI has not left the lab.

Market context matters when you benchmark. Supply is concentrated, so a UK competitor with three MLOps hires is relatively deeper than a US competitor with the same three.

| Market | Applied ML/AI/MLOps engineers | Ratio vs UK |
|---|---|---|
| United States | 15,269 | 6.2x |
| United Kingdom | 2,462 | 1.0x |

Counts are from Refolk's index; the ratio is derived (15,269 divided by 2,462). Reading hiring without a benchmark produces the wrong verdict; a modest absolute count can be a strong relative one in a thin market.

This is exactly the measurement that is slow to run by hand and fast with the right index. You do not want to hand-count titles across a competitor's team one profile at a time.

I ran this search: `MLOps and ML platform engineers at Databricks who joined in the last 18 months` - [see the full result list](https://www.refolk.ai/s/vdsmgjzbqr).

*Returns the named people in the scarcest, most load-bearing band, so you can see whether production plumbing actually exists rather than inferring it from marketing.*

[Refolk](/) lets you quantify the title mix directly, so the research-to-applied and MLOps-to-applied ratios become numbers you cite rather than impressions you form. Run the same query shape for the research band and you have both levers in minutes.

## Verifying the product surface

Release-stage labels are the clearest public signal that a capability is live, and the gap between beta and generally available is the gap between Piloting and Shipping. A feature behind login with a GA label that has passed major compliance, security, and performance tests is a shipped capability. An invite-only beta is a pilot.

The distinctions are public and precise. Beta makes a feature available for real use while its interface and behavior are being finalized, and beta features are functional end to end, not partial implementations. That matters in both directions: a working beta is not vaporware, but it is also not GA, because beta features are still evolving, with pricing not finalized and no guarantee they progress to general availability.

#### Product-surface test

1. **Find it** - Log in and locate the actual feature, not the marketing page
2. **Read the label** - Record beta, GA, or invite-only from the UI or docs
3. **Check cadence** - Confirm the changelog shows sustained releases, not one announcement
4. **Name the model** - Look in docs or network calls for a model or provider reference
5. **Run messy input** - Test on uncurated data to see if the demo survives reality

*Walk every claimed capability through these checks before you assign a band.*

Three checks confirm a live capability. First, is the feature behind login and GA, not invite-only beta? Second, does the changelog show sustained cadence rather than a single launch post that went quiet? Third, do the docs or network calls name a model or provider? GPT appears in 7.8 percent of AI-engineer job postings, and the same provider references surface in docs and network traffic, which both corroborates real integration and tells you whether the AI is a wrapper around a third-party API.

That last point is a trap worth stating plainly. A firm claiming proprietary AI while wrapping commercially available LLM APIs from a third party, and marketing the result as proprietary technology, is washing. The network call that names the provider settles it.

## The engineering-trace read after the 2026 inversion

Engineering traces corroborate real AI output, but the signal inverted and you must read it the new way. Contribution graphs, star counts, and follower counts stopped separating real engineers from prompt-typers once AI tooling could merge a pull request on its own.

The scale of the change is the reason. GitHub reports that 51 percent of committed code in early 2026 was AI-generated or AI-assisted, with some firms far higher. Volume metrics now produce false positives, so the signal moved from quantity of commits to merged pull requests in standard libraries.

**51% - Committed code that was AI-generated or AI-assisted in early 2026**

This is why green-square volume no longer proves skill; weight merged PRs to high-bar repositories instead.

What still carries weight is a merged pull request to a high-bar repository, because a maintainer reviewed and accepted it. Libraries such as vllm, pytorch, and huggingface transformers set the bar. For scale, vllm-project/vllm has more than 2,000 lifetime contributors, and a recent release shipped 411 commits from 212 contributors, 61 of them new. A competitor whose engineers appear in that contributor graph has demonstrated real AI engineering output that cannot be faked with trivial commits.

Model cards are the other durable trace. The model card is the de facto transparency standard, from the 2019 Mitchell et al. paper, and it is the standard document for AI model transparency. A firm that publishes version-pinned model cards is operating at a maturity that marketing-only firms do not reach.

> A merged pull request to vllm is reviewed by a maintainer; a green contribution graph is reviewed by no one.

## Governance, filings, and the honesty gap

Filings are honest where marketing is not, because companies face securities-fraud liability for investor-facing AI claims but not for homepage copy. The SEC did not need new AI-specific rules to bring its cases; it used existing anti-fraud provisions, the Marketing Rule, and fiduciary duty principles, and it can pursue AI-washing on negligent misrepresentation alone for registration statements.

The enforcement record tells you the claims are real risks, not theoretical. The SEC brought its first AI-washing actions in March 2024 against Delphia and Global Predictions, with Delphia agreeing to pay a 225,000 dollar penalty and Global Predictions 175,000 dollars. In January 2025 it charged Presto Automation, the first AI-washing action against a public company, alleging it overstated its drive-through product. In June 2024 it charged the founder of Joonko with defrauding investors of at least 21 million dollars over AI candidate-matching claims. The SEC created a Cyber and Emerging Technologies Unit in February 2025 focused on AI-related misconduct, and 12 AI-related securities class actions were filed in the first half of 2025, with companies taking an average 11.4 percent stock hit on announcement day.

> **Note:** The investor-versus-visitor gap is a finding
>
> Because homepage copy carries no fraud liability and filings do, a firm will often promise more to visitors than to investors. When the gap is large, score the lower claim and record the gap itself as evidence of washing.

For governance scale, look for documented model governance and responsible-AI references that match the size of the claimed operation. ISO 42001:2023 requires documented model governance, and the EU AI Act Annex IV requires technical documentation for high-risk systems from August 2, 2026. A firm claiming an AI platform with no governance footprint at all is claiming more than it runs.

## The scoring procedure

Run the seven steps in order, then combine the four signal classes into one verdict. Each step has an owner and a rough time budget so you can staff it, and each ends with a defined done-state so you know when to move on.

#### From claim inventory to verdict

1. **Scope the claim** - Collect every public AI claim from homepage, press releases, filings, and earnings calls into a dated list of specific capabilities, with filings flagged because they carry the most legal weight.
2. **Apply the turn-it-off test** - For each capability, ask what remains if the AI is removed, and tag the claim core or cosmetic.
3. **Read the hiring mix** - Count AI and ML roles and their composition across research scientists, MLOps, and applied engineers, producing ratios rather than a single headcount.
4. **Check the product surface** - Log in, find the feature, and record its release label, changelog cadence, and any model or provider reference, marking each capability vaporware, beta, or GA.
5. **Inspect engineering traces** - Look for named repos, model cards, and merged PRs to high-bar libraries, weighting merged PRs over contribution-graph volume, until each capability is corroborated or flagged absent.
6. **Check governance and filings** - Look for model governance and responsible-AI references at expected scale and compare investor claims against operational detail.
7. **Score and weight** - Combine hiring, engineering output, product surface, and infrastructure into one verdict of Marketing-only, Piloting, or Shipping, with each score citing a specific artifact.

Sources disagree on order. Washing-focused writers lead with messaging; talent-intelligence writers lead with hiring. In practice, start with the claim inventory because it defines what you are scoring, then let the hardest evidence class present decide ties.

To combine the four classes, place a competitor on two axes: how much operational substance exists and how loud the claim is. The quadrant tells you the band and what to do.

#### Substance against claim volume

Horizontal axis runs from Quiet claims to Loud claims. Vertical axis runs from Thin substance to Deep substance.

| Quadrant | What it means |
| --- | --- |
| Understated builder | Deep substance, quiet claims: score Shipping, watch as a real threat |
| Genuine leader | Deep substance, loud claims: score Shipping, the claim is earned |
| Not yet started | Thin substance, quiet claims: score Marketing-only or absent, low priority |
| AI washing | Thin substance, loud claims: score Marketing-only, record the gap |

*The dangerous quadrant is loud claims with thin substance, which is the washing signature.*

## How this read goes wrong

The failure modes are where most verdicts break, and each has a specific correction. Treat this section as the real work; a rubric that does not name its own false positives is not defensible.

- **Messaging mistaken for maturity.** Dense AI language with no operational detail reads as leadership but proves nothing. When the claim is much larger than the contribution, you are looking at AI washing. Correct with the turn-it-off test and by requiring at least one GA feature behind login.
- **Repo volume mistaken for skill.** Busy contribution graphs are now inflated by AI commits. A fully green graph does not mean someone is a great engineer, because some developers game graphs with trivial commits. Correct by reading merged PRs to high-bar repos, not green squares.
- **Demo mistaken for product.** The controlled demo is an optimized environment, and features shown may not exist in the shipped version. Correct by finding the GA label and running the feature on messy input.
- **Wrapper mistaken for proprietary AI.** Claiming proprietary AI while wrapping a commercial LLM API is a documented pattern. Correct by checking docs and network calls for provider names.
- **Beta counted as shipping.** An invite-only beta scored as production inflates the verdict. Beta features are evolving, with pricing not finalized and no guarantee they reach GA. Correct by reading the release label.
- **Hiring counted as production.** Open AI roles with generic requirements prove intent, not delivery. Correct by cross-checking hiring against the product surface and repos.
- **False negative on a mature private-stack firm.** A company with proprietary repos and no public case studies looks empty but may be shipping heavily. Correct by checking filings, model cards, GA features, and the MLOps and research title mix in Tables B and C.
- **Stale model card or changelog.** A card not updated after a claimed retrain, or a changelog gone quiet, signals the capability may not be live. Model cards should be updated whenever the model changes in a way that could affect performance, risk, or use, with triggers including retraining, fine-tuning, a new version, or a new deployment context. Correct by verifying version pins and cadence dates.

> **Watch out:** The private-stack false negative is the costliest error
>
> A mature competitor with closed repos and no case studies can look like a marketing-only firm. Before scoring low, confirm you checked filings, GA features behind login, and the MLOps title band. Underrating a real threat is worse than overrating a loud one.

## Before you call the verdict

Run this checklist before you sign off. It catches the common ways a read looks finished but is not defensible.

#### Verdict sign-off

- [ ] Every claimed capability is tagged core or cosmetic by the turn-it-off test.
- [ ] Each capability is marked vaporware, beta, or GA with the release label recorded.
- [ ] Hiring is reported as band ratios, not a single headcount, and benchmarked against the right market.
- [ ] The MLOps and research bands were checked against Tables B and C, not assumed.
- [ ] Engineering output is scored on merged PRs to high-bar repos, not contribution-graph volume.
- [ ] Investor-facing claims were compared against homepage claims and the gap recorded.
- [ ] Model cards and changelogs were checked for version pins and recent cadence.
- [ ] Every band assignment cites a specific, independently pullable public artifact.

## Keeping the read current

A maturity read is a snapshot, and the signals decay at different rates, so re-check on the mechanism rather than the calendar. Product announcements decay fastest because a demo or press release can be aspirational, so re-verify any product-surface claim whenever the competitor announces a new launch. Hiring patterns and sustained repository contribution move slowly and resist faking over quarters, so a quarterly re-pull of the title mix is enough. Model cards are triggered by retraining, so treat a claimed model update as the trigger to re-read the card.

**Competitor AI-maturity scorecard**

```
Capability:
Claim source (homepage / filing / earnings):
Turn-it-off result (core / cosmetic):
Hiring signal (research count / MLOps count / applied count):
Product surface (vaporware / beta / GA) + artifact:
Engineering trace (merged PR / model card / none) + link:
Filings and governance (consistent / gap / absent):
Verdict (Marketing-only / Piloting / Shipping):
Re-check trigger (next launch / quarterly / claimed retrain):
```

*Fill one row per claimed capability; the lowest-confidence class caps the verdict.*

The discipline that makes this read defensible is the same one that keeps it current: every line of the scorecard points to an artifact someone else can pull. When a competitor's story changes, re-pull the artifacts, not the narrative, and let the hardest evidence class decide the new band.

## Frequently asked questions

### How can I tell if a competitor actually uses AI or is just marketing it?

Apply the turn-it-off test to every claim: if the AI component were removed, what would be left? Then confirm with artifacts that are hard to fake. Look for a feature that is generally available behind login, a model card, merged pull requests to high-bar libraries, and MLOps hires. Heavy AI messaging with no operational detail surfacing anywhere public is the classic marketing-only signature.

### What is the single most reliable public signal of production AI use?

Production plumbing, visible through MLOps and ML-platform hiring. In Refolk's index there are 81 applied ML engineers for every MLOps or ML-infra engineer, making that band the scarcest and most load-bearing. A firm with applied engineers but near-zero MLOps is almost certainly piloting, not running models in production. Shipping AI to real users requires deployment infrastructure that someone has to build and run.

### Why can I no longer trust GitHub contribution graphs for this?

GitHub reports that 51% of committed code in early 2026 was AI-generated or AI-assisted, so star counts, follower counts, and green-square volume stopped separating real engineers from prompt-typers. The signal moved from quantity to quality: merged pull requests to high-bar repositories such as vllm, pytorch, and huggingface transformers. Those require maintainer review and are far harder to inflate.

### Are company filings more trustworthy than marketing pages for AI claims?

Yes. The SEC has treated false AI claims as securities fraud under existing anti-fraud rules, so investor-facing claims carry liability that homepage copy does not. The gap between what a firm tells investors and what it tells visitors is itself diagnostic. The SEC brought its first AI-washing actions in March 2024 and charged its first public company, Presto Automation, in January 2025.

### How do I avoid wrongly scoring a mature private-stack firm as empty?

A company with proprietary repositories and no public case studies can look empty while shipping heavily. Correct for this false negative by checking regulatory filings, model cards, generally available features behind login, and the MLOps and research title mix. If the MLOps band is present at expected scale and a GA feature exists, score it Shipping even without open repos.

### How long are these signals good for before I should re-check?

Signal shelf life varies and is not a fixed public standard. Product announcements decay fastest because a demo or press release can be aspirational, while hiring patterns and sustained repository contribution move slowly and resist faking over quarters. Model cards should be version-pinned and are triggered by retraining, so a stale card after a claimed retrain is itself a warning sign worth re-checking.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/competitor-ai-maturity-read*
