The Competitor AI-Maturity Read: Shipping, Piloting, or Marketing-Only
You will score any named competitor as Shipping, Piloting, or Marketing-only, citing the specific public artifacts that separate real production use from AI washing.
Deciding how far a competitor has actually gone with AI is a judgement call most teams get wrong because the public record is a sea of tool round-ups and vendor landing pages that treat a press release as a shipped capability. This guide is for strategy and research teams, talent-intelligence analysts, and operators who need to score a named competitor on a defensible AI-maturity verdict. It gives you a weighted rubric across four signal classes that produces the same verdict no matter who runs it, and it tells you which artifacts separate real production use from marketing claims.
The verdict has three values: Marketing-only, Piloting, or Shipping. The rest of this document is how to assign one and cite the artifact that earned it.
What the verdict means and why three bands
The three bands describe how far a capability has traveled from claim to production. Marketing-only means the claim exists but no operational substance does. Piloting means the capability runs for some users, often behind a beta label, with the plumbing half-built. Shipping means a generally available feature serves real users on real infrastructure, corroborated by hiring and engineering traces.
The reason to score bands rather than yes or no is that most of the gap between competitors lives in the middle. Adoption is broad but shallow by base rate. In Federal Reserve analysis, the AI share of job postings is only 1.6 percent across all firms but 8.6 percent among firms that have ever posted an AI role. The useful comparison is a competitor against AI-active peers, not against the whole economy.
The single most diagnostic test underneath all three bands is the simplest one. If you turned off the AI component of a given capability, what would be left? Everything that survives is cosmetic. A chatbot bolted onto unchanged infrastructure as a presentation layer survives the test with its real product intact, which tells you the AI is a veneer. A fraud model whose removal breaks the product does not survive it, which tells you the AI is load-bearing.
The four signal classes and what each one proves
Four classes of public evidence carry the verdict, and no single one is sufficient. Practitioners triangulate across all four because each proves something different and each lies in a characteristic way.
| Signal class | What it proves | How it lies |
|---|---|---|
| Hiring mix | Intent and depth of the AI team | Open roles with generic requirements and no shipped surface |
| Engineering output | Real engineering capability | Contribution-graph volume inflated by AI-generated commits |
| Product surface | A capability users can reach | A polished demo on curated data, or an invite-only beta |
| Infrastructure and filings | Governance and investor-grade honesty | Dense AI language in investor decks with no operational detail |
Hiring data proves intent and depth but not shipped product. Open repositories prove engineering capability but are now noisy. Case-study pages prove a named customer exists but not scale. Regulatory filings are the most falsifiable class because the SEC has treated false AI claims as securities fraud under existing anti-fraud rules, so investor-facing claims are the ones a company has the most incentive to keep honest.
Signal classes from softest to hardest
- Marketing copyHomepage and press releases, zero liability, discard on conflict
- Hiring mixOpen roles and title composition, proves intent and depth
- Product surfaceRelease labels and changelogs, proves reachable capability
- Engineering tracesMerged PRs and model cards, proves real output
- Filings and governanceInvestor claims under fraud liability, hardest to fake
When two classes disagree, the harder layer wins. A homepage that promises an AI platform while filings describe only a pilot tells you to score Piloting and note the gap. The gap between what a firm tells investors and what it tells visitors is itself a finding worth recording.
Reading the hiring mix for depth, not headcount
Depth hides in title composition, not in headcount. A firm claiming AI transformation with a dozen generically titled AI Engineers but no research scientists and no MLOps engineers is almost certainly running a single bolt-on team, not an AI operation.
Refolk's index gives you the bands to measure this against. The applied-engineer base is large, the research band is a meaningful minority, and the MLOps band is scarce enough to be the sharpest tell of all.
| Role band | Count | Share of applied engineers |
|---|---|---|
| Applied ML/AI/MLOps engineers | 15,269 | 100% |
| ML research / applied scientists | 3,535 | 23% |
| MLOps / ML-platform / ML-infra | 188 | 1.2% |
Counts are from Refolk's index of professional profiles; shares are derived against the applied-engineer base. The interpretation lever is direct. A higher research share signals depth of modeling capability. A near-zero MLOps share signals a bolt-on with no production plumbing, because models do not reach users without deployment infrastructure that someone builds and runs.
That scarcity is the core insight of the whole read. In Refolk's index there are 4.3 applied engineers per research scientist but 81 per MLOps engineer. The firm that has hired into the 81x band has paid for the hard, rare skill that production demands. The firm that has not is telling you, through its own hiring, that its AI has not left the lab.
Market context matters when you benchmark. Supply is concentrated, so a UK competitor with three MLOps hires is relatively deeper than a US competitor with the same three.
| Market | Applied ML/AI/MLOps engineers | Ratio vs UK |
|---|---|---|
| United States | 15,269 | 6.2x |
| United Kingdom | 2,462 | 1.0x |
Counts are from Refolk's index; the ratio is derived (15,269 divided by 2,462). Reading hiring without a benchmark produces the wrong verdict; a modest absolute count can be a strong relative one in a thin market.
This is exactly the measurement that is slow to run by hand and fast with the right index. You do not want to hand-count titles across a competitor's team one profile at a time.
Refolk lets you quantify the title mix directly, so the research-to-applied and MLOps-to-applied ratios become numbers you cite rather than impressions you form. Run the same query shape for the research band and you have both levers in minutes.
Verifying the product surface
Release-stage labels are the clearest public signal that a capability is live, and the gap between beta and generally available is the gap between Piloting and Shipping. A feature behind login with a GA label that has passed major compliance, security, and performance tests is a shipped capability. An invite-only beta is a pilot.
The distinctions are public and precise. Beta makes a feature available for real use while its interface and behavior are being finalized, and beta features are functional end to end, not partial implementations. That matters in both directions: a working beta is not vaporware, but it is also not GA, because beta features are still evolving, with pricing not finalized and no guarantee they progress to general availability.
Product-surface test
- Find itLog in and locate the actual feature, not the marketing page
- Read the labelRecord beta, GA, or invite-only from the UI or docs
- Check cadenceConfirm the changelog shows sustained releases, not one announcement
- Name the modelLook in docs or network calls for a model or provider reference
- Run messy inputTest on uncurated data to see if the demo survives reality
Three checks confirm a live capability. First, is the feature behind login and GA, not invite-only beta? Second, does the changelog show sustained cadence rather than a single launch post that went quiet? Third, do the docs or network calls name a model or provider? GPT appears in 7.8 percent of AI-engineer job postings, and the same provider references surface in docs and network traffic, which both corroborates real integration and tells you whether the AI is a wrapper around a third-party API.
That last point is a trap worth stating plainly. A firm claiming proprietary AI while wrapping commercially available LLM APIs from a third party, and marketing the result as proprietary technology, is washing. The network call that names the provider settles it.
The engineering-trace read after the 2026 inversion
Engineering traces corroborate real AI output, but the signal inverted and you must read it the new way. Contribution graphs, star counts, and follower counts stopped separating real engineers from prompt-typers once AI tooling could merge a pull request on its own.
The scale of the change is the reason. GitHub reports that 51 percent of committed code in early 2026 was AI-generated or AI-assisted, with some firms far higher. Volume metrics now produce false positives, so the signal moved from quantity of commits to merged pull requests in standard libraries.
What still carries weight is a merged pull request to a high-bar repository, because a maintainer reviewed and accepted it. Libraries such as vllm, pytorch, and huggingface transformers set the bar. For scale, vllm-project/vllm has more than 2,000 lifetime contributors, and a recent release shipped 411 commits from 212 contributors, 61 of them new. A competitor whose engineers appear in that contributor graph has demonstrated real AI engineering output that cannot be faked with trivial commits.
Model cards are the other durable trace. The model card is the de facto transparency standard, from the 2019 Mitchell et al. paper, and it is the standard document for AI model transparency. A firm that publishes version-pinned model cards is operating at a maturity that marketing-only firms do not reach.
A merged pull request to vllm is reviewed by a maintainer; a green contribution graph is reviewed by no one.
Governance, filings, and the honesty gap
Filings are honest where marketing is not, because companies face securities-fraud liability for investor-facing AI claims but not for homepage copy. The SEC did not need new AI-specific rules to bring its cases; it used existing anti-fraud provisions, the Marketing Rule, and fiduciary duty principles, and it can pursue AI-washing on negligent misrepresentation alone for registration statements.
The enforcement record tells you the claims are real risks, not theoretical. The SEC brought its first AI-washing actions in March 2024 against Delphia and Global Predictions, with Delphia agreeing to pay a 225,000 dollar penalty and Global Predictions 175,000 dollars. In January 2025 it charged Presto Automation, the first AI-washing action against a public company, alleging it overstated its drive-through product. In June 2024 it charged the founder of Joonko with defrauding investors of at least 21 million dollars over AI candidate-matching claims. The SEC created a Cyber and Emerging Technologies Unit in February 2025 focused on AI-related misconduct, and 12 AI-related securities class actions were filed in the first half of 2025, with companies taking an average 11.4 percent stock hit on announcement day.
For governance scale, look for documented model governance and responsible-AI references that match the size of the claimed operation. ISO 42001:2023 requires documented model governance, and the EU AI Act Annex IV requires technical documentation for high-risk systems from August 2, 2026. A firm claiming an AI platform with no governance footprint at all is claiming more than it runs.
The scoring procedure
Run the seven steps in order, then combine the four signal classes into one verdict. Each step has an owner and a rough time budget so you can staff it, and each ends with a defined done-state so you know when to move on.
From claim inventory to verdict
- Scope the claimCollect every public AI claim from homepage, press releases, filings, and earnings calls into a dated list of specific capabilities, with filings flagged because they carry the most legal weight.
- Apply the turn-it-off testFor each capability, ask what remains if the AI is removed, and tag the claim core or cosmetic.
- Read the hiring mixCount AI and ML roles and their composition across research scientists, MLOps, and applied engineers, producing ratios rather than a single headcount.
- Check the product surfaceLog in, find the feature, and record its release label, changelog cadence, and any model or provider reference, marking each capability vaporware, beta, or GA.
- Inspect engineering tracesLook for named repos, model cards, and merged PRs to high-bar libraries, weighting merged PRs over contribution-graph volume, until each capability is corroborated or flagged absent.
- Check governance and filingsLook for model governance and responsible-AI references at expected scale and compare investor claims against operational detail.
- Score and weightCombine hiring, engineering output, product surface, and infrastructure into one verdict of Marketing-only, Piloting, or Shipping, with each score citing a specific artifact.
Sources disagree on order. Washing-focused writers lead with messaging; talent-intelligence writers lead with hiring. In practice, start with the claim inventory because it defines what you are scoring, then let the hardest evidence class present decide ties.
To combine the four classes, place a competitor on two axes: how much operational substance exists and how loud the claim is. The quadrant tells you the band and what to do.
Substance against claim volume
How this read goes wrong
The failure modes are where most verdicts break, and each has a specific correction. Treat this section as the real work; a rubric that does not name its own false positives is not defensible.
- Messaging mistaken for maturity. Dense AI language with no operational detail reads as leadership but proves nothing. When the claim is much larger than the contribution, you are looking at AI washing. Correct with the turn-it-off test and by requiring at least one GA feature behind login.
- Repo volume mistaken for skill. Busy contribution graphs are now inflated by AI commits. A fully green graph does not mean someone is a great engineer, because some developers game graphs with trivial commits. Correct by reading merged PRs to high-bar repos, not green squares.
- Demo mistaken for product. The controlled demo is an optimized environment, and features shown may not exist in the shipped version. Correct by finding the GA label and running the feature on messy input.
- Wrapper mistaken for proprietary AI. Claiming proprietary AI while wrapping a commercial LLM API is a documented pattern. Correct by checking docs and network calls for provider names.
- Beta counted as shipping. An invite-only beta scored as production inflates the verdict. Beta features are evolving, with pricing not finalized and no guarantee they reach GA. Correct by reading the release label.
- Hiring counted as production. Open AI roles with generic requirements prove intent, not delivery. Correct by cross-checking hiring against the product surface and repos.
- False negative on a mature private-stack firm. A company with proprietary repos and no public case studies looks empty but may be shipping heavily. Correct by checking filings, model cards, GA features, and the MLOps and research title mix in Tables B and C.
- Stale model card or changelog. A card not updated after a claimed retrain, or a changelog gone quiet, signals the capability may not be live. Model cards should be updated whenever the model changes in a way that could affect performance, risk, or use, with triggers including retraining, fine-tuning, a new version, or a new deployment context. Correct by verifying version pins and cadence dates.
Before you call the verdict
Run this checklist before you sign off. It catches the common ways a read looks finished but is not defensible.
Verdict sign-off
- Every claimed capability is tagged core or cosmetic by the turn-it-off test.
- Each capability is marked vaporware, beta, or GA with the release label recorded.
- Hiring is reported as band ratios, not a single headcount, and benchmarked against the right market.
- The MLOps and research bands were checked against Tables B and C, not assumed.
- Engineering output is scored on merged PRs to high-bar repos, not contribution-graph volume.
- Investor-facing claims were compared against homepage claims and the gap recorded.
- Model cards and changelogs were checked for version pins and recent cadence.
- Every band assignment cites a specific, independently pullable public artifact.
Keeping the read current
A maturity read is a snapshot, and the signals decay at different rates, so re-check on the mechanism rather than the calendar. Product announcements decay fastest because a demo or press release can be aspirational, so re-verify any product-surface claim whenever the competitor announces a new launch. Hiring patterns and sustained repository contribution move slowly and resist faking over quarters, so a quarterly re-pull of the title mix is enough. Model cards are triggered by retraining, so treat a claimed model update as the trigger to re-read the card.
Capability: Claim source (homepage / filing / earnings): Turn-it-off result (core / cosmetic): Hiring signal (research count / MLOps count / applied count): Product surface (vaporware / beta / GA) + artifact: Engineering trace (merged PR / model card / none) + link: Filings and governance (consistent / gap / absent): Verdict (Marketing-only / Piloting / Shipping): Re-check trigger (next launch / quarterly / claimed retrain):
Fill one row per claimed capability; the lowest-confidence class caps the verdict.
The discipline that makes this read defensible is the same one that keeps it current: every line of the scorecard points to an artifact someone else can pull. When a competitor's story changes, re-pull the artifacts, not the narrative, and let the hardest evidence class decide the new band.
Questions practitioners ask
How can I tell if a competitor actually uses AI or is just marketing it?
Apply the turn-it-off test to every claim: if the AI component were removed, what would be left? Then confirm with artifacts that are hard to fake. Look for a feature that is generally available behind login, a model card, merged pull requests to high-bar libraries, and MLOps hires. Heavy AI messaging with no operational detail surfacing anywhere public is the classic marketing-only signature.
What is the single most reliable public signal of production AI use?
Production plumbing, visible through MLOps and ML-platform hiring. In Refolk's index there are 81 applied ML engineers for every MLOps or ML-infra engineer, making that band the scarcest and most load-bearing. A firm with applied engineers but near-zero MLOps is almost certainly piloting, not running models in production. Shipping AI to real users requires deployment infrastructure that someone has to build and run.
Why can I no longer trust GitHub contribution graphs for this?
GitHub reports that 51% of committed code in early 2026 was AI-generated or AI-assisted, so star counts, follower counts, and green-square volume stopped separating real engineers from prompt-typers. The signal moved from quantity to quality: merged pull requests to high-bar repositories such as vllm, pytorch, and huggingface transformers. Those require maintainer review and are far harder to inflate.
Are company filings more trustworthy than marketing pages for AI claims?
Yes. The SEC has treated false AI claims as securities fraud under existing anti-fraud rules, so investor-facing claims carry liability that homepage copy does not. The gap between what a firm tells investors and what it tells visitors is itself diagnostic. The SEC brought its first AI-washing actions in March 2024 and charged its first public company, Presto Automation, in January 2025.
How do I avoid wrongly scoring a mature private-stack firm as empty?
A company with proprietary repositories and no public case studies can look empty while shipping heavily. Correct for this false negative by checking regulatory filings, model cards, generally available features behind login, and the MLOps and research title mix. If the MLOps band is present at expected scale and a GA feature exists, score it Shipping even without open repos.
How long are these signals good for before I should re-check?
Signal shelf life varies and is not a fixed public standard. Product announcements decay fastest because a demo or press release can be aspirational, while hiring patterns and sustained repository contribution move slowly and resist faking over quarters. Model cards should be version-pinned and are triggered by retraining, so a stale card after a claimed retrain is itself a warning sign worth re-checking.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.