Refolk
TeardownInvesting and deal sourcing

Restating an AI Startup's Gross Margin Before You Price the Round

You can restate one AI company's gross margin with inference moved to COGS and credits removed, then rule whether the headline survives, compresses, or was propped up.

17 min readLast reviewed October 11, 2026Read as Markdown

Before you price an AI company's round, you need to know what its gross margin actually is once model inference sits in cost of goods sold and vendor credits are stripped out. This guide is for early-stage investors, platform and talent partners, and angels who have to defend a valuation. It carries one worked case all the way through the inference-to-COGS reclass and the credit-stripping math, with the real intermediate numbers, so you can follow along on your own deal.

The gross-margin line is where AI deals mislead. ARR pressure-tests, cap-table recasts, and traction reads are well-covered, but they stop at the top line. The margin itself gets quoted from a management deck and waved through. That is a mistake, because two accounting choices - where inference is booked and whether credits are netted - can swing the number by 18 to 20 points. A true 60 percent dressed up as a SaaS-like 78 percent is exactly the deception this guide is built to catch.

Why AI gross margins are structurally lower than SaaS

AI products run materially below classic SaaS, so a margin that looks SaaS-grade is the first thing to distrust. Bessemer's pricing playbook puts AI gross margins at 50 to 60 percent against 80 to 90 percent for traditional software, and a16z's framing lands in the same range.

The survey data agrees. ICONIQ's State of AI snapshot, covering roughly 300 software executives, restates the series as 41 percent in 2024, 45 percent in 2025, and a projected 52 to 53 percent in 2026. The spread inside that average is wide: Bessemer's "Supernovas" reached about $125 million of ARR in year two at gross margins near 25 percent, while "Shooting Stars" ran about 60 percent. So before you even open the data room, your prior should be that a credible AI company sits in the 45 to 60 percent band, and anything claiming 78 percent needs to earn it.

SegmentGross marginSource
Traditional SaaS80-90%Bessemer playbook
AI product average45% (2025), 52-53% (2026)ICONIQ State of AI
Application-layer resellers45%ICONIQ
Balanced-differentiation53%ICONIQ
Bessemer "Supernovas"~25%Bessemer State of AI

The reason the number is lower is structural, not temporary. Inference is the one COGS line that does not amortise with scale. It climbs from about 20 percent of AI product cost pre-launch to roughly 23 percent at scale, while talent's share falls from 32 percent to 26 percent, because adoption grows consumption faster than efficiency shrinks it. Mature SaaS COGS runs 10 to 25 percent of revenue; an AI company carries a usage-metered compute bill on top of that.

45%
Average AI product gross margin, 2025
ICONIQ's survey of roughly 300 software executives; projected to reach 52-53% in 2026 against 80-90% for traditional SaaS.

The worked case: Northwind, 72 percent on paper

Take one company all the way through. Call it Northwind, an application-layer AI product with a management deck showing 72 percent gross margin on $100 million of revenue. The CloudZero worked math gives us the shape: $20 million of traditional COGS and $8 million of AI COGS already disclosed, netting to a 72 percent gross margin. That is the headline I will try to break.

The first read is already a signal. A 72 percent margin on an application-layer company sits above the 45 percent ICONIQ prior for resellers and above the 53 percent balanced-differentiation mark. Either Northwind has genuine differentiation and in-house cost control, or the number is propped up. The rest of this guide is the method that tells the two apart.

Northwind gross margin as each adjustment lands

  1. Headline margin
    72%

    management deck

  2. After inference reclass
    ~64%

    ~8 pt swing

  3. After stripping credits
    ~55%

    ~9 pt swing

  4. After decile test
    55% blended, top decile negative

    power-user exposure

Each diligence step peels a layer off the headline until only the defensible number remains.

I will walk the two load-bearing adjustments - the inference reclass and the credit strip - with real point swings, then the decile test that exposes what the average hides.

Reclassifying production inference into COGS

Production inference belongs in COGS, and moving it there is usually the first high-single-digit hit to a flattered margin. Under US GAAP the test is nature and purpose. Inference, LLM API fees, serving GPUs, embeddings, and vector databases scale with usage like hosting and serve current customer traffic, so they sit above the line. Model training and platform-wide fine-tuning are R&D under ASC 730, or potentially capitalizable internal-use software under ASC 350-40.

The practitioner default settles the argument cleanly. An AWS bill for production EC2 serving customer traffic is COGS. An OpenAI bill for production inference serving the same workloads is also COGS. There is no accounting principle that distinguishes them. If a company has parked inference in R&D, it is not reading the standard, it is dressing the margin.

On Northwind, the deck already shows $8 million of "AI COGS", which is encouraging, but the reclass question is whether all customer-facing compute is actually in that line or whether a chunk is sitting in engineering. The documented size of this error is well-attested: one case moved hosting to COGS and dropped gross margin from 76 percent to 68 percent, an 8-point swing, because the cost had been hidden in the wrong line item. In a full GAAP recast during Series A diligence, one modeled case saw gross margin drop by 21.7 percent once the chart of accounts was recast.

For Northwind, assume the diligence finds $8 million of serving GPU and third-party product API spend tagged as R&D. Reclassifying it lifts COGS from $28 million to $36 million and drops the margin from 72 percent to 64 percent. That is the first layer off the headline.

What stays in R&D

Leave training and fine-tuning where they belong. In most real-world AI development scenarios, training data costs do not meet the criteria for capitalization and must be expensed as R&D. The one thing to check is whether fine-tuning was capitalized under ASC 350-40 to keep it out of operating costs entirely - that flatters both COGS and operating expense at once.

Stripping credits and restating at rate

Credits do not reduce COGS. Free cloud or model credits hide a cost, so the restated margin has to show economics before credits and after any contractual discount. This is the second big swing, and it runs in the same 8-to-10-point range as the inference reclass.

The mechanism is straightforward and the size is documented. One worked narrative moved gross margin from 67 percent to roughly 76 percent purely from discounted credits - a 9-point gift that evaporates the moment the credits run out. Promotional credits can improve reported cash and hide mature unit cost, which is why underwriting should show the margin both ways.

The method: rebuild from the company's own usage ledger. Apply contracted meter rates to telemetry, reconcile the aggregate to the invoice, and retain any difference as allocation variance until resolved. Then produce two numbers, because sources genuinely disagree on which rate to use.

Credit-strip waterfall
Line                         At list rate    At contracted rate
Metered usage (telemetry)    _______         _______
Less: committed-capacity     _______         _______
Less: negotiated discount    (excluded)      _______
Less: promotional credits    (excluded)      (excluded)
= Restated AI COGS           _______         _______
Reconcile to invoice:        variance _____  variance _____

Fill from the usage ledger, not the invoice. Run both the list-price and contracted-rate columns.

Run both. One view restates at list price because many find it clearer to model margin before any discount. Another restates at the contracted rate after negotiated discounts but before free credits. Neither lets promotional credits reduce COGS. For Northwind, assume the $8 million of booked AI COGS was net of $1.5 million in expiring provider credits; grossing that back pushes AI COGS to $9.5 million. Combined with the reclassed $8 million, total COGS reaches $45.5 million and the margin falls to roughly 55 percent. The headline 72 has become a defensible 55.

Once you have found the ledger-to-invoice reconciliation is the choke point, finding the person who can speak to it is its own task. Compute-talent depth is a proxy for in-house cost control, and it is searchable.

Refolk turns that into a named list in plain English, which is faster than reconstructing it from scattered public profiles. The depth matters: in Refolk's index, 2,952 US profiles carry LLM plus Inference skills against 448 UK profiles, a 6.6x ratio, concentrated at employers like OpenAI and Atlassian and heavily in the San Francisco Bay Area. A UK target is far less likely to have hired the inference-optimisation talent that defends margin, which makes its restated number more fragile.

MarketProfiles (LLM + Inference)ShareRatio vs UK
United States2,95286.8%6.6x
United Kingdom44813.2%1.0x

The decile test: what the blended average hides

A clean 55 percent blended margin can still sit on top of a negative top decile, so the decile plot is the load-bearing number, not the headline. Sort users by usage, group them into ten buckets, and compute gross margin per bucket. If the top decile is negative and the bottom decile is at 80 percent, pricing is not capturing value from power users.

This is not an edge case. Around 10 percent of users can generate 40 to 70 percent of the inference bill, and the top decile consumes 5 to 10 times the median. Under flat pricing, a single price covering all usage lets top users subsidize themselves straight into your loss column. The blended average is the deception vector precisely because it averages that away.

Running the decile test

  1. Build inputs
    Per-action cost at provider rates plus a 30% buffer, P50 and P90 monthly usage from the last 30 days of logs
  2. Bucket users
    Sort by usage into ten deciles and set a base price benchmarked to the P50 user
  3. Compute per bucket
    Gross margin for each decile; read the top decile and quantify the power-user loss
Three inputs turn a usage log into a per-decile margin curve that exposes power-user losses.

The inputs are specific: per-action cost at current provider rates with a 30 percent buffer, projected P50 and P90 monthly usage per active user from the last 30 days of logs, and a base price benchmarked to the P50 user. Watch for agent retries, which blow up cost per task - one ICONIQ-quoted workflow projected at $0.10 per run reached $1.50 or more once agents retried. Check the actual logged tokens, not the plan assumption.

For Northwind, assume the decile plot shows the top bucket running at negative 15 percent while the bottom sits near 80 percent. The company is one large power-user onboarding away from the blended 55 percent sliding further. That is the difference between "compresses" and "propped up".

The clean company average is the deception vector; the decile curve is where the loss actually lives.

The procedure, end to end

This is the full method, in order, from pulling the P&L to ruling the headline. Each step names who does it and what "done" looks like so you can run it on a live deal.

Rebuilding the restated gross margin

  1. Pull and normalise the P&L
    Get three years of management accounts plus current YTD and reconcile revenue and COGS to bank and cash. Done when you have a consistent chart of accounts and have reviewed hosting, third-party API expense, and support overhead.
  2. Classify every AI cost line
    Apply the test that COGS carries only the cost of serving current customers, excluding sales commissions, R&D engineering, product, and marketing. Done when each line is tagged COGS, R&D, or SG&A with a one-line rationale.
  3. Reclassify production inference into COGS
    Move customer-facing inference, serving GPUs, vector DB, embeddings, and third-party product APIs above the line, leaving training and fine-tuning in R&D. Done when you have a restated COGS with inference broken out.
  4. Strip credits and restate at rate
    Rebuild cost from the usage ledger, apply both list and contracted rates, and reconcile to the invoice. Done when you have a before-credits and after-contractual-discount waterfall, run both ways.
  5. Reconcile telemetry to invoice to GL
    Reproduce the margin for selected customers: start from contract and revenue, trace provider and tool calls, apply invoice rates, add direct service cost. Done when per-customer margin matches the aggregate within tolerance.
  6. Run the decile sensitivity test
    Sort users into ten usage buckets and compute gross margin per bucket using P50 and P90 inputs. Done when the top-decile margin is computed and any power-user loss is quantified.
  7. Cross-check compute dependency against public signals
    Compare implied inference-to-revenue against the application-versus-balanced split and any committed-spend or MD&A disclosures. Done when the dependency estimate is triangulated.
  8. Rule the headline
    State whether the restated margin survives, compresses, or was propped up, attributing the point swing to each adjustment. Done when you have one defensible restated number plus a bridge.

The point-swing arithmetic is what makes the ruling defensible. Each adjustment has a documented range, and your bridge should attribute the move explicitly rather than presenting one number out of thin air.

AdjustmentBeforeAfterSwing
Strip discounted credits67%76%9 pts
Hosting reclass to COGS76%68%8 pts
Full GAAP recast (modeled)n/an/a21.7%

Cross-checking when the invoices are withheld

When a company will not hand over its usage ledger, model-provider economics and public filings let you triangulate dependency anyway. Position in the stack is the strongest free prior: the closer a company is to just passing through a frontier API, the thinner the margin it keeps. The 45 percent application-layer versus 53 percent balanced-differentiation split is your anchor when invoices are hidden.

Three public reads back that up:

  • MD&A ratios. Public SaaS companies have started disclosing inference-related cost ratios, generally between 4 and 9 percent of revenue. A private target claiming materially lower is making a claim it has to defend.
  • Committed-spend announcements. These read directly on dependency and concentration. Under one neocloud deal, Anthropic is paying roughly $1.25 billion per month through May 2029 and Google approximately $920 million per month. For your target, any multi-year committed-compute contract tells you the cost floor is locked, not flexible.
  • Provider margin trends. Anthropic's inference margins rose from 38 percent to 70 percent in a single year per SemiAnalysis data. That gain does not automatically reach the app; falling token prices can be recaptured upstream by provider pricing power rather than passed through to the startup you are pricing.

One mechanic people reach for does not hold up here: there is no publicly established way to read inference dependency off a Form D filing. Use committed-compute announcements and MD&A ratios instead, and say plainly in your memo that the dependency estimate is triangulated rather than reconciled.

How this goes wrong

The failure modes below are where a restated margin gets faked or missed. Each has a tell and a check, and this is the part of diligence that actually protects the valuation.

Failure modeWhat the false positive looks likeHow to check
Inference parked in R&DAn 80%+ "SaaS-grade" marginDoes customer-paid traffic trigger the compute? If yes, it is COGS.
Credits masking cash costMargin healthy only while credits lastDemand a before-credits vs after-contractual-discount waterfall.
Blended-average blindnessA clean company averageRun the decile plot; the headline hides power-user losses.
Mid-year reclass whiplashMargin jumps 76% to 88%Check for consistent methodology across all periods.
Telemetry not tying to invoiceDashboards show high efficiencyReconcile provider dashboards to invoice, GL, and cash.
Capitalized trainingLow operating cost and clean COGSCheck whether fine-tuning was capitalized under ASC 350-40.

Two more deserve their own note. Agent retries exploding cost per task is the quiet killer: a $0.10 plan line can become $1.50 or more in production, so read the actual logged tokens, not the pricing assumption. And customer-success staff can be hidden in either COGS or sales and marketing to flatter whichever line the company wants to look good; review job descriptions and compensation plans, as venture associates do, to see where the people actually sit.

One counterintuitive tell closes the set. A margin that improves with scale is a flag, not a comfort. Inference does not amortise, so genuine operating leverage on the compute line is rare. If the trend is up, probe whether it reflects credits, a reclass, or capitalized training before you credit it to the model.

Before you call the number defensible

Run this check before the restated margin goes in the memo. It is the difference between a number you can defend to the partnership and one that falls apart on the first question.

Restated gross margin sign-off

  • Three years of management accounts plus YTD reconciled to bank and cash.
  • Every AI cost line tagged COGS, R&D, or SG&A with a written rationale.
  • Production inference, serving GPUs, vector DB, and product APIs sit in COGS.
  • Training and fine-tuning confirmed as R&D, not capitalized under ASC 350-40.
  • Margin waterfall shows both before-credits and after-contractual-discount, at list and contracted rates.
  • Per-customer margin reconciles to the aggregate within tolerance.
  • Decile plot computed; top-decile margin and power-user loss quantified.
  • Compute dependency triangulated against the stack split, MD&A ratios, or committed-spend data.
  • Bridge from headline to restated number attributes each point swing to a named adjustment.

Keeping the method current

The thresholds in this guide move, so re-check the mechanism rather than the number. The margin benchmarks are survey-driven and reset each ICONIQ and Bessemer edition; pull the latest band before you anchor a prior. Provider pricing is the volatile input - token prices fall, but provider margin gains like Anthropic's 38-to-70-percent jump show that the cost floor is set upstream, so re-read committed-spend announcements and MD&A inference ratios each time you price a deal. And billing models shift under you: when a dominant tool moves to usage-based billing, every downstream reseller's unit economics change at once. The procedure holds; the inputs do not, so run the reclass, strip the credits, and plot the deciles fresh on every term sheet.

Questions practitioners ask

Where does model inference go, COGS or R&D?

Production inference that serves current customer traffic belongs in COGS. Under US GAAP the test is nature and purpose: inference, LLM API fees, serving GPUs, and vector databases scale with usage like hosting, so they sit above the line. There is no accounting principle that distinguishes a production inference bill from a production EC2 bill. Model training and platform-wide fine-tuning are R&D under ASC 730 or potentially ASC 350-40.

How much can reclassifying inference move gross margin?

A reclass of production compute into COGS commonly costs high single digits. One documented case moved hosting to COGS and dropped gross margin from 76 percent to 68 percent, an 8-point swing, purely because the cost had been hidden in the wrong line item. In a full GAAP recast during Series A diligence, one modeled case saw gross margin drop by 21.7 percent.

Should I restate at list price or at the contracted rate?

Run both. Sources disagree: one view restates at list price because many find it clearer to model margin before any discount, while another restates at the contracted rate after negotiated discounts but before free credits. Credits do not reduce COGS. Showing both a before-credits and an after-contractual-discount number gives you the waterfall you need to see how much of the margin is real.

Why does a blended company average mislead on AI margin?

A clean blended average hides power-user losses. Around 10 percent of users can generate 40 to 70 percent of the inference bill, and the top decile can consume 5 to 10 times the median. If the top decile runs negative while the bottom runs at 80 percent, a flat-priced book is one heavy cohort away from turning red. The decile plot, not the headline, is the load-bearing number.

What public signals estimate compute dependency when invoices are withheld?

Use the stack-position split, MD&A ratios, and committed-spend announcements. Pure application-layer companies reselling someone else's model report around 45 percent margin versus 53 percent for balanced-differentiation firms, so position predicts the real number. Public SaaS companies have begun disclosing inference-related cost ratios in MD&A, generally 4 to 9 percent of revenue. Committed-compute announcements read directly on dependency and concentration.

Is a margin that improves with scale a good sign?

Not automatically. Inference is the one COGS line that does not amortise with scale: it climbs from roughly 20 percent of AI product cost pre-launch to about 23 percent at scale, because adoption grows consumption faster than efficiency shrinks it. A margin that improves with scale is itself a flag to probe, since it may reflect credits, a mid-year reclass, or capitalized training rather than genuine operating leverage.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next