Refolk
TeardownMarket and talent intelligence

Reconstructing a Competitor's Customer Base and Market Share

You can turn one named competitor into a deduplicated, date-stamped customer list and an at-least share estimate with a confidence band, without crossing into trade-secret territory.

16 min readLast reviewed August 24, 2026Read as Markdown

Key takeaways

  • A publicly reconstructed logo list captures only a fraction of a vendor's base, so the resulting share estimate is a lower bound and must be framed as at least, never as precise.
  • Job-postings that require experience with a product out-yield the vendor's own case studies: one search surfaces 50 to 100 user companies versus a handful of curated logos.
  • A missing Wayback snapshot is silence, not innocence; robots.txt hides pre-existing snapshots retroactively, so an undated relationship must never be scored as new.
  • The trade-secret line tracks effort and access: Seyfarth's five-column rule means enriching a public name with pricing, buying habits, or login-gated data can flip a fair-game list into protectable territory.
  • Employee-skill footprints scale the same signal to national level: 73,771 US professionals list Salesforce as a skill versus 8,280 in Germany, an 8.9x market-depth gap in Refolk's index.
  • SEC disclosure only triggers at 10 percent of revenue and identity is optional, so concentrated markets leak named whales while fragmented ones hide behind Customer A.

You have one named competitor and a question that will not go away: who do they actually sell to, and how much of the market do they hold? This guide is for strategy and research teams, talent-intelligence analysts, and operators sizing a market. It carries one method all the way through, using only public sources, to a deduplicated and date-stamped customer set plus a share estimate with a stated confidence band, and it draws the trade-secret line that lead-gen guides skip.

Most competitor teardowns read a roadmap from open roles or track talent outflow. Reconstructing the customer footprint is a different discipline: it is part archival forensics, part sampling, and part law. The single most important thing to hold in your head is that a public reconstruction only ever captures a fraction of the base. Every number you produce is a floor, not a level. Write "at least" on it and mean it.

What public sources actually name a competitor's customers?

Named customers leak from a small set of public channels, and each one is biased in a knowable direction. The vendor's own site is the usual first stop but the thinnest yield; the highest-yield sources sit outside the vendor's control.

Many B2B vendors run a case-study section or a logo wall, and a quick search gives a handful of names from testimonials, webinars, or videos. That is real, but it is a small and curated slice of the base. The marketing team chose those logos. The richer sources are the ones no one manages.

SourceWhat it yieldsThe bias to remember
Vendor case studies and logo wallsA few curated namesSelf-selected, small slice of the base
Review aggregators (reviewer firmographics)Company size, industry, title of actual buyersReviewers self-select; a fraction, not a census
Job postings requiring the product50 to 100 user companies per searchSkews to companies actively hiring
SEC 10-K major-customer notesCustomers at 10%+ of revenueIdentity optional; triggers only at 10%
Technology footprint lookupsSites running the product's scriptsFrontend inference can over-count

The counter-intuitive lesson: job postings out-yield the vendor's own site. A single search for "requires experience with the product" can surface 50 to 100 user companies, because hiring is decentralized and unmanaged and so leaks more than the controlled logo wall. Review aggregators are the next best, because every reviewer lists their company size, industry, and job title. Filtering reviews by company size shows you who the competitor actually sells to, self-reported by buyers.

How SEC disclosures expose whales but hide the long tail

Public companies must disclose a customer that is 10 percent or more of revenue, but they need not name it, so this source is powerful in concentrated markets and nearly silent in fragmented ones. Under ASC 280, if revenues from a single external customer reach 10 percent or more of a public entity's revenues, the entity must disclose that fact and the amount, but the customer's identity is optional.

That threshold makes public share reconstruction lopsided by market structure. Concentrated markets expose named whales. Fragmented ones stay hidden behind "Customer A." The table below shows the range across real filings.

FilerCustomer% of revenueFY
FabrinetLumentum Operations20%2019
JabilApple18%2014
DLH HoldingsDept of Health & Human Services49.8%2025
DLH HoldingsDept of Veterans Affairs33.8%2025
Commercial Vehicle GroupCustomer A (unnamed)13%2025

Two lessons follow. First, when your competitor is a supplier to public companies, run an EDGAR full-text search for the competitor's name inside 10-K filings: their customers may have disclosed the relationship even if the competitor is private. Second, remember the deduplication rule these filings encode. A group under common control counts as a single customer, and each government body (federal, state, local, foreign) counts as one customer. You will apply that same rule when you clean your own list.

The worked example: carrying one competitor through

Take a mid-market B2B SaaS competitor selling to companies in the 100 to 999 employee band. I will use the sizing figures from a real bottom-up exercise so the arithmetic is concrete, and I will flag the wrong turns as they happen.

The scope decision comes first: one competitor, one product, one geography. Skipping this is the most common early mistake, because a fuzzy denominator makes every downstream number unfalsifiable. Write the counting rules on one page before harvesting a single logo.

The harvest then runs in two waves. Wave one, the vendor site, returns maybe a dozen curated logos. Wave two, the third-party sources, is where the volume comes from. Here the sources disagree on order, and the disagreement matters. One school starts with the vendor site for a clean base; another starts with job postings and review aggregators because they yield more. Follow the higher-yield path: run the job-posting and reviewer searches first, then use the vendor site to fill gaps and confirm. This is the fork where a lot of teams waste a day scraping a logo wall that was never going to move the number.

From raw candidates to a defensible customer set

  1. Raw candidate logos
    312,000

    universe of segment entities

  2. After de-dup and shell removal
    287,000

    dormant shells removed

  3. Verified current customers
    identified fraction

    the at-least numerator

Every stage discards rows, and the survivors are still only a fraction of the true base.

The universe here starts at 312,000 entities in the 100 to 999 employee band drawn from commercial registries, of which 287,000 remain after de-duplication and removal of dormant shells. That 287,000 is your denominator. Your numerator is whatever fraction of it you can verify as current customers. If your harvest and verification produce, say, a few thousand verified logos, your share is "at least that fraction," never "that fraction exactly." The gap between what you found and what exists is invisible, so you report the floor and state the precision target.

287,000
Segment entities after de-duplication in the worked universe
Starts at 312,000; dormant shells and duplicates are removed before it becomes the denominator.

Dating each relationship without over-reading the archive

To date when a customer relationship became public, pull a domain's snapshot history and find the first snapshot in which the logo or case study appears; that snapshot date bounds when the relationship went public. The Internet Archive's CDX Server API returns the timestamp list you need, across over a trillion saved pages.

The Wayback Machine is an excellent resource for this, but its coverage gaps are severe and directional, and misreading them is the most expensive error in the whole method. The crawler is weighted toward popular sites, so coverage of any single domain is uneven. Low-traffic sites are crawled less than once a year, snapshot intervals are irregular, and a site can show dozens of captures in one month and none the next.

Two further traps corrupt dating. Robots.txt causes retroactive blackouts: if a site's current robots.txt blocks crawlers, the Wayback Machine hides even old snapshots made before that rule appeared. When a domain looks suspiciously empty, try Common Crawl (usable coverage from around 2013 to 2014, crawling roughly monthly) or a cache as a fallback. And redirects distort content: some CDX results correspond to 301 or 302 redirects rather than the final page, so confirm the snapshot actually renders the logo before you trust its date.

Where reconstruction crosses into trade-secret territory

A publicly reconstructed logo list is generally lawful; the line into misappropriation is crossed when you attach non-public detail or acquire information by improper means. The governing test under the DTSA and comparable state law is whether the information is generally known or readily ascertainable by proper means. Information can be a trade secret only if it is subject to reasonable secrecy measures and derives independent economic value from not being generally known or readily ascertained (18 U.S.C. 1839(3)).

The case law is clear that a simple list of names compiled from a telephone book or, more likely today, the internet, generally does not rise to a trade secret. The list flips into protectable territory when non-public detail is paired with the names: a customer's credit history, buying habits, specific pricing, or sales volume. Access matters too. Courts have found a customer list protectable where some of the information was not readily ascertainable because someone had to log in to a network to access it.

Improper means taints acquisition regardless of whether the underlying fact is public. Hacking, breaching an NDA, or scraping behind a login can poison an otherwise clean reconstruction. The safe posture is simple: only proper means, only public fields, and a documented source URL on every row so you can prove how each fact was obtained.

When the widening step is the bottleneck, Refolk collapses it into a single plain-English query instead of a day of manual searching across job boards and review sites. Ask for the companies hiring against a product, or the people who list it as a skill in a given segment, and you get a sourced list back rather than a scrape you have to clean by hand.

Verifying live customers and reading the skill footprint

A logo in an archive is not a current customer; set a written churn definition before counting and verify every survivor against present-day signals. Companies must define what qualifies as a lost customer and apply that definition consistently, because moving the goalposts obscures real trends. Standard treatment counts a customer as churned when the subscription terminates and is not reactivated in the period, and treats a downgrade to a free plan as churn.

Three public checks separate a current customer from a stale or aspirational one. Confirm the logo is on the current live site, not only in an archived snapshot. Check the case-study publication date. Check whether current job postings still require the product. Treat pilots and "selected" announcements as aspirational until a second source corroborates them.

For national-scale sizing, employee-skill footprints are a scalable proxy for the same signal. In Refolk's index of professional profiles, 73,771 US professionals list Salesforce as a skill and 69,210 list HubSpot, while only 8,280 Germany-based professionals list Salesforce. The Germany figure is about one-ninth of the US one, mirroring the market-depth gap that job postings show one company at a time.

SkillCountryProfessionals listing skill
SalesforceUnited States73,771
HubSpotUnited States69,210
SalesforceGermany8,280

Two derived ratios carry the warning. The US Salesforce footprint is 8.9x Germany's, which is a genuine market-depth signal. But HubSpot sits at 93.8 percent of the US Salesforce footprint, and that near-parity is a trap. HubSpot's self-serve base inflates head-count relative to contract value, so a skill count must be revenue-weighted before it enters a share estimate. Read a raw skill count as revenue share and you will over-credit the self-serve vendor every time.

Read a skill count as revenue share and you over-credit the self-serve vendor every time.

The procedure, end to end

Run these eight steps in order. Each produces a named artifact, and the whole chain converts one competitor into a defensible share number.

Reconstruct the base and estimate share

  1. Scope and define
    Name one competitor, one product, one geography, and write the counting rules for what is a customer and what excludes a churned or pilot logo. Done: a one-page rule sheet.
  2. Harvest named logos from the vendor site
    Pull case studies, logo walls, testimonials, and press releases. Done: a raw logo list with a source URL per entry.
  3. Widen with third-party sources
    Add reviewer firmographics, job postings requiring the product, technology-footprint lookups, and SEC 10-K notes. Done: a merged candidate list, each row carrying its source.
  4. Date each relationship
    Use the Wayback CDX history to find the first appearance of each logo and record the snapshot date; mark rows with no coverage as date unknown. Done: date-stamped rows.
  5. Verify live vs churned or pilot
    Confirm each logo against the current live site and recent signals; drop or flag stale and pilot logos per the rule sheet. Done: a cleaned, deduplicated set.
  6. Deduplicate under common control
    Merge subsidiaries and rebrands, since a group under common control is a single customer. Done: a canonical entity list.
  7. Size the market universe
    Pull entity counts for the defined segment from registries, de-dupe, and remove dormant shells. Done: a denominator with a stated source.
  8. Compute share two ways and set the band
    Divide verified customers by the universe for a bottom-up figure, run a top-down cross-check, and report the spread as the band with a precision target. Done: a share estimate with a band and an implied-share sense-check.

The share estimate as a two-method triangulation

  1. Bottom-up
    Verified customers divided by the cleaned universe
  2. Top-down
    Independent estimate from market size and adoption
  3. Triangulate
    The gap between the two becomes the confidence band
  4. Sense-check
    Compare implied share against known incumbent shares
Running both methods and reporting their spread is what turns a point guess into a defensible band.

Bottom-up sizing starts from real data points such as customer counts, unit prices, and transaction volumes and builds upward. It is slower but more reliable, and the most rigorous exercises use both top-down and bottom-up, then triangulate. Either approach works alone, but running both gives you a confidence interval. State the precision target explicitly, as practitioners do: plus or minus 10 percent at 80 percent confidence. Then sense-check the implied share against the shares of dominant players in your category and adjacent ones. Because a public reconstruction only ever captures a fraction of the base, the reconstructed count is a lower bound and the share must read as "at least."

How this goes wrong: the failure modes

Most reconstructions fail not in the harvest but in the reading, where a benign gap gets scored as a fact. The table below is the set to check against before you brief anyone.

Failure modeThe false positiveThe check
Empty Wayback stretch read as no relationshipA crawl gap looks like the logo was absentConfirm crawl density; treat silence as unknown
Robots.txt blackout read as never archivedRetroactively hidden snapshots look missingTry Common Crawl or a cache
Tech footprint over-countsA frontend script infers a product no longer usedCorroborate with a second source before counting
Case-study logo counted as currentA churned logo still shows on an old pageConfirm live site plus a recent job posting
Double-counting subsidiariesParent and acquired brand counted twiceMerge under common control per ASC 280
Reconstructed detail treated as safePricing or login-gated data attached to a rowKeep to public fields; apply the five-column line

Two more failures sit above the row-level ones and undermine the whole estimate. The first is stating a single-method share as precise. Run bottom-up and top-down, report the spread as the band, and sense-check against incumbents' known shares. The second is projecting selection bias onto the whole base. Reviewers on aggregators and featured logos on a vendor site are self-selected fractions, not a census. Say so in the brief. A reader who thinks your sample is representative will read your floor as a level, which is exactly the error the whole method exists to avoid.

Technology footprints deserve their own caution because they feel authoritative. Frontend detection often infers a backend product from a script that may appear late in a customer's lifecycle or not at all, so it both misses real customers and flags former ones. Use it as one corroborating source, never as a standalone census.

Before you ship: the verification checklist

Run this before the estimate leaves your desk. It is the difference between a number that survives a leadership review and one that gets picked apart in the first five minutes.

Ship-readiness for a reconstructed customer base and share

  • The rule sheet defines a customer and the churn exclusion, and it was written before counting.
  • Every candidate row carries a source URL obtained by proper means only.
  • No row is enriched with pricing, buying habits, or login-gated data.
  • Each logo is dated from a rendered Wayback snapshot, or explicitly flagged date unknown.
  • Undated relationships are never scored as new.
  • Every current logo is confirmed on the live site plus a second recent signal.
  • Subsidiaries and rebrands are merged under common control.
  • The denominator states its registry source and its de-dup and shell-removal step.
  • Share is computed two ways and reported as a band with a stated precision target.
  • The share is framed as at least, and selection bias is disclosed in the brief.

Keeping the reconstruction current

A customer footprint decays the moment you finish it, so treat the artifact as a living file with a re-check cadence rather than a one-time report. Logos churn, case studies get pulled, and new hires post new job requirements every week. Set a review interval tied to how fast the market moves, and on each pass re-run the two highest-yield sources first: job postings requiring the product and reviewer firmographics.

Version the file with a date stamp on every row so a reader always knows how fresh a given data point is, and keep the source URL against each one so you can re-verify without re-harvesting. When a relationship's date is unknown, leave it unknown until the archive fills in; do not let pressure to look complete turn a gap into a guess. The value of this work is that it is defensible, and defensibility comes from being honest about the floor, the band, and the silence in the archive.

Questions practitioners ask

Is it legal to reconstruct a competitor's customer list from public sources?

Yes, when you use only proper means and public fields. Case law holds that a bare list of names compiled from the internet generally is not a trade secret. The risk starts when you attach non-public detail such as specific pricing, buying habits, credit history, or anything behind a login, or when you acquire information by improper means like hacking or breaching an NDA. Keep to publicly ascertainable fields and you stay on the right side of the line.

How do I estimate market share when I can only see part of the customer base?

Treat your verified count as a lower bound and report share as at least. Divide verified, deduplicated customers by a defined market universe from commercial registries, then run an independent top-down cross-check. The spread between the two methods is your confidence band. State a precision target, such as plus or minus 10 percent at 80 percent confidence, and sense-check the implied share against known incumbent shares.

Why does a missing Wayback Machine snapshot not mean the customer was not there?

Because the Wayback Machine crawls selectively, weighted toward popular sites, so a single domain can show dozens of captures one month and none the next. A missing snapshot proves only that the crawler was not present. Worse, if a site's current robots.txt blocks crawlers, the Wayback Machine retroactively hides even old snapshots taken before that rule existed. Treat a gap as unknown, and try Common Crawl or a cache as a fallback.

What is a better first source than the vendor's case-study page?

Job postings that require experience with the product. Hiring is decentralized and unmanaged, so it leaks far more than a marketing team's curated logo wall. A single search for companies requiring the product often surfaces 50 to 100 user companies, versus the handful of names on a case-study section. Reviewer firmographics on aggregators are a strong second source because reviewers self-report company size, industry, and title.

How do I decide whether a logo is a current customer or a churned one?

Write the definition before you count and apply it consistently. Standard treatment counts a customer as churned when the subscription terminates and is not reactivated in the period, and treats downgrades to free plans as churn. Practically, confirm the logo is on the current live site, not only in an archived snapshot, check the case-study publication date, and see whether current job postings still require the product. Treat pilots and selected announcements as aspirational until a second source corroborates them.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next