# The Back-Channel Customer Check: Sourcing a Startup's Real Users Before You Wire

*You will be able to build a startup's real customer list yourself, reach the right user at each account without the founder, and run a call that surfaces churn.*

- Canonical URL: https://www.refolk.ai/guides/back-channel-customer-check
- Pillar: Investing and deal sourcing
- Format: Playbook
- Published: 2026-09-29
- Last reviewed: 2026-09-29
- Reading time: 15 min

Before you commit to a round, you need to talk to a startup's actual customers and users that the founder did not hand you, so you can verify the product is loved and not quietly churning. This playbook is for early-stage investors, platform and talent partners at funds, and angels. It gives you the end-to-end method to reconstruct a startup's real customer base from public signals, pick and reach the right individual user at each account without going through the founder, and run a structured call that surfaces churn risk and weak product-market fit.

Existing library guides cover founder off-list references and public traction signals. This one covers the step that sets valuation: independently sourcing the product's customers yourself and reaching real users without tipping the founder.

## Why the founder's customer list is structurally biased

The founder list is not occasionally skewed. It is engineered to be skewed. The target hand-picks its happiest, most loyal, most articulate customers, and no rational team offers up a customer that is evaluating a competitor or considering a switch. The incentive structure guarantees a distorted sample.

The gap has a measured size. Reference-call satisfaction scores average 30 to 40 percent higher than independently recruited customer interviews for the same company. That is the whole reason back-channel checks exist. Independent sourcing does not just widen the sample; it changes the answer.

**30-40% - How much higher founder-provided references score on satisfaction versus independent interviews of the same company**

The gap is why a curated list cannot verify retention on its own.

The back-channel is where conviction forms and where deals die quietly, as investors reach out to people the founder did not name for unfiltered answers. A typical deal historically absorbed 118 hours of diligence, about 10 reference calls, and 83 days to close. Spending those hours only on the founder's list is spending them on a sample designed to reassure you.

> **Rule:** Run both lists, expect a gap
>
> Always run founder-provided and independently found accounts as two labeled tracks. Treat any deal where the two tracks agree perfectly as a signal to keep digging, not to relax.

## What each public source proves, and where it lies

A product with real usage leaves fingerprints across the open web. Your job is to collect the accounts, knowing each source proves something narrow and hides something you care about. The discipline is to name what a source proves and what it misses before you trust it.

| Source | Proves | Misses |
|---|---|---|
| Case study / logo page | A signed relationship existed | Currency and satisfaction; it is curated |
| G2 / Capterra review | A named individual used it | Volunteer skew |
| Job post naming the tool | Active tooling in use | Whether it is loved or churning |
| 10-K customer disclosure | Material revenue concentration | Private companies are excluded |

Aggregators and review sites do most of the reconstruction. Case-study and customer aggregators pull company lists from websites, case studies, and videos; they show only a slice of the base, but absolute counts can be moderately high. Tech-stack lookups are the best of the tools that list companies running specific software. Review profiles name individuals directly: on one product the names of users and their companies have been published, handing you named customers. Company websites carry at least seven ICP signals worth mining: careers pages, leadership bios, press releases, tech-stack clues, pricing models, customer logos, and About Us language.

For consumer and app startups, the sources shift to app-store listings, third-party review sites, and public communities. Capterra lists Discord at 4.6 across 532 verified reviews, an example of a public, named-reviewer source you can work through one contributor at a time. Reddit and Discord communities add moderators and heavy users who will talk candidly.

For public companies, annual reports must name major customers whose loss could cause revenue problems, which gives you concentration truth for free. Private companies give you no such disclosure, so you rebuild it from the individuals.

#### The public customer footprint, outermost first

1. **Aggregators** - Case-study and customer-list sites naming companies at scale
2. **Tech-stack and hiring** - Software-usage lookups and job posts proving active tooling
3. **Review sites** - G2, Capterra, TrustRadius profiles naming individual users
4. **Communities** - Reddit, Discord, app-store reviewers posting under names

*Each layer names customers at a different resolution, from company down to named individual.*

## How reconstructable a customer base really is

The reason this method works at scale is that individual users are publicly addressable in large numbers. Any product with meaningful adoption produces thousands of people who list it on their own profiles.

In Refolk's index of professional profiles, 75,565 US professionals publicly list Salesforce as a skill and 62,652 list HubSpot. Those are not company logos; they are named individuals you can reach without the founder's introduction. If a portfolio-scale product leaves tens of thousands of addressable users, a Series A startup with real traction leaves hundreds to low thousands, which is far more than the five to eight calls you need.

| Tool (skill) | Market | People publicly listing it | Share of largest |
|---|---|---|---|
| Salesforce | United States | 75,565 | 100% |
| HubSpot | United States | 62,652 | 83% |
| Salesforce | United Kingdom | 8,582 | 11% |

Geography changes reachability, not just market size. UK Salesforce users number 8,582, about 11.4 percent of the US pool. A sampling plan built for a US startup does not transfer to a UK-focused one. Plan more candidate accounts per finding abroad, because the pool you draw from is an order of magnitude smaller.

> Any product with real usage leaves thousands of individually addressable users; the founder is a convenience, not a gatekeeper.

## The procedure, start to finish

Run these eight stages in order. The list-building work is front-loaded; the calls are where judgment happens.

#### The back-channel customer check

1. **Scope and trigger the check** - After the partner meeting but before the term sheet, decide check depth by stage and check size, and write down the exact claims you are verifying such as retention, that the product is loved, and that named logos are real. Done is a short written list of claims.
2. **Reconstruct the customer list** - Mine case-study and logo pages, G2, Capterra and TrustRadius profiles, tech-stack lookups, job posts naming the tool, LinkedIn, and 10-K customer disclosures for public firms. Done is 15 to 40 candidate accounts each with a source URL.
3. **Pick the right individual per account** - For B2B target the daily user, admin, or RevOps owner rather than the buyer who signed; for consumer target reviewers and community posters. Done is a named contact and a reach path per account.
4. **Segment on-list versus off-list** - Separate founder-provided contacts from independently found ones and plan to run both. Done is two labeled lists; expect the independent track to score lower given the 30 to 40 percent satisfaction gap.
5. **Reach out with neutral framing** - Lead with context and ask for a short informal chat about how they work, not a formal reference. Done is five to ten calls booked over one to two days.
6. **Run the structured call** - Anchor on the last time they did the job, probe embeddedness and switching costs, and apply the one-thing-to-improve and Sean Ellis tests. Done is notes capturing specifics, not vibes, in 20 to 30 minutes each.
7. **Detect coaching and triangulate** - Flag rehearsal patterns and require five to eight independent voices including at least one struggling account. Done is a pattern read across accounts.
8. **Score and decide** - Convert notes into a retention and product-market-fit verdict that feeds valuation, and check for revenue concentration. Done is a go or no-go memo.

There is genuine disagreement on timing. Most sources run customer diligence before the term sheet, but in competitive situations investors shift more of it to after the term sheet to move faster. Decide which world you are in during step one, because it sets how aggressively you can reach out.

### Where the call volume should land

Right-size the effort to your check. The historical average VC made about 10 reference calls per deal; top firms run 10 to 20 or more; angels run 7 to 12 with only 2 to 3 from the founder.

| Investor type | Reference calls per deal | Founder-provided share |
|---|---|---|
| Average VC (historical study) | ~10 | not stated |
| Top-tier VC | 10-20+ | mixed named and unnamed |
| Angel | 7-12 | 2-3 of them |

For customer calls specifically, aim for five to eight independent conversations, including at least one account that is struggling. One fund's stated target is five to seven conversations including independently found back-channels and at least one struggling account.

## Reaching the right person without the founder

The person who signed the contract is the wrong person to call. The buyer praises the deal they championed; the daily user is the one who churns. Reach the admin, the power user, or the RevOps owner who lives in the product, and for consumer startups reach the reviewers and community posters who use it by choice.

This is where sourcing the individual, not just the account, decides the quality of your read. Given a customer logo, you want the person whose job touches the product every day, at a company the founder did not put in front of you.

I ran this search: `US customer success and operations leaders who mention switching from a legacy CRM to a newer sales tool` - [see the full result list](https://www.refolk.ai/s/gvrqe88jzv).

*Returns named operators who have lived through a switch, the exact voices who can tell you whether switching costs are real or whether people drift back.*

Reaching them is a research problem, not a networking problem. Rather than mining review sites by hand and guessing at who still uses the tool, [Refolk](/) lets you ask for the people directly across GitHub, LinkedIn, and the open web, so you can pull daily users at accounts outside the founder's list in one pass.

When you write, use neutral, non-attributed framing. Ask for a short informal chat about how they use the tool, not a formal reference, and make it low-stakes to say yes.

**Neutral back-channel outreach**

```
Subject: Quick question about how you use [product]

Hi [name] - I saw you work with [product] at [company]. I am doing some
research on the category and trying to understand how teams actually use
it day to day. Not a formal reference, just 15 minutes on what works and
what you would change. Open to a quick chat this week?
```

*Swap the product name and role for your target. Keep it short and never call it a reference.*

Most people will talk if you ask the right questions. The framing matters because it lowers the odds the contact loops in the founder before you speak, and because a candid user answers a curiosity question more openly than an interrogation.

## Running the call so it surfaces churn

Ask about the last time, never about the next time. Self-reported switching reasons matched the actual behavioral driver only 54 percent of the time, and across one vendor's 723 churn interviews the exit-survey reason matched the root cause less than a third of the time. Stated reasons are unreliable; past behavior is factual.

So anchor every probe on an event. "Walk me through the last time you did X" beats "would you use X" every time. Then push on three things that predict retention better than sentiment does.

- **Embeddedness and replacement behavior.** Listen for parallel processes. If the customer still exports to spreadsheets, reroutes work to email, or keeps a manual workaround alive, switching costs are weak and the product is not load-bearing. Embeddedness beats NPS, because customers describe churn as low consequence rather than dissatisfaction. If nobody would notice the tool disappearing, the risk is existential even with a decent score.
- **The one thing they would improve.** Every customer has one thing they wish were better. If the answer is "nothing," the call is now suspect. Candor produces a specific miss plus a retention signal, such as "reporting is weak, we built a workaround, and everything else held for 14 months." A frictionless rave is the anomaly to investigate, not the reassurance it feels like.
- **The Sean Ellis test.** Ask how disappointed they would be if the product went away tomorrow. Companies where 40 percent or more of users say "very disappointed" showed strong, sustainable growth. Under 40 percent across your independent sample is a real product-market-fit flag, regardless of what the logo page says.

> **Watch out:** Do not accept "everything's great"
>
> A rave with no named flaw is the most common false positive in customer diligence. Force a specific miss. If you cannot get one, downgrade the call, do not upgrade the deal.

You can also add a size read. Asking to whom the customer would recommend the product works as a qualitative proxy for market size and for how vigorously it spreads.

## How this goes wrong

Most bad customer diligence is not lazy; it is fooled. These are the failure modes that produce a confident yes on a churning product, each with the check that catches it.

1. **Trusting the founder's list.** Five glowing calls prove the founder chose well, not that the product is loved. Run independent accounts and expect the 30 to 40 percent satisfaction gap. If your independent track matches the curated track, keep sourcing.
2. **Mistaking stated reasons for real ones.** Customers rationalize. Stated switching reasons matched behavior only 54 percent of the time. Anchor on past events and probe replacement behavior instead of asking why.
3. **Accepting "everything's great."** A "nothing to improve" answer means the call is suspect, not that the product is flawless. Force a specific miss before you record the call as positive.
4. **Missing a coached reference.** Coaching shows as deflection to generalities, struggling under specific follow-ups, and excessive consistency. Answering a practice question identically three times running indicates over-rehearsal. Ask an off-script scenario and watch for identical phrasing across references.
5. **Reaching the wrong individual.** The buyer praises; the daily user churns. Confirm you are talking to the admin or power user, not the executive who signed.
6. **Treating a stale logo as a live customer.** A case study can be two years old. Confirm current usage through recent job posts or tech-stack data before you count the account.
7. **Over-indexing on one voice.** A single glowing or poor read is rarely enough; the value is spotting trends across trusted voices. Require five to eight independent contacts before you draw a conclusion.
8. **Ignoring concentration.** One happy whale can hide fragility. If 40 percent of ARR sits with a single customer on an annual contract with a 60-day termination window, that is a material risk warranting protective deal terms, no matter how happy that customer sounds.

#### Sentiment versus embeddedness

Horizontal axis runs from Low embeddedness to High embeddedness. Vertical axis runs from Low sentiment to High sentiment.

| Quadrant | What it means |
| --- | --- |
| Vocal but detached | Investigate; may churn on any friction |
| Loyal and locked in | Strongest signal; verify it repeats across accounts |
| Quietly at risk | Highest churn risk; nobody would notice it disappear |
| Grumbling but stuck | Retained by switching cost, not love; check contract terms |

*A high score means little if the product is not load-bearing; embeddedness is the axis that predicts churn.*

### Spotting the coached reference

Structure is the anti-coaching mechanism. Because coached references break under specific follow-ups, a fixed probe sequence converts polished delivery into detectable rehearsal, which is exactly why unstructured calls let coaching pass. Genuine, specific enthusiasm stands out immediately from coached talking points. When you hear hesitation on specifics, or hear the same phrasing twice across two references, treat it as a flag and route around that account with an independently sourced voice.

## Scoring and keeping the read current

Turn your notes into a verdict that feeds valuation, not a folder of transcripts. The output of this work is a retention and product-market-fit read the partnership can price against, plus a concentration check.

#### Before you call the customer check done

- [ ] I reconstructed 15 to 40 candidate accounts from public sources, each with a source URL.
- [ ] I ran founder-provided and independently found accounts as two labeled tracks.
- [ ] I reached the daily user or admin at each account, not only the buyer who signed.
- [ ] I completed five to eight independent conversations including at least one struggling account.
- [ ] Every call anchored on a past event and produced at least one specific improvement, not a rave.
- [ ] I applied the Sean Ellis 40 percent test across the independent sample.
- [ ] I checked for revenue concentration and any single terminable contract above 40 percent of ARR.
- [ ] I flagged and re-sourced any account where the reference showed coaching or identical phrasing.

Score against the claims you wrote in step one. If the independent track lands within a few points of the founder track and clears the 40 percent disappointment bar, you have a defensible loved-product read. If the independent track sits 30 to 40 percent below the curated one and shows replacement behavior, the founder's traction narrative is inflated and your valuation should reflect it.

Keep the read current by mechanism, not by date. Recheck usage through fresh job posts and tech-stack signals near close, since logos go stale, and re-run one or two calls if the deal slips past the window where your original conversations were fresh. The market context also moves: when down rounds run near 19 to 20 percent of rounds rather than the 10 to 12 percent norm, the cost of overpaying on inflated traction rises, so weight the independent track harder. The method holds; the numbers you plug into it are the part you refresh.

## Frequently asked questions

### How do I find a startup's customers without asking the founder?

Reconstruct the list from public signals. Mine the startup's case-study and logo pages, review profiles on G2, Capterra and TrustRadius that name individual users, tech-stack lookups, job posts that mention the tool, and 10-K customer disclosures for public firms. In Refolk's index, tens of thousands of professionals publicly list common tools as skills, so any product with real usage leaves individually addressable users. Aim for 15 to 40 candidate accounts, each with a source URL, before you reach out.

### When in diligence should back-channel customer calls run?

At most top firms reference and customer checks happen after the partner meeting but before the term sheet, and some funds run customer calls two to four weeks into due diligence. In competitive situations investors sometimes shift more diligence to after the term sheet. Start reconstructing the customer list as soon as you have conviction to spend the 118 hours a typical deal absorbs, so the calls do not become the bottleneck near close.

### How many customer reference calls do I actually need?

Plan for five to eight independent voices per deal, including at least one struggling or churned account. A historical study put the average VC at about 10 reference calls, top-tier firms run 10 to 20 or more, and angels do 7 to 12 with only 2 to 3 founder-provided. A single glowing or poor read is rarely enough; the value is in the pattern across trusted, independent voices.

### Will contacting a startup's customers tip off the founder and damage the deal?

That specific claim is not established publicly, so treat it as a risk to manage rather than a rule. Use neutral, non-attributed framing: ask for a short informal chat about how they use the tool, not a formal reference, and reach the daily user rather than the executive most likely to call the founder. Back-channel norms hold that you should route through people whose judgment you trust and keep the conversation low-key.

### How do I tell a coached reference from a genuine one?

Coached references deflect toward general assessments, struggle with follow-up questions that probe specifics, and show excessive consistency that suggests rehearsal. Answering a practice question identically three times running is a tell. Structure is the countermeasure: run a fixed probe sequence, ask an off-script scenario, and force a specific miss. Genuine, specific enthusiasm stands out immediately from polished talking points, and a 'nothing is wrong' answer should make the whole call suspect.

### What questions actually surface churn and weak product-market fit?

Anchor on past behavior with 'walk me through the last time you did X,' because past behavior is factual and stated reasons matched real drivers only 54 percent of the time. Probe replacement behavior: if the customer still exports to spreadsheets or maintains parallel processes, switching costs are weak. Ask what one thing they would improve, and apply the Sean Ellis test, since 40 percent or more of users saying they would be very disappointed without the product signals sustainable growth.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/back-channel-customer-check*
