Customer Reference Calls: Verifying Traction Before You Wire
You can source an independent customer list, run the calls in the right order, and turn the answers into a defensible verdict on demand and retention.
You have a term sheet out on a seed or Series A deal and the traction slide looks good. Before you wire, you need to know whether the demand and the retention are real, and the only primary source that can tell you is the customer. This guide is the end-to-end method for running customer reference calls: sourcing an independent list, sequencing and running the calls without burning the founder's relationships, and converting the answers into a defensible verdict on demand and churn.
This is not a founder character check and it is not a public-signal traction read. Those tests ask whether the founder is who they say they are and whether the metrics are internally consistent. The customer call tests something neither can reach: does the buyer actually want this, and will they pay again. It is the last and most invasive step, and it is the one that most often kills a deal at the wire.
Why the customer call is the test the other checks cannot run
The customer reference call is the only diligence step that puts you in direct contact with the person deciding whether the startup's revenue renews. Founder references test character, public signals test consistency, and the customer call tests demand itself.
The distinguishing feature is that the founder is absent. Reference calls are often the first time in the process that the management team is completely removed and cannot course-correct in the middle of the action. That absence is the whole point. Everything the founder controls has already been presented; the customer call is where you hear the version the founder did not stage.
It is also the most invasive step you take before pulling the trigger. Every call asks a customer to put their relationship with the startup on the line for a deal that has not closed. That risk is why the call sits at the end of diligence and why doing it badly does lasting damage. Handled well, one round of calls settles the question that the deck cannot: is the traction load-bearing.
When to make the call in the diligence sequence
Hold customer calls until the end of diligence, after the partner meeting, when a term sheet is issued or clearly imminent. Calling earlier burns the base for a deal that may never close.
The classic framing is direct: the last two due diligence items a VC wants before offering a term sheet are personal references and customer references, and they should be done last because it is easy to burn out the person giving the reference. The reason is not politeness. Customer references create real obligations, because asking someone to put their customer relationships on the line for a deal that has not closed carries genuine risk. You reserve that ask for the moment your interest is serious and a term sheet is likely.
There is a real market shift worth naming. Not too long ago most VCs asked for reference calls only after issuing a term sheet; in a tighter market there is far more diligence done upfront, and some funds now call earlier. Both positions are defensible. What does not change is the invasiveness of the ask. If you front-load, front-load the off-list back-channel work first, and keep the founder-supplied calls for the term-sheet window when your intent is credible.
Building two lists: managed intros and an independent set
Build two lists, not one. The founder supplies a small managed set of two to three introductions per firm; you build an independent off-list set from your own network, LinkedIn overlap, and public buyer signals, and it must include at least one churned or lost-deal account.
The managed number is a stated norm: two to three customer introductions per VC firm should be enough. Beyond that, the VC can do the work of reaching out to potential customers in their own network to ask their take. Off-list sourcing is explicitly endorsed - investors who seek off-list references by leveraging their networks get a broader range of insights than on-list references and a more holistic evaluation.
For traction specifically, the strongest independent signal comes from the negative side. Independent outreach to lost deals and churned accounts tells you more than any managed reference call. Managed calls test satisfaction among hand-picked champions. Churned and lost-deal accounts test demand elasticity: did the customer leave because the problem stopped mattering, or because a competitor was better. That is the traction question, and no on-list call will answer it.
How hard the independent list is to build depends on the buyer persona. A product bought by procurement has a far denser pool of reachable off-list contacts than a product owned by a customer success team. In Refolk's index, there are roughly 7.9 times more US procurement decision-makers than US customer success leads, so dev-tool and CS-led deals demand more sourcing effort per independent reference.
Table A - Findability of the customer-side reference contact, US vs UK (Refolk's index)
| Persona | US | UK | US:UK ratio |
|---|---|---|---|
| Head/VP Customer Success | 1,249 | 821 | 1.52x |
Table B - Which buyer persona is easiest to source independently, US (Refolk's index)
| Buyer persona | Count | Multiple of CS base |
|---|---|---|
| Procurement (Head/Mgr/VP) | 9,893 | 7.9x |
| Head/VP Engineering | 3,996 | 3.2x |
| Head/VP Customer Success | 1,249 | 1.0x |
Read these as a sourcing forecast. If the startup sells into procurement, your independent list is easy to build and there is no excuse for an on-list-only sample. If it is a CS-owned account or a dev tool sold to engineering leaders, the pool is thinner and you should budget more time to find real off-list contacts before you start dialing.
The step-by-step procedure
The job runs in seven stages, from confirming your stage to writing the verdict. Follow them in order; the sequencing is what protects the base and controls your own bias.
Running customer reference calls end to end
- Confirm you are at the right stageDo not request customer calls until after the partner meeting and other diligence. Done when you have issued or are about to issue a term sheet and have told the founder these are among the last items.
- Build the two listsTake the founder's two to three managed intros, then separately build an independent list from your own network, LinkedIn overlap, and public buyer signals. Done when you have named on-list and off-list contacts, including at least one churned or lost-deal account.
- Sequence the callsRun off-list and back-channel calls first to form an unfiltered baseline before the curated calls. Done when the order is set so managed intros confirm or contradict your view rather than seed it.
- Prepare the question setScript buying-decision, alternatives, renewal-intent, expansion, and who-initiated questions. Done when the same core questions run across every call so answers are comparable.
- Run each callKeep the founder absent and let the reference talk without cutting them off. Done when you have captured verbal renewal intent for each account.
- Protect the baseCap intros per firm, avoid re-calling the same accounts, and consider the group-call format for structured rounds. Done when no customer has been contacted more than necessary.
- Pattern the answersCompare across calls for recurring concerns and benchmark retention against the segment norm. Done when you have a written verdict grounded in the pattern, not one testimonial.
Why sequence off-list first
Running off-list calls before the curated ones is a bias control, not just an efficiency trick. References come back positive by default, so a mixed reference is usually a bad signal. If you hear the champions first, their enthusiasm anchors your read of everyone else. Form the unfiltered baseline from independent and churned accounts, then let the managed intros confirm or contradict it. This mirrors hiring practice, where back-channel references are started early precisely to avoid wasting time on a candidate the market already knows is weak.
The call sequence that controls bias
- Off-list back-channelApproach independent contacts and lost deals to form an unfiltered baseline
- Churned accountsTest why customers left, the true demand-elasticity signal
- Managed introsLet founder-supplied calls confirm or contradict, not seed, your view
- Pattern readCompare across all calls and benchmark retention before writing the verdict
The question set that separates a buyer from a soft reference
The high-value content of any customer call is the buying decision itself: how they chose to work with the company, what competitors they considered, and how the experience has been. A soft reference praises the product; a genuine buyer can reconstruct the decision that led them to pay.
Keep the specific script yours, but structure it around five comparable themes so answers line up across calls. The load-bearing question is renewal intent, because it is the one answer the founder cannot pre-coach away. A founder can prep a champion to speak warmly, but a dated, binary "yes, I plan to renew" is not something you can rehearse into existence. The gap between a claimed 95% retention rate and a customer saying they are not renewing is the single most common late-stage deal killer.
1. WHO INITIATED. Walk me through how this purchase started. Who inside your company pushed for it, and why then? 2. ALTERNATIVES. What else did you evaluate? What almost made you choose someone other than them? 3. EXPERIENCE. What has actually been better or worse than you expected since you started using it? 4. RENEWAL INTENT. Do you plan to renew? When is that decision, and who makes it? (Push for a dated verbal yes.) 5. EXPANSION. Are you using more or less of it than a year ago? Would you buy the next product they build?
Ask the same five themes on every call so the answers are comparable. Adapt wording to the buyer, not the structure.
Do not phrase any question as a request to endorse your investment. Asking a customer to bless the deal makes customers cautious about sounding too enthusiastic, and you will misread their caution as a weak product. You are diligencing the business, not auditioning the customer as a promoter. Keep the framing about their experience, and let renewal intent do the scoring.
Reading the pattern: real retention versus churn risk
Score the pattern across accounts, never a single testimonial. VCs are more focused on patterns of negative feedback emerging than on any one-off comment, and a consistent pattern of concerns or areas that need to be managed across many references is the signal to walk.
The retention-specific test is verbal renewal intent. Treat every account as a churn risk until you get renewal confirmation, a plain "yes, I plan to renew." Enthusiasm is not renewal. A customer can love a product and still be cutting the line item next quarter, so a warm call with no dated renewal commitment scores as unconfirmed, not as a win.
Then benchmark against the segment, not the founder's framing. A 92% gross retention rate sounds strong until you learn the segment norm is 95%. Pull comparables from public benchmarking sources for the specific stage and category before you judge the number. The founder will always frame a figure to sound good; your job is to know what good looks like for this segment.
Scoring an account after the call
Table C - Call cadence and thresholds (published sources)
| Parameter | Value | Source note |
|---|---|---|
| Call length | 15-30 min | Standard reference call length |
| Intros per firm | 2-3 | Stated founder-supply norm |
| Renewal notice window | 120 days | A window for timing renewal-intent questions |
| Retention norm example | 95% segment | Benchmark before judging any figure |
Enthusiasm is not renewal; a warm call with no dated commitment scores as unconfirmed, not as a win.
How this goes wrong: failure modes and false positives
Most botched customer diligence fails in one of seven predictable ways, and each has a false positive that feels like a good result. Learn the tell for each.
- Calling too early. The false positive is a founder who complies eagerly. Check: have you passed the partner meeting and signaled term-sheet intent? If not, you are burning the base for a maybe.
- Curated-only list. The false positive is three glowing calls. Check: are all three on the cap table or hand-picked champions? Three references drawn from the current cap table get discounted by partners for a reason.
- Mistaking enthusiasm for renewal. The false positive is "we love it." Check: did they give a dated verbal yes to renewing? Treat every account as a churn risk until you get that confirmation.
- Single-call verdict. The false positive is one strong testimonial. Check: is the same concern surfacing across calls? Weight patterns of negative feedback over one-off comments.
- No churned accounts called. The false positive is a 100% positive sample. Check: if the win-loss work only talked to wins, it is not win-loss work.
- Reference feels on the hook. The false positive is muted enthusiasm read as a weak product. Check: did you ask the customer to endorse the investment decision? That framing makes customers cautious about sounding too enthusiastic.
- Retention with no benchmark. The false positive is a 92% figure that sounds great. Check: what is the segment norm? Without the comparable, the number means nothing.
The deepest failure mode does not appear on any single call. Over-calling actively creates the churn it is meant to detect. Each redundant call plants doubt about whether the startup will have funding to follow through and whether it will be around in a year, and if the investor passes, it is embarrassing and potentially damaging to the relationship.
kind: warning
title: Over-calling damages the asset you are diligencing
A twenty-fund process turns into the same three customers taking six investor calls; five funds pass, and by the time the round closes the customers have drawn their own conclusions. Cap intros per firm, do not re-call accounts, and use the group-call format for structured rounds.
The mitigations are concrete. Reserve calls for term-sheet-stage funds only. Cap introductions at two to three per firm. Avoid re-calling accounts another fund already burned. For a highly structured financing process, consider the group-call format, where one investor leads the Q&A while others submit questions, so the customer takes one call instead of six.
Before you call the job done
Verify the sample, the confirmations, and the benchmark before you write the verdict. A clean-looking set of calls with no churned accounts and no dated renewals is not a diligence result; it is a testimonial reel.
Pre-verdict checklist
- You are at term-sheet stage and told the founder these are among the last items.
- You have both a managed on-list set and an independently sourced off-list set.
- At least one churned or lost-deal account is in the sample.
- You ran off-list and back-channel calls before the founder's intros.
- The same five core questions ran on every call.
- Every account has a captured renewal-intent answer, and none is warm words without a dated yes.
- No customer was contacted more than necessary, and intros were capped per firm.
- Retention figures are benchmarked against a public segment norm, not the founder's framing.
- The verdict rests on a pattern across accounts, not a single testimonial.
Keeping the work current
Customer diligence is not a one-time artifact; the conditions that make a good call change with the market and with the deal. Two things drift and are worth re-checking each time you run this.
First, the timing norm. The classic rule holds customer calls until after the term sheet, but in tighter markets more diligence moves upfront and some funds call earlier. Do not treat either as fixed. Before each deal, decide where in your own sequence the calls belong, and if you front-load, front-load only the off-list back-channel work and keep the founder-supplied calls for the moment your intent is credible.
Second, the independent list. The pool of reachable off-list contacts differs sharply by buyer persona, and the startup's go-to-market may not match the last deal you diligenced. Re-scope the sourcing effort for each company: a procurement-bought product gives you a deep independent pool, while a CS-owned or engineering-sold product will need real work to reach genuine off-list references. When the pool is thin, plain-language sourcing across public profiles is where Refolk removes the friction - ask for former customers who switched to a competitor, or the procurement leads who evaluated the category, and you get a starting off-list set without waiting on the founder. Rebuild that list per deal, benchmark the retention numbers against fresh comparables, and the method stays defensible.
Questions practitioners ask
How many customer reference calls should I run for a seed or Series A deal?
There is no fixed number, but the shape matters more than the count. Founders typically supply two to three managed introductions per firm, and you should add independent off-list contacts on top, including at least one churned or lost-deal account. Aim for enough calls that a pattern can emerge across accounts rather than resting on a single testimonial. Five to eight total, split between on-list and off-list, is a reasonable working target for a seed round.
When in diligence should I ask for customer references?
The classic norm is to hold customer calls until the end, after the partner meeting and when a term sheet is likely. Asking someone to put their customer relationships on the line for a deal that has not closed carries real risk, so deeper reference requests are reserved for serious interest. Some funds now front-load more diligence and call earlier, but the invasiveness argument still holds: call late enough that your intent is credible.
What is the best question to catch inflated traction?
Ask for dated, binary renewal intent: does the customer plan to renew, and when is the decision. This is the one answer a founder cannot pre-coach away, which is why the gap between a claimed 95% retention rate and a customer saying they are not renewing is the single most common late-stage deal killer. Treat every account as a churn risk until you get a verbal yes.
Are back-channel customer references better than founder-supplied ones?
For testing traction specifically, yes. Managed calls test satisfaction among hand-picked champions, while independent outreach to lost deals and churned accounts tests demand elasticity. Because references default to positive, off-list contacts give you a more holistic and less anchored read. Run them first so the champions do not seed your view before you have an unfiltered baseline.
How do I avoid damaging the founder's customer relationships during diligence?
The customer interview is the most invasive step you do before wiring, and over-calling actively creates the churn it is meant to detect. Reserve calls for term-sheet-stage funds, cap introductions at two to three per firm, avoid re-calling the same accounts, and use the group-call format for structured rounds. Each redundant call plants doubt about the startup's survival.
What retention number should I trust from a customer call?
Trust dated verbal renewal intent over any headline percentage, and always benchmark against the segment norm rather than the founder's framing. A 92% gross retention rate sounds strong until you learn the segment norm is 95%. Pull comparables from public benchmarking sources for the specific stage and category before you decide whether the number is good.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.