Refolk
PlaybookSales and go-to-market

The Look-Alike Account List: From Closed-Won to Ranked Net-New

You will extract the win-predicting attributes from your closed-won cohort and ship a ranked, CRM-deduplicated look-alike list with a match reason on every row.

17 min readLast reviewed August 25, 2026Read as Markdown

Your closed-won accounts are the best training data you own, and most look-alike workflows waste them. They feed the domains into a vendor button, get back companies of similar size and industry, and hand a rep a list that looks right and behaves wrong. This guide is for founders selling their own product, account executives, SDR leads and partnerships teams who want to do the part the tool comparisons skip: extract the attributes that actually predicted the win, turn them into a searchable ICP, and produce a ranked, CRM-deduplicated list of net-new companies with a stated match reason on every row.

The reasoning between "here are my ten wins" and "here is a ranked list I can hand a rep" is where the value lives. That is what this playbook covers, in order, with what to do at each stage and what a good result looks like.

What a look-alike list actually is, and what it is not

A look-alike list is a set of net-new companies that match the pattern of your closed-won wins, ranked, deduplicated against your CRM, with a defensible match reason on every row. It is not an industry code plus a headcount band, and it is not a signal of intent.

Two distinctions decide whether the list is worth a rep's time. First, similarity is not intent. ICP fit answers "could this company buy from us?" It does not answer "is this company buying now?" A perfect-fit account that just renewed with a competitor scores identically to one three weeks from a budget cycle if you blend the two. Second, firmographics are the shape of your customer base, not the substance. Two firms can match on size and industry yet sit in completely different buying situations. The attributes that separate a win from a loss are usually technographic and buying-role signals, not the size band.

The practical consequence: every row on a good look-alike list carries a reason that is more specific than "matches size and industry." If a row's only justification is firmographic, it is a filter output wearing a look-alike costume.

Firmographics give you the shape of your customer base. The win pattern is in the substance.

How many closed-won accounts you need

You need 5 to 20 ranked closed-won accounts to read a human-readable pattern, or roughly 100 or more to let software model win and loss statistically. The disagreement between sources here is a method fork, not a contradiction, and it decides your whole approach.

Practitioner guides say to rank closed-won by revenue, retention, expansion and sales-cycle length and take the top 10 to 20% as your seed set. Lookalike workflows commonly accept 5 to 100 best-customer domains as input. But AI-driven pattern extraction states a much higher floor: at least 100 closed-won accounts and a defined sales motion, because below that there is not enough signal for the model to find statistically meaningful patterns.

MethodSeed countWhat you get
Human pattern-reading5 to 20 ranked winsA written, searchable ICP you can defend
Lookalike domain input5 to 100 domainsNet-new rows sharing firmographic DNA
Statistical win/loss modeling~100+ closed-wonModeled patterns with a defined sales motion

If you have fewer than 100 closed-won deals, do not pretend to model. Read the pattern by hand, name the attributes, and write them down. The named-attribute path is more defensible anyway, because it forces you to state why each attribute belongs.

Which attributes to capture beyond firmographics

Capture six categories, not one: behavioral, intent, firmographic, technographic, first-party, and deal-level signals. The categories most sources under-weight are technographic and the buying role, and those are usually where the win pattern actually lives.

From each closed-won record, capture the state of the account at the time of purchase, not today. That means firmographics, the technologies they ran, who signed, and the signals that fired before the account entered pipeline, such as hiring activity, funding, and news events. Technographic capture is especially useful when it records a stack shift: adopting a complementary tool, churning from a competitor, or adding a platform that creates a natural integration opportunity.

The "who signed" attribute is worth isolating. The buyer persona and the stack together narrow a look-alike far faster than an industry code. In Refolk's index of professional profiles, among US Revenue Operations profiles, HubSpot outnumbers Salesforce as a listed skill by 51 to 20, a 2.55x ratio in this slice. That is a concrete technographic seed split: it tells you the modal RevOps buyer in this cohort skews mid-market HubSpot, which is a far sharper match reason than "SaaS, 50 to 500 employees."

Skill listedCount in Refolk's indexShare of the two
HubSpot5172%
Salesforce2028%

This is a narrow title-string slice, directional rather than a census, but the direction is the point: a technographic attribute gives every downstream row a reason.

Extract the win pattern

Extracting the win pattern means finding attributes that are far more common in your wins than in your losses, then writing them down as a searchable ICP with a match-reason vocabulary. This is the step vendor buttons skip, and it is the difference between a describable ICP and a black box.

Work across all four judgement categories: firmographic, technographic, behavioral and buying-role. For each candidate attribute, ask a single question: does this appear in my wins much more than in my losses? An attribute that is common in both proves nothing. Aim for 5 to 7 named attributes. Fewer, and the list is too broad; more, and you will match almost nobody.

The output is not a mental model. It is a written sentence a rep could read aloud, something like: "US Series A to B SaaS, runs HubSpot, hired their first RevOps person in the last six months, signed by a VP of Revenue Operations." Each clause becomes a match-reason token you can attach to rows later.

From closed-won cohort to searchable ICP

  1. Rank seeds
    Order closed-won by revenue, retention, expansion, cycle length
  2. Enrich as-of-purchase
    Append firmographic, technographic, hiring, funding, who-signed
  3. Compare win vs loss
    Keep attributes far more common in wins
  4. Write the ICP
    5 to 7 named attributes plus a match-reason vocabulary
The pattern-extraction path that turns wins into a describable, searchable ICP rather than a vendor button output.

At this point you have the hard part done: a pattern you can search against and defend. The friction now is finding net-new companies that fit it and confirming a real buyer at each one.

Because Refolk reads across public LinkedIn records, the public GitHub graph and Refolk's own index, you can ask for the attribute combination you extracted in plain English rather than reconstructing it filter by filter, and each result comes back tied to the reason it matched.

Generate candidates, then deduplicate and suppress

Generating candidates means running similarity against your seed domains or filtering a fresh source with your extracted ICP, so you get a candidate universe larger than the seed set with a match reason on every row. Deduplication and suppression then strip out everything that is not genuinely net-new.

Deduplication is where lists quietly poison themselves. Use root domain as the primary key for accounts and company name only as a secondary matching signal. Name-only dedupe causes false merges: "Acme Consulting" and "Acme Manufacturing" are different companies, and merging them either loses a real prospect or lets a current customer slip back onto the cold list. Run fuzzy-match logic on domain, name and, at the contact level, email address, and flag conflicts for review rather than auto-merging.

Suppression is separate from deduplication and follows a precedence rule. Broad, permanent suppression takes precedence over campaign inclusion. A rep should not be able to bypass a permanent objection or a current-customer flag simply by adding the account to a new list.

Suppression categoryWhat it removesPrecedence
Current customersAccounts you already sell toPermanent, wins
CompetitorsDirect competitorsPermanent, wins
Legal / do-not-contactRestricted or opted-out recordsPermanent, wins
Campaign / pipeline exclusionOpen opportunities, already-workedCampaign-level

The scale of the cleanup is easy to underestimate. One company cut its CRM accounts from 650,000 to 180,000 by deduplicating, and an estimated 91% of CRM data is incomplete, stale, or duplicated annually. If you skip this step, a meaningful fraction of your "net-new" list is neither net nor new.

2.55x
HubSpot to Salesforce ratio among US RevOps profiles in Refolk's index
A concrete technographic seed split; the modal RevOps buyer in this slice skews mid-market HubSpot.

Enrich contacts and verify before send

Contact enrichment means finding the buyer persona at each account and verifying their work email before anything enters a sequence. This is the most common place a list built with care still fails, because stale records degrade the model at the source and inflate bounces.

The buyer persona is the equivalent of "who signed" your best deals. Size the persona pool before you commit to a market. In Refolk's index, there are 143 US profiles matching "VP of Revenue Operations" or "RevOps Manager" versus 17 in the UK, an approximately 8.4x US-to-UK ratio. If your win pattern requires a RevOps buyer and you plan to expand into the UK, that ratio tells you the pool is thin and you may need to adjust the persona rather than force it.

Persona sliceUnited StatesUnited KingdomUS:UK ratio
VP of Revenue Operations / RevOps Manager143178.4x

Verification is not optional, and it is time-sensitive. B2B contact data decays at roughly 22 to 25% a year, so a list enriched weeks ago is already worse than it looks. Verify immediately before send and hold to the bounce thresholds below.

MetricValueSource
Acceptable hard bounce<0.5%emailaddress.ai
Tested campaign bounce3.4%overloop.com
User-reported bounce (large DB)15%+amplemarket.com
User-reported bounce (another DB)20-30%amplemarket.com
Annual B2B data decay22-25%emailaddress.ai

Note what the table proves: a large database is not a clean one. One vendor's own verified-email filter cut its database from 275M to 96M, a 65% reduction, and independent testing found a 200M-contact database at under 3% bounce outperforms a 1.7B-contact database at 25% bounce. Match-reason-per-row beats raw size, because the reason forces a verified, defensible attribute.

Score fit, signal and timing separately, then gate

Score fit and intent as two independent axes, layer timing on top, and apply a threshold before any account enters a sequence. Never collapse them into one number, because a blended score can hide a zero-intent account inside an average.

The documented gate is Fit, Signal (Intent) and Timing. Fit checks whether the company matches your ICP. Intent measures the strength of the signal, tiered: Tier 1 is funding, a new VP, headcount growth; Tier 2 is competitor research and pricing-page visits; Tier 3 is content downloads. Timing accounts for how recent the signal is. One common approach multiplies the three into a composite and pushes accounts above a threshold, typically 7 out of 10, to outreach. Sources disagree on whether to multiply all three or to score fit first and then rank by timing; the non-negotiable is that fit and intent stay visible as separate values.

Fit versus timing, the two-axis judgement

High fitLow fit
High fit, low intent
Nurture; do not burn a rep on a dead-fit account
High fit, high intent
Work now; these clear the threshold to sequence
Low fit, low intent
Drop from the list
Low fit, high intent
Watch, but fit gates the invite
Low intent / stale timingHigh intent / fresh signal
A perfect-fit account with no active signal is not the same as a perfect-fit account near a budget cycle; keep the axes separate.

Why this matters in numbers: trigger presence roughly doubles win rate, at 37% versus 19% for cold outreach, and intent-prioritized accounts have closed at 21.3% versus 8.4% for firmographic-only, a 2.5x difference. Yet only 5 to 15% of your TAM is actively buying at any given time. A look-alike list without a timing gate is statistically about 85% mistimed even when fit is perfect. That is the single strongest argument for the timing axis.

Where this goes wrong

The failure modes below are the reason look-alike lists disappoint even when the tooling works. Each has a false positive that looks fine on inspection, so each needs an explicit check rather than a glance.

  • Weak seeds. Ranking by deal size alone smuggles in one-time whales. The false positive is a "premium" seed that churned in six months and still shapes your pattern. Check: re-rank by net revenue retention, not signature deal size.
  • Firmographic-only similarity. The list looks right on size and industry but behaves wrong, because two firms of the same size sit in different buying situations. Check: require a technographic or behavioral attribute on every row.
  • Similarity mistaken for intent. A perfect-fit account that just renewed with a competitor scores identical to one near a budget cycle. Check: score fit and timing on separate axes and never let a blended score hide a zero-intent account.
  • Single-signal chasing and signal noise. Most intent stacks capture thousands of signals a week and most are worthless; a content download inflates a score without predicting a meeting. Check: cut any signal category that fails a controlled meeting-book test.
  • Unenriched list handed to reps. Stale records degrade the model at the source and inflate bounces; a catch-all address verifies as valid and still bounces. Check: verify emails and confirm hard bounce inside the 0.5 to 2% caution range before sending.
  • Name-only dedupe. "Acme Consulting" and "Acme Manufacturing" merge, or a customer slips back onto the cold list. Check: root domain is the primary key, name is secondary.
  • Suppression bypass. A rep re-adds a suppressed or existing-customer account to a new list and it sequences. Check: enforce precedence so permanent suppression overrides campaign inclusion.
  • Stale "zombie" Tier 1. An account scored Tier 1 in January looks identical in April after the VP left and hiring froze. Check: re-score weekly for intent, refresh firmographics quarterly.

The through-line is decay. At 22 to 25% annual data decay, plus 30% of B2B firmographic data going stale each year, the largest avoidable driver of a bad list is not matching quality, it is age. The "stop at an unenriched CSV" mistake compounds faster than any of the others.

The procedure, start to finish

Run the eight stages below in order. Time estimates assume a mid-sized closed-won cohort and one RevOps operator or founder doing the work.

Build the look-alike list

  1. Select and rank seeds
    Pull closed-won from the last 12 to 24 months and rank by revenue, net revenue retention, expansion and cycle length. Take the top decile to quintile, 5 to 20 named accounts, and remove churned or bad-fit accounts. (1 to 2 hours)
  2. Enrich seeds on every dimension
    Append firmographic, technographic, hiring, funding and who-signed data as it stood at the time of purchase. Every seed ends with a full attribute row. (1 to 2 hours)
  3. Extract the win pattern
    Find attributes far more common in wins than losses across firmographic, technographic, behavioral and buying-role categories. Write a searchable ICP with 5 to 7 named attributes. (2 to 4 hours)
  4. Generate net-new candidates
    Run similarity against seed domains or filter a fresh source with the ICP, so the candidate universe is larger than the seed set and each row is tied to a match reason. (minutes to hours)
  5. Deduplicate and suppress
    Match on root domain then name; remove existing customers, open opportunities, already-worked and do-not-contact records; apply precedence so permanent suppression wins. (1 to 2 hours)
  6. Enrich contacts and verify
    Find the buyer persona at each account and verify work emails before send, flagging bounce risk. (hours)
  7. Score fit, signal and timing, then gate
    Score fit and intent as separate axes, layer timing, and apply a threshold before an account enters a sequence. (1 to 2 hours)
  8. Refresh on cadence
    Re-score weekly for intent and refresh firmographics quarterly, treating the list as decaying rather than fixed. (ongoing)

Here is a match-reason vocabulary you can copy and adapt so every row states why it belongs. Reasons force discipline: if you cannot write one, the row should not ship.

Match-reason row schema
account_domain | persona_title | verified_email | fit_score | intent_tier | timing_days_since_signal | match_reason
example.com | VP Revenue Operations | yes | 8/10 | Tier 1 (raised Series B) | 12 | HubSpot stack + first RevOps hire in last 6mo, matches top-decile win pattern

One row per account. Fill the reason from your extracted ICP attributes; leave no reason blank.

Keep the list current

A look-alike list is a decaying asset, not a deliverable you finish. Re-score weekly for intent and refresh firmographics quarterly, because the signals that made an account worth working fade on a clock.

The decay is measurable and fast. Funding signals decay 50% every 60 days, topic surge 50% every 30 days, and response to trigger events drops by 80% after five days. A Tier 1 account from January can look identical in April on paper while the VP who signed has left and hiring has frozen. Weekly intent re-scoring catches that; a quarterly firmographic refresh catches the slower drift, since roughly 30% of firmographic data goes stale each year.

Before you hand the list to a rep

  • Seeds are ranked by net revenue retention, not deal size, with churn removed
  • Every row carries a technographic or behavioral match reason, not just size and industry
  • The list is deduplicated on root domain first, company name second
  • Current customers, open opportunities and do-not-contact records are suppressed with precedence enforced
  • Work emails are verified with projected hard bounce below the 2% critical threshold
  • Fit and intent are scored on separate axes, with timing layered and a threshold applied
  • A weekly intent re-score and quarterly firmographic refresh are scheduled

One last discipline. When you build multi-channel into the follow-up, remember email plus LinkedIn plus phone produces about 40% higher engagement than single-channel, but that lift only pays off on a list where every row already has a verified contact and a defensible reason. The reason column is the spine of the whole exercise. If you keep it honest, the list stays useful long after the first send.

Questions practitioners ask

How many closed-won accounts do I need to build a lookalike list?

It depends on your method. For a human-readable pattern you can read and search against, 5 to 20 ranked closed-won accounts is enough, and most lookalike workflows accept 5 to 100 seed domains. For statistical win/loss modeling where software finds the patterns for you, the documented floor is roughly 100 closed-won accounts, because below that there is not enough signal for meaningful patterns. Pick the method your deal count actually permits.

How is a lookalike list different from an ICP filter?

An ICP filter is usually firmographics: industry code plus a headcount band. A lookalike list starts from the accounts that already bought and extracts the attributes that predicted the win, which almost always includes technographic and buying-role signals firmographics miss. The test is whether every row carries a defensible match reason. If a row's only reason is size and industry, it is a filter output, not a lookalike.

Why should I score fit and intent separately instead of one blended number?

Because fit and intent are independent. Fit answers whether a company could buy from you; intent answers whether it is buying now. A perfect-fit account that just renewed with a competitor and one three weeks from a budget cycle can produce the identical blended score, which sends a rep to work a dead-fit account. Score fit and timing on separate axes so a zero-intent account can never be hidden inside an average.

What is an acceptable bounce rate for a lookalike outreach list?

Below 0.5% hard bounce is acceptable for most B2B email programs. Between 0.5% and 2% is a caution range, and above 2% is critical, because bulk sender requirements flag persistent hard bounce rates above 2%. Since B2B contact data decays at roughly 22 to 25% a year, verify emails immediately before send rather than trusting an enrichment done weeks earlier.

How do I keep a suppressed or existing customer from slipping onto the cold list?

Deduplicate on root domain as the primary key and company name only as a secondary signal, so 'Acme Consulting' and 'Acme Manufacturing' do not falsely merge. Then enforce a precedence rule: broad, permanent suppression overrides any campaign inclusion. A rep should not be able to bypass a do-not-contact flag or a current-customer flag by adding the account to a new list.

How often does a lookalike account list need refreshing?

Treat it as decaying, not fixed. Re-score weekly for intent, because trigger signals fade fast: funding decays 50% every 60 days and topic surge 50% every 30 days, and response to trigger events drops by 80% after five days. Refresh firmographics quarterly, since roughly 30% of B2B firmographic data goes stale each year. A one-off CSV is worth less every week you leave it alone.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next