Refolk
PlaybookProcess, data, and compliance

The Enrichment Waterfall Playbook: Provider Order and Blended Cost

You will be able to order an enrichment cascade by cost per valid result, gate out dead data, protect good records on write-back, and score each source monthly.

18 min readLast reviewed October 2, 2026Read as Markdown

Key takeaways

  • A waterfall costs the path each record walks, not the sum of its providers, so ordering the cheapest accurate source first can halve spend on the expensive specialist tier.
  • Full email verification with an SMTP mailbox check catches 95 to 99 percent of bad addresses versus 70 to 90 percent for syntax-only, which is what forces a dead result to cascade instead of counting as a match.
  • Incremental coverage collapses from 15 to 25 percent at the second provider to 3 to 5 percent at the fourth, so three sources is the practical optimum and the fourth rarely pays for itself.
  • A documented waterfall reaches an effective cost of $0.16 to $0.47 per valid record against $0.15 to $1.00 for single-source, with match rates moving from roughly 45 to 50 percent to over 80 percent.
  • Replace a phone source when connect-or-wrong-person outcomes fall below about 30 percent, because advertised 95 percent accuracy usually means line validity, not person ownership.
  • In Refolk's index there are 1,446 RevOps professionals in the US and 250 in the UK, so this cascade is usually owned by one person and the scorecard must stay light enough for one person to run.

Chaining several data sources so you fill the most contact and company fields at the lowest cost is a sequencing problem, not a shopping problem. This playbook is for the RevOps or data owner answerable for how the data was gathered, and it carries the ordering math, the validation-gate logic, the write-back rules, and the monthly scorecard in one place. Work it start to finish and you will have a cascade that cannot let dead data count as a match or let automated writes overwrite good records.

Every page ranking for this job is written by a data vendor and ordered around its own product. This one names no tools. It tells you how to decide the order, where to put the gate, and when to fire a source.

What a waterfall enrichment setup actually is

A waterfall is a cascade of data providers where a record stops at the first confident result and is billed only for the lookup that returned it. The expensive providers near the end only ever see records that every cheaper source above them failed to find.

That one mechanic changes everything about cost. A waterfall does not cost the sum of its providers. It costs the path each record walks. The first provider sees every record and consumes the most credits. The last sees only the leftovers. This is why single-source match rates of roughly 35 to 50 percent for contacts give way to 80 percent and above once you sequence several sources, and why independent testing of 1,000 B2B records found single-source email matching at 62 percent against 98 percent for a waterfall.

98%
Email match rate from a waterfall vs 62% single-source
From an independent 2026 test of 1,000 B2B records; the lift comes from sequencing, not from any one provider.

The job splits into four decisions you make once and one you repeat monthly: the order of sources, the gate between them, the stopping rule, and the write-back policy, then the scorecard that tells you which source to swap. Get those right and "multi-source contact enrichment" stops being a credit sink and becomes a predictable cost per valid record.

How to order data providers by cost per valid result

Order the cheapest accurate provider first and the premium specialist last, because the expensive tier only processes records everyone above it missed. This is a budget lever, not a quality lever, and it is the single highest-leverage decision in the whole setup.

The worked logic is blunt. Run your list through a low-cost source at roughly $0.01 to $0.03 per record before a premium source at $0.15 to $0.50 per record. The cheap source matches 50 to 60 percent of records at a fraction of the cost, and the premium source then only processes the 40 to 50 percent that remain, cutting its spend in half. You pay the high per-record rate on half as many records.

Records reaching each tier of a three-source cascade

  1. Tier 1 (cheapest)
    10,000

    sees every record

  2. Tier 2 (mid)
    4,000

    only tier-1 misses

  3. Tier 3 (specialist)
    1,500

    only tier-1 and tier-2 misses

Each tier sees only what the tier above it missed, so credit consumption falls sharply down the chain.

There is a real dissent here worth stating plainly. One camp argues accuracy-first, because the first provider to return a value wins and populates the field. Lead with a cheap-but-weak source and your best records come from your worst vendor. The reconciliation most practitioners reach is to order cheapest-first for cost but run a single terminal verifier at the end, so quality does not depend on who answered first. You decouple the two decisions: cost scales with how many records reach each tier, and quality scales with the verifier.

How many providers to chain

Three is the practical optimum. Incremental coverage collapses as you add sources, so the stopping rule is set by the diminishing-returns curve, not by vendor count.

PositionIncremental liftCumulative (starting 45%)
Provider 1~45%45%
Provider 215-25%~60-70%
Provider 38-12%~70-80%
Provider 43-5%~75-83%

The second provider recovers a meaningful chunk. The third still earns its place. The fourth adds 3 to 5 percent, which for most lists is not worth the per-record cost it adds to your blended rate. Add a fourth source only when a record is worth enough that any additional coverage pays, and even then, cap it.

The blended cost math: what you should actually pay

A well-built three-source waterfall lands at an effective cost of $0.16 to $0.47 per valid record, which is the lowest cost per valid record for most use cases. The number that matters is cost per valid record, not cost per query, because queries that return dead data are pure waste.

MethodCost per query% returning dataEffective cost per valid record
Manual research$5-15n/a$5-15
Single-source API$0.10-0.5050-70%$0.15-1.00
Waterfall$0.15-0.4085-95%$0.16-0.47

Read the right-hand column, not the left. A single-source API looks cheap per query but can cost up to $1.00 per valid record once you account for the half of queries that come back empty. A documented three-tier cost-per-enriched-record benchmark sits at $0.05 to $0.15 when the cascade is tight. Manual hand-verification is the floor of quality and the ceiling of cost: one vendor test took 143 hours to verify 10,000 contacts at 91 percent accuracy. You are automating to avoid that, not to match it exactly.

$0.16-0.47
Effective cost per valid record from a waterfall
Against $0.15 to $1.00 single-source. The spread is driven almost entirely by how many queries return usable data.

The validation gate that forces a dead result to cascade

The gate between tiers is email verification, and its job is to treat a syntactically valid but undeliverable result as a non-match so the record keeps cascading. Without the gate, a pattern-guessed address that passes syntax counts as a filled field and your coverage number lies to you.

Full email verification runs four checks and returns a status in one to three seconds:

  • RFC 5322 syntax confirms the address is well-formed. Proves nothing about deliverability on its own.
  • Domain MX records confirm the domain can receive mail. Still says nothing about the specific mailbox.
  • SMTP handshake connects to the mail server and asks whether that exact mailbox exists. This is the check that forces a cascade.
  • Catch-all detection flags domains that accept any address, which makes the SMTP check falsely succeed.

Full verification catches 95 to 99 percent of bad addresses, against 70 to 90 percent for syntax-only. The difference is the SMTP mailbox-existence check. The credit-saving variant is to verify each returned email and, when a provider sends an invalid one, charge no credit and continue the cascade until a valid address is found.

For phone, the equivalent gate is a line and ownership check, and the honest verdict is far messier than email. More on that in the scorecard section.

Write-back governance: the rules that keep the cascade from corrupting the CRM

Default every field to fill-empty-only, protect any value that is both verified and recently checked, and stage incoming data to diff it before writing anything live. The governance layer, not the provider, is where most CRMs get corrupted, so this is where a careless hour does the most damage.

The documented standard is a field-level policy matrix with four states. Assign every enriched field to exactly one of them:

StateWhat it doesUse for
OverwriteExternal source replaces CRM valueFields where the vendor is more authoritative and freshness matters
Fill onlyWrites when blank, never replacesDefault for most contact fields
FlagSurfaces the difference for a humanHigh-stakes fields where disagreement needs judgement
NeverCRM is the sole authorityIdentity and match keys like email and website

The concrete protection rule to encode first: do not overwrite if the existing email status is verified and it was last verified within 90 days. Layer confidence-tiered write-back on top: high-confidence matches write automatically when all protection rules pass, medium-confidence matches go to a review queue without overwriting, and low-confidence matches make no change and log the failed match. Records that do not match are left untouched, and identity fields used for matching are never overwritten.

The payoff is measurable. In one test of 1,200 accounts, verified email coverage moved from 58 percent to 71 percent after adding a second provider with do-not-overwrite rules, which also halved the bounce rate. The win came from the governance logic, not only from the extra vendor.

The governance layer, not the provider, is where most CRMs quietly get corrupted. </pull> Staging matters because it lets you diff against what you already hold rather than writing straight over live fields. Refresh into a staging layer, apply the field-level rules, and never let an unverified incoming value replace a verified existing one. If you want the suppression and governance logic handled in plain English rather than hand-built rules, [Refolk](/) lets you describe the people you want and keeps the identity keys stable across GitHub, LinkedIn, and the open web so your match layer stays honest. ## The step-by-step waterfall enrichment setup Run these eight steps in order. The first five you do once; the last three are the operating rhythm. Time estimates assume a single RevOps owner, which is realistic given how scarce the role is.

steps title: Build and run the cascade step: Run a 500-record baseline test :: Run the same 500-record ICP sample through two or three candidate providers independently and record which records each found that the others missed. Done means you know per-provider match rate and overlap on your own list, not a vendor's. step: Order the cascade by cost per valid result :: Place your cheapest accurate provider first and your premium specialist last, so the expensive tier only sees records everyone above it missed. Done means a documented sequence with per-tier cost per record written down. step: Insert the validation gate between tiers :: Route every returned email through SMTP mailbox and catch-all verification, and treat any non-valid verdict as a miss that cascades to the next tier. Done means an undeliverable result triggers the next provider instead of stopping the record. step: Set stopping and cost-cap rules :: Set a per-record cost cap and a tier limit so a low-value record stops after two providers rather than running the full chain. Done means no record can exceed its own enrichment budget. step: Configure write-back governance :: Build the field-level policy matrix of Overwrite, Fill-only, Flag, and Never, stage incoming data and diff it before writing, and protect verified-and-recent fields. Done means a test run writes nothing to live fields until every protection rule passes. step: Run the full list and sample-verify :: Execute the cascade across the target list, then run a random sample back through email verification to measure true accuracy. Done means coverage lift, blended cost per valid record, and bounce rate are all recorded. step: Run the monthly scorecard and swap :: Track match rate, cost, ROI, and validation pass rate per provider, then demote or replace any source that misses its threshold. Done means one keep-or-replace decision recorded for every source. step: Schedule re-runs by field :: Set a quarterly bulk refresh as the floor and layer trigger-based refreshes on job-change and engagement signals for high-value accounts. Done means refresh cadence is tuned per field, with titles and emails checked most often.


On the cost cap in step four: set a per-record ceiling so a record that is not worth more than $0.30 to enrich stops the waterfall after two providers instead of four. The cap and the three-source optimum work together to keep your blended rate flat as the list grows.

The path one record walks through the cascade

  1. Tier 1 lookup
    Cheapest source attempts a match
  2. Validation gate
    SMTP and catch-all check; valid stops, dead cascades
  3. Tier 2 or 3 lookup
    Only runs if the gate rejected the prior result
  4. Terminal verifier
    Every accepted value re-checked before write
  5. Write-back policy
    Fill-only, flag, or hold per field matrix
A record only advances when the gate rejects what the current tier returned, so spend tracks failures, not lookups.

How this goes wrong: the failure modes

The cascade fails quietly. Coverage looks fine on the dashboard while the underlying data rots. These are the specific ways it happens and the check that catches each.

  • First match is not the correct match. Sequential waterfalls stop at the first returned value, not the freshest. A cheaper vendor queried first can populate a field even when a better provider would have returned better data. Check: run every accepted value through one terminal verifier.
  • A dead email counts as a win. A pattern-guessed address passes syntax but bounces. The false positive looks exactly like a filled field. Check: require SMTP mailbox existence, not syntax or MX alone.
  • Catch-all domains mask failure. They accept any address, so the SMTP check reports success. Check: three-probe catch-all detection; flag as risky and cascade.
  • Silent overwrite of rep-entered data. A rep spends 20 minutes finding a direct number, enters it, and enrichment replaces it with the company switchboard. Check: fill-empty-only plus the 90-day verified lock.
  • Multiple writers on one field. A sales tool writes an industry, marketing automation replaces it, enrichment changes it again. Reports drift and nobody sees a failure. Check: one authoritative writer per field.
  • Fake coverage from shared databases. Three providers licensing the same stale underlying database add no real coverage no matter how you configure the order. Check: test independence by overlap of records found in the baseline step.
  • Trusting vendor confidence scores. A vendor's score may not reflect your field definitions or risk tolerance. Check: validate against your own verified sample.
  • "95 percent accurate" means line-valid, not person-valid. A vendor advertising 95 percent accurate phone data almost always means the line is valid, not that the person owns it. Check: measure connect-or-wrong-person rate and replace below 30 percent.

The monthly per-source scorecard that says which source to replace

Track four metrics per provider every month and hold each source to a numeric threshold: match rate, cost, ROI measured as data found divided by cost, and validation pass rate measured as the portion of found data that verifies. Keep it light, because this is usually run by one person.

MetricWhat it measuresReplace-or-demote trigger
Position-one match rateShare of list the lead source clearsNo longer clears the bulk at top of chain
Phone connect rateConnect-or-wrong-person outcomesBelow ~30%
Email bounce rateAccepted emails that bounceAbove 2%, or above 0.3% for Gmail/Yahoo
ROIData found / costFalls below peer sources at same position

The thresholds are not arbitrary. Phone-verified mobile providers typically show 40 to 60 percent connect rates on cold calls, so a source producing fewer than 30 percent connected-or-wrong-person outcomes is likely pulling from outdated data and should be pulled. For email, bounce above 2 percent is the ceiling, but under current Gmail and Yahoo bulk-sender rules, bounce above 0.3 percent can hurt your sender reputation and push mail to spam, so treat 0.3 percent as the real bar for high-volume sending.

The position-one match rate is the key ordering KPI. A tier-one source that no longer clears the bulk of your list has lost the economic argument for sitting first, since everything below it is now doing the work. Demote it, or replace it.

This is a one-person job for a reason. In Refolk's index there are 1,446 current RevOps professionals in the United States and only 250 in the United Kingdom, a 5.8x gap. For context, there are 3,587 Rust engineers in the US, 2.48x the number of RevOps professionals who run cascades like this.

MarketRevOps professionalsIndex vs other market
United States1,4465.8x UK
United Kingdom250baseline
US Rust engineers (context)3,5872.48x US RevOps

Because the role is scarce and concentrated, the scorecard above has to be something one owner runs in an hour or two a month, not a committee deliverable.

Keeping the cascade current: re-run cadence by field

Set a quarterly bulk refresh as the floor and layer trigger-based refreshes on top, because data decays continuously and a list verified once is out of spec within weeks. B2B contact data decays about 2.1 percent per month, compounding to roughly 22.5 percent a year, and email decays faster at about 3.6 percent per month.

Refresh by field rather than on a single global schedule. Titles and emails decay fastest and justify frequent checks on records due for outreach, while industry, domain, and location change rarely and can ride the quarterly bulk. The acceptable bounce bar is tightening faster than data quality is improving, so there is no such thing as a one-time clean: the cadence is the hygiene.

Here is the pre-publish check for the cascade itself, before you point it at a live list.

Before you run the cascade on live data

  • Baseline test on 500 records shows real per-provider overlap, not inherited shared-database coverage
  • Sources are ordered cheapest-accurate-first with per-tier cost documented
  • A terminal verifier runs on every accepted value regardless of source
  • The validation gate uses SMTP mailbox existence, and catch-all addresses are flagged risky and cascaded
  • A per-record cost cap and tier limit are set so low-value records stop early
  • Every enriched field is assigned exactly one of Overwrite, Fill-only, Flag, or Never
  • Verified-and-recent values are locked, and identity keys are set to Never overwrite
  • A staging-and-diff pass confirms nothing writes to live fields until policies pass
  • The monthly scorecard has thresholds set for match rate, bounce, phone connect, and ROI
  • Re-run cadence is scheduled by field, quarterly bulk minimum plus triggers on high-value accounts
Enrichment waterfall configuration record
CASCADE: [name / ICP segment]
TIER 1: [source] | cost/record: $___ | baseline match %: ___
TIER 2: [source] | cost/record: $___ | incremental lift %: ___
TIER 3: [source] | cost/record: $___ | incremental lift %: ___
TERMINAL VERIFIER: [verifier] | catches % bad: ___
GATE RULE: accept email only if SMTP=valid AND catch-all=false
COST CAP: stop after tier ___ if record value < $___
WRITE-BACK: default = Fill-only | locked if verified AND last_verified <= 90d
IDENTITY FIELDS (Never overwrite): email, website, [add]
RE-RUN: bulk = quarterly | triggers = job-change, engagement
SCORECARD TRIGGERS: phone connect < 30% | bounce > 2% (0.3% bulk) | pos-1 match below bulk

Fill one per cascade and revisit it at the monthly scorecard. Keep it in the same place as the CRM admin notes.

Run this once, keep the configuration record open, and the only recurring decision is the monthly keep-or-replace call on each source. That is the whole job: order by cost, gate out the dead data, protect the good records, and let the scorecard tell you when a source has stopped earning its place.

Questions practitioners ask

Should I order enrichment providers cheapest-first or most-accurate-first?

Order cheapest accurate provider first for cost, because only unmatched records cascade to the expensive tiers, which can halve spend on a premium source. The accuracy risk that cheapest-first introduces, where a weaker source wins the first match, is solved not by reordering but by running every accepted value through one terminal verifier at the end of the chain. That decouples the cost decision from the quality decision.

How many data providers should a waterfall have?

Three is the practical optimum for most lists. The second provider recovers 15 to 25 percent of records the first missed, the third adds 8 to 12 percent, and the fourth adds only 3 to 5 percent. Because blended cost per valid record rises once lift collapses that far, a fourth source rarely pays for itself unless you are enriching very high-value accounts where any additional coverage justifies the cost.

What counts as a valid match versus a dead result in enrichment?

A valid email match passes full verification: RFC 5322 syntax, domain MX records, an SMTP mailbox-existence handshake, and catch-all detection. A syntactically correct but undeliverable address is a dead result, not a match, and must cascade to the next provider. Full verification catches 95 to 99 percent of bad addresses against 70 to 90 percent for syntax-only, so skipping the SMTP step lets dead data count as a win.

How do I stop enrichment from overwriting good CRM data?

Default every field to fill-empty-only, lock any value whose status is verified and whose last-verified date is within 90 days, and never overwrite identity or match keys like email and website. Stage incoming data in a separate layer and diff it against what you hold before writing anything live. Use a four-state policy matrix per field: Overwrite, Fill-only, Flag for human review, or Never.

When should I replace a data source in the waterfall?

Replace a phone source when connect-or-wrong-person outcomes fall below about 30 percent, since that signals outdated underlying data regardless of advertised accuracy. Replace or demote an email source if accepted addresses push bounce above 2 percent, or above 0.3 percent under Gmail and Yahoo bulk-sender rules. For position-one sources, demote any that no longer clears the bulk of your list at the top of the cascade.

How often should I re-run enrichment on my list?

Run a quarterly bulk refresh as the minimum, because B2B contact data decays about 2.1 percent per month, roughly 22.5 percent a year. Email decays faster, around 3.6 percent per month, so verify addresses close to the point of outreach. Refresh by field rather than on one global schedule: titles and emails change fastest, while industry, domain, and location change rarely and need far less attention.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next