The Data-Source Onboarding Playbook: Sample Test to Verified Backfill
You can run a new data feed from a locked sample test through a staged, segment-by-segment backfill with a signed-off match rate, field-ownership map, and dedup rules before cutover.
Key takeaways
- The average single enrichment vendor matches about 49% of companies and 56% of contacts, while a waterfall across three or four providers reaches 80 to 95%, so a source built on one feed silently runs on half your data.
- A match only means the vendor holds a record, not that it is correct, so you must hand-verify a random 50 against a primary source and report accuracy separately from match rate.
- Salesforce merge is irreversible, skips the Recycle Bin, keeps Chatter from the principal record only, and inherits the oldest Created Date, so export a full backup and dry-run every merge in a sandbox first.
- Blended match rate is the number that hides the failure: a stack can score 85% overall while missing an entire region or function, so gate on segment-level coverage lift, not the headline figure.
- For EU and UK personal data you cannot write a defensible Legitimate Interest Assessment without a per-record provenance document, and Article 14 forces you to name the source at first contact.
- In Refolk's index only 13 US profiles pairing a data-operations title with a Data Quality skill exist, so the person who owns this playbook is rarer than the tooling suggests.
You bought a new people-and-company data feed and now have to bring it into your system of record without creating duplicates, overwriting good fields, or importing records you cannot lawfully use. This playbook is for the RevOps, recruiting-ops, or data owner answerable for how the data was gathered. It runs the new source from a controlled sample test through a staged, segment-by-segment backfill, with a written field-ownership map, dedup rules, and a measured match-rate and coverage lift signed off before you touch a live record.
Every other operations guide handles data that already lives in your system: refreshing it, grading it, merging an acquired book. This is the distinct one-time job of standing up a brand-new external feed from first sample to full cutover. Do it in order and you never improvise over production data.
Why onboarding a new feed is its own job, not just an import
Onboarding a new data source is a gated procedure that ends in a signed go or no-go, not a file upload. The reason it needs its own method is that a feed's quality is unknowable from vendor marketing and only becomes visible when you test it against your own list and measure two numbers that trade off against each other.
Practitioners agree on this point: the only trustworthy test is a bake-off on your own records. Run the same 1,000-record list through five vendors and you get five different winners, because no provider tops both quality and coverage at once. Broad databases carry stale records, which lifts match rate and drops accuracy. Verified databases drop coverage to keep accuracy high. That trade-off is by design, which is why the sign-off gate must measure match rate and accuracy separately and never as a composite.
The other reason this is a real job is that the damage from doing it casually is irreversible. A Salesforce merge cannot be undone, the deleted records skip the Recycle Bin, and the merged record keeps Chatter from the principal only and the Created Date of the oldest record. A backfill that "worked" can still erase engagement history. So the work is front-loaded into tests and dry runs, and production is touched last, one segment at a time.
The onboarding stack, outermost gate first
- Compliance and provenanceDPA, per-record source, transfer mechanism, LIA, suppression cross-check
- Field ownershipper-field survivorship and null-handling rules
- Stagingvalidation against a permanent intermediate table, never production
- Dedup and mergenormalize keys, dry-run in show-duplicates mode, back up first
- Segmented backfillone segment cut over, measured, then widened
What "good" looks like: the numbers that clear each gate
A good result is a measured before-and-after on your own sample, not a vendor benchmark. There is no universal pass line to borrow, because a truly independent, methodology-disclosed third-party test of enrichment match rates barely exists. You set thresholds per source against your baseline.
Use published benchmarks only as reference points for what is achievable. The gap between them is the whole argument for staging carefully.
| Approach | Company match | Contact match |
|---|---|---|
| Single vendor (avg) | ~49% | ~56% |
| Waterfall (best case) | up to 94% | up to 83% |
| Single-source (range) | 50-70% blended | 50-70% blended |
| Waterfall (range) | 80-95% blended | 80-95% blended |
The lift from about 49% to up to 94% company match is the difference between scoring half your database and nearly all of it. Models, routing, and scoring built on a single feed silently run on half the data. That is the case for a waterfall - an enrichment stack that cascades unmatched records from one provider to the next and stops at the first match to save credits. But the range is blended, and the blended number is exactly where the failure hides, which the failure-modes section covers in full.
Measure five numbers at the gate and treat each as independent: match rate by record type, accuracy from a hand-verified sub-sample, coverage lift by segment, cost per verified record, and duplicate rate. The duplicate rate matters because more than 25% of records in the average database already contain duplicate data, so you need a pre-import baseline to detect regression.
Build the control sample before you talk price
The control sample is the single most important artifact in this playbook, and it must over-weight the worst part of your database. Vendors are strongest on clean, US, mid-market records, so a sample skewed toward those flatters everyone and tells you nothing.
Pull 100 to 1,000 records, sized to your patience and how many vendors you are comparing. The composition is what counts. If you want to enrich only leads missing phone number and job title, that scope is your baseline: pull a sample of leads that have no phone and no title. Then add a known-truth seed set of your own colleagues, customers, and contacts whose accuracy you can confirm by hand. Without the seed set you can measure match rate but not accuracy.
Lock the CSV so every vendor gets the identical file, and document the critical fields you are scoring. Freezing the file is what makes the comparison valid.
Run the procedure in order
Run these eight stages in sequence. Compliance clearance can run parallel to the bake-off, but staging, dedup, and backfill are strictly ordered because each depends on the artifact the previous stage produces.
Sample test to verified backfill
- Scope and build the control samplePull 100 to 1,000 records skewed toward the messy segment you want to fix, plus a known-truth seed set of your own colleagues and customers. Done: a locked CSV and a documented list of critical fields.
- Run a blind bake-offSend the identical file to each shortlisted vendor with the free-test requirement in the RFP. Done: returned files scored on match rate, accuracy from a hand-checked 50, and cost per verified record.
- Clear compliance and provenanceCollect DPA, provenance doc, hosting location, and transfer mechanism, write the LIA, and cross-reference the vendor list against your suppression list. Done: signed DPA, dated LIA, and suppression cross-check log.
- Write the field-ownership mapAssign per-field ownership as source-wins, preserve-manual, or most-recent-verified-wins, with a conflict rule and null-handling per field. Done: a written field map naming winner logic and null behaviour.
- Stage the dataLoad into a staging table or sandbox with batch-ID, source-file, and processed metadata, then run schema, format, date, picklist, and dedup validation, flagging failures. Done: validation report with zero blocking errors on a sample review.
- Dry-run dedup and design the mergeNormalize phone and email, configure matching keys on name, email, and company, and run in show-duplicates mode. Back up before any merge. Done: duplicate report reviewed, history-preserving merge rules confirmed in sandbox.
- Backfill one segment at a timeCut over a single segment first, measure before and after, then widen. Confirm the vendor supports a phased rollout. Done: segment loaded with before and after metrics captured.
- Run the sign-off gate or roll backCompare measured match rate, coverage lift, duplicate rate, and verified accuracy against pre-set thresholds. Done: a written go or no-go with the numbers attached and a rollback plan for any missed threshold.
A tight version of the bake-off fits in three days: day one, pull 500 real records; day two, run them through two free tiers; day three, score match rate and per-field fill rate. Insert a vendor match-rate report step first if you want their self-reported number to compare against your measured one.
Score the bake-off: match, accuracy, and cost per verified record
Score three separate numbers, because a match on its own tells you almost nothing. A match only means the vendor holds a record. It does not mean the record is any good.
From returned records to verified records
- 500Sample sent
identical file to every vendor
- 280Records matched
~56% contact match on a single vendor
- 250Fields filled
match does not guarantee the field you bought
- 200Hand-verified correct
random 50 checked against a primary source, extrapolated
The three numbers to record for every vendor:
- Match rate, by record type, so company and contact are scored separately rather than blended.
- Accuracy, from a random 50 records hand-checked against a primary source. This is the number that separates a database that holds a record from one that holds a correct record.
- Cost per verified record, which divides the price by the accurate matches, not the raw matches.
Check email specifically. A provider that returns an email without validating syntax, domain, and mailbox hands you bounces, so confirm the vendor verifies rather than merely appends. Phone is the weakest field across the board and where cheap APIs fall apart fastest, so if phone is what you are buying, weight the sample and the score toward it.
Clear compliance before the data lands, not after
Provenance is inherited liability, not a vendor checkbox. The moment you connect an enrichment feed, the vendor's compliance posture becomes yours, so compliance clearance is a gate that stands before the import, not a footnote after it.
For EU and UK personal data the load-bearing document is a Legitimate Interest Assessment, a documented evaluation of three tests - purpose, necessity, and balancing - that you complete before relying on legitimate interest as your lawful basis. It takes roughly two to three hours to produce and it cannot be written without knowing where each record came from. That is why you require a per-record provenance document from the vendor. Article 14 of GDPR forces you to inform any person whose data you obtained from a third-party source at the time of first contact, so "we do not know the source" is not a defensible answer to an access request.
Collect this evidence from the vendor before onboarding:
- A signed Data Processing Agreement.
- A clear data provenance document showing where records originate.
- Confirmation of hosting location.
- The cross-border transfer mechanism for any EU or UK records - whether the provider relies on Standard Contractual Clauses or an adequacy decision.
Two retention facts to encode in the load. ICO and CNIL set a maximum retention of three years from last contact for prospecting data, so stamp an acquisition date on every imported record. And GDPR data subject requests must be actioned within 30 days, extendable to 90, which only holds if you can trace a record back to its source. Refolk gives every returned record a stated origin across public LinkedIn, the public GitHub graph, and the open web, which is the provenance a legitimate interest assessment needs to stand up.
Map field ownership before you stage a single row
Assign survivorship per field, not per record, because a new feed will hold some fields better than yours and some worse. Modern master data management assigns ownership by field, not by record, and platforms that operationalize the golden record - the single reconciled version of a record - do it with field-level rules. The golden record is the version that survives conflict resolution.
Pick one winner rule for each field. The common inputs are source priority, completeness, recency, and verification status. Three rules cover most fields:
| Rule | When it wins | Null handling |
|---|---|---|
| Source-wins | The feed is authoritative for this field | Skip-null so blank source never clears production |
| Preserve-manual | A human authored the value | Manual value always overrides consolidated |
| Most-recent-verified-wins | Freshness matters and both sources are trusted | Only a verified update survives |
The rule that saves you is preserve-manual: manually authored values override consolidated ones, so a hand-corrected phone number is never clobbered by a bulk feed. The default in tools that document this is that manually added changes are prioritized, otherwise the most recent value wins. Add one safeguard on top: the most recent verified update survives, so a low-quality or stale source cannot beat a trusted value.
Overwrite mode has no way to express write NULL, so a blank source field silently erases a good production value.
Choose the staging merge mode that matches the source. Use upsert plus overwrite for a source sending a mix of new and changed records where empty source fields should not clear production values. Use overwrite-with-sentinel only when a user genuinely needs to blank a field, because default overwrite cannot express a deliberate NULL.
Stage, validate, and dry-run the merge
Everything runs against a staging table first, and no data touches production until it has passed. A staging table is a permanent intermediate table where raw data lands before it is validated, transformed, and loaded. Rows that fail validation get flagged with an error code so you can see exactly which checks failed rather than losing the batch.
Give the staging table metadata columns for batch ID, source file, and a processed flag, then run schema, format, date, currency, picklist, and dedup validation. The Salesforce and Zoho sandboxes mirror production for this, though note a Zoho constraint: records added in the Sandbox cannot be deployed to production, so use the sandbox to prove the logic, not to stage the real load.
Then design the merge carefully, because merge is where a clean onboarding quietly destroys value. Salesforce out-of-the-box matching considers name, email, and company for leads and contacts. Before any merge:
- Normalize phone and email so the matching keys actually match.
- Run in show-duplicates mode and review the duplicate report without writing anything.
- Export a full backup, because merge is irreversible and deleted records skip the Recycle Bin.
- Remember leads and contacts cannot be merged cross-object, so convert a lead to a contact before merging it with one.
- Know that the merged record inherits the oldest Created By and Created Date, and keeps Chatter from the principal record only.
The Salesforce UI merges up to three contacts at a time, so bulk cleanup needs tooling that preserves relationships and an audit trail rather than the native screen.
How this goes wrong: the failure modes to test against
The failure modes below are where a technically successful onboarding still leaves you worse off. Each has a false positive that looks like success and a specific check that catches it.
| Failure mode | Looks like success | Check that catches it |
|---|---|---|
| Flattering sample | 90%+ match in test | Build the sample from records missing the exact fields you buy |
| Match counted as accuracy | High fill rate, wrong values | Hand-verify 50 against a primary source |
| Email appended, never verified | Full inbox of enriched emails | Confirm syntax, domain, and mailbox verification |
| Blended rate hides a dead segment | 85% overall | Break match rate down by segment, not blended |
| Overwrite wipes manual data | Clean, fully populated fields | Skip-null and preserve-manual rules with a low-quality safeguard |
| Irreversible merge destroys history | Deduped, tidy record count | Export a backup and dry-run merges in sandbox |
| New records on match failure | Import "succeeded" | Normalize keys and dedup in staging before load |
| Missing provenance or suppression | A "clean" file | Per-record provenance and suppression cross-check before import |
The one that fools experienced teams is the blended match rate. A stack can score 85% overall while missing an entire region or function. Agencies defaulting a US-strong provider onto EU lists watch coverage fall from around 55% to about 30%, invisible behind the headline number. Gate on segment-level lift, which is why the backfill is staged one segment at a time: you measure each segment on its own before you widen.
Sign off, roll back, and keep the source honest
The sign-off gate is a written go or no-go with the numbers attached, and a rollback plan for any missed threshold. The definition of done is that the critical workflow inputs reach the required completeness and freshness threshold before your workflow engine acts on them.
Before you call the source live
- Sample was built from records missing the exact fields being bought, plus a known-truth seed set
- Match rate is reported by record type, never blended, and broken out per segment
- Accuracy comes from a random 50 hand-verified against a primary source
- Signed DPA, dated LIA, hosting location, and transfer mechanism are on file
- Suppression list was cross-referenced against the full file before any load
- Field-ownership map names a winner rule and null-handling for every field
- Data was validated in staging with zero blocking errors before touching production
- A full backup exists and merges were dry-run in sandbox before any write
- Coverage lift, duplicate rate, and cost per verified record are recorded against baseline
- A rollback plan is written for the segment being cut over
Keeping the source honest after cutover is its own discipline. Re-run the segment-level match and accuracy check on a fresh sample each quarter, because a feed that cleared the gate can drift as its underlying database ages. Watch the duplicate rate against your pre-import baseline as the first sign of a dedup rule that stopped working.
One last point on ownership. This competency is scarce even inside operations teams. In Refolk's index only 13 US profiles pairing a data-operations title with a Data Quality skill exist, so the person who owns this playbook is rarer than the tooling suggests.
| Market | RevOps Manager profiles | Index (US=100) |
|---|---|---|
| United States | 927 | 100 |
| United Kingdom | 168 | 18 |
The US pool is roughly 5.5 times the UK pool, and within the narrow US RevOps and marketing-ops slice, HubSpot appears about 4.3 times more often than Salesforce as a listed skill - directional given the small sample, but a reason to name the CRM this person will actually own before you hire. Make the data-quality ownership explicit in the role, because the field map, the survivorship rules, and the sign-off gate all depend on one accountable owner rather than a committee.
Questions practitioners ask
How big should my sample file be for testing an enrichment vendor?
Somewhere between 100 and 1,000 records, weighted toward the messy segment you are actually paying to fix rather than clean US mid-market accounts. If you are buying phone and job title for leads that lack both, your sample must be leads with no phone and no title. Add a known-truth seed of your own colleagues and customers so you can hand-check accuracy against values you already know.
What match rate should clear the sign-off gate?
There is no universal published pass line, because a truly independent, methodology-disclosed third-party test of enrichment match rates barely exists. Set the threshold per source against your own before-and-after baseline. As reference points, a single vendor averages about 49% company and 56% contact match, and a waterfall commonly reaches 80 to 95%. Gate on segment-level coverage lift and hand-verified accuracy, not a single blended number.
How do I backfill a CRM without creating duplicates?
Load into a staging table or sandbox first, never straight into production. Normalize phone and email, configure matching keys on name, email, and company, then run a dedup pass in show-duplicates mode before any write. Import-without-dedup spawns new records on match failure instead of updating existing ones, which is the most common way a backfill inflates a database instead of enriching it.
What stops a new feed from overwriting good manual data?
A per-field survivorship map. Assign each field a rule such as source-wins, preserve-manual, or most-recent-verified-wins, and use skip-null or upsert-plus-overwrite modes so empty source fields never clear production values. Add a safeguard so a low-quality or stale source cannot beat a trusted verified value. Modern MDM assigns ownership by field, not by record, for exactly this reason.
What compliance evidence do I need before importing a purchased list?
For EU and UK personal data, require a signed DPA, a per-record provenance document, the hosting location, and the cross-border transfer mechanism, then write a Legitimate Interest Assessment covering the purpose, necessity, and balancing tests. Article 14 forces you to name the data source at first contact. Cross-reference the whole file against your suppression list before any record enters a sequence.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.