The Sourced-Database Fitness Standard: Fit, Remediate, or Quarantine
You will grade any sourced list across six dimensions against fixed numeric thresholds and return one of three verdicts that two graders reach independently.
Key takeaways
- A sourced list is fit to use only when accuracy is 95% or higher, deal-critical field completion is 90% or higher, email validity is 93% or higher, and the duplicate rate is under 2%; anything worse than the middle band quarantines the list.
- The verdict must gate on the worst dimension, not the mean, because most B2B databases score 45 to 60 out of 100 on their first audit and a weighted average lets one strong dimension mask a fatal one.
- Email is its own dimension, not part of accuracy: email fields decay at about 3.6% per month, roughly 1.7 times faster than the 22.5% annual blended contact-record rate.
- A verifier saying 90% valid can still produce a 40% bounce because catch-all, role, and disposable addresses pass verification; bucket catch-all separately and never count it toward the 93% validity pass.
- Exact-key deduplication misses roughly 30 to 40% of real duplicates, so a reported 2% duplicate rate can actually be 10% or higher until you run fuzzy matching on a sample.
- Only about 22% of organisations ever hit a 1% duplicate rate while untreated databases run 10 to 30%, which is why 5% is the realistic quarantine line rather than 1%.
Before a freshly sourced or enriched list of people and companies goes to reps or into a campaign, someone has to decide whether the whole dataset is good enough to trust. This standard is for the RevOps and recruiting-operations owner who is answerable for that call. It gives you six dimensions, fixed numeric thresholds, and a checklist so that two people grading the same export land on the same verdict: fit to use, remediate, or quarantine.
Most operations guides stop at listing metrics. This one fixes the pass lines and the order of operations so the grade is reproducible and you can adopt it as team policy. The evidence here is practitioner benchmark data, not opinion, and where the number is genuinely not established I say so and tell you what to check locally.
What the six data-quality dimensions actually measure
The dominant model uses six measurable criteria that determine whether data is fit for business use: accuracy, completeness, consistency, timeliness, validity, and uniqueness. These are the axes you grade against, and each one answers a different question about the same export.
- Accuracy is the degree to which data represents real-world things, events, or an agreed source. A title of "VP Sales" is accurate only if that person still holds it.
- Completeness is the share of records with the fields you require populated. It is measured per field, not per record, because a missing email matters more than a missing website.
- Validity means the data conforms to the syntax, format, type, and range of its definition. An email that fails a format check is invalid before anyone even tries to send to it.
- Uniqueness counts distinct real-world entities against records. A database showing 520 records for 500 real people (Fred Smith and Freddy Smith stored twice) has a uniqueness of 500/520 x 100 = 96.2%.
- Timeliness is the recency of information and its availability for use, which is what decay erodes.
- Consistency is agreement between copies of the same fact across systems and fields.
The classification is not universally agreed. Some frameworks add currency, conformity, integrity, and precision, and ISO/IEC 25012 defines a comprehensive model on characteristics such as accuracy, completeness, and credibility. For a go/no-go grade on a sourced list, six is the right resolution: enough to catch the failure modes, few enough that two graders can hold the whole rubric in their heads.
The verdict thresholds every list gets graded against
A list is fit to use only when it clears the fit column on every dimension below; it quarantines the moment any single dimension enters the quarantine column, regardless of the others. This is the core of the standard, so it sits up front.
The pass lines come from converging practitioner benchmarks. Most B2B operations should target 95% accuracy or higher; 90 to 94% is acceptable but signals room to improve; below 90% is critical and forecasts become unreliable. A composite 2026 CRM benchmark holds that good data quality means 93% or higher email validity, under 5% duplicate rate, 80% or higher field completion, and under 2% bounce. The verdict bands below are my synthesis of those benchmarks into three actionable columns.
| Dimension | Fit to use | Remediate | Quarantine |
|---|---|---|---|
| Accuracy | 95%+ | 90-94% | <90% |
| Field completion (deal-critical) | 90%+ | 80-89% | <80% |
| Email validity | 93%+ | 85-92% | <85% |
| Duplicate rate | <2% | 2-5% | >5% |
| Hard bounce (proxy) | <2% | 2-3% | >3% |
Read the table as a gate, not an average. Field completion applies only to deal-critical fields you named in your rubric, because flagging a database on 80% completeness of a field nobody sequences on wastes remediation time. The academic threshold for an email-address field, for reference, marks it unacceptable below 80% complete; I hold deal-critical completion higher, at 90%, because a list that cannot be sequenced is not fit to hand a rep.
The composite score gives you a headline, and the bands are: above 90 is excellent, 80 to 90 acceptable with a documented improvement plan, and below 80 requires immediate remediation. But the composite is a summary, not the verdict. The verdict gates on the worst dimension.
Why the verdict gates on the worst dimension, not the mean
Score by dimension and by segment, then let the worst result decide, because averaging hides the exact failures that kill a campaign. A weighted composite is useful for tracking trend over time; it is dangerous as a go/no-go gate.
Three facts force this rule. First, a single job change invalidates three or four fields at once - title, email, and direct-dial together - so a dead record can still push a high field-completion average while being entirely unreachable. Second, if three or more dimensions land in poor or critical, you have an urgent hygiene problem, and an 85 composite can sit on top of one quarantine-grade segment. Third, email decays far faster than the rest of the record, so it must be scored as its own dimension and never folded into a blended accuracy figure.
Composite score versus worst dimension
The masked-failure quadrant, top-right, is the one this standard exists to catch. A list can average well and still carry a segment where email validity is 60%. Score by source and segment, and the failing slice reveals itself instead of hiding inside the mean.
A weighted average lets one strong dimension pay the ransom for a fatal one. Gate on the worst.
How fast the data decays, by field
Data quality is not a fixed property of a list; it erodes on a clock, and different fields erode at different speeds. This is why timeliness is a scored dimension and why the verdict has a shelf life. Plan against published decay rates rather than assuming a graded list stays graded.
| Field type | Published decay | Source basis |
|---|---|---|
| Email address | ~43%/yr (3.6%/mo) | Field-level email decay data |
| Contact record (blended) | 22.5%/yr | MarketingSherpa/HubSpot decay simulation |
| Job title / role | 25-35%/yr changed | Job-change data |
| Firmographic | 20-30%/yr | Dun & Bradstreet estimate |
The decisive comparison is email against the blended rate. An email field decaying 3.6% per month erodes roughly 1.7 times faster annually than the 22.5% blended contact rate. That gap is why a database can pass overall accuracy while the email column alone triggers quarantine, and why email gets its own row in the threshold table.
Timing, not tool choice, is the lever. A 90-day-old unverified list carries an estimated 6 to 9% invalid rate. At 40 emails per day, an 8% invalid list produces 3.2 hard bounces per inbox, above the 2% danger threshold that damages deliverability. Verify at the point of use, because decay is monthly and a grade from last quarter is already stale.
The scoring procedure, step by step
Run these eight steps in order. The first two are policy set once; the rest execute against every export. The point of fixing the order is that two graders who follow it produce the same verdict.
Grading a sourced database, freeze to verdict
- Freeze a snapshotExport the full segment to CSV with a fixed snapshot date so both graders work from identical rows. Done: a dated, immutable file.
- Define required fields and weightsFix deal-critical fields and dimension weights in writing before touching data. Deciding it live guarantees inconsistency. Done: a written rubric.
- Baseline the automatable dimensionsRun field completion and duplicate reports across all objects to set a current-state benchmark. Done: completeness, duplicate, and validity percentages per field.
- Sample for accuracyPull 100+ random records and check email, phone, title, and company against reality. Done: an accuracy percentage with the sample size stated.
- Verify emails before scoringRun the list through a verifier, treating catch-all and risky as separate buckets. Done: counts of valid, invalid, catch-all, and risky.
- Compute the composite scoreApply the weighted formula to combine sub-scores into one number. Done: a 0-100 score plus every per-dimension sub-score.
- Assign the verdict against fixed thresholdsMap the composite and each sub-dimension to fit, remediate, or quarantine, gating on the worst. Done: one verdict with written reasons.
- Set the re-verification cadenceTier records by value and territory and assign an interval to each. Done: cadence documented per segment.
For the computation itself, the generic score practitioners use is: Data Quality Score = [1 - (bad records / total records)] x 100. For a per-rule score, Microsoft Purview uses passed / (passed + failed + miscast + empty), which correctly treats empty and miscast records as failures rather than dropping them from the denominator. To combine dimensions into a single composite, weight them and take a weighted average.
Weights: 40 30 20 10 Scores: 88 93 91 72 Composite = SUMPRODUCT(Weights, Scores) / SUM(Weights) Example = (40*88 + 30*93 + 20*91 + 10*72) / 100 = 87.9 Verdict rule: Composite >= 90 AND every sub-dimension in "fit" -> FIT TO USE Composite 80-90 OR any sub-dimension in "remediate" -> REMEDIATE Composite < 80 OR any sub-dimension in "quarantine" -> QUARANTINE
Set the weights row to your rubric; the example uses 40/30/20/10 across completeness, accuracy, consistency, timeliness.
Note the verdict rule uses AND for fit and OR for the worse outcomes. That asymmetry is deliberate: fitness requires everything to clear, while a single failure is enough to demote. Keep the audit (finding and documenting problems) separate from the cleanup (fixing them). Keeping them separate gives you a reviewable work queue and a baseline you can measure the cleanup against.
The order has legitimate variants. Some workflows run the baseline before decay analysis; others score all dimensions in a single pass. Both are fine as long as audit stays separate from cleanup and the thresholds are fixed in advance.
Once the rubric and cadence are policy, the recurring cost is finding the records most likely to have moved since the list was built. That is precisely the query I would run to pressure-test a stale export.
For pulling the practitioners who own this audit, or the source experts who can pressure-test your catch-all and deliverability thresholds, Refolk resolves the roles directly - data quality analysts, RevOps managers on specific stacks, or founders of email verification startups - without you building and cleaning yet another list to audit the auditing.
How the grade goes wrong: failure modes and false positives
Every dimension has a way of reading better than it is, and a standard that ignores them produces confident wrong verdicts. This is the most valuable section to internalise: each item below is a false positive that makes a bad list look fit.
- Catch-all inflates valid. A catch-all server responds "accepted" even when the mailbox was never set up, and standard tools mark it valid. Check: bucket catch-all separately and never count it toward the 93% validity pass.
- Valid does not equal deliverable. A tool can report 90% valid, then the list bounces at 40%, because a valid rate can still include role accounts, disposable domains, and catch-alls. Check: send seed tests to a sample of valid addresses before trusting the score.
- Exact-match dedup understates duplicates. Exact-key matching alone misses roughly 30 to 40% of the duplicates that actually exist in a typical CRM. False positive: a 2% duplicate verdict that is really 10% or more. Check: run fuzzy matching on a sample.
- Auto-merge at low thresholds destroys records. Automated merging at low confidence creates false positives, and one bad bulk merge can erase years of conversation history. Check: back up first and auto-merge only the 95%+ confidence tier.
- Completeness does not equal correctness. A field can be 100% populated and 100% wrong, as with a stale title. Check: sample-verify populated fields rather than only counting fill rate.
- Whole-database averages hide dead segments. An 85 composite can mask a single quarantine-grade segment. Check: score by segment and by source, never only in aggregate.
- Quarterly cleanup lags decay. By the time you finish a quarterly clean, 15 to 20% has decayed again. Check: verify at point of use for anything entering a sequence.
- Inflows re-seed duplicates. More than 45% of all new records entered into CRMs are duplicates, per a study of 12 billion records. Check: track duplicate creation rate weekly, not just the stock rate.
The pattern across all eight is the same: aggregate numbers and single-pass checks flatter a list. The defence is to sample, segment, and separate the risky buckets out. The 1% duplicate standard is aspirational for the same reason - only about 22% of organisations ever hit it, while untreated databases run 10 to 30% - which is why the quarantine line sits at a realistic 5%, not 1%.
Setting the re-verification cadence after the verdict
A verdict is valid only until the next decay cycle, so every fit or remediate grade ships with a re-check interval tiered by what the record is worth. Cold prospects re-verify before every send; active-deal contacts re-verify monthly.
| Tier | Cadence | Basis |
|---|---|---|
| Active-deal contacts | Monthly | Highest value, worst cost if wrong |
| Target accounts | 60-90 days | Standard flagging window |
| Cold prospects | Before each send | Highest decay exposure at point of use |
Full comprehensive audits should happen at least twice a year, ideally quarterly, while duplicate creation rate, email bounce rate, and field completion get monitored weekly or monthly through a dashboard. Flag records not updated in 90 or more days as a reasonable starting threshold; fast-moving teams tighten to 60. At roughly 1% monthly contact decay, a 1,000-contact live-deal set yields about 31 records going wrong each quarter, which is the volume your monthly re-check is sized to catch.
One number this standard deliberately does not assert: a single stabilised accuracy level a database settles at. The sources give target bands - 95% or higher accuracy, sub-2% duplicates - rather than a defined stabilisation point, so I will not invent one. If you need a local equivalent, track your composite score across four consecutive audits and treat the plateau as your working stabilised level, then re-check it annually.
Before you call the verdict final
- The export is a dated, immutable snapshot both graders read from.
- Deal-critical fields and dimension weights were written down before scoring began.
- Accuracy is stated with its sample size (100+ records).
- Email validity excludes catch-all and risky, which are counted in separate buckets.
- Duplicate rate was measured with fuzzy matching, not exact-key only, on a 500-1,000 record sample.
- The list was scored by segment and by source, not only in aggregate.
- The verdict gates on the worst dimension, and every sub-score is recorded.
- A re-verification cadence is assigned per tier, with active-deal contacts on monthly.
- The audit findings and the cleanup queue are recorded separately.
Keeping the standard current
Adopt the threshold table as written, then re-check three things on a schedule so the standard does not drift out of date. The mechanisms below matter more than any single value, because the values move.
First, re-check your bounce and validity thresholds against your own sends, not against the published benchmark. If your seed tests show valid addresses bouncing above 2%, your verifier's definition of valid is looser than the standard assumes, and you tighten the email-validity pass line until the two agree. Second, re-check decay rates against your own list by re-sampling 100 records from an older cohort each quarter; if your measured decay runs above the 22.5% annual blended figure, shorten your cadence tiers accordingly. Third, re-check the duplicate quarantine line against your inflow: with more than 45% of new records arriving as duplicates, a source that pours in fast can breach 5% between audits, so watch the creation rate weekly.
The financial reason to hold the line is documented. Poor data quality costs organisations at least $12.9 million per year on average, and in one CRM data-health survey 44% of respondents lost more than 10% of annual revenue to it. Under the 1-10-100 rule, each stale record costs roughly $100 in wasted rep time, failed outreach, and deliverability damage. A quarantine verdict is not the expensive outcome. Handing a rep a list that reads 85 and hides a dead segment is.
Where this guide names a benchmark, it is a practitioner figure, not a law of nature. Treat the thresholds as your starting policy, instrument your own numbers against them, and let the local evidence move the pass lines. A standard that never gets re-checked is just an old opinion with a table.
Questions practitioners ask
What data quality score counts as good enough to use a prospect list?
Treat a composite above 90% as fit to use, 80 to 90% as remediate with a documented plan, and below 80% as quarantine. But the composite alone is not enough. Gate on the worst dimension too: accuracy must be 95% or higher, deal-critical completion 90% or higher, email validity 93% or higher, and duplicates under 2%. If any single dimension falls into the quarantine band, the whole list quarantines regardless of the average.
What is an acceptable duplicate rate threshold for a CRM?
Best-in-class organisations run below 2%, and the aspirational benchmark is about 1%, hit by only around 22% of organisations. Above 5% you have a significant problem and above 10% it actively distorts operations. Because untreated databases run 10 to 30% and exact-match deduplication misses 30 to 40% of real duplicates, I set the quarantine line at 5% measured with fuzzy matching, not 1%.
How do I measure data decay on a list I built months ago?
Pull 100 random records and check email, phone, job title, and company against current reality; if 30 of 100 have at least one stale field, your decay over that window is 30%. For planning, use published rates: contact records decay about 2.1% per month or 22.5% per year, and email addresses decay faster at about 3.6% per month. A 90-day-old unverified list carries an estimated 6 to 9% invalid email rate.
Why does my verifier say the list is valid but emails still bounce?
A verifier saying 90% valid can still produce a 40% bounce because valid rates include catch-all domains, role accounts, and disposable addresses. A catch-all server accepts anything, so tools mark it valid even when the mailbox does not exist. Bucket catch-all and risky separately, never count them toward the 93% validity pass, and send seed tests to a sample of valid addresses before trusting the score.
How often should I audit sourced data quality?
Run a full comprehensive audit at least twice a year, ideally quarterly, and monitor duplicate creation rate, bounce rate, and field completion weekly or monthly through a dashboard. Re-verify records by tier: active-deal contacts monthly, target accounts every 60 to 90 days, and cold prospects before every send. Quarterly cleanup alone lags decay, so verify at point of use for anything entering a sequence.
Should I weight the six dimensions or treat them equally?
Weight them for the composite score, for example 40% completeness, 30% accuracy, 20% consistency, and 10% timeliness, combined with a weighted average. But do not let the weighted number be the only gate. Keep every sub-dimension visible and quarantine the list if any one lands in its quarantine band, because a single strong dimension can otherwise mask a fatal one.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.