The Database-Health Metric Reference: Formula, Benchmark, Blind Spot
After reading, you can compute each data-quality metric the same way every time, state its benchmark, and name how a healthy-looking number hides a problem.
Key takeaways
- Deliverability is the metric most likely to lie upward: in one documented case 31% of records flagged valid sat on catch-all domains and another 9% were undeliverable, dropping the true deliverable share from about 96% to 71%.
- Phone connect benchmarks are useless unless they specify per-dial versus per-prospect: the same list reads 9.9% per dial or 24.5% per prospect, and the healthy benchmark of 15-25% is a per-prospect figure.
- Firmographic fields match at 85-97% while contact fields match at 70-85%, so a single blended accuracy number hides decaying contact data behind stable company data.
- A composite health score is only comparable month over month if you freeze and publish the weights and the denominator, because identical raw data yields different scores under different weightings.
- B2B contact data decays at 20-30% per year, so a database of 100,000 contacts can lose 25,000 to 35,000 valid records annually without any visible change in record count.
Your sourced database of people and companies is only as trustworthy as the numbers you use to describe it. This reference gives revenue operations, recruiting operations, and anyone answerable for data provenance one formula, one published benchmark, and one named failure mode for each data-quality metric, so the health report you hand leadership is the same every month and comparable across teams. It covers the recurring measurement layer, not a one-time fitness gate: the metrics you recompute on a cadence so a trend line actually means something.
Use it as a lookup. Jump to the metric you are computing, read its formula and benchmark, then read the blind spot before you report the number. Every threshold and figure here comes from documented sources; where the evidence is thin, I say so and tell you what to check locally instead.
What dimensions does a sourced database need measured?
The widely cited core set is six: completeness, accuracy, consistency, timeliness or freshness, validity, and uniqueness. For a sourced people-and-company database, that set collapses into a short list of metrics you can actually compute from your own rows plus a seed send.
The six dimensions are a taxonomy, not a dashboard. What you report is narrower: completeness, deliverability, bounce, phone connect, duplicate rate, accuracy, and freshness, then a single composite that rolls them up. Each one has a documented formula and a benchmark. The discipline is computing each the same way every time and attaching the denominator, because the number is worthless to a month-over-month comparison if the method drifted.
Two facts shape everything below. First, B2B contact data decays at 20-30% per year by one estimate and 25-35% by another, so a database of 100,000 contacts can lose 25,000 to 35,000 valid records annually without the record count changing. Second, company data and contact data fail in different ways, so you measure them separately or you will average a real problem out of sight.
Metric, formula, benchmark, failure mode
Each metric below has exactly one formula, one benchmark, and the primary way the number lies. This is the table to keep open while you compute.
| Metric | Formula | Benchmark | Failure mode |
|---|---|---|---|
| Completeness | non-null / expected x100 | 80%+ | masks accuracy; field vs record level |
| Deliverability | (sent - bounced) / sent x100 | 95%+ | catch-alls inflate it |
| Bounce | bounced / sent x100 | under 2% | denominator swapped to delivered |
| Phone connect | answered / dials x100 | 15-25% per-prospect | per-dial ~9.9% looks like decay |
| Duplicate | duplicates / total x100 | under 5% | entity-resolution threshold too loose |
Completeness is non-null values divided by total expected values, times 100, with 80%+ the benchmark for active prospect lists. Deliverability is sent minus bounced over sent, times 100, with 95%+ healthy and below 90% a sign of significant data quality problems. Bounce is bounced over sent, times 100, with under 2% healthy; above that, mailbox providers start filtering you.
Phone connect deserves care. The benchmark of 15-25% for a well-maintained database is a per-prospect figure, meaning the share of prospects who eventually pick up across multiple attempts. The per-dial rate is much lower: 9.9% across a large 2025 study, versus a per-prospect rate of 24.5% on the same kind of list. Duplicate rate is duplicates over total, times 100, benchmarked under 5%; its mirror image, uniqueness, is 100 minus that, so a run of 500 unique rows in 520 total reads 96.2%.
Contact records and company records fail differently
Firmographic fields match higher than contact fields by design, so you report their accuracy as two numbers, never one blended figure. Company name, employee count, and industry code are stable; email, phone, and title decay.
| Field type | Good match rate | Real-world delivered |
|---|---|---|
| Firmographic (company) | 85-97% | higher, limited by consistency |
| Contact (email/phone/title) | 70-85% | 70-85% delivered vs 90-98% claimed |
The reason to split them is that a blended accuracy number hides decaying contact data behind stable company data. If your firmographic fields sit at 92% and your contact fields at 74%, an average of 83% looks acceptable and conceals the fact that one in four contact records is wrong at the point of outreach.
Company records fail mainly on consistency and deduplication, which is entity resolution: "IBM," "International Business Machines," and "IBM Corp" create three records for one company. Contact records fail mainly on decay and deliverability. Independent testing found phone accuracy among B2B providers ranged from 63% to 91% with coverage from 26% to 92%, so even the accuracy of a single field varies enormously across sources. Treat a provider's headline accuracy claim as a demo-dataset figure: most providers claiming 90-98% accuracy delivered only 70-85% on real contact lists.
What rolls up into one database-health score
- Composite 0-100 scoreweighted, normalised, weights frozen and published
- Dimension metricscompleteness, deliverability, bounce, connect, duplicate, accuracy, freshness
- Field-level measuresfill rate and match rate per field, split contact vs firmographic
- Frozen sample200-500 rows, date-stamped, recomputed the same way each month
The composite health score and why weights must be frozen
A composite data quality score is a weighted average of the normalised dimension metrics, and it only earns month-over-month trust when the weights and denominators are frozen and published. The general form is the sum of each dimension's pass rate times its weight.
One documented contact-database weighting is worth copying as a starting point:
Health Score = (Completeness x 0.25) + (Accuracy x 0.25) + (Freshness x 0.20) + (Duplicate Penalty x 0.10) + (Consistency x 0.10) + (Engagement x 0.10) Where each component is normalised to a 0-100 scale, and Duplicate Penalty = 100 - duplicate percentage.
Normalise each component to 0-100 first. Invert duplicate rate (100 minus duplicate percentage) so higher is always better. Adjust the weights to your use case, then freeze and publish them.
The critical caveat: two organizations with the same raw data can report different scores if they weight the dimensions differently. That is not a flaw to fix; it is a reason to freeze your own weighting and never silently change it. The moment you re-weight, your trend line breaks and last month's number is no longer comparable to this month's. If you must change weights, recompute history under the new weighting so the series stays consistent.
A composite score is a promise about method, not a fact about data. Change the method and you have broken the promise.
Normalisation matters too. Each component must be scaled to 0-100 before combining, and duplicate percentage must be inverted so that higher always means better. Skip the inversion and a clean database with a low duplicate rate will drag its own composite down.
The people who own this recurring measurement are revenue operations: monitoring completeness, accuracy, and decay rate, and building the feedback loops that let sales and marketing flag inaccurate records. In Refolk's index of professional profiles there are 2,522 people in the US with a Revenue Operations Manager title, against 207 in the UK, roughly a 12-to-1 ratio. When you need the person who already owns a CRM data health program, that is the population to search.
The procedure: from raw table to a defensible score
Run this sequence once to establish a baseline, then on a monthly cadence against a fresh frozen sample. The order matters in one place: practitioner write-ups insist the seed send precedes trusting any deliverability figure, even though some sources compute the composite first and treat the seed send as optional. Trust the seed send.
Compute the database-health score
- Define required fields and the denominatorDecide which fields are mandatory per use case and write what 100% complete means as a business rule. Done is a written field spec naming the denominator for every ratio.
- Pull a frozen random sampleDraw a random sample of 200 to 500 rows and snapshot it with a date stamp. Done is a frozen sample you can recompute against without the table shifting underneath.
- Compute completeness and uniquenessRun non-null counts per field and apply your duplicate match rules. Done is a per-field fill rate plus a single duplicate rate.
- Validate emails through SMTP with catch-all flaggingValidate syntax and MX records, then run an SMTP handshake to flag each address valid, invalid, or catch-all. Done is a deliverable share with catch-alls in their own bucket.
- Seed-send to confirm real bounceSend a small seed batch, around 150 rows, and record the hard-bounce rate against emails sent. Done is a measured bounce figure you trust over tool green-lights.
- Measure accuracy against an authoritative sourceSpot-check titles and phones against a reliable source and record the sample size. Done is an accuracy rate with sample size noted.
- Normalise, weight, and compute the compositeNormalise each metric to 0-100, invert duplicate rate, then apply your fixed weights. Done is a single 0-100 score with per-dimension contributions shown.
- Band, log, and compare month over monthBand the score, log it with the denominator and weighting, and plot the trend. Done is a dashboard trend line comparable across months because the formula is frozen.
The seed send is the step people skip and then regret. A documented case ran SMTP verification with catch-all flagging, found that catch-alls and undeliverables cut the true deliverable share to about 71%, and a 150-row seed send confirmed the real hard-bounce rate at 6.2%. Non-validated datasets produce 5-7% bounce rates while properly verified data holds sub-1%, so the seed send is the difference between a number you can defend and one that collapses in production.
Deliverable share survives the validation stack
- 100Flagged valid by tool
before SMTP catch-all analysis
- 69Not on catch-all domains
31% sat on catch-all domains
- 60Also truly deliverable
another 9% were undeliverable
How a healthy-looking number hides a problem
This is the section to read before you report anything. Each failure mode below is a specific way a green number lies, with the check that exposes it.
Deliverability inflated by catch-alls
A list reads 96% valid and the true deliverable share is about 71%. Catch-all domains accept mail for any address and return a 250 OK, so the validation tool sees a green light but the user never gets the email. Treating catch-alls as valid without further analysis inflates deliverable rates and gives a false sense of security. Check with SMTP catch-all flagging plus a seed send, and report catch-alls as their own bucket.
Bounce denominator swapped
Some tools report bounce rate against emails delivered rather than emails sent, which produces a different, usually lower number that is not benchmark-comparable. The under-2% benchmark assumes the denominator is sent. Check which denominator your tool uses before you compare to any threshold.
Completeness masking accuracy
A dataset can be perfectly unique and completely inaccurate, with every field populated and every value wrong. Completeness is the cheapest metric to game, because populating a field raises the number without making it true. Check accuracy separately by sampling against an authoritative source; never infer accuracy from fill rate.
Match rate masking coverage
A provider might be 95% accurate on the 40% of your list they could find, which is less useful than one that is 85% accurate on 90% of your list. A single accuracy figure hides how much of the list was never matched. Check by reporting accuracy and coverage as two numbers, not one.
Phone connect misread
A per-dial connect rate near 9.9% looks like dead data when read against the 15-25% benchmark, which is per-prospect. The same list reads 24.5% per prospect across attempts. Check that your connect metric is defined per-prospect across attempts before comparing to the benchmark, or you will condemn healthy data.
Composite score drift
Identical raw data yields different scores under different weightings, so a composite that moved may reflect a weight change rather than a data change. Check by freezing and publishing the weighting and the denominator, and recompute history if you ever change them.
Freshness as a proxy for accuracy
A record verified 90 days ago can predate a job change, so a high freshness score does not mean the data is right. Freshness is a timeliness signal, not an accuracy one. Check by validating with live bounce and reply signals rather than trusting the verification timestamp.
Company dedupe false negatives
Name variants survive as separate records, so the duplicate rate reads clean while three rows describe one company. Check with entity-resolution match rules, not string equality, especially on firmographic data.
Keeping the measurement current
Freshness is commonly re-measured on a rolling 90-day window: the share of records verified or updated within the last 90 days, benchmarked at 60%+, with anything unverified for six or more months flagged for re-verification before outreach. That window is the one widely published cadence; the exact recompute cadence for every other dimension is not established as a single public standard, so set your own and document it.
A workable default: recompute the composite and deliverability monthly against a fresh frozen sample, track freshness continuously on its 90-day window, and run a full accuracy audit quarterly. Because decay runs at 20-30% per year, a quarterly audit catches roughly a quarter of the annual drift before it reaches your senders.
Before you publish any monthly number, run this check.
Before you report the score
- The denominator for every ratio is written down and unchanged from last month.
- Catch-all addresses are flagged separately and not counted as deliverable.
- Bounce rate is computed against emails sent, confirmed by a seed send.
- Contact accuracy and firmographic accuracy are reported as two numbers.
- Phone connect is defined per-prospect across attempts, not per-dial.
- The composite weighting is frozen and published alongside the score.
- Freshness uses the 90-day window with the 60%+ benchmark noted.
- The sample is random, 200-500 rows, and date-stamped.
The last point of discipline is ownership. Someone has to recompute this on a cadence, defend the denominator, and refuse to quietly re-weight the composite when the number looks bad. That is a revenue operations job, and the depth of that bench is measurable: Refolk shows 2,522 US Revenue Operations Managers and 17,458 US senior professionals carrying both Data Quality and Salesforce skills, so the person who can own this well is findable by name and skill rather than by guesswork. Fix the formula, freeze the weights, and the trend line finally means what it says.
Questions practitioners ask
What is a good CRM data health score?
There is no universal pass mark because the score depends on your weights. The dossier's documented weighting is Completeness 0.25, Accuracy 0.25, Freshness 0.20, Duplicate 0.10, Consistency 0.10, Engagement 0.10, each normalised to 0-100. A defensible target is to meet each component benchmark: completeness 80%+, deliverability 95%+, bounce under 2%, duplicates under 5%, freshness 60%+ verified within 90 days. Freeze the weights so the composite is comparable over time.
How do you calculate email deliverability rate?
Deliverability rate equals (emails sent minus bounced emails) divided by emails sent, times 100. The healthy benchmark is 95% or higher, and below 90% signals significant data quality problems. The trap is catch-all domains that return a green light in validation tools but never deliver, so confirm the number with an SMTP handshake that flags catch-alls plus a small seed send.
What is a healthy duplicate rate for a contact database?
The duplicate rate benchmark is under 5%, computed as duplicates divided by total records times 100. For people records, string equality usually catches duplicates. For company records the harder failure is false negatives, where IBM, International Business Machines, and IBM Corp survive as three records for one entity. Use entity-resolution match rules, not raw string matching, or the metric will look clean while hiding duplicates.
How often should I recompute data quality metrics?
Freshness is commonly re-measured on a rolling 90-day window with a 60%+ benchmark, and records not verified in six or more months should be flagged before outreach. The exact recompute cadence per dimension is not established as a single public standard. A practical cadence is monthly for the composite and deliverability, with freshness tracked on its 90-day window and accuracy validated indirectly through live bounce, reply, and connect signals.
Why does my deliverability look high but replies are low?
Catch-all domains are the usual cause. The SMTP server accepts mail for any address and returns 250 OK, so validation tools mark the record valid even though no mailbox exists. In one documented case 31% of valid-flagged records were catch-alls and another 9% were undeliverable, cutting true deliverable share to about 71%. Separate catch-alls into their own bucket and confirm with a seed send.
Should contact and company accuracy be one number or two?
Two. Firmographic fields such as company name, employee count, and industry code match at 85-97%, while contact fields such as email, phone, and job title match at 70-85%. A blended accuracy figure hides decaying contact data behind stable company data. Report firmographic accuracy and contact accuracy separately so a drop in contact quality is visible instead of averaged away.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.