# The Database-Health Metric Reference: Formula, Benchmark, Blind Spot

*After reading, you can compute each data-quality metric the same way every time, state its benchmark, and name how a healthy-looking number hides a problem.*

- Canonical URL: https://www.refolk.ai/guides/database-health-metric-reference
- Pillar: Process, data, and compliance
- Format: Reference
- Published: 2026-10-04
- Last reviewed: 2026-10-04
- Reading time: 14 min

## Key takeaways

- Deliverability is the metric most likely to lie upward: in one documented case 31% of records flagged valid sat on catch-all domains and another 9% were undeliverable, dropping the true deliverable share from about 96% to 71%.
- Phone connect benchmarks are useless unless they specify per-dial versus per-prospect: the same list reads 9.9% per dial or 24.5% per prospect, and the healthy benchmark of 15-25% is a per-prospect figure.
- Firmographic fields match at 85-97% while contact fields match at 70-85%, so a single blended accuracy number hides decaying contact data behind stable company data.
- A composite health score is only comparable month over month if you freeze and publish the weights and the denominator, because identical raw data yields different scores under different weightings.
- B2B contact data decays at 20-30% per year, so a database of 100,000 contacts can lose 25,000 to 35,000 valid records annually without any visible change in record count.

Your sourced database of people and companies is only as trustworthy as the numbers you use to describe it. This reference gives revenue operations, recruiting operations, and anyone answerable for data provenance one formula, one published benchmark, and one named failure mode for each data-quality metric, so the health report you hand leadership is the same every month and comparable across teams. It covers the recurring measurement layer, not a one-time fitness gate: the metrics you recompute on a cadence so a trend line actually means something.

Use it as a lookup. Jump to the metric you are computing, read its formula and benchmark, then read the blind spot before you report the number. Every threshold and figure here comes from documented sources; where the evidence is thin, I say so and tell you what to check locally instead.

## What dimensions does a sourced database need measured?

The widely cited core set is six: completeness, accuracy, consistency, timeliness or freshness, validity, and uniqueness. For a sourced people-and-company database, that set collapses into a short list of metrics you can actually compute from your own rows plus a seed send.

The six dimensions are a taxonomy, not a dashboard. What you report is narrower: completeness, deliverability, bounce, phone connect, duplicate rate, accuracy, and freshness, then a single composite that rolls them up. Each one has a documented formula and a benchmark. The discipline is computing each the same way every time and attaching the denominator, because the number is worthless to a month-over-month comparison if the method drifted.

**$12.9M - Gartner's estimate of the annual cost of poor data quality per organization**

The cost is driven by data that looks complete but is inaccurate or decayed, which is exactly what these metrics are built to expose.

Two facts shape everything below. First, B2B contact data decays at 20-30% per year by one estimate and 25-35% by another, so a database of 100,000 contacts can lose 25,000 to 35,000 valid records annually without the record count changing. Second, company data and contact data fail in different ways, so you measure them separately or you will average a real problem out of sight.

## Metric, formula, benchmark, failure mode

Each metric below has exactly one formula, one benchmark, and the primary way the number lies. This is the table to keep open while you compute.

| Metric | Formula | Benchmark | Failure mode |
|---|---|---|---|
| Completeness | non-null / expected x100 | 80%+ | masks accuracy; field vs record level |
| Deliverability | (sent - bounced) / sent x100 | 95%+ | catch-alls inflate it |
| Bounce | bounced / sent x100 | under 2% | denominator swapped to delivered |
| Phone connect | answered / dials x100 | 15-25% per-prospect | per-dial ~9.9% looks like decay |
| Duplicate | duplicates / total x100 | under 5% | entity-resolution threshold too loose |

Completeness is non-null values divided by total expected values, times 100, with 80%+ the benchmark for active prospect lists. Deliverability is sent minus bounced over sent, times 100, with 95%+ healthy and below 90% a sign of significant data quality problems. Bounce is bounced over sent, times 100, with under 2% healthy; above that, mailbox providers start filtering you.

Phone connect deserves care. The benchmark of 15-25% for a well-maintained database is a per-prospect figure, meaning the share of prospects who eventually pick up across multiple attempts. The per-dial rate is much lower: 9.9% across a large 2025 study, versus a per-prospect rate of 24.5% on the same kind of list. Duplicate rate is duplicates over total, times 100, benchmarked under 5%; its mirror image, uniqueness, is 100 minus that, so a run of 500 unique rows in 520 total reads 96.2%.

> **Rule:** Attach the denominator to every ratio
>
> A deliverability or bounce number is only comparable month over month if the denominator is fixed and documented. Report bounce against emails sent, not emails delivered, or the figure will drift lower and stop matching the published benchmark.

## Contact records and company records fail differently

Firmographic fields match higher than contact fields by design, so you report their accuracy as two numbers, never one blended figure. Company name, employee count, and industry code are stable; email, phone, and title decay.

| Field type | Good match rate | Real-world delivered |
|---|---|---|
| Firmographic (company) | 85-97% | higher, limited by consistency |
| Contact (email/phone/title) | 70-85% | 70-85% delivered vs 90-98% claimed |

The reason to split them is that a blended accuracy number hides decaying contact data behind stable company data. If your firmographic fields sit at 92% and your contact fields at 74%, an average of 83% looks acceptable and conceals the fact that one in four contact records is wrong at the point of outreach.

Company records fail mainly on consistency and deduplication, which is entity resolution: "IBM," "International Business Machines," and "IBM Corp" create three records for one company. Contact records fail mainly on decay and deliverability. Independent testing found phone accuracy among B2B providers ranged from 63% to 91% with coverage from 26% to 92%, so even the accuracy of a single field varies enormously across sources. Treat a provider's headline accuracy claim as a demo-dataset figure: most providers claiming 90-98% accuracy delivered only 70-85% on real contact lists.

#### What rolls up into one database-health score

1. **Composite 0-100 score** - weighted, normalised, weights frozen and published
2. **Dimension metrics** - completeness, deliverability, bounce, connect, duplicate, accuracy, freshness
3. **Field-level measures** - fill rate and match rate per field, split contact vs firmographic
4. **Frozen sample** - 200-500 rows, date-stamped, recomputed the same way each month

*The composite is the top layer, but it is only trustworthy when every layer beneath it uses a fixed formula and denominator.*

## The composite health score and why weights must be frozen

A composite data quality score is a weighted average of the normalised dimension metrics, and it only earns month-over-month trust when the weights and denominators are frozen and published. The general form is the sum of each dimension's pass rate times its weight.

One documented contact-database weighting is worth copying as a starting point:

**CRM Health Score weighting**

```
Health Score =
  (Completeness x 0.25)
+ (Accuracy      x 0.25)
+ (Freshness     x 0.20)
+ (Duplicate Penalty x 0.10)
+ (Consistency   x 0.10)
+ (Engagement    x 0.10)

Where each component is normalised to a 0-100 scale,
and Duplicate Penalty = 100 - duplicate percentage.
```

*Normalise each component to 0-100 first. Invert duplicate rate (100 minus duplicate percentage) so higher is always better. Adjust the weights to your use case, then freeze and publish them.*

The critical caveat: two organizations with the same raw data can report different scores if they weight the dimensions differently. That is not a flaw to fix; it is a reason to freeze your own weighting and never silently change it. The moment you re-weight, your trend line breaks and last month's number is no longer comparable to this month's. If you must change weights, recompute history under the new weighting so the series stays consistent.

> A composite score is a promise about method, not a fact about data. Change the method and you have broken the promise.

Normalisation matters too. Each component must be scaled to 0-100 before combining, and duplicate percentage must be inverted so that higher always means better. Skip the inversion and a clean database with a low duplicate rate will drag its own composite down.

The people who own this recurring measurement are revenue operations: monitoring completeness, accuracy, and decay rate, and building the feedback loops that let sales and marketing flag inaccurate records. In Refolk's index of professional profiles there are 2,522 people in the US with a Revenue Operations Manager title, against 207 in the UK, roughly a 12-to-1 ratio. When you need the person who already owns a CRM data health program, that is the population to search.

I ran this search: `Revenue operations managers at US B2B SaaS companies who own CRM data quality and have Salesforce admin experience` - [see the full result list](https://www.refolk.ai/s/gxdncdstpc).

*Returns RevOps managers with hands-on CRM data ownership, so you can staff or benchmark the recurring measurement rather than guessing who should own it.*

**17,458 - US senior professionals listing both Data Quality and Salesforce skills (Refolk's index)**

Of these, 10,909 are at director level, a senior-to-director ratio of roughly 1.6 to 1, which tells you how deep the owner bench runs for this work.

## The procedure: from raw table to a defensible score

Run this sequence once to establish a baseline, then on a monthly cadence against a fresh frozen sample. The order matters in one place: practitioner write-ups insist the seed send precedes trusting any deliverability figure, even though some sources compute the composite first and treat the seed send as optional. Trust the seed send.

#### Compute the database-health score

1. **Define required fields and the denominator** - Decide which fields are mandatory per use case and write what 100% complete means as a business rule. Done is a written field spec naming the denominator for every ratio.
2. **Pull a frozen random sample** - Draw a random sample of 200 to 500 rows and snapshot it with a date stamp. Done is a frozen sample you can recompute against without the table shifting underneath.
3. **Compute completeness and uniqueness** - Run non-null counts per field and apply your duplicate match rules. Done is a per-field fill rate plus a single duplicate rate.
4. **Validate emails through SMTP with catch-all flagging** - Validate syntax and MX records, then run an SMTP handshake to flag each address valid, invalid, or catch-all. Done is a deliverable share with catch-alls in their own bucket.
5. **Seed-send to confirm real bounce** - Send a small seed batch, around 150 rows, and record the hard-bounce rate against emails sent. Done is a measured bounce figure you trust over tool green-lights.
6. **Measure accuracy against an authoritative source** - Spot-check titles and phones against a reliable source and record the sample size. Done is an accuracy rate with sample size noted.
7. **Normalise, weight, and compute the composite** - Normalise each metric to 0-100, invert duplicate rate, then apply your fixed weights. Done is a single 0-100 score with per-dimension contributions shown.
8. **Band, log, and compare month over month** - Band the score, log it with the denominator and weighting, and plot the trend. Done is a dashboard trend line comparable across months because the formula is frozen.

The seed send is the step people skip and then regret. A documented case ran SMTP verification with catch-all flagging, found that catch-alls and undeliverables cut the true deliverable share to about 71%, and a 150-row seed send confirmed the real hard-bounce rate at 6.2%. Non-validated datasets produce 5-7% bounce rates while properly verified data holds sub-1%, so the seed send is the difference between a number you can defend and one that collapses in production.

#### Deliverable share survives the validation stack

| Stage | Figure | Note |
| --- | --- | --- |
| Flagged valid by tool | 100 | before SMTP catch-all analysis |
| Not on catch-all domains | 69 | 31% sat on catch-all domains |
| Also truly deliverable | 60 | another 9% were undeliverable |

*A list that reads 96% valid in a tool can fall to roughly 71% deliverable once catch-alls and undeliverables are separated out.*

## How a healthy-looking number hides a problem

This is the section to read before you report anything. Each failure mode below is a specific way a green number lies, with the check that exposes it.

### Deliverability inflated by catch-alls

A list reads 96% valid and the true deliverable share is about 71%. Catch-all domains accept mail for any address and return a 250 OK, so the validation tool sees a green light but the user never gets the email. Treating catch-alls as valid without further analysis inflates deliverable rates and gives a false sense of security. Check with SMTP catch-all flagging plus a seed send, and report catch-alls as their own bucket.

### Bounce denominator swapped

Some tools report bounce rate against emails delivered rather than emails sent, which produces a different, usually lower number that is not benchmark-comparable. The under-2% benchmark assumes the denominator is sent. Check which denominator your tool uses before you compare to any threshold.

### Completeness masking accuracy

A dataset can be perfectly unique and completely inaccurate, with every field populated and every value wrong. Completeness is the cheapest metric to game, because populating a field raises the number without making it true. Check accuracy separately by sampling against an authoritative source; never infer accuracy from fill rate.

### Match rate masking coverage

A provider might be 95% accurate on the 40% of your list they could find, which is less useful than one that is 85% accurate on 90% of your list. A single accuracy figure hides how much of the list was never matched. Check by reporting accuracy and coverage as two numbers, not one.

### Phone connect misread

A per-dial connect rate near 9.9% looks like dead data when read against the 15-25% benchmark, which is per-prospect. The same list reads 24.5% per prospect across attempts. Check that your connect metric is defined per-prospect across attempts before comparing to the benchmark, or you will condemn healthy data.

### Composite score drift

Identical raw data yields different scores under different weightings, so a composite that moved may reflect a weight change rather than a data change. Check by freezing and publishing the weighting and the denominator, and recompute history if you ever change them.

### Freshness as a proxy for accuracy

A record verified 90 days ago can predate a job change, so a high freshness score does not mean the data is right. Freshness is a timeliness signal, not an accuracy one. Check by validating with live bounce and reply signals rather than trusting the verification timestamp.

### Company dedupe false negatives

Name variants survive as separate records, so the duplicate rate reads clean while three rows describe one company. Check with entity-resolution match rules, not string equality, especially on firmographic data.

> **Watch out:** Blended accuracy is the quiet killer
>
> Reporting one accuracy number for contacts and companies together lets stable firmographic data (85-97% match) hide decaying contact data (70-85% match). The average looks fine while your outreach bounces. Always split the two.

> **Tip:** Validate accuracy indirectly between full audits
>
> You cannot spot-check every record every month. Between audits, use bounce, reply, and connect signals as live accuracy proxies: a rising bounce rate on a segment is accuracy decay you can act on before the next full sample.

## Keeping the measurement current

Freshness is commonly re-measured on a rolling 90-day window: the share of records verified or updated within the last 90 days, benchmarked at 60%+, with anything unverified for six or more months flagged for re-verification before outreach. That window is the one widely published cadence; the exact recompute cadence for every other dimension is not established as a single public standard, so set your own and document it.

A workable default: recompute the composite and deliverability monthly against a fresh frozen sample, track freshness continuously on its 90-day window, and run a full accuracy audit quarterly. Because decay runs at 20-30% per year, a quarterly audit catches roughly a quarter of the annual drift before it reaches your senders.

Before you publish any monthly number, run this check.

#### Before you report the score

- [ ] The denominator for every ratio is written down and unchanged from last month.
- [ ] Catch-all addresses are flagged separately and not counted as deliverable.
- [ ] Bounce rate is computed against emails sent, confirmed by a seed send.
- [ ] Contact accuracy and firmographic accuracy are reported as two numbers.
- [ ] Phone connect is defined per-prospect across attempts, not per-dial.
- [ ] The composite weighting is frozen and published alongside the score.
- [ ] Freshness uses the 90-day window with the 60%+ benchmark noted.
- [ ] The sample is random, 200-500 rows, and date-stamped.

The last point of discipline is ownership. Someone has to recompute this on a cadence, defend the denominator, and refuse to quietly re-weight the composite when the number looks bad. That is a revenue operations job, and the depth of that bench is measurable: [Refolk](/) shows 2,522 US Revenue Operations Managers and 17,458 US senior professionals carrying both Data Quality and Salesforce skills, so the person who can own this well is findable by name and skill rather than by guesswork. Fix the formula, freeze the weights, and the trend line finally means what it says.

## Frequently asked questions

### What is a good CRM data health score?

There is no universal pass mark because the score depends on your weights. The dossier's documented weighting is Completeness 0.25, Accuracy 0.25, Freshness 0.20, Duplicate 0.10, Consistency 0.10, Engagement 0.10, each normalised to 0-100. A defensible target is to meet each component benchmark: completeness 80%+, deliverability 95%+, bounce under 2%, duplicates under 5%, freshness 60%+ verified within 90 days. Freeze the weights so the composite is comparable over time.

### How do you calculate email deliverability rate?

Deliverability rate equals (emails sent minus bounced emails) divided by emails sent, times 100. The healthy benchmark is 95% or higher, and below 90% signals significant data quality problems. The trap is catch-all domains that return a green light in validation tools but never deliver, so confirm the number with an SMTP handshake that flags catch-alls plus a small seed send.

### What is a healthy duplicate rate for a contact database?

The duplicate rate benchmark is under 5%, computed as duplicates divided by total records times 100. For people records, string equality usually catches duplicates. For company records the harder failure is false negatives, where IBM, International Business Machines, and IBM Corp survive as three records for one entity. Use entity-resolution match rules, not raw string matching, or the metric will look clean while hiding duplicates.

### How often should I recompute data quality metrics?

Freshness is commonly re-measured on a rolling 90-day window with a 60%+ benchmark, and records not verified in six or more months should be flagged before outreach. The exact recompute cadence per dimension is not established as a single public standard. A practical cadence is monthly for the composite and deliverability, with freshness tracked on its 90-day window and accuracy validated indirectly through live bounce, reply, and connect signals.

### Why does my deliverability look high but replies are low?

Catch-all domains are the usual cause. The SMTP server accepts mail for any address and returns 250 OK, so validation tools mark the record valid even though no mailbox exists. In one documented case 31% of valid-flagged records were catch-alls and another 9% were undeliverable, cutting true deliverable share to about 71%. Separate catch-alls into their own bucket and confirm with a seed send.

### Should contact and company accuracy be one number or two?

Two. Firmographic fields such as company name, employee count, and industry code match at 85-97%, while contact fields such as email, phone, and job title match at 70-85%. A blended accuracy figure hides decaying contact data behind stable company data. Report firmographic accuracy and contact accuracy separately so a drop in contact quality is visible instead of averaged away.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/database-health-metric-reference*
