# The Duplicate-Company Verdict: Merge, Link, or Keep Separate

*You will score any candidate company pair, band it, and assign it to auto-merge, review, link-as-hierarchy, or keep-separate with a reason another grader can reproduce.*

- Canonical URL: https://www.refolk.ai/guides/duplicate-company-verdict
- Pillar: Process, data, and compliance
- Format: Framework
- Published: 2026-09-06
- Last reviewed: 2026-09-06
- Reading time: 15 min
- Keywords: merge or keep separate company records, duplicate account match confidence, parent child account hierarchy crm, subsidiary vs duplicate account, company deduplication match rules, when to auto merge accounts

## Key takeaways

- A duplicate-company decision has four outcomes, not two: auto-merge, human review, link as hierarchy, and keep separate. A contact-shaped merge/no-merge cutoff manufactures false merges on subsidiaries and regional entities.
- Identifiers should carry the merge, not names. A VAT match alone usually confirms identity in most jurisdictions, so a conflicting legal or tax ID is a stronger do-not-merge signal than name similarity is a merge signal.
- Domain fails in both directions: it over-matches on agency and coworking addresses and under-matches when one company holds several domains, which is exactly why subsidiaries must be linked rather than merged.
- Only 42 of 1,072 US Revenue Operations Managers list both Salesforce and Data Quality skills, roughly 3.9% of the pool, so documented reproducible rules matter more than tribal knowledge.
- A Salesforce merge sends non-master records to the Recycle Bin for 15 days and is otherwise permanent, so any verdict pushed to production without a backup is effectively irreversible within two pay cycles.
- Survivorship is a governance mechanism, not a merge algorithm. If identity is wrong, confident survivorship destroys account history faster than no deduplication at all.

You have two company records in front of you that look like the same account, and you have to decide what to do with them. This guide is for RevOps, data governance, and anyone answerable for how account data was consolidated. It gives you a way to score the pair across weighted signals, place it in a confidence tier, and assign one of four verdicts that another grader would reach independently.

Most deduplication guides treat companies and contacts the same and stop at a single merge-or-not cutoff. That cutoff is fine for people. It is dangerous for companies, because the company-specific traps - subsidiaries, regional entities, rebrands, and post-acquisition brands - need outcomes a binary rule cannot produce. This guide scores the pair on the signals that actually separate a true duplicate from a related-but-distinct entity, then maps the score to a defensible verdict instead of an auto-merge threshold that quietly destroys account history.

## Why a company pair needs four outcomes, not two

A company decision has four verdicts because domain, the field contact rules lean on, fails in both directions. Contacts get merged or they do not. Companies also get linked, and sometimes held.

Domain over-matches when a marketing agency, a consulting firm, or a shared coworking space puts many genuinely unrelated companies under one address. Domain under-matches when one real company legitimately holds several domains across product lines and regions. A domain-driven binary rule therefore manufactures false merges in one direction and false separations in the other. That is precisely why subsidiaries must be linked, not merged: a regional subsidiary and its parent may share a domain but represent distinct go-to-market relationships, and logic that flattens them destroys territory and routing data that was correct.

The four outcomes are the frame for everything that follows.

- **Auto-merge.** Deterministic identifier match with no do-not-merge gate. Safe to run in volume.
- **Human review.** Strong-and-possible, but no confirming identifier. A steward looks at the related records and decides.
- **Link as hierarchy.** Two real entities, one relationship. Parent and child, kept distinct.
- **Keep separate.** Related-looking but distinct. No link, no merge.

#### The company-pair verdict map

Horizontal axis runs from Weak fuzzy signal to Strong fuzzy signal. Vertical axis runs from No identifier match to Identifier match or conflict.

| Quadrant | What it means |
| --- | --- |
| Keep separate | Low similarity, no shared ID; leave both records alone. |
| Human review or link | Similar name and domain but no confirming ID; a steward checks for subsidiary structure. |
| Keep separate | Conflicting legal or tax ID; distinct entities even if names rhyme. |
| Auto-merge or link as hierarchy | Matching ID confirms one entity to merge; conflicting ID with shared parent means link. |

*A confirming identifier decides whether a high fuzzy score means merge or hierarchy.*

## The signals that discriminate, and the ones that lie

Rank your signals by whether they carry identity or merely suggest it. Identifiers carry identity. Names and domains suggest it, and both lie in predictable ways.

The asymmetry at the center of this framework: because a VAT match alone usually confirms identity in most jurisdictions, a *conflicting* VAT or registration number is a stronger do-not-merge signal than name similarity is a merge signal. You can merge with confidence on an identifier. You should almost never merge on a name.

| Signal | What a match proves | What it looks like when it lies |
|---|---|---|
| VAT / registration no. | Same legal entity, near-certain | Rarely lies; a conflict is a hard do-not-merge |
| Legal name + domain paired | Strong duplicate candidate | Two subsidiaries share a domain under one parent |
| Domain alone | Weak; one signal among several | Agency or coworking address pools unrelated firms |
| Company name (normalized) | Suggestive, needs a second signal | Short names under 3-4 chars, like ABC, match many |
| Shared Group / Holdings token | Suggestive at best | Common word inflates score without real match |

Two rules follow directly. First, normalize before you compare: strip legal suffixes, lowercase, and standardize the domain, because duplicates can only be detected once names are comparable. Second, weight tokens by frequency. Plain normalization cannot tell a meaningful shared word from a real name match, so a shared "Group" or "Holdings" token needs frequency-aware weighting - rare tokens count more than common ones - alongside normalization.

> **Rule:** Identifiers override names, always
>
> A conflicting legal or tax identifier is a do-not-merge gate that overrides any name or domain score. If the VAT numbers disagree, the records are distinct entities no matter how alike the names read.

Short names deserve their own caution. A company name under three or four characters creates real false-positive risk, so treat it as conservative: require an additional signal like domain or location before it contributes anything to the score.

## Scoring the pair on a weighted 0-1 number

Combine each field score into a single weighted number between 0 and 1, then read the number against pre-declared bands. The weights depend on your domain, but the shape is stable: identifiers dominate, paired signals matter, single soft signals count for little.

Published sources converge on a three-band model but the exact cutoffs vary, which is why you calibrate against a labeled sample rather than adopting a number blind.

| Source | Auto-merge floor | Review band | Block below |
|---|---|---|---|
| Supportbench | 0.95 | mid-range | - |
| Dynamics 365 checklist | 0.90 | 0.70-0.89 | 0.70 |
| Primentra (MDM) | 0.92 | 0.75-0.92 | 0.75 |

Pick a floor you can defend, then hold it against a sample of hand-labeled pairs and adjust. The point of a well-calibrated model is not to merge more aggressively. It is triage: a model that scores confidence well lets you safely auto-merge a larger high-confidence tier and shrink the manual-review queue to genuine edge cases.

**3.9% - US RevOps Managers who list both Salesforce and Data Quality**

Only 42 of 1,072 in Refolk's index, which is why the rules must be written down, not carried in one person's head.

That scarcity is the case for reproducibility. In Refolk's index, only 42 of 1,072 US Revenue Operations Managers pair Salesforce with Data Quality as listed skills. The people who own the CRM mostly do not list the exact competency this judgment requires, so a documented ruleset another grader can execute beats tribal knowledge every time.

## The do-not-merge and link-as-hierarchy gates

Some pairs must never auto-merge regardless of score. Run these gates before the score decides anything.

Documented hard blocks: shared mailboxes, holding companies, and records with conflicting legal or tax IDs are do-not-merge. A shared registered address is a flag, not a merge signal, because coworking and agency addresses pool unrelated firms.

The subsidiary case is where the third outcome earns its keep. A subsidiary and its parent may share a domain but represent distinct go-to-market relationships. Merging the subsidiary into the parent destroys local billing and contract accuracy; leaving them unlinked hides the enterprise-wide relationship. Linking as parent and child is the only outcome that preserves both.

The distinguishing test is legal-entity independence plus separate branding. Most subsidiaries have their own leadership and may operate under a different name, and they typically maintain their own balance sheets even when financials consolidate into the parent's reports. YouTube, Waze, and Nest all operate as independently-branded subsidiaries under Alphabet, and none should be merged into it.

#### Deciding merge, link, or hold on a suspected acquisition

1. **Same legal entity?** - If yes and only the name changed, it is a rebrand.
2. **Rebrand path** - Merge with an old-name-to-new-name alias so history resolves.
3. **Distinct entity, own brand?** - If the acquired brand still trades under its own name, link as hierarchy.
4. **Status unclear?** - Hold for a status check against a public registry before deciding.

*The acquired brand's trading status, not the acquisition itself, decides the verdict.*

Rebrands are the mirror image and need an alias map. A pure rebrand is the same entity under a new name, so it is a merge-with-alias: create a mapping from old names to new, the way Facebook resolves to Meta, and include common old names in lookup rules so historical references still match. Without that map, a rebrand looks like a distinct account and you split one company in two.

Acquisitions split by trading status. When the acquired brand still trades under its own name, link as hierarchy. P&G acquired Gillette in a deal worth $57 billion, and Gillette still trades as its own brand, so that pair links rather than merges. When the acquired brand is fully absorbed and retired, treat it like a rebrand with an alias.

I ran this search: `Data stewards and master data management leads in the US who have run CRM account merge or golden-record projects` - [see the full result list](https://www.refolk.ai/s/t3r7wpe5m8).

*Returns people who have owned exactly this verdict in production, with the survivorship and audit-trail scars to show for it.*

Naming the owner is half the battle, and the pools are deeper than most teams assume. Use [Refolk](/) to find the person who has run a golden-record project before, rather than handing the ruleset to whoever happens to own the CRM. In Refolk's index, 1,316 people in the US hold Data Governance, Data Steward, or Master Data Manager titles, a slightly larger pool than the 1,072 Revenue Operations Managers who tend to own the system itself.

| Role family | US count | Share of the two pools |
|---|---|---|
| Data Governance / Steward / MDM | 1,316 | 55.1% |
| Revenue Operations Manager | 1,072 | 44.9% |

## The procedure, from ruleset to prevention

Run these eight steps in order. The first is a writing job, not a data job, and skipping it is where most cleanups go wrong.

#### Duplicate-company verdict procedure

1. **Define the ruleset and four outcomes** - Write down thresholds, survivorship, and the four verdicts before touching data. Done: a signed matching contract that names who owns each decision.
2. **Normalize the match fields** - Strip legal suffixes, lowercase, normalize phone to E.164, standardize domains. Done: comparable fields, because duplicates can only be detected once names are comparable.
3. **Process deterministic matches first** - Clear exact identifier matches on VAT, registration number, or external ID. Done: near-certain duplicates resolved, shrinking the pool before fuzzy scoring.
4. **Score fuzzy candidates on weighted signals** - Combine field scores into one weighted 0-1 number with frequency-aware weighting. Done: every remaining pair carries a score and a band.
5. **Apply do-not-merge and hierarchy gates** - Flag conflicting legal or tax IDs, shared addresses, holding companies, and subsidiaries as hold or link. Done: edge cases routed away from auto-merge.
6. **Human review of the gray band** - Route strong-and-possible matches to a steward who can see related records and match evidence. Done: a documented verdict per pair.
7. **Execute merge in staging with survivorship** - Back up, choose master, apply pre-declared survivorship, verify related records reparent. Done: master validated against backup. Merge accounts before contacts.
8. **Post-merge audit and prevention** - Sample decisions on a schedule, log field-level survivorship, enable server-side duplicate prevention on every channel. Done: duplicate rate tracked toward sub-1%.

Two sequencing notes worth flagging. Sources agree you merge accounts before contacts, so the account decision comes first. And there is genuine disagreement on step 3: some teams process deterministic matches first because they are high-confidence and safe to clear in volume, while others enrich thin records before matching so the fuzzy phase has more to work with. If your records are sparse, enrich first; if they are reasonably complete, clear the deterministic matches first to shrink the queue.

## How this goes wrong: failure modes and false positives

The failure modes below are where account history dies. Each one has a signature and a check that catches it before the merge runs.

**Domain over-match.** Two unrelated clients on an agency or coworking domain score high. The false positive reads "matched on domain." Check: require a second identifier, a VAT or registration number, before any merge.

**Subsidiary flattened into parent.** A blind auto-merge on fuzzy logic collapses two real subsidiaries into one account, and you will not notice until a renewal goes missing. The false positive is shared domain plus similar name. Check: a conflicting legal or tax ID or a separate registered address forces link-as-hierarchy.

**Shared Group or Holdings token.** Normalization cannot tell a meaningful shared word from a real name match. Check: weighted, frequency-aware matching alongside normalization, so common tokens carry less weight.

**Closed Won opportunity destroyed on merge.** Commissions and forecasts break silently because merging Closed Won opportunities breaks historical data. Check: exclude Closed Won from the merge, or reconcile line items manually first.

**Survivorship overwrites good data.** The master's blank field wins over the duplicate's populated value. Check: backfill NULLs from the non-survivor rather than trusting the master wholesale. A stable pattern is to elect the oldest record as survivor for stability, then backfill only NULL fields from the newest for recency.

**Cleanup without prevention.** The duplicate rate rebounds after a clean one-time report. Check: enable server-side duplicate prevention across imports, workflows, and APIs, not just the form.

**Rebrand treated as new company.** The same entity under a new name looks distinct. Check: maintain an old-name-to-new-name alias map before scoring.

**Two graders disagree on the same pair.** There is no recorded evidence to reconcile. Check: store the match features and a field-level decision log for every verdict.

> **Watch out:** The Recycle Bin is your only net, and it is short
>
> A Salesforce merge sends non-master records to the Recycle Bin for 15 days and is otherwise permanent. Push a verdict to production without a backup and you have made an irreversible change inside two pay cycles.

> A wrong merge with confident survivorship destroys history faster than no deduplication at all.

## Survivorship and the audit trail that makes a verdict reproducible

Survivorship is a governance mechanism, not a merge algorithm, and if identity is wrong, survivorship is dangerous. Decide field-by-field which value survives before you run anything.

Survivorship selects and combines the most reliable attribute values from the duplicate records. Organizations typically use source priority, completeness, recency, verification status, and conditional overwrite rules to decide which values win. Declare those rules in the matching contract so the run executes a decision you already made, not one a steward improvises live.

Know your platform's mechanics before you commit. In native Salesforce, all related records - opportunities, contacts, cases, activities - move to the master, non-master records go to the Recycle Bin recoverable for 15 days, and the master retains its original Record ID. Salesforce limits merges to three records at a time. In Apex merges, master field values supersede the others and external ID fields cannot be used with merge.

The audit trail is what lets a second grader reproduce your verdict. Make the evidence visible: store the exact match features behind each suggestion so reviewers can see why a pair was flagged, and keep field-level decision logs showing which values won and lost in every merge and why. That record is also what settles disputes when two graders read the same pair differently.

**Field-level survivorship and verdict log (one row per pair)**

```
pair_id: ACC-00812 / ACC-01199
verdict: link_as_hierarchy
score: 0.88
gates_fired: conflicting_registration_number
identity_signal: separate VAT, shared domain, distinct GTM region
survivorship_rule: oldest record survives; backfill NULLs from newer
excluded: Closed Won opportunities (reconcile manually)
reviewer: [name]
reason: Regional subsidiary of parent; link, do not merge, to preserve local billing.
```

*Fill one before executing. This is the contract the run and any future audit both read.*

**Merge-run confidence bands**

```
auto_merge: identifier match AND no gate fired
review: score in review band OR shared domain without confirming identifier
link_as_hierarchy: conflicting legal/tax ID with shared parent or ownership
keep_separate: low score AND no identifier match
hold_for_status_check: suspected acquisition or rebrand, status unconfirmed
```

*Adapt the floors to your own labeled sample; do not adopt a number blind.*

## Keeping the standard current

A duplicate-company standard is only as good as its last calibration, so treat it as a recurring job, not a one-time cleanup. After go-live, sample AI-assisted or automated decisions on a set schedule to catch threshold drift before it snowballs, and track the duplicate rate against a target: organizations without an active data-quality program often carry duplicate rates between 10 and 30 percent, while best-in-class organizations keep that below 1 percent.

Two external mechanisms change under you and are worth re-checking rather than memorizing. Identifier schemes evolve - the US federal government replaced DUNS with the Unique Entity ID on SAM.gov, and the LEI, a 20-character code under ISO 17442, is now required or recommended by over 300 regulations - so revisit which identifier you trust as your merge anchor in each jurisdiction. Free public registries let you confirm legal existence without a paid vendor: GLEIF for LEIs, and jurisdiction registries such as Companies House in the UK and the Handelsregister in Germany.

Before you call any run done, walk the checklist.

#### Before you execute a duplicate-company verdict

- [ ] The ruleset, thresholds, and four outcomes are written and signed before any data was touched.
- [ ] Match fields are normalized: suffixes stripped, names and emails lowercased, domains standardized.
- [ ] Deterministic identifier matches were cleared before fuzzy scoring ran.
- [ ] Conflicting legal or tax IDs, holding companies, and subsidiaries were gated to hold or link.
- [ ] Every gray-band pair has a documented verdict with stored match features.
- [ ] A backup exists and survivorship rules were declared field by field.
- [ ] Closed Won opportunities were excluded or reconciled before the merge.
- [ ] Server-side duplicate prevention is enabled across imports, workflows, and APIs.
- [ ] A recurring sample and duplicate-rate target are scheduled toward sub-1%.

The verdict you can defend is the one another grader reaches from the same evidence. Write the rules first, gate on identifiers, link the subsidiaries, and log every field that survived. That is the difference between deduplication that cleans the CRM and deduplication that quietly deletes its history.

## Frequently asked questions

### When is it safe to auto-merge company accounts?

Only when a deterministic identifier matches, such as a VAT number, company registration number, or external ID, and no do-not-merge gate fires. Published confidence bands put the auto-merge floor between 0.90 and 0.95, but a score alone is not enough for companies. A high fuzzy score with a shared domain and no confirming identifier belongs in human review, not auto-merge, because that is the exact profile of a subsidiary.

### How do I tell a subsidiary from a duplicate account?

Test for legal-entity independence and separate branding. A subsidiary keeps its own registration number, often its own balance sheet, and frequently trades under a different name for a distinct market. A shared domain and similar name with a conflicting or separate legal ID means link as parent and child, not merge. Duplicates share identity; subsidiaries share ownership but represent distinct go-to-market relationships you must not flatten.

### What should I do with a company that rebranded?

Treat a pure rebrand, meaning the same legal entity under a new name, as a merge-with-alias. Maintain an old-name-to-new-name map so historical references still resolve, the way Facebook resolves to Meta. This is different from an acquisition where the acquired brand still trades under its own name, which is a link-as-hierarchy case, not a merge.

### Why not just use a single merge threshold like contact dedup does?

Because domain, the field most contact rules lean on, fails in both directions for companies. It over-matches on agency and coworking addresses and under-matches when one company holds several domains. A single cutoff cannot express the two company-specific outcomes you actually need: link as hierarchy for subsidiaries and hold-for-status-check for suspected acquisitions or rebrands.

### How do I stop a merge from destroying account history?

Pre-declare survivorship, back up before every run, and keep a field-level decision log. In Salesforce, related records move to the master, non-master records sit in the Recycle Bin for only 15 days, and the master keeps its Record ID. Exclude Closed Won opportunities from merges because they can break commissions and forecasting silently. Survivorship is governance, not an algorithm, so a wrong identity with confident survivorship is worse than no dedup.

### Who should own the duplicate-company verdict?

A named data steward or master-data owner runs the verdict, and RevOps owns the ruleset. In Refolk's index, 1,316 people in the US hold Data Governance, Data Steward, or Master Data Manager titles against 1,072 Revenue Operations Managers, so the pools are comparable in size. The scarce combination is RevOps who also list data-quality skill, so write the rules down rather than relying on one person's judgment.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/duplicate-company-verdict*
