Refolk
StandardProcess, data, and compliance

The AI-Inferred Field Standard: Trust, Flag, or Reject

You can grade any AI-inferred field on a sourced record as Trust, Flag, or Reject using a measured accuracy bar, a provenance log, and a sampling check.

17 min readLast reviewed October 7, 2026Read as Markdown

Key takeaways

  • Calibration, not raw accuracy, decides auto-write: a field at 85% accuracy with low Expected Calibration Error is safer to threshold than one at 90% accuracy whose confidence scores are untrustworthy.
  • Title data is the quiet exception - it decays 25-35% a year yet tests above 89% accuracy across vendors because titles are re-published on profiles, making it Trust-eligible while mobile numbers at 41-67% match are not.
  • The judge is a bigger risk than the data: the same architecture spans Cohen's kappa 0.04 to 0.85, and one deployed judge over-flagged 77.1% of human-faithful records, which would route an entire batch to humans and destroy the economics of automation.
  • GDPR turns inference into a routing rule: because an inferred field is an opinion and significant-effect actions need human involvement, person-level inferred fields default to Flag regardless of confidence.
  • Manual verification does not scale as a fallback - one test took 143 hours to hand-verify 10,000 contacts at 91%, which is why acceptance sampling verifies a statistically sized sample of 52 to 315 units, not the whole lot.
  • Vendor headline accuracy of 91% to 98% collapses to roughly 65-70% in real use with 20-30% bounce, so grade on your own measured accuracy, never the marketed number.

You ran your sourced people and company list through AI enrichment, and now every record carries fields a model inferred rather than retrieved from a source. This standard is for the RevOps, recruiting ops, or data lead answerable for that data, and it gives you a gradeable rule - Trust, Flag, or Reject - that two teammates would apply the same way, backed by a measured accuracy bar, a required provenance log, and a pre-publish sampling check you can adopt as team policy.

Most enrichment guides stop at provider waterfalls, conflicting values, and staleness. None of them sets a pass/fail bar for data an LLM made up rather than looked up. That is the gap this fills. The top search results argue AI enrichment is either magic or dangerous without ever telling you which field to trust today. This document tells you which field, on what evidence, and what to do when the evidence is thin.

What counts as an AI-inferred field, and why it needs its own standard

An AI-inferred field is any value a model generated by reasoning over evidence rather than copying it from a named source. A retrieved email is a fact lifted from a record. An inferred seniority, a guessed company stage, or a derived "likely decision-maker" flag is an opinion the model produced, and the distinction changes everything about how you treat it.

The reason this needs a separate standard is legal as much as operational. Published data-lineage practice treats provenance and lineage as distinct: provenance tracks the origin and creation of data, lineage tracks its transformation across its life cycle. For an inferred field, both matter, because you need to know both what grounded the inference and what transformed it. The ICO accuracy checklist is explicit that you must record the source of collected data and keep records that clearly identify any matters of opinion. An inferred field is a matter of opinion by definition, so it falls under that requirement whether or not you planned for it to.

The second legal hook is automated profiling. Under GDPR, a data subject has the right not to be subject to a decision based solely on automated processing, including profiling, that produces legal or similarly significant effects. If an inferred field drives such a decision, human involvement is not optional. This is why the standard below routes person-level inferences differently from company-level ones.

How field type and source set the baseline grade

Field type is the first filter, because measured accuracy and decay vary enormously across fields, and the grade follows the numbers. Email and mobile fields are high-decay and high-risk; job titles are the quiet exception; company names are the most durable.

Here is the measured picture. Email addresses decay about 43% a year and tested at 78% accuracy for one vendor and 84% for another in an independent 1,000-lead test - both below the numbers each vendor markets. Mobile phone match landed at just 41% to 67% in the same test. Job titles decay fast too, at 25-35% a year, yet tested above 89% for both vendors, because titles are frequently re-published on public profiles, so high observability offsets high churn.

Field typeAnnual decayMeasured accuracyGrade implication
Email~43%/yr78-84%Flag, verify before use
Job title25-35%/yrabove 89%Trust-eligible if calibrated
Mobile phone~20-25%/yr41-67% matchFlag or Reject
Company nameslowestnot isolated in testTrust-eligible

Read this table as the floor, not the verdict. A field type being Trust-eligible means it can earn a Trust grade once calibration is measured, not that you auto-write it blind. Title is the clearest case of why the two-axis read matters: it decays faster than email but tests more accurately, so decay alone would mis-grade it.

143 hours
time to hand-verify 10,000 contacts at 91% accuracy in one vendor test
Manual verification does not scale as a fallback, which is the whole reason acceptance sampling exists.

The second axis is source: person-data versus company-data. A company's inferred stage affects no individual's rights. An individual's inferred seniority does, and GDPR makes that difference a routing rule rather than a footnote. Tag both axes at field level, not table level, before you grade anything.

Why calibration, not accuracy, decides auto-write

The number that decides whether you can auto-write a field is not its accuracy rate but its calibration: how closely the model's stated confidence tracks its real accuracy. The confidence score is what routes each individual field, so a model whose scores lie is dangerous even when its average accuracy is high.

Calibration is measured by Expected Calibration Error, the gap between stated confidence and actual accuracy. The documented automation pattern is selective: high-confidence predictions process automatically, uncertain cases get flagged for human review. That pattern only works if the confidence number is honest. By default it is not - LLMs often exhibit misaligned confidence scores, usually overestimating the reliability of their own predictions. One study across seven models found self-reported confidence achieved the best calibration of the methods tested, at an average ECE of 0.166 versus 0.229 for self-consistency, but even the best was far from perfect.

The practical consequence: a model at 85% accuracy with low ECE is safer to threshold than one at 90% accuracy with high ECE, because the second one produces confident wrong answers you cannot filter out. Auto-writing a title the model reported at 0.95 confidence and got wrong is exactly the failure a high ECE hides.

The calibration-accuracy quadrant

Well calibrated (low ECE)Poorly calibrated (high ECE)
Poorly calibrated, low accuracy
Reject the field type; neither right nor filterable.
Poorly calibrated, high accuracy
Flag all; confident wrong answers slip through thresholds.
Well calibrated, low accuracy
Flag; low confidence correctly routes most to humans.
Well calibrated, high accuracy
Trust-eligible; selective auto-write on high-confidence slice.
Low accuracyHigh accuracy
Accuracy tells you how often the field is right; calibration tells you whether you can believe the confidence score that routes it.

The Trust, Flag, Reject grading rule

Grade every inferred field into one of three buckets with a written rule keyed to measured accuracy, calibration, and legal risk. Trust means auto-write. Flag means hold in staging for human verification before use. Reject means discard the inferred value entirely.

Here is the rule, stated so two teammates apply it identically:

  • Trust (auto-write): Company-data fields only, where field-type accuracy is measured above your written bar, ECE is low on your holdout, and the confidence score clears your calibrated threshold. A provenance record with grounding evidence and timestamp must exist.
  • Flag (verify before use): Any field where accuracy is marginal (roughly 70-89%), calibration is mediocre, or the value is person-data acted on with significant effect. Email and mobile default here. The field sits in staging until a human confirms it.
  • Reject (discard): Any field below your minimum accuracy bar, any field with no grounding evidence, any provenance record missing a timestamp, and any mobile-match type performing near the 41% floor.

The person-data override is absolute: an inferred field about a person that drives a significant-effect action routes to Flag by default, regardless of confidence, because Article 22 requires human involvement and the ICO requires opinions be flagged. Confidence can move a company field from Flag to Trust. It cannot do the same for a person-level inference that triggers a consequential decision.

Field grading rubric (paste into your data policy)
GRADE = TRUST  when: field is company-data
                 AND measured accuracy >= 90% on our holdout
                 AND ECE <= 0.10 for this field type
                 AND confidence >= calibrated threshold
                 AND provenance record has evidence + timestamp
GRADE = FLAG   when: field is person-data with significant effect (always)
                 OR measured accuracy 70-89%
                 OR ECE 0.10-0.20
                 OR confidence below threshold
GRADE = REJECT when: measured accuracy < 70%
                 OR no grounding evidence linkable to a source
                 OR provenance record missing timestamp or model version

Set the accuracy bars to your own measured numbers; the defaults below match the dossier's test figures.

Writing a rubric is one thing; staffing it is another. In Refolk's index, "Revenue Operations Manager" returns 1,041 profiles in the US against 207 in the UK, and the "Data Governance" skill returns 57,667 US profiles against 35,113 for "Data Quality". Those pools tell you who you are hiring into the owner role and how thinly the governance skill is spread. Refolk lets you ask for exactly that owner profile in plain English and get the list back, instead of reverse-engineering a boolean string.

SegmentCountDerived ratio
RevOps Manager, US1,041baseline
RevOps Manager, UK207UK is ~20% of US
"Data Governance" skill, US57,667baseline
"Data Quality" skill, US35,113~61% of governance pool

Running the standard: the procedure

The procedure moves a batch from raw enrichment output to graded, logged write-back. It runs in order, though in practice a cheap LLM pre-screen (step 6) often runs before human sampling (step 5) so you sample only the survivors.

From enriched batch to graded write-back

  1. Classify each field by type and source
    Tag every field as inferred versus retrieved and as person-data versus company-data. Field-level granularity is required, not table-level.
  2. Attach a provenance record to every inferred field
    Each inferred value carries grounding evidence, a timestamp, model name and version, the transformation applied, and a confidence score. The ICO requires the source be recorded and the opinion flagged.
  3. Measure calibration before trusting any confidence score
    Compute Expected Calibration Error per field type on your own labelled holdout. Set a threshold only where confidence tracks accuracy; do not inherit a vendor threshold untested.
  4. Set Trust, Flag, and Reject bars per field type
    Write a rule two teammates apply identically, keyed to measured accuracy and legal risk, with person-data significant-effect actions routing to human review by default.
  5. Run pre-publish acceptance sampling on each batch
    Draw a random sample sized by ISO 2859-1 for the lot, verify independently against ground truth, and accept or reject by the AQL number.
  6. Add a Pass, Fail, or Unable-to-verify verdict layer
    Measure the judge's agreement with human gold in kappa, precision, and recall before trusting it. Route borderline and unable-to-verify cases to humans.
  7. Write back and log every decision
    Auto-write only Trust-graded fields, hold Flag in staging, discard Reject. Log every decision and its provenance for audit and rectification.
  8. Re-verify on a schedule
    Re-sample decayed fields at least quarterly and keep a rectification path open for data-subject challenges.

The grading pipeline

  1. Enrich
    Model produces inferred values with raw confidence
  2. Log provenance
    Attach evidence, timestamp, model version per field
  3. Grade
    Apply Trust/Flag/Reject rubric against measured accuracy and ECE
  4. Sample + judge
    ISO 2859-1 sample plus validated judge verdict
  5. Write back
    Trust auto-writes, Flag to staging, Reject discarded, all logged
Each field passes through provenance capture, calibration-gated grading, sampling, and a verdict layer before any auto-write.

How to size the pre-publish sampling check

Size your pre-publish check with acceptance sampling, which draws a documented sample by lot size and acceptable quality limit rather than a flat percentage. ISO 2859-1 (equivalently ANSI-ASQ Z1.4) gives you the sample size and the accept/reject numbers, so the check is reproducible and defensible.

At General Inspection Level II, a 1,500-unit lot yields a sample of 125, an 8,000-unit lot yields 200, and a 15,000-unit lot yields 315. The acceptable quality limit (AQL) sets how strict you are: an AQL of 2.5% means lots with 2.5% or fewer defective items are accepted roughly 95% of the time. A worked Minitab plan at AQL 1.5 and RQL 10 with alpha 0.05 and beta 0.1 gives a sample of 52: accept if 2 or fewer are defective, reject at 3 or more.

Lot sizeSample size (Level II)Note
1,500125Code K
8,000200Code L
15,000315Code M
52 (Minitab AQL 1.5 plan)52Accept at 2, reject at 3

The reason you sample rather than verify everything is cost. One team spent 143 hours hand-verifying 10,000 contacts to confirm 91% accuracy. A statistically sized sample of 52 to 315 gives you a defensible accept/reject verdict in hours. Draw the sample randomly, verify each item against ground truth independently of the model, and record the verdict against the lot.

You verify a statistically sized sample, not the lot, because 143 hours per 10,000 contacts is not a fallback, it is a bottleneck.

The verdict layer: when the judge is the bigger risk

Before an LLM judge gets to Pass or Fail your records, measure the judge itself, because a bad judge does more damage than bad data. The same judge architecture spans a Cohen's kappa of 0.04 to 0.85 depending on validation, and an un-validated one can wrongly route your entire batch to humans.

The spread is not academic. A validated judge hit 92.31% raw agreement with human gold, Cohen's kappa 0.85, precision 1.00, recall 0.87, and F1 0.93 on a 52-item set. A deployed judge agreed with gold at only kappa 0.04 on a disagreement-enriched set and over-flagged 77.1% of human-faithful cases. That second judge looks like a quality crisis in your data. It is actually a broken judge destroying the economics of automation by sending correct records to humans.

SystemAgreementPrecisionOver-flag rate
Validated judge92.31%1.00low
Deployed judgekappa 0.04not reliable77.1%

A production verdict scheme should offer three outcomes, not two: Pass, Fail, and Unable-to-verify (insufficient support). The judge should enumerate each claim, check it against the supporting passage, verify entity attribution, and flag truncated or missing context as insufficient support rather than guessing. Route Fail and Unable-to-verify to humans.

There is a reproducibility trap underneath all of this. The exact same trace can receive a Pass on a Tuesday and a Fail on a Wednesday because of token sampling. A standard that claims two people grade the same case the same way cannot rest on a single judge call. Mandate a multi-sample jury and require agreement, or the verdict layer violates the standard's own premise.

How this goes wrong: failure modes and false positives

The failures below are the most valuable part of this standard, because each one produces an output that looks correct and is not. For every signal that can lie, here is what the lie looks like and how to catch it.

  • Trusting vendor accuracy claims. Headline figures of 91% to 98% collapse to roughly 65-70% real accuracy with 20-30% bounce in use. Check: mail a sample and measure the actual hard-bounce rate against your own data, not the marketed number.
  • Trusting raw confidence. Models are overconfident, so a high score with high ECE means confident wrong answers. The false positive is auto-writing a 0.95-confidence wrong title. Check: compute ECE on a labelled holdout before any threshold is live.
  • Judge over-flagging. A deployed judge flagged 77.1% of correct cases, which reads as a data crisis but is a broken judge. Check: measure judge-versus-human kappa before acting on any verdict.
  • Judge non-determinism. The same record earns different verdicts across runs. Check: use a multi-sample jury and require agreement before the verdict counts.
  • Catch-all emails. They validate as deliverable yet still bounce on send. The false positive is a "valid" email that fails. Check: segment catch-all domains and verify them separately.
  • Treating inference as fact. An inferred seniority is an opinion, and GDPR requires opinions be flagged. Check: the provenance record must carry grounding evidence, not just a value.
  • Format-polished hallucination. Outputs that look authentic but invent content, such as hallucinated citations. Check: require grounding evidence linkable to a real source passage.
  • Stale "verified" flag. A once-accurate field with no re-verification date drifts silently. Check: reject any provenance record that lacks a timestamp.

Two of these deserve extra weight. The first, vendor accuracy inflation, is why the whole standard grades on your measured numbers rather than any marketed figure. The third, judge over-flagging, is the one most likely to make you distrust clean data and tear up a working pipeline. When quality metrics suddenly crater, suspect the judge before the data.

The pre-publish checklist and keeping the standard current

Before you call any batch done, run this checklist. It is the gate between graded output and a live system of record, and it is written so a reviewer can tick each item or stop the batch.

Pre-publish gate for an enriched batch

  • Every field is tagged inferred versus retrieved and person versus company, at field level.
  • Every inferred field carries a provenance record with evidence, timestamp, model name and version, transformation, and confidence.
  • Expected Calibration Error has been measured per field type on our own labelled holdout.
  • The Trust/Flag/Reject bars are written down and keyed to our measured accuracy, not a vendor's.
  • Person-data fields driving significant-effect actions are routed to Flag regardless of confidence.
  • An ISO 2859-1 sample was drawn by lot size and the accept/reject verdict is recorded.
  • The judge's agreement with human gold is measured in kappa and precision before its verdicts were trusted.
  • Verdicts ran as a multi-sample jury, not a single call.
  • Only Trust-graded fields auto-wrote; Flag sits in staging; Reject is discarded.
  • Every decision and its provenance are logged, and a rectification path is open for data-subject challenges.

Keeping the standard current is mechanical, not one-off. Decay is the clock: B2B contact data decays about 2.1% a month, compounding to roughly 22.5% a year, and ZeroBounce's analysis of over 11 billion emails found at least 23% of a typical list degrades per year. Email at 3.6% a month and titles at 2-3% a month decay faster still. Re-sample decayed fields at least quarterly and re-grade them, because a Trust grade is a snapshot, not a permanent status.

Two mechanisms will shift the bars over time without any calendar date attached. First, your measured accuracy and ECE change as your model and vendors change, so re-run the holdout whenever you swap either. Second, your judge's kappa drifts as prompts and models update, so re-validate it against a fresh gold set on the same cadence. When a data subject challenges an inferred field, GDPR Article 5(1)(d) requires you rectify or erase it without delay, which is why the provenance log and the staging buffer are not optional - they are the machinery that makes rectification possible. Re-check these three things, the holdout, the judge, and the decay clock, and the standard stays honest.

Questions practitioners ask

When can I trust AI enriched CRM data enough to auto-write it?

Auto-write only when the field type has measured accuracy on your own data and a low Expected Calibration Error, so the confidence score actually tracks accuracy. Company fields and calibrated job titles are usually Trust-eligible. Person-level inferences acted on with significant effect default to Flag under GDPR regardless of confidence. Never trust a vendor's headline accuracy, which commonly drops from 91-98% marketed to 65-70% in real use.

How do I validate LLM inferred fields before I act on them?

Measure the judge, not just the data. Run the field through a validated judge whose agreement with human gold you have scored in Cohen's kappa, precision, and recall. One deployed judge over-flagged 77.1% of correct records at kappa 0.04, while a validated one hit 92.31% agreement at kappa 0.85. Because a single judge call can flip Pass to Fail run to run, require a multi-sample jury, then human-sample the survivors with acceptance sampling.

What does GDPR require for automated profiling and inferred fields?

An inferred field is a matter of opinion, not fact, so the ICO accuracy checklist requires you record its source and flag it as an opinion. Article 22 restricts solely-automated decisions with legal or similarly significant effects and lets the data subject obtain human intervention. In practice this means person-level inferences that drive significant-effect actions route to human review by default, and Article 5(1)(d) requires inaccurate personal data be rectified or erased without delay.

How big a sample do I need for a pre-publish quality check?

Use ISO 2859-1 acceptance sampling, which sizes the sample by lot size and acceptable quality limit rather than by percentage. At General Inspection Level II a 1,500-unit lot needs 125 samples, 8,000 needs 200, and 15,000 needs 315. A worked Minitab plan at AQL 1.5 and RQL 10 gives a sample of 52: accept if 2 or fewer are defective, reject at 3. This is why you sample rather than hand-verify a whole lot, which took one team 143 hours for 10,000 contacts.

Why is calibration more important than accuracy for deciding to trust a field?

Because the confidence score, not the accuracy rate, is what routes each field. A model at 85% accuracy with low Expected Calibration Error is safer to threshold than one at 90% accuracy with high ECE, since the latter produces confident wrong answers you cannot filter. Self-reported confidence averaged ECE 0.166 against 0.229 for self-consistency across seven models, but calibration is unreliable by default, so you must measure ECE on your own labelled holdout before any threshold is safe.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next