# The Briefing Readiness Standard: When Talent and Market Findings Can Ship

*You can grade every line of a talent-and-market briefing as fact, estimate, or assumption, score its source and freshness, and reach a repeatable ship-or-hold call.*

- Canonical URL: https://www.refolk.ai/guides/briefing-readiness-standard
- Pillar: Market and talent intelligence
- Format: Standard
- Published: 2026-08-20
- Last reviewed: 2026-08-20
- Reading time: 16 min
- Keywords: how to label analyst confidence in a report, distinguishing fact from estimate in research, market intelligence report review checklist, when is a competitive intelligence report done, sourcing claims in a market map

## Key takeaways

- Grade every briefing claim on two separate axes: source reliability (A to F) and information credibility (1 to 6), because collapsing them lets a reliable source launder a weak claim.
- People-field claims decay far faster than aggregate data: job titles change at 65.8% per year against a 22.5% aggregate contact-record rate, so a 'current title' line is the most fragile assertion in any talent brief.
- Corroboration only counts when the second source is independent in origin, not the same primary fact republished; two citations tracing to one press release rate as single-source.
- Never combine a confidence level and a likelihood term in the same sentence, per ICD 203, because it hides what exactly is uncertain.
- Score reviewer agreement on pre-discussion labels only; consensus reached after a conversation measures politeness, not repeatability.
- In Refolk's index the UK CI pool of 24 is roughly one-tenth the US pool of 246, so a market-sizing claim built on UK headcount rests on a thinner base and warrants an automatic downgrade.

Before a competitor-and-talent-pool briefing goes in front of leadership, you need to know every claim in it will hold up under scrutiny. This is a claim-level standard for strategy, research, and talent-intelligence teams: a way to grade a briefing line by line so that fact is separated from estimate, every source is characterized, confidence is stated in fixed language, and two reviewers reach the same ship-or-hold decision on the same draft.

Most research standards cover a single deliverable - a market-size estimate, a pay range, an adoption read. Public competitive-intelligence material stops at program-level monitoring checklists. Neither gives you a definition of done for a briefing that mixes people assertions and company assertions in the same document. This standard adapts structured analytic tradecraft - fact-versus-judgment separation, source characterization, and calibrated confidence language - into a commercial grading rubric you can apply before you hit send.

## What "briefing readiness" means as a definition of done

A briefing is ready when every factual claim in it has been graded on four axes - type, source strength, corroboration, and freshness - and a second reviewer reaches the same ship-or-hold call working from the pre-discussion labels. Readiness is not "the deck looks finished." It is a property of each individual claim, verified twice.

The standard rests on three separations that professional analysts treat as non-negotiable, and that most commercial briefings quietly collapse:

- **Fact versus judgment.** ICD 203, the US Intelligence Community's analytic-standards directive first issued in 2007 and amended on January 21, 2022, requires analysts to separate underlying information from assumptions and judgments. A briefing that states an estimate in the grammar of a fact fails before it starts.
- **Source reliability versus information credibility.** The Admiralty code, used across defence and security, rates these on two independent scales precisely so a trustworthy source cannot silently upgrade a shaky claim.
- **Likelihood versus confidence.** ICD 203 explicitly prohibits combining a confidence level and a likelihood term in the same sentence, because mixing them confuses the reader about what exactly is uncertain.

If your briefing honors those three separations for every claim, and a second reader independently agrees, it is ready. If it does not, it is a draft.

> **Rule:** The ship gate
>
> Ship only if no linchpin claim rests on a single low-reliability source or a decayed field. A linchpin claim is one that, if wrong, changes the recommendation.

## The two-axis source grade you apply to every claim

Grade each claim on two separate scales: how reliable the source is, and how credible the specific information is. The Admiralty code combines them into a two-part grade such as B2 or F6, and the two halves must be assigned in separate passes.

Source reliability runs from A, a history of complete reliability, down to E, a history of invalid information, with F reserved for a source without sufficient track record to judge. Information credibility runs from 1, confirmed, to 5, improbable, with 6 for information whose reliability cannot be evaluated. The two axes are independent by design: a reliable source can still pass on bad information, and a questionable source can still deliver intelligence later confirmed.

The failure this prevents is real and documented. Analysts tend to grade on the diagonal - A1, B2, C3 - collapsing two judgments into one. When that happens, a reliable source auto-upgrades a weak claim, which is exactly the laundering the two-axis system exists to stop.

> **Watch out:** Grade the two axes in separate passes
>
> If you find your grades clustering on A1, B2, and C3, you are collapsing the two scales into one. Grade every claim's source reliability first, then start over and grade every claim's information credibility. The diagonal is a known bias, not a coincidence.

Under STANAG 2511, a claim earns credibility rating 1, "confirmed by other sources," only when the reported information can be stated with certainty to originate from a source other than the existing information on the same subject. That is a rule about origin, not citation count. Two footnotes that both trace to one press release are one source republished, and the claim stays single-sourced.

#### Reliability versus credibility, and what each corner means

Horizontal axis runs from Low information credibility to High information credibility. Vertical axis runs from Low source reliability to High source reliability.

| Quadrant | What it means |
| --- | --- |
| Reliable source, weak claim | Do not let the source upgrade the claim; hold until corroborated. |
| Reliable source, strong claim | Ship-ready on this axis; check freshness next. |
| Weak source and weak claim | Cut it or label it an assumption with a linchpin note. |
| Weak source, strong claim | Corroborate the origin independently before you trust it. |

*A claim's two-part grade tells you which corner it sits in and what to do before it ships.*

## Labeling fact, estimate, and assumption

Every claim gets exactly one of three type labels before anything else happens: fact, estimate, or assumption. This is the first pass, and it decides how hard the later passes have to work.

- **Fact** - an underlying observation with a retrievable citation. "This company lists 40 open engineering roles on its public careers page." Facts still get a source grade and a freshness check, but they are not judgments.
- **Estimate** - a judgment built from evidence. "The addressable pool of competitive-intelligence analysts in this market is roughly 250." Estimates carry calibrated confidence language and must expose the evidence base underneath.
- **Assumption** - something accepted without evidence so the analysis can proceed. Every linchpin assumption needs a "what changes if wrong" note, because ICD 203 requires assumptions to be stated, not smuggled in as facts.

The most common way this pass fails is an assumption dressed as a fact - a load-bearing belief stated flatly with no flag. The reader then treats a guess as an observation. Splitting people-claims from company-claims at inventory time helps here, because people-claims are far more likely to be stale estimates masquerading as current facts.

> An assumption stated as a fact is not a small error. It is the reader inheriting your guess as their evidence.

## The freshness downgrade, and why people-claims decay fastest

Apply an automatic confidence downgrade to any claim resting on data past its decay threshold, and set the threshold by field type, not by one blanket rule. People-fields decay roughly an order of magnitude faster than the aggregate figure most teams have in their heads.

The headline aggregate - B2B contact records decaying about 22.5% per year, or 2.1% per month compounding - hides the real risk. Roughly 70.8% of business contacts see at least one data change within 12 months, and job titles change at 65.8% annually, making title the single fastest-decaying field. Email addresses go bad at about 23% per year, and about 30% of professionals change jobs each year, invalidating their contact details.

**65.8% - Annual decay rate of job titles, the fastest-moving people field**

Against a 22.5% aggregate contact-record rate, a "current title" line is the most fragile claim in any talent brief.

That gap is the practical core of the freshness rule. A "current title" claim is the most fragile line in a talent brief, so it needs the tightest threshold. The downgrade logic is deliberately mechanical so two reviewers apply it identically:

- People-field claim older than about **90 days**: downgrade confidence one band.
- Title or role claim older than **6 to 12 months**: downgrade, and reclassify from fact to estimate.
- Every people-claim carries a **"captured on" date**, or it cannot ship.

| Field | Annual decay | Source |
|---|---|---|
| Aggregate contact record | ~22.5% | HubSpot/MarketingSherpa |
| Job title | 65.8% | Landbase (via SalesHandy) |
| Email address | ~23% | ZeroBounce |
| Job change (person moves) | ~30% | Apollo/Cleanlist |

Trigger-based re-verification beats a calendar because decay concentrates in a few fields: titles and phone numbers move fast, physical addresses slowly. If you can re-pull the people-claims cheaply at review time rather than trusting a months-old capture, you remove the whole downgrade problem instead of managing it. Re-running a live query - for example, [Refolk](/) can resolve "competitive intelligence analysts and managers at US enterprise software companies, currently in role" against its index on the day you review - is often faster than auditing a stale export field by field.

I ran this search: `Competitive intelligence analysts and managers at US enterprise software companies, currently in role.` - [see the full result list](https://www.refolk.ai/s/q05pkvjkhz).

*Returns current-in-role practitioners you can use to refresh a people-claim rather than shipping a months-old title capture.*

## Calibrated confidence language you can standardize

Assign confidence with a fixed lexicon of estimative-probability terms, and use those words verbatim. Vague hedges - "may suggest," "could indicate" - are weasel words that carry no shared meaning and must be banned from a graded briefing.

The Kent schema maps words to probability midpoints with bands, and it is the cleanest table to adopt as team policy. Kesselman's updated list keeps seven such terms, grouped in 15% bands except a 10% middle category, and never reaching complete certainty or impossibility - but the underlying discipline is identical: one term, one meaning, applied consistently.

| Term | Midpoint | Band |
|---|---|---|
| Almost certain | 93% | ±6% |
| Probable | 75% | ±12% |
| Chances about even | 50% | ±10% |
| Probably not | 30% | ±10% |
| Almost certainly not | 7% | ±5% |

The one rule that trips teams up: never put a likelihood word and a confidence level in the same sentence. "High confidence it is likely" reads as strong evidence when it may be neither. Say what you assess ("probable"), then, if needed, separately state how confident you are in that assessment. Kesselman's research is a warning here - across 50 words used in national estimates over five decades, only 13 were statistically distinguishable in how readers interpreted them. Most hedge words do not mean what the writer thinks.

**Claim grade-sheet row**

```
Claim ID | Claim text | People/Company | Type (fact/estimate/assumption) | Source grade (A-F / 1-6) | Corroboration (single / confirmed) | Captured-on date + decay flag | Confidence term (from Kent list) | Linchpin? (Y/N) + what-changes-if-wrong
</template>

## The grading procedure, step by step

Run these nine steps in order. Steps 1 through 6 are the first analyst's work; steps 7 and 8 bring in the second reviewer; step 9 is the joint decision. For a 20-claim brief, budget one to two hours for the inventory and about an hour for the second review.
```

*One row per claim. Adapt column names to your tooling, but keep all seven fields.*

steps
title: Grading a briefing from inventory to ship-or-hold
step: Inventory every claim :: Extract each discrete factual assertion into its own row, splitting people-claims from company-claims. Done when every sentence carrying a factual assertion has a dedicated row.
step: Label type: fact, estimate, or assumption :: Apply the ICD 203 distinction between underlying information, assumptions, and judgments. Done when every row carries one label and every linchpin assumption has a "what changes if wrong" note.
step: Characterize each source :: Assign an Admiralty-style two-part grade, source reliability A to F and information credibility 1 to 6, in two separate passes. Done when each fact or estimate has a grade and a retrievable citation.
step: Count corroboration :: Mark each claim single-source or "confirmed by other sources," verifying the second source is genuinely independent in origin. Done when every claim shows a corroboration count.
step: Apply the freshness downgrade :: Flag people-field claims older than ~90 days and title or role claims older than 6 to 12 months, then downgrade confidence. Done when each claim has a "captured on" date and a decay flag.
step: Assign calibrated confidence language :: Map each estimate to the fixed lexicon and never put a likelihood word and a confidence level in one sentence. Done when all wording is drawn only from the approved list.
step: Run an independent second review :: A second reviewer re-grades type, source strength, and confidence blind to the first reviewer's grades. Done when both grade sheets exist before any discussion.
step: Reconcile and score agreement :: Compute percent agreement on the pre-discussion labels, then hold a consensus meeting to resolve disagreements and update the codebook. Done when disagreements are resolved and definitions clarified.
step: Make the ship-or-hold decision :: Ship only if no linchpin claim rests on a single low-reliability source or a decayed field. Done when both reviewers sign off against that criterion.
```

## Making two reviewers agree, and measuring whether they did

Two reviewers grade the same draft independently, then you measure how consistently they applied the criteria before they talk. Agreement measured after discussion is not reliability - it is consensus, and consensus overwrites how each reviewer originally read the evidence.

Borrow the mechanics from systematic-review and qualitative-coding practice. Each reviewer independently classifies each claim - include, exclude, or uncertain, in review terms; for a briefing, that maps to ship, cut, or hold. The reliability calculation must use the independent pre-discussion decisions, because a final consensus decision cannot reveal how consistently reviewers first applied the criteria.

The procedure that makes this repeatable:

1. **Practice-code a sample together** first, so both reviewers interpret the labels the same way before the real grading.
2. **Grade independently and blind.** Neither reviewer sees the other's sheet.
3. **Score the pre-discussion labels.** Report percent agreement with the sample size n. Cohen's Kappa is an option, but sources disagree on whether Kappa fits interpretive judgments, so treat percent agreement plus n as the safer minimum.
4. **Hold the consensus meeting** to resolve disagreements and update the codebook - after scoring, never before.

#### The two-reviewer readiness loop

1. **Practice-code** - Both reviewers grade a shared sample to align on definitions.
2. **Grade blind** - Each reviewer grades the full draft without seeing the other's sheet.
3. **Score pre-discussion** - Compute percent agreement plus n on the independent labels.
4. **Consensus meeting** - Resolve disagreements, update the codebook, then decide ship or hold.

*Agreement is scored on the blind labels, then consensus fixes the draft and the codebook.*

## How this goes wrong: the failure modes

Most briefings fail readiness in predictable ways. Each of these has a false positive - a specific way it makes the briefing look stronger than it is - and a check that catches it. This section is the part of the standard worth rereading.

| Failure mode | How it lies | The check |
|---|---|---|
| Diagonal grading | A reliable source auto-upgrades a weak claim | Assign the two axes in separate passes |
| Mixing likelihood and confidence | Reader thinks a claim is stronger than the evidence | One axis per sentence, per ICD 203 |
| Stale people-data as fact | A title captured 8 months ago is likely wrong | Require a "captured on" date; downgrade past threshold |
| Fake corroboration | Two citations from one press release read as "confirmed" | Verify independent origin before rating credibility 1 |
| Consensus masking disagreement | Post-discussion agreement hides real inconsistency | Score pre-discussion labels only |

Two more deserve their own note because they are subtle:

- **Assumption smuggled as fact.** A linchpin assumption stated flatly, with no flag, so the reader inherits your guess as evidence. The check is procedural: ICD 203 requires you to state the assumption and what changes if it is wrong. If a claim would flip the recommendation and has no evidence under it, it is an assumption and must be labeled one.
- **Kappa theater.** Reporting "substantial agreement" without the coefficient or the sample size. It signals rigor while hiding whether the reviewers actually agreed. The check: report the percent agreement figure and the number of claims it was computed over, or report nothing.

> **Tip:** When the base is thin, downgrade by default
>
> Small talent pools change the confidence math. In Refolk's index the UK competitive-intelligence pool is 24 people against 246 in the US, roughly one-tenth. A market-sizing estimate built on the smaller base rests on a far thinner foundation, so apply an automatic confidence downgrade for claims drawn from small populations, regardless of how the arithmetic looks.

That last point deserves its own emphasis, because it is where the people-side and market-side of a briefing collide. A sizing estimate is only as strong as the population it is derived from, and headcount ratios expose the difference immediately.

| Segment | Total people | Derived ratio |
|---|---|---|
| CI Analyst/Manager - US | 246 | baseline |
| CI Analyst/Manager - UK | 24 | 0.10x US |
| Market Intelligence Analyst/Manager - US | 337 | 1.37x US CI |

**10x - How much larger the US CI analyst pool is than the UK pool in Refolk's index**

246 in the US against 24 in the UK, so a UK-based sizing claim warrants an automatic confidence downgrade.

## The readiness checklist

Run this before you send. If any item fails, the briefing holds.

#### Ship-or-hold verification

- [ ] Every factual assertion has its own row, with people-claims split from company-claims.
- [ ] Every row is labeled fact, estimate, or assumption.
- [ ] Every linchpin assumption has a "what changes if wrong" note.
- [ ] Every fact and estimate has a two-part source grade assigned in separate passes.
- [ ] Every claim shows a corroboration count, with second sources verified independent in origin.
- [ ] Every people-claim has a "captured on" date; title claims older than 6 to 12 months are downgraded.
- [ ] All confidence wording comes from the approved lexicon, with no likelihood-plus-confidence sentences.
- [ ] A second reviewer graded the draft blind, and pre-discussion percent agreement is recorded with n.
- [ ] No linchpin claim rests on a single low-reliability source or a decayed field.

## Keeping the standard current

A standard is only useful if it stays anchored to its sources, and two anchors in this one are load-bearing enough to re-check periodically. First, the exact high, moderate, and low confidence band definitions in the IC directive were not fully retrieved for this guide - if your team adopts a three-band scheme, verify the band language against the directive text itself before you make it policy. Second, the decay figures drive the freshness downgrade, and decay benchmarks shift as labor markets move; re-pull the aggregate and per-field rates rather than trusting a number that ages like the data it describes.

The freshness rule contains its own maintenance instruction. Because job titles decay at 65.8% a year, the cheapest way to keep a talent briefing current is not to re-audit stale captures but to re-run the underlying people-query at review time and grade the live result. When the population you are sizing is small, treat that live re-pull as mandatory, not optional. A briefing that was ready three months ago is a draft again today, and this standard is the fastest way to prove it either way.

## Frequently asked questions

### How do I label analyst confidence in a report without sounding vague?

Pick a fixed lexicon and use it verbatim. The Kent schema maps words to probability midpoints: 'probable' means about 75% and 'almost certain' about 93%. Assign one term per estimate, and never combine a confidence level with a likelihood word in the same sentence, because that hides which part is uncertain. Restricting wording to the approved list is what makes two reviewers land on the same reading.

### What is the difference between a fact and an estimate in a research briefing?

A fact is an underlying observation you can point to with a retrievable citation, such as a company's stated headcount on a public record. An estimate is a judgment you build from evidence, like a projected pool size. An assumption is something you accept without evidence to make the analysis work. ICD 203 requires you to separate the three so the reader can see where observation ends and judgment begins.

### When is a competitive intelligence report actually done?

It is done when every claim has been graded by type, source, corroboration, and freshness, and a second reviewer reaches the same ship-or-hold call on the pre-discussion labels. The hard gate: ship only if no linchpin claim rests on a single low-reliability source or a decayed field. Anything short of that is a draft, not a briefing.

### How fast does people data go stale in a talent brief?

Fast, and unevenly. Aggregate contact records decay at about 22.5% per year, but job titles change at 65.8% per year, making a 'current title' claim the single most fragile line in a talent brief. About 30% of professionals change jobs annually. Require a 'captured on' date and downgrade any title claim older than 6 to 12 months.

### How do two reviewers grade the same briefing consistently?

Both grade independently and blind to each other, then compute agreement on those pre-discussion labels using percent agreement, and only afterward hold a consensus meeting. Measuring agreement after a conversation overwrites how each reviewer originally applied the criteria, so it tells you nothing about repeatability. Practice-code a small sample together first to align interpretations.

### Do two citations count as corroboration?

Only if their origins are independent. Two sources that both trace back to the same press release are one source republished, so the claim is still single-sourced. Under the Admiralty code, a claim rates as 'confirmed by other sources' only when the second report demonstrably originates elsewhere. Check origin, not citation count.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/briefing-readiness-standard*
