# The Decision-Grade Intelligence Standard: Brief It, Caveat It, or Hold

*You will be able to grade any single intelligence finding as decision-grade, directional, or hold, using a rubric explicit enough that two analysts score it the same way.*

- Canonical URL: https://www.refolk.ai/guides/decision-grade-intelligence-standard
- Pillar: Market and talent intelligence
- Format: Standard
- Published: 2026-09-23
- Last reviewed: 2026-09-23
- Reading time: 18 min

You have a finding about a company and its people, and someone is about to act on it. This standard tells you whether that single finding is solid enough to put in front of a decision-maker, and it is written for market-intelligence analysts, talent-intelligence teams, and operators sizing a market. It grades one finding at the moment before it enters a brief, using a rubric explicit enough that two analysts score the same finding the same way.

Most readiness standards grade a whole deliverable: a competitor profile, a named-candidate list, a market map. But analysts do not get burned by whole deliverables. They get burned by one finding inside an otherwise sound brief - a stale job posting read as live demand, a headcount gap that was definitional all along, three citations that turned out to be one press release. This standard borrows the intelligence community's source-reliability and analytic-confidence discipline and adapts it to findings built from public company and people signals, so you get a concrete rubric instead of "verify your sources" or an abstract military code.

## What "decision-grade" means, and the two axes it rests on

A finding is decision-grade when a decision-maker can act on it without you standing next to them to caveat it. That threshold rests on two independent judgements the intelligence community keeps deliberately separate: how reliable the source is, and how credible this specific claim is.

The Admiralty code, codified in NATO AJP-2.1, rates these on two axes and notates them as a pair such as B2. Reliability runs A to F and asks about the source's history. Credibility runs 1 to 6 and asks about the piece of information in front of you right now. Each descriptor is considered in isolation so that the reliability of the source does not influence the assessed accuracy of the report.

| Axis | Scale | The question it answers |
|---|---|---|
| Source reliability | A to F | What is this source's track record, independent of what it says now? |
| Information credibility | 1 to 6 | Is this specific claim plausible, consistent, and corroborated? |

Reliability A means completely reliable, running down through B usually reliable, C fairly reliable, D not usually reliable, E unreliable, to F cannot be judged. Credibility 1 means confirmed, then 2 probably true, 3 possibly true, 4 doubtfully true, 5 improbable, and 6 cannot be judged. A regulator filing feed or an issuer's own 10-K is a high-reliability source; a paywalled trade publication is mid; anonymous social is low. But reliability is about the channel, not the claim - which is why the two passes stay separate.

> **Rule:** Score the two axes in separate passes
>
> Rate reliability without looking at the specific claim, then rate credibility without re-litigating the source. A trusted source can carry a weak claim, and the pair B4 tells the reader exactly that.

The commercial version of this discipline is thin. Practitioner CI guides tend to say "cross-verify" and stop. This standard makes the pair explicit because the pair is where two analysts either converge or diverge.

## How many independent sources it takes to reach "confirmed"

Exactly one. Under the Admiralty scale, a single additional independent source lifts a finding to the top credibility grade of 1, provided that grade rests on independent corroboration from a separate reporting chain, internal consistency, and alignment with known information.

That is the single most useful fact in this guide, and it cuts against instinct. Adding a fourth citation that traces to the same origin buys nothing. Adding one citation from a genuinely separate collection path changes the grade. Credibility rewards independence, not volume.

> The jump from single-source to confirmed is exactly one independent chain, not a pile of citations.

The journalistic standard is stricter in practice. The New York Times imposes a minimum that a fact not attributed to a single named speaker be verified by at least two independent sources. For a finding that will drive a real decision - a hire, a spend, a market entry - I hold to the stricter bar: at least two genuinely independent chains before I call something decision-grade, and I reserve the single-corroboration grade for directional.

A worked example of a credibility-1 finding: a report that a competitor is planning to exit a segment reaches a 1 when separate, unrelated sources point to the same conclusion through different channels - a trade publication citing insiders, a job-posting pattern showing hiring freezes, and a supplier reporting canceled orders. Three channels, three collection paths, one conclusion. That is corroboration. Three outlets reprinting one press release is not.

**2 - Independent sources for a fact not tied to a named speaker**

The New York Times minimum, and my recommended bar for a decision-grade label.

## The circularity check, as its own gate

Circular reporting is information that looks multi-sourced but in reality traces to a single origin. It is the failure that produces a confident, wrong grade, and it deserves its own step because it is the exact way "cross-verify" fails.

Because circular reporting can happen inadvertently, extra care is needed to confirm that multiple sources actually are independent rather than interconnected in some obscure manner. The mechanism is quiet: a startup publishes a blog post, a newsletter summarizes it, an analyst cites the newsletter, and a second analyst cites the first analyst. Four "sources," one origin.

The check is mechanical. Trace each supposedly independent source to its root - the actual URL, dataset, filing, or person it came from. Collapse any that share a root. Then count what remains. If three citations collapse to one press release, you have one source and a credibility grade no higher than what a single origin earns.

#### The circularity trace

1. **List sources** - Write out every citation supporting the claim
2. **Trace to root** - Find the original filing, dataset, post, or person behind each
3. **Collapse shared roots** - Merge any citations that trace to the same origin
4. **Count chains** - The remaining count is your true independent-source number

*Every finding earns its credibility grade only after sources are traced to root and collapsed.*

Many practitioner CI workflows fold this into a generic cross-verify step and never separate it. Treat it as its own gate. A finding that passed cross-verification but never had its roots traced is not decision-grade; it is untested.

## Shelf-life: when a true finding goes stale

A finding can be correctly graded and still be wrong by the time it is briefed, because the signal underneath it expired. Freshness is not a recency nicety - it is a credibility signal, because for signals like job postings, age proxies for whether the event is even real.

Job postings decay fastest. A posting under seven days old is a strong signal, but one older than 14 days should be treated as weak, and most genuinely urgent roles are filled or substantially progressed within three to four weeks. The quality split is stark: about 52% of postings from always-hiring companies are ghost jobs that never fill, while about 83% of fresh postings under seven days old from mid-market companies represent genuine hiring intent. The same "they are hiring" finding flips grade purely on the age of the posting.

| Signal | Fresh window | Re-verify / weak-signal threshold |
|---|---|---|
| Job posting | Under 7 days | Weak after 14 days |
| Urgent role open | - | Likely filled by 3-4 weeks |
| Corporate-event hiring | - | Window is 30-90 days |
| Market/CI report cadence | - | Quarterly refresh |

Corporate events carry a longer but bounded window: acquisitions, product launches, and tech-stack changes drive hiring surges within the same 30 to 90 day window after the announcement. A hiring burst - three or more new postings for the same role within 14 days - is a signal in its own right, but only when it has breadth as well as freshness.

> **Watch out:** One fresh data point is not a trend
>
> A single six-day-old burst treated as acceleration is a false positive. Require breadth plus freshness, not volume alone. Three postings for one role can be one requisition reposted, not three hires.

Record a last-verified date separately from the source date. The source date tells you when the signal was created; the last-verified date tells the reader when you last confirmed it was still live. Set a refresh cadence - monthly or quarterly for most markets - and the shelf-life stamp becomes an auditable field rather than a memory.

## The disclosures a number must carry

Before you cite a headcount or a posting count, the finding must disclose how the number was built, because headcount disagreements are usually definitional, not factual. Two analysts can "disagree" on a company's size without either being wrong.

The recurring gap is workforce definition. Some sources exclude contractors and count only direct employees; other platforms include anyone who lists the company. A firm with 800 full-time staff and 200 contractors can read as 800 in one source and 1,000 in another. Compare the contractor-inclusive number against a filings FTE number and you manufacture a growth spike that never happened.

A citable number needs five disclosures:

- **Period.** The date range the count covers.
- **Geography.** Which countries or regions are included.
- **Deduplication method.** How records for the same real-world entity were merged across messy or slightly different sources through entity resolution.
- **Occupation-mapping basis.** What the occupational breakdown anchors on - commonly BLS data for the US, extrapolated to other countries using ILO statistics.
- **Contractor and staffing-firm treatment.** Whether contractors are counted, excluded, or unknown.

There is also a recency correction to know about: because people are slow to update their profiles after a job change, some sources run a nowcasting model that estimates recent inflows and outflows. A number that already includes a nowcast is not directly comparable to a raw count, so that too belongs in the disclosure line.

> **Rule:** No number enters a brief without its five disclosures
>
> Period, geography, dedup method, occupation-mapping basis, and contractor treatment. A number missing any of these is directional at best, and never comparable to a number built differently.

## The analytic-confidence layer beyond the source grade

Source reliability and credibility grade the evidence. Analytic confidence grades your analysis of it, and it is a separate note. The Peterson Table of Analytic Confidence Assessment, a 2008 Mercyhurst thesis, weighs six factors: the information and analysis method used, source reliability, expertise, collaboration through peer review, task complexity, and time pressure to produce the analysis.

For a single public-source finding, four of these carry the weight. Task complexity is usually low for one claim, and source reliability is already captured in your A-to-F pass. The load-bearing four are the method you used, your expertise in the domain, whether the finding was peer-reviewed, and how much time pressure produced it.

#### The three layers of a graded finding

1. **Label** - Decision-grade, directional, or hold, with one confidence sentence
2. **Analytic confidence** - Method, expertise, peer review, time pressure
3. **Evidence grade** - Source reliability A-F and information credibility 1-6

*A decision-grade label sits on top of both an evidence grade and an analysis grade, not one alone.*

Analytic-confidence thinness has a supply dimension that is easy to miss. In Refolk's index of professional profiles, the pool of people holding a competitive-intelligence family title is far thinner in some markets than others, which caps the defensible confidence a lone analyst can claim in a market with few peers to review the work.

| Market | CI-family title holders | Share of the two-market total |
|---|---|---|
| United States | 52 | 90% |
| United Kingdom | 6 | 10% |

The US pool is roughly 8.7 times the UK pool. The mechanism matters for grading: a UK market read rests on a much smaller expert base, so the same method yields lower defensible confidence there. Peer review is scarcer, corroboration is harder, and the "collaboration" factor in Peterson's table is structurally weaker. Do not paper over thin populations with a confident label.

I ran this search: `UK-based competitive intelligence managers at telecom or media companies, London or Oxford area.` - [see the full result list](https://www.refolk.ai/s/mh4ee2q9j7).

*Returns named CI practitioners you can route a finding to for the peer-review pass, which is the factor most often skipped on a thin-population market read.*

When you need a second reviewer for the collaboration factor and your own bench is short, finding the right named analyst to route a finding to is exactly the friction [Refolk](/) removes: you ask for the people you want in plain English and get them back.

## The components a decision-grade finding must carry

The intelligence-community baseline for what a finding must contain is ICD 203's tradecraft standards. Four of them are load-bearing for a single finding, and a finding missing any one is not decision-grade regardless of how clean its source grade looks.

- **Sourcing.** Identify the underlying sources and methodologies, describing factors affecting quality and credibility including accuracy, possible denial and deception, currency, and the source's access, motivation, and bias.
- **Distinctions.** Clearly distinguish the underlying information from your assumptions and judgments.
- **Alternatives.** Identify and assess at least one plausible alternative hypothesis.
- **Linchpin assumptions.** When an assumption serves as the linchpin of the argument or bridges a significant gap in the evidence, state it explicitly and explain what would change if that assumption turned out to be wrong.

The linchpin test is the one that saves careers. Force the "what would change this finding" line before you label anything. If you cannot name the single premise that, if false, collapses the finding, you have not surfaced your assumptions - you have smuggled one in as fact.

ICD 203 also enforces a wording discipline: do not combine a confidence level and a likelihood term in the same sentence, because mixing them confuses the reader about what exactly is uncertain. "Highly likely and high confidence" reads as emphasis but says nothing - the reader cannot tell whether you are unsure about the event or unsure about your own analysis. Separate them. Say the likelihood in one sentence and your confidence in it in another.

The commercial CI analogue rounds this out: a best-in-class finding answers what it means for stakeholders, what to do next, and what is likely to happen, not just what has happened. A finding with a clean grade but no "so what" is verified trivia.

## The procedure

Run these eight steps in order. The first six a single analyst can do; the last two pull in a reviewer.

#### Grading one finding, end to end

1. **Frame the finding as one claim** - Reduce it to a single testable statement with an implied action. Write the claim, the decision it feeds, and the cost of being wrong in one line.
2. **Rate source reliability A to F** - Judge each contributing source's track record independently of the content. Give every source a letter and a one-line justification.
3. **Rate information credibility 1 to 6** - Score the specific claim on plausibility, internal consistency, and corroboration. A 1 requires at least one confirmation from a separate collection path.
4. **Run the circularity check** - Trace each independent source to origin and collapse any that share a root such as one press release or one dataset. Record the count of genuinely independent chains.
5. **Apply shelf-life** - Stamp the finding with its signal type's re-verify window. Record the last-verified date separately from the source date.
6. **Apply the Peterson factors** - Score method used, expertise, peer review, and time pressure. Write an analytic-confidence note distinct from the source grade.
7. **Check the decision-grade components** - With a reviewer, confirm surfaced assumptions, the linchpin what-would-change-this line, methodology disclosure for any number, and one alternative hypothesis.
8. **Grade and label** - Assign decision-grade, directional, or hold. Write one confidence sentence that does not mix a confidence term with a likelihood term, and sign off.

The output of the whole procedure is small: one label, one confidence sentence, and an audit trail behind it. That compactness is the point. A decision-maker reads the label; a challenger reads the trail.

## How this goes wrong

The most valuable part of any standard is its list of failure modes, because these are the ways a finding passes a lazy review and fails in the field. Each one below has a false positive it produces and a specific check that catches it.

| Failure mode | The false positive it produces | The check |
|---|---|---|
| Fake corroboration (circularity) | Credibility-1 grade on a single origin | Trace each source to its root URL; collapse shared roots |
| Reliability bleeding into credibility | A-source, therefore 1-rated | Score the two axes in separate passes |
| Stale posting read as live | "They are hiring aggressively" off a ghost job | Last-verified date plus a still-active check, not posting date |
| Headcount definitional mismatch | A fabricated growth spike | Confirm contractor treatment and period on both numbers first |
| Assumption smuggled as fact | A confident finding that collapses on one unstated premise | Force the what-would-change-this line before labeling |
| Confidence/likelihood word-blur | Reader cannot tell what is uncertain | Enforce the ICD 203 separation rule |

Two of these deserve extra weight. **Fake corroboration** is the deadliest because it earns a top grade rather than exposing a weak one - the finding looks its strongest at the exact moment it is most wrong. **The headcount mismatch** is the most common in day-to-day market sizing, because it does not feel like an error; it feels like two sources disagreeing, and the analyst "resolves" it by picking the number that fits the story.

#### Reliability against credibility

Horizontal axis runs from Low credibility (claim weak) to High credibility (claim corroborated). Vertical axis runs from Low reliability (source unproven) to High reliability (source with track record).

| Quadrant | What it means |
| --- | --- |
| Hold: unproven source, weak claim | Discard or re-collect from a better source |
| Directional: unproven source, corroborated claim | Brief with a caveat; find a reliable channel |
| Hold: reliable source, weak claim | Do not let the source's history inflate the grade |
| Decision-grade: reliable source, corroborated claim | Brief it, with one confidence sentence |

*The label a finding earns depends on both axes, and a strong source with a weak claim is not decision-grade.*

The matrix makes the reliability-bleed trap visible. The bottom-right quadrant - reliable source, weak claim - is where analysts most often over-grade, because the source's reputation feels like it should count for the claim. It does not. Only the top-right quadrant is decision-grade.

## The verification checklist

Run this before you attach a label. It is the definition of done, phrased so two analysts checking the same finding tick the same boxes.

#### Before you label the finding

- [ ] The finding is stated as one testable claim with the decision it feeds and the cost of being wrong
- [ ] Every source carries an A-to-F reliability letter with a one-line justification
- [ ] The claim carries a 1-to-6 credibility grade scored in a separate pass from reliability
- [ ] Every source has been traced to root and shared roots are collapsed, with the true independent-chain count recorded
- [ ] A credibility-1 grade rests on at least one confirmation from a separate collection path
- [ ] The finding carries a last-verified date separate from the source date, within the signal's shelf-life window
- [ ] Any cited number discloses period, geography, dedup method, occupation-mapping basis, and contractor treatment
- [ ] An analytic-confidence note records method, expertise, peer review, and time pressure
- [ ] The linchpin assumption is stated with an explicit what-would-change-this line
- [ ] At least one alternative hypothesis is named and assessed
- [ ] The confidence sentence does not combine a confidence level with a likelihood term
- [ ] A reviewer has signed off on the label

## Keeping the standard current

Adopt this as team policy and it will drift unless two things are maintained. First, the shelf-life table is the part that ages: re-check the posting-decay and ghost-job figures against your own funnel each quarter, because a signal's decay rate is specific to the markets you cover. If your urgent roles fill in two weeks rather than three to four, tighten the window.

Second, the analytic-confidence layer depends on having reviewers, and reviewer supply is uneven across markets. In Refolk's index, the pool of experienced-but-not-yet-director CI analysts is the scarce band: at Director seniority in the US, 8,754 people list competitive intelligence as a skill, against 4,109 at Senior - the Director band is 2.13 times larger, an inverted pyramid driven by skill inflation at senior titles. When you need a genuine second pass on a finding, the constraint is finding a qualified peer, not finding a title.

| Seniority band | People with CI skill (US) | Ratio to Senior |
|---|---|---|
| Senior | 4,109 | 1.0x |
| Director | 8,754 | 2.13x |

**2.13x - US Director-band CI-skill supply versus Senior band**

The scarce reviewer is the experienced Senior analyst, not the Director; plan peer review around that.

Treat the standard itself as a finding subject to re-verification. The rubric holds; the numbers inside it have shelf-lives too.

## Frequently asked questions

### What is the difference between source reliability and information credibility?

Reliability asks about the source's track record over time and is rated A to F. Credibility asks about the specific claim in front of you right now and is rated 1 to 6. They are scored on two independent axes and notated as a pair, such as B2, so that a trusted source's history does not inflate the grade of a weak individual claim. Score them in separate passes to keep one from bleeding into the other.

### How many independent sources make a finding decision-grade?

Under the Admiralty scale, exactly one additional independent source lifts a claim to the top credibility grade of 1, provided that source comes from a genuinely separate collection path and the claim is internally consistent. The New York Times standard is stricter in practice, requiring at least two independent sources for a fact not attributed to a single named speaker. The key word is independent: three citations that trace to one press release count as one source.

### How do I know if my market finding is fresh enough to brief?

Match the finding to its signal type and its shelf-life window. A job posting under seven days old is a strong signal, but one older than 14 days should be treated as weak, and about 52% of postings from always-hiring companies are ghost jobs that never fill. Corporate-event hiring surges cluster in a 30 to 90 day window after announcement. Record a last-verified date separate from the source date so freshness is auditable.

### What must I disclose before citing a headcount or posting-count number?

Disclose the period, the geography, the deduplication or entity-resolution method, the occupation-mapping basis, and how contractors and staffing firms are treated. A firm with 800 direct employees and 200 contractors can read as 800 in one source and 1,000 in another purely from workforce definition. Without these disclosures, two numbers are non-comparable and a fabricated growth spike can appear real.

### What is circular reporting and how do I detect it?

Circular reporting is information that looks multi-sourced but in reality traces to a single origin. It produces fake corroboration, which is the most dangerous false positive because it earns a top credibility grade on one source. Detect it by tracing each supposedly independent source to its root URL or dataset and collapsing any that share a root. Extra care is needed because circularity can be inadvertent, with sources interconnected in ways that are not obvious.

### Why can two good analysts grade the same finding differently?

Usually because a step was skipped or a definition was left undisclosed, not because the evidence is genuinely ambiguous. The most common causes are an unstated headcount definition, an unsurfaced linchpin assumption, and folding the circularity check into a generic cross-verify step. Running reliability and credibility as separate passes and forcing the what-would-change-this line before labeling collapses most of the disagreement.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/decision-grade-intelligence-standard*
