# The Buying-Signal Yield Score: Keep, Demote, or Retire a Trigger

*You can score any buying signal you already chase on six measured dimensions and rule it Keep, Demote, or Retire against thresholds two people would apply identically.*

- Canonical URL: https://www.refolk.ai/guides/buying-signal-yield-score
- Pillar: Sales and go-to-market
- Format: Framework
- Published: 2026-09-01
- Last reviewed: 2026-09-01
- Reading time: 17 min
- Keywords: which buying signals actually convert, signal to opportunity conversion rate, back-test buying signals against closed won, win rate by signal type, retire low performing buying signals

## Key takeaways

- Signal-to-opp rate must be computed per signal type, not blended, because a decorative trigger at 2% hides inside a healthy-looking 12% portfolio average.
- Signal-based outreach clears generic by roughly 5x - a signal-based midpoint near 11% against a generic midpoint near 2% - because it fires inside the intent window.
- Correlation inflates a signal's apparent value by about 3x: one documented holdout showed 40% claimed contribution against 14% true incremental lift.
- Attribution is decided before outreach, not at quarter-end; the signal-source field most teams skip is what makes the yield score computable at all.
- In Refolk's index the UK SDR-to-RevOps ratio is 35.6 to 1 versus 29.7 to 1 in the US, so the person who instruments attribution is scarcer per rep in the UK.
- A signal needs two clocks: monthly review to catch alert fatigue and quarterly recalibration to catch weight drift against closed-won.

This guide is for founders selling their own product, account executives, SDR leads, and partnerships teams who already act on buying signals and need to decide which ones still earn rep time. It scores one signal you already chase on six measured dimensions from your own pipeline and turns the argument into a verdict: Keep, Demote, or Retire. The library already catalogs which triggers to watch. This one measures whether the trigger you chase produces closed revenue in your CRM.

The problem is not a shortage of signals. It is that teams keep chasing signals long after the data stopped supporting them, because nobody ever computed the yield. A blended dashboard number keeps decorative triggers alive, and rep time bleeds into work that never closes. The fix is a scoring method precise enough that two people looking at the same numbers reach the same verdict.

## What the Yield Score decides

The Yield Score decides whether one buying signal earns rep time, should be watched at lower priority, or should be dropped. It is a per-signal judgement, applied to a trigger you already act on, resolved against numeric thresholds rather than opinion.

A buying signal is any observable event that suggests an account might buy: a champion changing jobs, a pricing-page visit, a funding round, a category intent surge, an executive hire into a buying role. The score answers one question about one of them at a time: does this signal, in my pipeline, produce opportunities and closed revenue at a rate that justifies the rep hours it consumes?

The verdict has three states, and they are not synonyms for good, medium, and bad.

- **Keep**: the signal beats your cold baseline on conversion and survives a lift check. Reps should action it inside its tier window.
- **Demote**: the signal has real but weak yield, or high volume with thin lift. Watch it, but rank it below Keep signals and never let it displace them.
- **Retire**: the signal converts at or below your cold baseline, or its apparent value collapses under a holdout. Stop spending rep time on it.

The reason to formalise this is that only about 5% of a market is in-market at any time. Rep attention is the scarce input, and every hour spent on a decorative signal is an hour not spent on a predictive one. The score exists to reallocate that hour on evidence.

## The six dimensions, and what each proves

Score the signal on six dimensions, each with a numeric threshold, so the verdict is reproducible. Each dimension proves one thing, and each has a way it lies.

| Dimension | What it proves | What it looks like when it lies |
| --- | --- | --- |
| Reply lift | The signal beats generic outreach on relevance | Inflated by Apple MPP pixel opens, not human replies |
| Signal-to-opp rate | The signal produces pipeline, not just replies | A healthy blend hiding a 2% signal inside it |
| Win rate vs cold | Signal-sourced deals actually close | Borrowed from industry average, not your CRM |
| Time-to-decay | The signal fires inside a usable window | Scored as a footnote while stale alerts pile up |
| Volume | Enough events to matter and to measure | A perfect rate on five accounts, statistically noise |
| Incremental lift | The signal caused the conversion | Raw conversion mistaken for causation, ~3x inflated |

**Reply lift** measures whether outreach referencing this signal beats your generic baseline. Signal-based outreach clears generic by roughly 5x, so a signal that does not lift reply rate is not behaving like a signal at all. The trap is counting opens: Apple Mail Privacy Protection inflates open rates by 15 to 20 points and by early 2025 accounted for nearly half of tracked opens, so opens are dead as a metric. Count human replies only.

**Signal-to-opp rate** is the spine. It is opportunities created from the signal divided by accounts actioned on it, times 100. It must be tracked per signal type. Aggregation is precisely where decorative signals survive, because a champion job change and a content download can sit inside one blended number.

**Win rate versus cold** asks whether signal-sourced opportunities close at a better rate than cold-sourced ones. This is where you derive weights from your own closed-won and closed-lost data. A signal that converts in a vendor's benchmark but not in your CRM should not survive on borrowed evidence.

**Time-to-decay** captures how fast the signal's value bleeds after the moment of intent. The reply-rate gap is a relevance effect, and relevance is time-bound, so decay is a scored dimension, not a caveat.

**Volume** guards against overfitting to noise. A signal that converts perfectly across five accounts tells you nothing you can act on. Set a floor - a minimum count of actioned accounts - below which you defer the verdict.

**Incremental lift** is the one that separates a real signal from a coincidence. It measures whether the signal caused the conversion or merely correlated with accounts that were already going to buy.

## Does the signal beat baseline

A signal earns nothing until it beats your cold baseline on reply and conversion. Here are the published reference rates by outreach type, which set the shape of what "beating baseline" looks like before you plug in your own numbers.

| Outreach type | Avg reply rate | Source |
| --- | --- | --- |
| Generic templated | 1-3% | GTME benchmarks |
| Personalized cold (enrichment) | 4-8% | GTME benchmarks |
| Signal-based (job change, funding) | 8-14% (top quartile 16-22%) | GTME benchmarks |
| Signal-specific (Instantly 2026) | ~18% | Autobound |

The signal-based midpoint near 11% is roughly 5.5x the generic midpoint near 2%. Instantly's 2026 report puts signal-specific personalization at about 18% response, roughly 5.2x the 3.4% generic average. On conversion, expect 4 to 10% signal-to-meeting and 30 to 40% lower cost per qualified meeting versus cold. Stacked signals, meaning two or three firing on one account, are reported to convert at 5 to 10x cold outreach.

Use these as the shape you expect, not the numbers you record. Your own baseline is the only comparison that decides a verdict. If your signal replies at 4% while your cold sequences reply at 4%, the signal is decorative in your pipeline regardless of what it does in a benchmark table.

**~11% - Signal-based reply midpoint versus a ~2% generic midpoint**

Roughly 5.5x, and the gap is a relevance effect from firing inside the intent window - not better copy.

> **Watch out:** Opens are not a signal metric
>
> Apple Mail Privacy Protection inflates open rates by 15 to 20 points and by early 2025 accounted for nearly half of tracked opens. Score reply lift on human replies only, or the dimension will report engagement that never happened.

## Scoring time-to-decay

Time-to-decay measures how fast a signal loses value after the moment of intent, and it should carry real weight because relevance is time-bound. The cleanest quantified curves come from inbound speed-to-lead research, which is an adjacent proxy for trigger-event decay, not an identical one.

| Elapsed time | Effect on qualification/conversion | Source |
| --- | --- | --- |
| Within 1 minute | up to +391% conversion lift | Kixie speed-to-lead |
| 5 min vs 30 min | 21x more likely to qualify | Kixie speed-to-lead |
| Per minute, first 5 min | ~10% conversion probability lost | Strolid automotive dataset |
| By 30 minutes | ~75% of conversion opportunity lost | Strolid automotive dataset |

Treat these figures with care. They come from inbound lead response, drawn from datasets like Strolid's 2.3M-lead, 847-dealership record, where the buyer just raised a hand. External trigger events decay more slowly, and a published decay curve specific to trigger events is not established. What transfers is the mechanism: value falls fast right after the moment, so tier the signal and set a window.

- **Tier 1**: high-decay signals like a live demo request or pricing visit. Respond inside 5 minutes.
- **Tier 2**: slower-decay signals like a champion job change or funding round. Respond inside 24 hours.

For account scoring, apply recency decay so stale signals do not inflate scores. Sources offer two conventions: halve the score every 30 days, or decay it 15% per day. Pick one and apply it consistently. The point is that a 90-day-old trigger should not score like a fresh one.

#### How a signal earns its verdict

1. **Beats reply baseline** - Human replies clear your cold rate, opens stripped out
2. **Produces pipeline** - Per-signal signal-to-opp rate computed, not blended
3. **Closes deals** - Win rate by signal type beats cold in your own CRM
4. **Survives lift** - Holdout or pilot shows the signal caused the conversion
5. **Verdict** - Keep, Demote, or Retire against pre-agreed thresholds

*A signal moves left to right; it can be ruled Retire at any gate it fails.*

## The scoring procedure

Follow this procedure in order; each step produces an input the next one needs, and skipping instrumentation makes everything after it guesswork. The steps below match the machine-readable version submitted with this guide.

#### Score one signal end to end

1. **Define the one signal and its window** - Write a single unambiguous trigger definition and assign a tier: tier 1 respond in 5 minutes, tier 2 in 24 hours. Done when two people would tag the same account identically.
2. **Instrument the CRM before counting** - Add a signal-source field, signal date, and outreach date to every actioned account, populated during enrichment not backfilled. Done when attribution survives to the opportunity record.
3. **Pull the denominator and numerator** - Isolate accounts actioned on the signal and opportunities created from it, then compute signal-to-opp rate. Done when the rate is per this signal alone.
4. **Back-test against closed-won and closed-lost** - Compute win rate by signal type from your own closed deals and derive weights from that history. Done when you have a win rate against your cold baseline.
5. **Score the six dimensions** - Score reply lift, signal-to-opp rate, win rate vs cold, time-to-decay, volume, and incremental lift against the rubric. Done when every dimension has a number.
6. **Run a controlled comparison for lift** - Run a concurrent holdout or a 3 to 5 rep signal-vs-non-signal pilot over 30 to 90 days. Done when you have a lift number net of market drift.
7. **Rule Keep, Demote, or Retire** - Assign the verdict against pre-agreed thresholds. Done when two people would reach it independently.
8. **Set the re-review cadence** - Put the signal on both a monthly alert-hygiene review and a quarterly recalibration calendar. Done when it carries a next review date on each clock.

Step 2 is the one teams skip and the one that decides everything. When a signal-sourced contact converts to an opportunity, a custom source field populated during enrichment is what carries the attribution into the CRM. Most teams skip that field and then cannot answer the renewal question of what their tooling actually generated. Attribution is decided before outreach, not at quarter-end.

> **Rule:** No source field, no score
>
> If the signal-source field was not populated at first outreach, do not compute a Yield Score. Backfilled attribution is a guess about which signal drove the deal, and a guess cannot support a Keep verdict. Instrument first, count second.

Finding the accounts to action against a signal is its own bottleneck, and it is where a plain-English search removes the most friction. Instead of building a saved search per trigger, describe the stacked signal you want to test and pull the population directly.

I ran this search: `Companies that closed a Series B in the last 90 days and are hiring their first VP of Sales` - [see the full result list](https://www.refolk.ai/s/jmbkf4k39j).

*Returns accounts carrying a funding-plus-hiring stacked signal, ready to action and then back-test against your own conversion.*

[Refolk](/) lets you express a signal as a sentence and get the accounts back, so the denominator in step 3 is a list you can action rather than a query you have to maintain. For stacked signals in particular, where two or three triggers must co-occur on one account, describing the combination in plain English is faster than assembling it from separate filters.

## Isolating true lift from correlation

The sixth dimension, incremental lift, is the one that decides whether a Keep verdict is honest. It asks whether the signal caused the conversion or merely marked accounts that were already going to buy. Raw conversion cannot tell the difference, and it inflates a signal's apparent value by roughly 3x.

The documented example is stark: one holdout case found vendor attribution claiming 40% contribution when the true incremental lift was 14%. If you rule Keep on raw conversion, you will over-retain, and you will keep spending rep hours on signals whose real effect is a fraction of what the dashboard shows.

The technique is a randomized holdout, imported from marketing measurement. Hold out 5 to 10% of the target audience from signal-based exposure entirely, run the test concurrently so external trends hit both groups equally, and compare exposed against held-out after a full purchase cycle. A valid test needs four things: randomization, matched cohorts, isolation, and a measurement window.

> Correlation inflates a signal's value by about 3x, so a Keep verdict without a holdout will over-retain.

There is a rigor-versus-effort tradeoff here, and geography shapes it. The full holdout is the rigorous path. The pragmatic path is a 30-day pilot comparing 3 to 5 reps who action the signal against reps who do not. Sources disagree on which is required, and the answer depends partly on who you have to run it.

**9.9x - US-to-UK ratio of Revenue/Sales Operations profiles in Refolk's index**

1,530 in the US against 155 in the UK, so the person who must instrument attribution and run a holdout is far scarcer in the UK.

In Refolk's index of professional profiles, the US holds 45,413 SDR/BDR profiles against 5,523 in the UK, and 1,530 RevOps profiles against 155. That works out to a US SDR-to-RevOps ratio of 29.7 to 1 and a UK ratio of 35.6 to 1. The person who instruments attribution and runs the lift test is stretched thinner per rep in the UK, so the sensible default there is the 30-day pilot, not the full holdout. Match the rigor to the RevOps capacity you actually have.

## How this goes wrong

Most bad signal decisions come from a small set of repeatable errors, and each one has a check. This section is the most valuable part of the method, because a score that survives these traps is one you can defend.

**Aggregating signal-to-opp across signal types.** A blended 12% rate looks healthy while a decorative signal at 2% hides inside it. Recompute per signal type and demote anything below your cold baseline.

**Counting reply rate on inflated opens.** The dashboard shows engagement that is really Apple MPP pixel-loading, not humans. Strip auto-loads and count human replies only.

**No signal-source field at first outreach.** Quarter-end attribution ends up guessing which signal drove the opportunity. Audit that the source field was populated during enrichment, not backfilled.

**Deriving weights from industry averages.** You keep a signal that converts elsewhere but not in your CRM. Recalibrate against your own closed-won and closed-lost data.

**Mistaking correlation for lift.** Signal accounts convert well because they were already high-intent, the 40% claimed versus 14% true gap. Run a concurrent holdout before trusting any Keep.

**Reading a lift test too early.** An on/off difference flips a week later. Run at least one full purchase cycle inside the measurement window.

**Overfitting the score to noise.** A many-condition model looks perfect on history and fails live. Keep few parameters and re-test on a fresh quarter before ruling Keep.

**Never retiring anything.** Low-value signals accumulate into alert fatigue. Enforce a monthly rule: retire any signal that has not moved pipeline.

> **Tip:** Watch for the model that is too good
>
> If a signal's historical fit looks perfect, suspect overfitting, not brilliance. A score built on many conditions will match past data and break on live accounts. Fewer parameters that survive a fresh quarter beat a complex model that never leaves the spreadsheet.

## Thresholds and the verdict rubric

Set the thresholds before you look at the numbers, so the verdict is a lookup, not a debate. The rubric below is a starting template; adapt the cutoffs to your own baseline, but agree them in advance.

**Keep / Demote / Retire rubric**

```
KEEP - all of the following:
  Reply lift: signal reply rate >= 2x your cold reply rate (human replies only)
  Signal-to-opp rate: >= your cold signal-to-opp rate
  Win rate vs cold: signal-sourced win rate > cold-sourced win rate
  Volume: >= your agreed minimum actioned accounts
  Incremental lift: positive after holdout or 30-day pilot

DEMOTE - real but weak:
  Beats cold baseline on reply OR opp rate but not both, OR
  High volume with lift at or just above zero
  Action: rank below all Keep signals, never displace them

RETIRE - any one of the following:
  Signal-to-opp rate at or below cold baseline
  Win rate vs cold not positive
  Incremental lift collapses under holdout (raw conversion was correlation)
  No pipeline moved across a full monthly review
```

*Replace the baseline placeholders with your own cold reply and win rates before scoring, then agree the cutoffs as a team.*

The two-clock cadence is deliberate, because monthly and quarterly reviews catch different failure modes. Monthly review catches alert fatigue and retires signals that have stopped moving pipeline. Quarterly recalibration catches weight drift against closed-won and lets you subtract points for negative signals like champion departures or competitor contract renewals. A team needs both clocks, not a choice between them.

#### Where a signal lands

Horizontal axis runs from Low volume to High volume. Vertical axis runs from Low yield to High yield.

| Quadrant | What it means |
| --- | --- |
| Rare but strong | Keep, but confirm volume floor before trusting the rate |
| Predictive core | Keep and protect rep time for it |
| Noise | Retire, converts at or below cold baseline |
| Decorative | Demote or retire, high volume masking thin lift |

*Plot each signal on conversion yield against event volume to see the verdict at a glance.*

The bottom-right quadrant is the dangerous one. A high-volume signal with thin lift feels productive because reps are always busy actioning it, but it is exactly the decorative signal that a blended dashboard protects. The matrix forces you to see that busy is not the same as predictive.

## Keeping the score current

A Yield Score is a snapshot, and signals decay, so the score has to be maintained on a calendar rather than run once and filed. Put every scored signal on both the monthly and quarterly clocks the day you rule it, and treat the verdict as provisional until it survives a fresh quarter.

Before you call any single signal scored, run this check.

#### Before you rule Keep, Demote, or Retire

- [ ] The signal has one unambiguous trigger definition and a tier
- [ ] The signal-source field was populated at first outreach, not backfilled
- [ ] Signal-to-opp rate is computed for this signal alone, never blended
- [ ] Reply lift counts human replies with opens stripped out
- [ ] Win rate is derived from your own closed-won and closed-lost, not industry averages
- [ ] Volume clears your agreed minimum actioned-account floor
- [ ] A holdout or 30-day pilot has produced an incremental lift number
- [ ] The lift test ran at least one full purchase cycle
- [ ] The verdict follows the pre-agreed thresholds, not the discussion in the room
- [ ] The signal is on both a monthly and a quarterly review calendar

When you re-review, the fastest way to refresh the denominator is to re-pull the population that carries the signal and re-run the arithmetic against fresh closed-won data. If you are benchmarking your signal-to-opp rates against peers, Refolk can surface SDR leaders at comparable companies to compare notes with, which turns a private number into a calibrated one. The measurement discipline is the same every cycle: instrument first, count per signal, isolate lift, then let the thresholds make the call. That is what stops a team arguing from anecdote and moves rep time from decorative signals to predictive ones.

## Frequently asked questions

### Which buying signals actually convert?

Only the ones your own closed-won data confirms. Industry benchmarks show signal-based outreach replying at 8 to 14% versus 1 to 3% for generic, and the highest-correlation starting signals are demo requests, pricing visits, champion job changes, category intent surges, and executive hires into buying roles. But a signal that converts elsewhere can be dead in your CRM, so compute signal-to-opp rate per signal type before trusting any external list.

### How do I calculate signal-to-opportunity conversion rate?

Divide the opportunities created from a signal by the accounts you actioned on that signal, then multiply by 100. Track it per signal type, never in aggregate, because a champion job change converts very differently from a generic content download. Aggregation is exactly where a 2% decorative signal hides inside a healthy-looking 12% blended number, so the per-signal denominator is the whole point.

### How do I back-test a buying signal against closed-won?

Isolate every account you actioned on the signal, then check which of those became opportunities and which of those closed. Compute win rate by signal type from that history and compare it to your cold baseline. Derive your weights from this data, not from vendor averages, since the failure mode is keeping a signal that converts in someone else's market but not yours.

### Why does correlation overstate a signal's value?

Because signal accounts often convert well since they were already high-intent, not because the signal drove anything. One documented holdout showed vendor attribution claiming 40% contribution when true incremental lift was 14%, roughly a 3x overstatement. Any Keep verdict resting on raw conversion instead of a holdout will over-retain, which is why the sixth dimension is measured lift, not measured conversion.

### How often should I review buying signals?

Use two clocks. Review alerts monthly to catch alert fatigue and retire any signal that has not moved pipeline. Recalibrate weights quarterly against closed-won and closed-lost, subtracting points for negative signals like champion departures. These cadences catch different failure modes, so a team needs both rather than choosing one, and each signal should carry a next review date on each calendar.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/buying-signal-yield-score*
