The Source Attribution Standard: One Tag, Two Graders, No Drift
You can write and enforce a source-tagging rule specific enough that two people credit the same record to the same source, so source reports reconcile.
Key takeaways
- A source tag is only trustworthy if two people grade the same record the same way; the borrowable bar is Cohen's kappa at or above 0.60, with 0.80 preferred.
- Split origin from latest: an original-source field must be write-once and fixed at creation, while a latest-source field updates continuously. Confusing the two is where most source reporting drifts.
- Source of hire should reflect where the relationship that led to the placement started, not the application intake point, because job boards produce volume while sourcers often produce the actual hire.
- In Refolk's index of professional profiles, account-side revenue-operations roles outnumber candidate-side recruiting-operations roles roughly 3.5 to 1 in the US (1,385 versus 398).
- Double-counting comes from the taxonomy, not the tool: when a journey maps to multiple picklist values, channel percentages sum above 100. Enforce mutually exclusive values upstream.
- Watch the kappa paradox: near-100 percent raw agreement can hide low kappa when one source dominates volume, so report both numbers together.
Every source-of-hire and lead-source report rests on one assumption: that the source tag on each record means the same thing to everyone who reads it. This standard is for the recruiting-operations or revenue-operations lead who has to defend those reports when leadership challenges them. It gives you a definition of done for a source tag, a tie-break rule for the hard cases, and a two-grader test so you can prove that two people credit the same record to the same source.
Most guides on this topic cover matching, cleanup, or freshness, or they list vendor metrics. None of them tell you what a correct source tag is or how to grade one. That gap is why reports drift. This document closes it for both candidate and account records, not one channel's dashboard.
Why source reports drift even when the data looks clean
Source reports drift because one field is being asked two questions at once. A tag that records where a relationship began is a different question from a tag that records where the last interaction happened, and a single mutable field cannot answer both.
The mechanism is simple and unforgiving. If your source field updates every time a contact re-engages, then every re-engagement silently rewrites acquisition history. The record still looks clean - high fill rate, valid picklist value - but the value no longer means what your report assumes it means. This is the single most common way source-based reporting fails, and no amount of downstream cleanup fixes it, because the original information is already gone.
The second driver is taxonomy, not tooling. Salesforce and HubSpot both hand you a picklist. The arithmetic breaks when a single journey maps to more than one value: the same campaigns apply 100 percent of amounts, multiple contact roles double the totals, and channel percentages sum above 100. A mutually exclusive taxonomy is a policy choice, not a system default. You have to make it.
Reports drift because one field is being asked two questions and nobody chose which one it answers.
There is a staffing reality behind why this standard has to cover both record types. In Refolk's index of professional profiles, account-side revenue-operations roles outnumber candidate-side recruiting-operations roles by roughly 3.5 to 1 in the US. RevOps practice and CRM attribution tooling matured first, so recruiting-ops inherits attribution norms built for accounts. A standard that covers only one side leaves the other running on borrowed, mismatched rules.
What a correct source tag must prove
A source tag is correct when it names one mutually exclusive channel, records the origin of the relationship rather than the intake point, and lives in a field whose update behavior matches the question it answers. State all three and two graders can agree; skip any one and they will not.
Break that into criteria you can grade. Each criterion tells you what it proves and what it looks like when it lies.
| Criterion | What it proves | What it looks like when it lies |
|---|---|---|
| One value per record | The channel split will sum to 100% | Percentages over 100; same record under two channels |
| Origin, not intake | Credit follows the relationship, not the last click | Job board credited for a hire the sourcer closed |
| Write-once origin field | Acquisition history is preserved | Value changed since creation; re-engagement overwrote it |
| No placeholder values | The field is genuinely populated | "N/A", "unknown", "test" counted as filled |
| Date-stamped taxonomy | Categories mean the same thing across time | Values from six months ago no longer map to today's list |
The load-bearing criterion is origin over intake. Source of hire should reflect where the relationship that led to the placement actually started, not just where the application came in. Many candidates apply via a job board but convert to hires through later sourcer engagement. First-touch attribution on the intake point credits the board; the driver was the sourcer. The same logic holds on the account side, where a lead's original touch and its closing channel routinely differ.
First touch, last touch, or multi touch: pick by cycle length
Choose your attribution model by how long and how many-stepped the journey is, not by which one sounds most sophisticated. First-touch fits short, single-channel journeys; position-based fits long B2B cycles; last-touch fits a single closing channel.
| Model | Credit rule | Documented fit condition |
|---|---|---|
| First-touch | 100% to first campaign | Sales cycle under 30 days |
| Last-touch | 100% to final touch | Single closing channel |
| Position-based 40/20/40 | Split first, middle, last | B2B cycle over one month |
First-touch is honest when the cycle is under 30 days and acquisition runs through one channel. It fails on multi-step B2B journeys with 10 to 15 touchpoints, where it credits the intake and hides everything after. For those, position-based 40/20/40 is a reasonable default, and linear is fine while you are still learning the shape of your journeys.
One vendor claims a 22 percent budget-efficiency gain and 1.7x faster revenue growth from switching to multi-touch. That is a single-vendor figure and I treat it as unverified; do not put it in a report you have to defend. The defensible claim is narrower: pick the model whose documented fit condition matches your cycle, and write down why.
Which attribution model fits your journey
Whichever model you choose, it sits on top of the origin-versus-latest field split. The model decides how you distribute credit across touches; the field discipline decides whether those touches were recorded truthfully in the first place. Get the fields wrong and no model can save the report.
The write-once field pattern in practice
The origin field is stamped once, at record creation, and never edited. The latest field evolves with every interaction. Enforce this in the platform, not in a written wish, because a field that is merely defaulted will still get overwritten.
HubSpot ships this pattern natively: Original Source is fixed at contact creation and cannot be edited, while Latest Source keeps changing. The catch is that both picklists share the same fixed values and cannot be changed to custom ones, so if your taxonomy needs values HubSpot does not offer, you build custom properties alongside them.
In Salesforce, you implement the same discipline with Flow. On the first campaign add, populate custom Campaign Source and Channel Source fields. Once populated, those fields are not overwritten, and they become the source of truth for acquisition. Salesforce's default of one primary campaign per opportunity is why native reporting undercounts multi-touch journeys unless you turn on Customizable Campaign Influence.
On the ATS side, the vocabulary differs but the principle holds. Greenhouse distinguishes Source Strategy from Source Name for reporting, and treats sourced people as "Prospects" who convert into job-attached "Candidates." That prospect-to-candidate conversion is exactly the moment where source of applicant and source of hire diverge, so protect the origin tag through it.
If you need to find the person who already runs this discipline - the one who has built the Flow, protected the field, and lived through a reconciliation - describing them in plain English is faster than filtering a database by title. Refolk returns people by what they have actually done, across public LinkedIn records and the open web, so a query like the one above surfaces operators with hands-on attribution experience rather than everyone who lists a keyword.
The procedure: from taxonomy to a report that reconciles
Run these eight steps in order. The first four build the standard, the last four prove it holds. Owners and rough durations are indicative; adjust to your team.
Build and enforce the source-tagging standard
- Define the taxonomyDraft a mutually exclusive, collectively exhaustive picklist for both candidate and account records, with one plain-English decision rule per value. Done: a written list where no journey fits two values.
- Split origin from latestDesignate one write-once original-source field and one updating latest-source field per object. Done: the original field is locked at creation and cannot be edited.
- Write the tie-break ruleState which touch wins when intake and close differ, defaulting to relationship origin. Done: a rule that resolves the applicant-versus-hire case the same way every time.
- Enforce at entryUse validation, hidden form fields, or Flow to auto-populate and protect the field. Done: manual free-text entry into the source field is impossible.
- Calibrate two gradersTwo people independently tag the same 30 to 50 records; compute agreement before rollout. Done: kappa at or above your chosen threshold.
- Baseline auditRun a baseline audit in under two hours to find which dimensions drag quality down. Done: a documented starting accuracy score.
- Set audit cadencePick one cadence - monthly spot checks plus quarterly deep audits, or twice-yearly cleanse with entry-point validation - and write it down. Done: a scheduled review with a named owner.
- Reconcile and reportRe-run the two-grader test on a fresh sample each cycle; publish only if agreement passes. Done: source-of-hire and lead-source reports reconcile within tolerance.
The tie-break rule in step three is the one people skip, and it is the one that decides whether two graders agree. Write it as a concrete decision, not a principle. Below is a skeleton you can adapt.
When intake channel and closing channel differ, credit SOURCE OF HIRE to the channel where the relationship that led to the placement began. Apply in this order: 1. If a sourcer contacted the person before they applied, source = Direct outreach (sourced-after-apply flag = true). 2. Else if the person was already in the database from a prior relationship, source = Existing database. 3. Else if a named employee referred them, source = Referral. 4. Else if they applied via a job board, source = that Job board. 5. Else, source = Website/inbound. Latest-source field records the most recent touch and is not used for source-of-hire reporting.
Replace the channel names with your own MECE picklist values. Keep the ordering; it is the deterministic part.
Origin survives; latest evolves
- Record createdOrigin source stamped, then locked write-once
- Re-engagementLatest source updates; origin untouched
- Prospect to candidateSourced-after-apply flag set if applicable
- Placement or closeSource-of-hire read from origin plus tie-break rule
- ReconcileTwo graders re-tag a sample; report if agreement passes
The two-grader test: how to know the tag is applied the same way
Two people independently tag the same sample of records, and you measure how often they agree beyond chance using Cohen's kappa. This is the test that turns "we have a source field" into "our source data is defensible." There is no recruiting-specific published threshold, so the standard is borrowed from inter-rater reliability statistics, and you must pick your bar explicitly.
| Kappa range | Landis and Koch label | Use as a bar? |
|---|---|---|
| 0.01 to 0.20 | Slight | No |
| 0.21 to 0.40 | Fair | No |
| 0.41 to 0.60 | Moderate | Marginal |
| 0.61 to 0.80 | Substantial | Common minimum |
| 0.81 to 1.00 | Almost perfect | Preferred |
The passing bar is contested and you should know the argument. A common minimum is kappa at or above 0.60, described as substantial. The stricter published view, from McHugh, treats any kappa below 0.60 as inadequate and notes many texts want 80 percent or more raw agreement before trusting a rating. The 0.80 mark is preferred where the stakes are high. Choose one, write it into the standard, and hold every audit to it.
Applying kappa to source tagging as a two-grader reconciliation test is my synthesis. The threshold itself is publicly established; the recruiting application is not. If you want to check the fit locally, run the calibration in step five on 30 to 50 records and see whether the disagreements cluster on a specific taxonomy value - that tells you which decision rule is ambiguous.
Report raw agreement alongside kappa, always. When one source dominates volume, expected chance agreement is very high, which leaves little room for kappa to move - the kappa paradox. You can see 95 percent raw agreement and a low kappa on the same sample, and neither number alone tells the truth. Together they do.
How this goes wrong: failure modes and false positives
Most of these failures produce data that looks clean. That is what makes them dangerous, and why a standard that only checks fill rate is worthless. Here is what actually breaks and how to catch it.
- Double-counting in reports. Looks like: channel percentages sum above 100. Cause: the same campaigns apply 100 percent of amounts while multiple contact roles double totals. Check: filter to one primary source per record before scoring.
- Original field silently overwritten. Looks like: clean, fully populated data. Reality: acquisition history is gone. Check: confirm the field is truly write-once, not merely defaulted, by attempting an edit on a test record.
- Self-reported source treated as truth. Looks like: high fill rate. Reality: wrong values, because candidates are thinking about the job, not your reporting. Check: cross-reference against the actual first touch.
- First-touch masking the real driver. Looks like: a decisive board attribution. Reality: the sourcer closed it. Check: cross-check hires flagged sourced-after-apply.
- Taxonomy drift over time. Looks like: a populated field. Reality: categories used six months ago do not match today's. Check: version and date-stamp the picklist.
- Placeholder pollution. Looks like: a filled field. Reality: "N/A", "unknown", or "test" - in most CRMs 5 to 10 percent of "filled" fields are placeholders. Check: filter these out before any scoring.
- Kappa paradox. Looks like: near-100 percent raw agreement. Reality: low kappa because one source dominates. Check: report raw agreement and kappa together.
- Passing calibration once, then never re-testing. Looks like: a solved problem. Reality: drift returns silently. Check: re-run the two-grader test each audit cycle.
The through-line: a high fill rate is not a quality signal. It is often the opposite, because placeholders and self-reported guesses fill fields fast. Grade on origin accuracy, protection, and grader agreement instead.
Verify before you publish
Run this checklist before any source report leaves your hands. It is the difference between a report that survives a leadership challenge and one that gets picked apart in the room.
Source report readiness
- The taxonomy is mutually exclusive - no journey maps to two values - and date-stamped with a version.
- The origin-source field is write-once and I have confirmed it by trying to edit a test record.
- The latest-source field is separate and is not used for source-of-hire reporting.
- The tie-break rule is written as an ordered decision that resolves the applicant-versus-hire case deterministically.
- Manual free-text entry into the source field is impossible.
- Two graders re-tagged a fresh 30 to 50 record sample this cycle and cleared the chosen kappa bar.
- Raw agreement is reported alongside kappa.
- Placeholder values (N/A, unknown, test) were filtered before scoring.
- The audit sample size fits the database: 100 to 200 records, or 10,000 to 20,000 stratified for databases over 100,000.
- Source-of-hire and lead-source reports reconcile within tolerance.
Keeping the standard current
Set the cadence to the faster of two clocks: data decay and human-entry drift. CRM data decays at roughly 30 percent per year, which argues for auditing at least twice a year. But source-tag drift from manual entry can outpace decay, which is why the tighter cadence - monthly spot checks on critical fields plus quarterly deep audits - is the safer default whenever any part of the tagging is manual.
The sources genuinely disagree here, and both cite the same 30 percent decay figure. The decay rate is a floor, not a schedule. Pick the cadence that matches how much of your tagging a human touches, name an owner, and put it on the calendar. An audit with no owner does not happen.
Re-run the two-grader calibration every cycle, not just at rollout. Passing once and never re-testing is its own failure mode: drift returns silently, and a taxonomy that was clean six months ago is exactly the kind that no longer matches today's channels. Treat the kappa test as a recurring gate on publication, the same way you treat the placeholder filter. If a fresh sample fails the bar, you hold the report and find the ambiguous decision rule before you ship a number anyone will act on.
Questions practitioners ask
What is the difference between source of applicant and source of hire?
Source of applicant records where an application came in; source of hire records where the relationship that led to the placement actually started. They differ because many channels produce high applicant volume with low hire conversion, while others produce few applicants but high hire rates. A candidate may apply via a job board but convert through later sourcer engagement, so first-touch attribution on the intake point often credits the wrong channel.
Should original source and last touch be the same field?
No. Keep two separate fields: an original-source field that is write-once and fixed at the moment the record is created, and a latest-source field that updates continuously as interactions unfold. Confusing the two is where most source-based reporting goes wrong, because a single mutable field silently rewrites acquisition history every time someone re-engages.
What kappa score means two people tag sources consistently enough to trust?
There is no recruiting-specific published threshold, so the borrowable standard is Cohen's kappa. A common minimum is 0.60, described as substantial agreement, with 0.80 (almost perfect) preferred. McHugh treats any kappa below 0.60 as inadequate. Pick your bar explicitly and report raw agreement alongside kappa, because near-100 percent raw agreement can hide a low kappa when one source dominates the volume.
How often should I audit source tags, and on how many records?
Sources disagree on cadence: monthly spot checks on critical fields plus quarterly deep audits, versus a twice-yearly cleanse with real-time validation at entry between. Both cite roughly 30 percent annual data decay. For sample size, pull 100 to 200 random records to verify manually. For databases over 100,000 records, pull 10,000 to 20,000, stratified by lead source and creation date to avoid bias.
Why do my channel percentages add up to more than 100 percent?
Because a single journey maps to multiple picklist values, so the same record gets counted under several channels. This shows up when campaigns each apply 100 percent of an amount and multiple contact roles double the totals. The fix is upstream: enforce one primary source per record with mutually exclusive taxonomy values, rather than filtering the double-counts out downstream.
Can I trust candidate self-reported source?
Not on its own. Self-identification produces high fill rates but wrong values, because candidates are thinking about the job, not your reporting. Treat it as one input, protect the field so free text cannot pollute it, and cross-check hires flagged as sourced-after-apply so a first-touch board answer does not mask the sourcer who actually closed the relationship.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.