# The Lawful-Sourcing Standard: What You May Collect, Keep, and Send

*You will be able to set a sourcing data policy the legal team signs off on and recruiters actually follow, with a lawful basis, retention rule, and transfer tool for every flow.*

- Canonical URL: https://www.refolk.ai/guides/lawful-sourcing-standard
- Pillar: Process, data, and compliance
- Format: Standard
- Published: 2026-08-22
- Last reviewed: 2026-08-22
- Reading time: 17 min
- Keywords: gdpr recruiting compliance, candidate data policy, lawful basis sourcing, recruitment data retention, cross-border candidate data transfer

## Key takeaways

- For active sourcing and matching candidates to roles, the correct lawful basis is legitimate interest under Article 6(1)(f), not consent - consent is reserved for keeping data in a talent pool.
- Retention guidance converges on 6 to 12 months for unsuccessful candidates unless you hold explicit consent, and the SAR deadline is one calendar month from receipt, not 30 days.
- The load-bearing document is the necessity analysis: proof you rejected less-intrusive options, which must exist before collection and cannot be reconstructed after.
- GDPR fines reach EUR 20 million or 4% of global turnover, and cumulative fines since 2018 stand at EUR 6.8B, yet the biggest recruitment exposure on record was a security lapse - 12,000 records left open to the internet.
- In Refolk's index, Germany shows 202 DPO-type profiles against 150 sourcers, a compliance-dense market where a policy assuming heavy oversight will not transfer to the US, which shows 4,192 sourcers.
- EU-to-US transfers need a valid tool: the Data Privacy Framework for self-certified US recipients or Standard Contractual Clauses plus a Transfer Impact Assessment for anywhere else.

This is a policy standard for setting how your team sources, stores, and moves candidate data lawfully. It is for the recruiting operations and revenue operations leads who own the candidate CRM, and for the DPO or legal partner who has to sign the result. It gives you a definition of done for each processing activity, the procedure to get there, and the failure modes that turn a clean policy into an enforcement case.

The point of a standard is that two people grade the same case the same way. A candidate data policy fails when a recruiter and a lawyer disagree about whether a given collection was allowed. Everything below is written so that disagreement has a documented answer.

## What lawful basis applies to sourcing candidates?

For active recruitment - sourcing candidates and matching them to open roles - legitimate interest under Article 6(1)(f) is the standard basis, not consent. Most agencies think they need consent for everything. They do not.

Every piece of candidate data you process needs one of six lawful bases under GDPR Article 6. Two of the six matter for recruiting: legitimate interest and consent. You have a legitimate business interest in finding candidates for your clients, and candidates have a reasonable expectation that recruiters will contact them about relevant opportunities. That is the ground legitimate interest stands on.

Legitimate interest is not a free pass. It requires a documented three-part balancing test: the interest must be real and present, the processing must be necessary, and the candidate's rights must not override your interest. The middle test is the one that bites. Processing must be necessary and not merely useful, meaning there is no other reasonable, equally effective, less-intrusive option.

Consent is reserved for a narrower case: keeping candidate data beyond a specific recruitment process, for example adding someone to a talent pool for future roles. Consent must be freely given, specific, informed, and unambiguous. A pre-ticked checkbox on a careers page does not qualify.

> **Rule:** Legitimate interest for sourcing, consent for the pool
>
> Use legitimate interest with a documented LIA for active sourcing and matching. Use consent only to hold data beyond the role it was collected for.

The consent reflex is the single most common error, and it makes teams less compliant, not more. Invalid consent collapses the moment it is tested, and a team relying on it has no fallback. Legitimate interest with a documented Legitimate Interest Assessment survives an audit; a pre-ticked box does not.

## Why the necessity analysis is the load-bearing document

The document that carries the most weight in an audit is the necessity analysis, and it must exist before collection, not after. Regulators accept legitimate interest, but they require proof that less-intrusive options were considered and rejected.

This is a discipline you cannot reconstruct retroactively. If you decide after a complaint that scraping a profile was necessary, you are writing fiction. The necessity showing has to be dated before the collection it justifies. That is why the Legitimate Interest Assessment sits so early in the procedure.

There is a genuine sequencing tension in the sources worth flagging. Practical LIA guidance for recruiters places the assessment before any sourcing starts. The EDPB web-scraping guidance stresses that the necessity documentation must exist before collection begins. Both point the same way: document first, collect second. If your team is already sourcing, the honest move is to write the LIAs now and treat everything before them as legacy data to review against the same test.

> The necessity analysis cannot be written after the complaint arrives; if it is late, it is fiction.

Public availability is a specific trap here. Consent is off the table for indiscriminate collection, and the EDPB is explicit that publishing data online is not consent to scrape it. A visible LinkedIn or GitHub profile is not permission. Legitimate interest remains the realistic route, subject to the three cumulative conditions, and where a profile captures Article 9 data - health, ethnicity, beliefs, and similar - you need both an Article 6 basis and an Article 9(2) derogation.

## How long may you keep candidate data?

No fixed statutory period exists. Guidance converges on 6 to 12 months for unsuccessful candidates unless you hold explicit consent to keep them longer.

GDPR names no exact retention period for candidate data, so the standard has to set one. Guidance from the European Data Protection Supervisor and the ICO is consistent: retain unsuccessful candidate data no longer than 6 to 12 months, or longer only with explicit consent. Practitioners commonly automate a 2 to 3 year deletion rule for wider CRM records, but the tighter window applies to people who applied and did not get the role.

Retention is not a filing preference. ICO enforcement against recruiters starts with a person, not an audit. The mechanism is a candidate complaint, usually an ignored deletion request or continued emails after an objection, which then triggers the regulator to demand evidence. So the retention and deletion policy is the first thing tested, which is why it must be automated rather than left to a recruiter's memory.

**6.8B - Cumulative GDPR fines in euros since 2018**

The statutory ceiling is EUR 20 million or 4% of global annual turnover, whichever is higher.

The deadline for a data subject access request is separate and firm: one calendar month from the day you receive the request, extendable by two further months for complex requests only if you tell the requester within the original month. Two details cause most misses. "One month" is a calendar month, not 30 days - a request received on 15 March is due 15 April regardless of how many days sit between. And the clock starts on receipt, not when you locate the data. Log the receipt date the moment a request lands.

## What can you lawfully send across borders?

EU-to-US transfers need a valid transfer tool. The two you will actually use are the Data Privacy Framework for self-certified US recipients and Standard Contractual Clauses with a Transfer Impact Assessment for everywhere else.

The Data Privacy Framework is a self-certification program. A US organisation commits to a set of privacy principles, registers with the US Department of Commerce, and appears on a public list the European Commission recognises as adequate. Once certified, that company can receive EU personal data without separate SCCs for that flow. The European Commission adopted the DPF adequacy decision on 10 July 2023 under Article 45 GDPR, and certification requires annual re-certification to stay on the list.

SCCs are broader. They can be used to transfer personal data from the EU to any non-EU country, subject to a Transfer Impact Assessment. The 2021 SCCs use a modular structure covering four scenarios: controller-to-controller, controller-to-processor, processor-to-processor, and processor-to-controller. Pick the module that matches the flow.

| Tool | Scope | Extra step required |
|---|---|---|
| Data Privacy Framework | US recipients only, self-certified | Annual re-certification |
| Standard Contractual Clauses | Any non-EU country | Transfer Impact Assessment |
| Binding Corporate Rules | Intra-group | DPA approval, ~18 to 24 months |

Do not treat the DPF as permanently safe. The PCLOB quorum was lost in January 2025, a development that affects the DPF's redress mechanism and is being monitored by the EDPB. The practical rule is to keep an SCC fallback ready for any flow you route through the DPF, so a single adequacy shock does not strand your data.

> **Watch out:** The DPF is not a set-and-forget tool
>
> Relying on the DPF alone, with no SCC fallback, is a false sense of safety. Monitor the redress-mechanism status the EDPB has flagged and keep clauses on the shelf.

## The procedure: from data map to signed policy

Set the policy in eight moves, each with a named owner and a done condition. The order matters: you cannot assign a lawful basis to data you have not mapped, and you cannot run a valid LIA after collection has already started.

#### Set the lawful-sourcing policy

1. **Map where candidate data lives** - Build a written inventory of every system holding candidate data, with an access owner and an export/delete method for each. Done when the map is complete before any request arrives.
2. **Assign a lawful basis to each activity** - Tag every processing activity as legitimate interest or consent, with talent-pool retention flagged as consent-based. Active sourcing and matching carry legitimate interest.
3. **Run and document a Legitimate Interest Assessment** - For each activity relying on legitimate interest, complete and sign an LIA passing the purpose, necessity, and balancing tests. Necessity must show less-intrusive options were rejected before collection.
4. **Screen for special category and excess data** - Limit collection to fields that contact or evaluate a candidate, and strip anything sensitive under Article 9. Recruiters can name why each field is held.
5. **Run a DPIA for any AI or automated screening** - Complete a Data Protection Impact Assessment before deploying any tool that scores, ranks, or screens EU candidates. Vendor compliance does not substitute for your own DPIA.
6. **Set retention and automated deletion** - Configure automated deletion of unsuccessful-candidate data at 6 to 12 months, longer only with explicit consent. Answer requests within one calendar month of receipt.
7. **Put processor agreements and transfer tools in place** - Sign a DPA with every vendor holding candidate data, and cover each cross-border flow with the DPF or SCCs plus a TIA. No data leaves the region without a named tool.
8. **Publish a privacy notice and train the team** - Put a one-page plain-English notice live and train everyone who touches candidate data. Training covers CV sharing, retention limits, and the one-month deadline.

Rough timings from the field: mapping runs one to two weeks, basis assignment about a week, each LIA a few days, a DPIA two to four weeks, retention configuration about a week, and processor and transfer paperwork two to six weeks. The DPIA and the vendor paperwork are the long poles, so start them the day you commit to the project.

#### The eight-step policy build

1. **Map** - Inventory every system holding candidate data
2. **Assign basis** - Tag each activity legitimate interest or consent
3. **Assess** - LIA and DPIA before collection or deployment
4. **Retain** - Automated 6 to 12 month deletion
5. **Cover transfers** - DPA plus DPF or SCC on every flow
6. **Publish and train** - One-page notice, whole team briefed

*Each step gates the next, so mapping and basis assignment must finish before any assessment begins.*

## How this goes wrong: failure modes and false positives

The most valuable part of a standard is knowing where it breaks. Each failure below has a false positive - the thing that looks compliant but is not - and a check that catches it.

| Failure mode | What it looks like when it lies | The check |
|---|---|---|
| Defaulting to consent for sourcing | A pre-ticked box that looks compliant | Is active sourcing recorded as legitimate interest with an LIA? |
| Treating a public profile as permission | Assuming a visible profile equals consent | Is there a dated necessity analysis rejecting less-intrusive options? |
| Starting the SAR clock late | Belief the month begins at discovery | Log the receipt date, not the discovery date |
| Counting one month as 30 days | A calendar tool set to 30 days | Use the same date next month |
| Data scattered across systems | A "complete" SAR that misses a tool | Reconcile every response against the data map |
| AI screening without a DPIA | Assuming vendor compliance covers you | A DPIA dated before go-live |

Two of these deserve extra weight. First, SAR failures are architecture failures, not deadline failures. The month rarely runs out because drafting is slow; it runs out because a candidate's data sits in the ATS, the CRM, a team inbox, and an AI screening tool, and no one mapped all four. Many teams discover their data is scattered across a dozen systems with no single view only when a request forces the count. The data map, built before any request arrives, is the fix.

Second, collecting special category data by accident. A CV harvest or a scraped profile can capture health, ethnicity, or beliefs without anyone intending it. Data that does not help the team contact or evaluate a candidate, or that is sensitive under Article 9, is not related to recruiting and generally should not be collected. The check is a single question per field: does this genuinely help contact or evaluate the candidate?

> **Tip:** Reconcile the SAR against the map, not against memory
>
> When a request lands, walk the data map system by system. A response that feels complete but was assembled from memory is how the AI screening tool or shared inbox gets missed.

There is a harder lesson under all of this. The biggest recruitment data exposure on record was not a lawful-basis problem at all. A recruiter left a network storage container of 12,000 records covering 3,000 workers open to the internet, where a security researcher accessed it. A perfect lawful-basis policy is undone by one misconfigured storage bucket. The standard has to cover access controls and encryption, not just the paperwork of consent and legitimate interest.

## Sizing the compliance function against your sourcing volume

Match your oversight to your sourcing volume, and do not assume a policy tuned for a compliance-dense market transfers to a lean one. The staffing data shows how sharply this varies.

In Refolk's index, the United Kingdom holds 691 people with a Data Protection Officer-type title against 202 in Germany. That is a heavier concentration of named data-protection roles in the UK market. But the ratio that matters for policy design is compliance headcount against sourcing headcount, and there Germany looks unusual.

| Market | Sourcers | DPO-type profiles | Ratio sourcers to DPO |
|---|---|---|---|
| Germany | 150 | 202 | 0.74 : 1 |
| United States | 4,192 | not queried | n/a |

Germany shows more DPO-type profiles than sourcers - a compliance-dense market where a policy can lean on heavy oversight per recruiter. The United States shows 4,192 sourcing specialists in the same index, an order of magnitude more sourcing headcount. A policy that works in Germany because a DPO reviews nearly every activity cannot assume equivalent oversight in a market where each compliance person covers dozens of sourcers. In lean markets, automation and clear self-service rules carry the load that a DPO carries elsewhere.

**4,192 - Sourcing specialists in the US in Refolk's index**

Against 150 in Germany, which shows more DPO-type profiles than sourcers.

This is also where you decide who signs the policy. If you do not have a DPO or a data-protection lawyer with recruitment experience, you need one on the review, either on staff or external. Finding those people is a sourcing problem, and it is one I built [Refolk](/) to solve.

I ran this search: `Data Protection Officers at UK recruitment agencies and staffing firms` - [see the full result list](https://www.refolk.ai/s/6rsczem06r).

*Returns the people who own and sign off sourcing policy, so you can route the review to someone who has done it before.*

## A copy-pasteable LIA skeleton and privacy notice

Adopt a fixed template for the two documents recruiters skip most: the Legitimate Interest Assessment and the candidate privacy notice. Fixing the format is what makes the standard reproducible across a team.

**Legitimate Interest Assessment - one per processing activity**

```
Activity: [what you do with the data, e.g. sourcing engineers from public profiles]
Purpose test - the interest:
  - What is the interest? (real and present business need)
  - Who benefits and how?
Necessity test:
  - Why is this processing necessary, not just useful?
  - What less-intrusive alternatives were considered and rejected, and why?
Balancing test:
  - What is the candidate's reasonable expectation?
  - What is the impact on their rights?
  - Do their rights override the interest? (if yes, stop)
Safeguards applied: [retention limit, minimisation, opt-out route]
Decision: proceed / do not proceed
Owner and date: [DPO or recruiting lead], [date before collection]
```

*Fill each section before collection begins. If necessity cannot be shown, do not proceed on legitimate interest.*

**Candidate privacy notice - one page, plain English**

```
Who we are: [company], acting as [controller/processor] for candidate data.
Why we hold your data: to source and match you to roles. Our lawful basis is legitimate interest; we use consent only to keep your data for future roles.
What we collect: contact details and information that helps us evaluate you for a role. We do not collect sensitive data (health, ethnicity, beliefs) as part of recruiting.
How long we keep it: unsuccessful applications are deleted within 6 to 12 months unless you ask us to keep you on file.
Your rights: access, correction, deletion, and objection. We respond within one calendar month. Contact: [DPO email].
Where your data goes: named processors under contract; any transfer outside the region uses an approved transfer tool.
```

*Publish where candidates first encounter you. Swap the bracketed details for your own.*

## What to verify before you sign the policy

Run this checklist before the DPO signs and before recruiters start working to the new rules. Every item is a statement you can mark true or false, not a topic to think about.

#### Definition of done for the lawful-sourcing policy

- [ ] A written data map lists every system holding candidate data, its access owner, and its export/delete method.
- [ ] Every processing activity is tagged legitimate interest or consent, with talent-pool retention flagged as consent.
- [ ] A signed, dated LIA exists for each legitimate-interest activity, and each passes the necessity test with rejected alternatives named.
- [ ] No field is collected that fails the "does this contact or evaluate the candidate" test, and no Article 9 data is captured by accident.
- [ ] A DPIA dated before go-live exists for every AI or automated screening tool touching EU candidates.
- [ ] Automated deletion is configured at 6 to 12 months for unsuccessful candidates, with longer retention only where consent is recorded.
- [ ] The SAR process logs receipt date, treats one month as a calendar month, and reconciles responses against the data map.
- [ ] A signed DPA covers every vendor, and every cross-border flow names a valid transfer tool with an SCC fallback where the DPF is used.
- [ ] A one-page privacy notice is live, and everyone who touches candidate data has been trained on the rules.
- [ ] Access controls and encryption are in place so no candidate store is reachable from the open internet.

## Keeping the standard current

Treat the policy as a living document with three review triggers, because the ground under it moves. The lawful-basis logic is stable, but the transfer tools and the scraping guidance are not.

First, monitor the transfer tools. The DPF adequacy decision depends on a functioning US redress mechanism, and the PCLOB quorum loss in January 2025 is being watched by the EDPB. If the DPF's status changes, your SCC fallback needs to be ready to activate, not written from scratch. Re-check the public DPF list at each annual re-certification cycle for the vendors you rely on.

Second, watch the scraping and AI guidance. The EDPB adopted Guidelines 03/2026 on web scraping for generative AI, open for consultation until 30 October 2026, and issued Opinion 28/2024 on AI models in December 2024. The ICO audited AI recruitment tool providers and published findings in November 2024, uncovering considerable room for improvement on fairness and minimisation. When this guidance firms up, revisit your LIAs and DPIAs for any tool that scrapes or scores.

Third, re-run the data map on a schedule. New tools enter the stack quietly - a trial inbox, a new screening vendor, a spreadsheet a recruiter built. A map that was complete a year ago is the exact failure that blows the SAR deadline. Rebuild it at least annually and after any new tool goes live. The map is the cheapest insurance in the whole standard, and it is the first thing that goes stale.

## Frequently asked questions

### Do I need consent to source candidates under GDPR?

No. For active recruitment - sourcing candidates and matching them to open roles - the correct lawful basis is legitimate interest under Article 6(1)(f), supported by a documented Legitimate Interest Assessment. Consent is reserved for a narrower case: keeping candidate data beyond a specific recruitment process, such as adding someone to a talent pool. A pre-ticked checkbox is not valid consent, and defaulting to consent for everything usually makes teams less compliant, not more.

### How long can I keep candidate data after a role closes?

No fixed statutory period exists, but guidance from the ICO and the European Data Protection Supervisor converges on 6 to 12 months for unsuccessful candidates, longer only with explicit consent. Many teams automate a 2 to 3 year deletion rule for wider CRM records. The practical safeguard is automated deletion, because ICO enforcement against recruiters usually starts with an ignored deletion request, which makes retention the first policy tested.

### What is the deadline to answer a data subject access request?

One calendar month from the day you receive the request, extendable by two further months for complex requests only if you tell the requester within the original month. Note that 'one month' is a calendar month, not literally 30 days: a request received on 15 March is due 15 April. The clock starts on receipt, not when you locate the data, so log the receipt date immediately.

### Can I scrape public LinkedIn or GitHub profiles?

Public availability is not consent. The EDPB is explicit that publishing data online does not mean permission to scrape it. The realistic route is legitimate interest under Article 6(1)(f), which requires a documented necessity analysis showing less-intrusive alternatives were rejected before collection. If the profile contains Article 9 special category data, you also need an Article 9(2) derogation on top of your Article 6 basis.

### When is a DPIA mandatory for recruitment?

A DPIA is required before deploying systematic, extensive profiling with significant effects based on automated processing, and for large-scale processing of Article 9 special category data or criminal-conviction data. In practice, any AI tool that scores, ranks, or automatically screens EU candidates needs a documented DPIA dated before go-live. Vendor compliance does not cover you; the DPIA is your own accountability record.

### What transfer tool do I need to send candidate data to the US?

You need a valid transfer tool. If the US recipient is self-certified under the EU-US Data Privacy Framework, that flow is covered without separate clauses, though it requires annual re-certification. Otherwise use Standard Contractual Clauses plus a Transfer Impact Assessment, which work for any non-EU country. Keep an SCC fallback ready, since the DPF's redress mechanism has been under monitoring after the PCLOB quorum loss in January 2025.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/guides/lawful-sourcing-standard*
