Refolk
PlaybookProcess, data, and compliance

The Scheduled Purge Playbook: Deleting Sourced Contact Data When Its Basis Expires

You can build a per-record-type retention schedule with clock triggers, then run a repeatable purge that deletes expired records and leaves a regulator-grade audit trail.

16 min readLast reviewed September 7, 2026Read as Markdown

Key takeaways

  • The widely repeated "6 to 12 months" is a UK and German claim-limitation artefact, not an EU floor; France caps an unconsented candidate file at 2 years from last contact under the CNIL framework.
  • Data refresh is the silent over-retention engine: the CNIL fined a data scraper 240,000 euros for resetting a 5-year clock on every job change across a 160-million-contact database.
  • Backups, not production, are where purges fail audits; the EDPB found backup deletion gaps flagged by half of participating supervisory authorities.
  • Legitimate interest is self-expiring and consent is self-revoking, so a purge must read the lawful-basis field, not just the record date.
  • Statutory floors override the purge only where their condition applies: HMRC PAYE records must be kept at least 3 years, NMW and VAT records at least 6 years.
  • The work is under-staffed for its risk: Refolk's index shows 17 UK Recruiting-Ops professionals against 698 UK Data Protection Officers, roughly 41 DPOs per recruiting-ops owner.

This playbook is for the recruiting-ops, RevOps, and data owners answerable for how long sourced candidate and contact data sits in your systems. The job is narrow and unglamorous: work out how long you may keep each type of sourced record, then actually delete the ones past their window without breaking an open requisition or a live deal. What you get by the end is a per-record-type retention schedule with named clock triggers and a repeatable quarterly purge that leaves an audit trail a regulator would accept.

Most guidance on this stops at a vague "6 to 12 months" and moves on. That number is a UK and German artefact, and it governs one record type in two jurisdictions. It does not tell you what to do with a French talent pool, a cold-outreach contact list, or a placed-candidate file, and it says nothing about how you run the deletion at scale. This playbook covers the proactive, scheduled side of storage limitation. It sits alongside the Lawful-Sourcing Standard, which sets policy, and the Data Access and Deletion Request Runbook, which handles one subject at a time when they ask. Here I am deleting at scale, on a clock, before anyone asks.

Why "6 to 12 months" is the wrong answer to start from

The single number everyone repeats is a claim-limitation artefact, not an EU retention floor, and it breaks the moment your dataset spans more than one country. It derives from the periods in which a rejected applicant can bring a legal claim, so it is really a defence window, not a purpose window.

Trace where it comes from. In Germany, a discrimination claim under Section 15(4) of the AGG must be asserted within two months, with a further three months to file under the Labour Court Act. German practice therefore lands on keeping rejected-applicant data for a maximum of around six months so the employer can defend an AGG claim. In the UK, the Equality Act limitation for recruitment claims is commonly cited as six months, which is the usual basis for six-month CV retention. Both numbers are downstream of litigation risk.

France breaks the pattern entirely. The CNIL reference framework caps an unconsented candidate file at two years from the last contact with the candidate, and applies the same two-year outer limit to talent-pool and CV-database records, provided the profile remains of genuine interest and the candidate has not objected. That is not a claim window. It is a purpose window with a different clock. One number cannot govern a multi-jurisdiction dataset, and any schedule that pretends otherwise is wrong somewhere.

2 years
CNIL cap on an unconsented candidate file, from last contact
The same two-year outer limit applies to French talent-pool and CV-database records.

The practical consequence: your schedule is a matrix of jurisdiction against record type, not a single field. Build it that way from the start.

The retention schedule: window, trigger, and basis per record type

A retention schedule maps every sourced record type to a defensible window and the exact event that starts its clock. The two variables that decide the window are jurisdiction and lawful basis; the trigger is a named event, never "whenever we last touched the record."

Here is the core of the schedule, drawn from the published frameworks. Read each row across the clock trigger, which is the column that actually drives the purge.

JurisdictionUnsuccessful applicantConsented talent poolClock trigger
France (CNIL)2 years2 yearsLast contact
Germany (AGG practice)~6 monthsUp to 3 yearsReceipt of rejection
UK (ICO)End of claim period (~6 months)Duration of consentVacancy close

Two structural points sit behind this table. First, lawful basis changes both the window and the trigger. Legitimate interest, such as defending a recruitment claim, supports only a short window tied to the claim period; once the role is filled and that window closes, the basis lapses and deletion is forced. Consent supports a longer talent-pool window, but it is revocable at any time, and withdrawal itself forces deletion. As the ICO's reasoning goes, legitimate interest does not cover keeping a candidate in your database indefinitely after the original purpose has ended. Once the role is filled and the claim window has passed, the legitimate interest has expired with it, and you need either a new basis or a deletion.

Second, the trigger event is the thing that most schedules get wrong, so name it explicitly. "Last contact" and "receipt of rejection" are original events. "Last update" is not, and treating it as one is how you keep active people forever.

Statutory floors that override the purge

Some records carry a hard legal floor that forbids early deletion, and these must be encoded as overrides so the purge cannot touch them. The floors are narrow, they attach to payroll and tax records rather than to sourced candidate data, and they only apply where their triggering condition is met.

RecordFloorBasis
PAYE/payroll3 years from end of tax yearSI 2003/2682 reg.97
NMW records6 yearsNMW Regs 2015
VAT records6 yearsHMRC

The ICO confirms that floors defend retention: if you keep personal data to comply with a legal requirement such as income-tax or audit records, you will not be treated as having kept it for longer than necessary. That is the useful half of the rule. The dangerous half is invoking a floor where it does not actually apply. In the PAP case, the company kept ten-year contract data below the monetary threshold that would have triggered the obligation, so the floor it relied on did not exist for those records. Before you let a floor stop a deletion, verify that the floor's condition is met.

The critical point for a sourcing dataset: sourced applicant CVs, interview notes, and cold-outreach marketing contacts have no statutory floor at all. For those record types, the purpose-limited clock is the only control. There is nothing holding them back, so if the schedule says delete, you delete.

Run the purge: the eight-step procedure

The purge is a repeatable quarterly process that moves from inventory to a signed-off deletion to an immutable audit log. Run it in this order. The sequencing matters most at the front: inventory before you set windows, because the CNIL and ICO both require the schedule to reflect your documented purpose, not a generic template.

The quarterly purge, start to finish

  1. Inventory and classify the sourced dataset
    List every record type, and for each record its lawful basis, source, and system location. Produce a register naming each type against those attributes.
  2. Assign each type a window and a named clock trigger
    Map each type to a defensible period and record the exact trigger event. End with one "keep until X, then purge" rule per type.
  3. Encode statutory floors as overrides
    Flag records with hard floors so the purge cannot delete them early, and verify each floor's condition applies. Reconcile the override list against the schedule.
  4. Build the purge query and a dry-run report
    Generate the list of expired records without deleting. Produce a reviewable candidate-for-deletion export with counts by type and reason.
  5. Human review against active searches and deals
    Suppress false positives such as open reqs, live deals, and fresh consent. Sign off the deletion list and log every exception with a reason.
  6. Execute hard delete or irreversible anonymization
    Delete from production and schedule the records out of backups. No soft-delete or pseudonymization alone.
  7. Produce the audit artefact
    Write an immutable log of what was deleted, the count, rule triggered, method, operator, timestamp, and backup-disposition note.
  8. Schedule the next run and review the schedule
    Calendar the next quarterly purge and book an annual review of the schedule itself.

Rough timings, so you can staff it: inventory takes one to two weeks the first time, the schedule three to five days, floors two to three days, the query build about a week, human review two to three days, execution one to two days plus a backup cycle, and the audit artefact the same day as execution. After the first run the front half collapses, because the register and schedule already exist and you are only maintaining them.

The purge pipeline

  1. Dry-run export
    Query flags expired records, deletes nothing, counts by type and reason
  2. Human review
    Recruiting and RevOps suppress open reqs, live deals, and fresh consent
  3. Execute deletion
    Hard delete or irreversible anonymization across production and backups
  4. Audit artefact
    Immutable log of what, how many, which rule, what method, who, when
Expired records move from a dry-run export through human review to verified deletion and an immutable log.

Step one is where the sourcing tools earn their place. You cannot schedule a purge for records you cannot find, and the person who knows where sourced contact data actually lives is rare. If your inventory needs to reach the people who own candidate data and ATS hygiene, Refolk finds them by asking in plain English rather than guessing at titles.

What compliant deletion actually means

Compliant deletion is a hard delete or a genuinely irreversible anonymization, proven in a way an auditor can verify. Two things that look like deletion do not qualify, and both fail audits regularly.

The first is pseudonymization. Replacing identifying fields with tokens while retaining the token map leaves the data as personal data under GDPR. The EDPB's erasure findings are blunt on this: applying pseudonymization in response to an erasure request and then treating the obligation as fulfilled is a compliance failure. True anonymization must be irreversible, meaning no reasonable means of re-identification exists in the hands of the controller, any joint controller, or any reasonably foreseeable third party. If a re-identification key survives anywhere, you have anonymized nothing.

The second is soft-delete. Article 17 requires verifiable erasure, not just functional deletion. An API that returns no records can mean either deletion or suppression, so soft-delete provides no real proof. Soft-deleted records remain recoverable at the storage layer, which is precisely what erasure is meant to prevent.

Levels of "gone," weakest to strongest

  1. Soft-delete
    Record flagged deleted in the app, still present at the storage layer
  2. Pseudonymization
    Identifiers tokenised, token map retained, still personal data
  3. Irreversible anonymization
    No re-identification key exists anywhere, data no longer personal
  4. Hard delete
    Record removed from production and disposed of through the backup cycle
Only the bottom two layers satisfy an erasure obligation; the top two only look like it.

Proof is the part people skip. Because verifiability is the legal standard, the audit artefact is not paperwork you produce for its own sake; it is the evidence that the deletion happened. Capture the count, the rule that triggered each deletion, the method, the operator, the timestamp, and a note on backup disposition. Store it where it cannot be edited afterwards.

A clean production database is not evidence of compliance if the backups still hold the records.

How this goes wrong: failure modes and false positives

Most purges fail in predictable ways, and nearly all of them produce a false positive that makes the system look compliant when it is not. Treat this section as the checklist behind the checklist.

  • Pseudonymization mistaken for deletion. The dashboard shows "anonymized," but a token map still exists somewhere, and the data remains personal. Check: confirm no re-identification key survives in any system, backup, or third-party hand.
  • Soft-delete leaves ghosts. An API returning nothing can mean suppression, not erasure. Check: query the storage layer and backups directly, not the application.
  • Clock reset on data refresh. Re-scraping or a job-change update silently restarts the window. This is the exact pattern the CNIL fined: a data scraper kept contact details for five years from each data update, which typically happens when a person changes job, so active people were kept far too long across a database of roughly 160 million contacts. Check: the trigger must be original last-contact, never last-update.
  • Statutory floor applied where it does not exist. A floor invoked without its condition being met, as in the PAP case, is over-retention dressed up as compliance. Check: verify the floor's condition actually applies before using it as a reason to keep.
  • Backups excluded from the purge. Production reads clean while backups still hold the records; half of participating supervisory authorities flagged this gap. Check: confirm the backup rotation disposes of purged records within its cycle.
  • Consent treated as permanent. Consent is revocable, and a stale opt-in from years ago is not a valid basis. Check: log the consent date and re-confirm before relying on it.
  • Manual purge that never runs. A written policy exists but no job executes it. This is the most common cause of over-retention. Check: produce the audit log from the last quarterly run as evidence the process is live, not just documented.
240,000 euros
CNIL fine for resetting a retention clock on every job change
The five-year window renewed on each data update, keeping active people indefinitely across roughly 160 million contacts.

The through-line: a schedule that exists on paper does nothing. Regulators now read retention schedules as maximums rather than minimums, and the failure to sort and delete stale records is itself the violation. The recent Free Mobile case, a 42-million-euro fine, turned on exactly that failure to sort and delete former-subscriber data. "Keep everything just in case" is no longer a neutral default; it is the finding.

Who owns this, and why it usually has no operator

The single biggest structural risk is that no one owns the purge, because the person who understands where sourced data lives is genuinely scarce. This is not a staffing footnote; it is why the manual-purge-that-never-runs failure mode is the most common one.

The scale of the mismatch shows up in the population itself. In Refolk's index, the UK holds 698 Data Protection Officer titleholders against just 17 people with a current Recruiting or Recruitment Operations title, roughly 41 DPOs for every recruiting-ops owner. The DPO can write the policy, but the person who knows which system holds the scraped contacts and how the ATS clocks a rejection is rare, and the purge usually lands on no one's desk.

SegmentCountDerived
Recruiting Ops, US111baseline
Recruiting Ops, UK17US has ~6.5x more
Data Protection Officer, UK698~41x the UK Recruiting-Ops pool

The fix is to name an operator, not a policy owner, and to give that operator the audit log as their deliverable. Refolk is useful here twice: to inventory where sourced data sits at step one, and to find the RevOps and privacy-engineering people who can build and run the automation. A search for privacy engineers who have built GDPR data-deletion or retention automation surfaces the exact profile most teams are missing, without you having to guess which title they carry.

Before you call the run done

Verify each of these before you sign off a purge. Every item maps to a failure mode above, and every unchecked item is a way the run looks complete while leaving data behind.

Purge sign-off checklist

  • The schedule maps every record type to a jurisdiction, a window, a lawful basis, and a named clock trigger.
  • The clock trigger is an original event (last contact, receipt of rejection, vacancy close), never last-update.
  • Statutory floors are encoded as overrides, and each floor's triggering condition has been verified to apply.
  • The dry-run export was reviewed by recruiting and RevOps against open reqs and live deals, with exceptions logged.
  • Consent-based records were re-confirmed against a logged consent date, not treated as permanent.
  • Deletion was a hard delete or irreversible anonymization, with no surviving re-identification key anywhere.
  • Purged records are confirmed disposed of from backups within the rotation cycle, verified at the storage layer.
  • An immutable audit log records what was deleted, count, rule, method, operator, timestamp, and backup disposition.
  • The next quarterly run is calendared, and the annual schedule review is booked.

Keeping the schedule current

The schedule is a living document, and the fastest way to fall out of compliance is to let it drift while the purge keeps running against stale rules. Review the schedule itself annually, and re-check the mechanism whenever a trigger changes rather than trusting a number you wrote once.

Three things reliably move. Jurisdictional guidance shifts, so re-read the CNIL framework and the ICO storage-limitation guidance on your annual review rather than assuming the windows held. Your own purposes change: a new sourcing channel or a new lawful basis adds a record type that needs its own window and trigger, and that only enters the schedule if inventory is repeated. And your systems change, because an ATS migration or a new backup provider can quietly break the storage-layer deletion you verified last year.

Retention schedule row
Record type:        [e.g. cold-outreach contact]
Jurisdiction:       [e.g. UK]
Lawful basis:       [legitimate interest | consent]
Window:             [e.g. end of claim period, ~6 months]
Clock trigger:      [original event, e.g. vacancy close]
Statutory floor:    [none | floor + verified condition]
System location:    [where this record type lives]
Deletion method:    [hard delete | irreversible anonymization]
Last reviewed:      [date]

One row per record type per jurisdiction. Fill every field; a blank clock trigger is a purge waiting to fail.

Keep the audit logs from every run in one place. They are your evidence that the process is live, they show a regulator the schedule is enforced rather than aspirational, and they are the first thing you will be asked for. A purge that runs quarterly and logs each run is a defensible position. A policy document with no run behind it is the finding waiting to happen.

Questions practitioners ask

How long can I keep sourced candidate data?

It depends on jurisdiction and lawful basis, not a single number. Under the CNIL framework an unconsented candidate file caps at two years from last contact. German practice keeps rejected-applicant data around six months, tied to the AGG claim window. The UK ICO says keep unsuccessful-applicant records only to the end of the statutory claim period, commonly cited as six months, absent a clear business reason.

Does anonymizing candidate data count as deleting it?

Only if the anonymization is genuinely irreversible. The EDPB confirmed that pseudonymization, where identifying fields are swapped for tokens but the token map is retained, remains personal data under GDPR and does not satisfy an erasure obligation. True anonymization means no reasonable means of re-identification exists in the hands of the controller, any joint controller, or any foreseeable third party. If a re-identification key survives anywhere, you have not deleted.

Do I have to delete data from backups too?

Yes, and this is the most common audit failure. A clean production database is not evidence of compliance if backups still hold the records. The EDPB found backup deletion gaps flagged by half of participating supervisory authorities. Confirm your backup rotation disposes of purged records within its cycle and note that disposition in the audit log for the run.

What restarts the retention clock?

In France the trigger is last contact with the candidate; in Germany it is receipt of rejection. Statutory claim periods can extend the clock. The dangerous reset is data refresh: re-scraping or a job-change update that silently restarts the window. The CNIL fined a data scraper 240,000 euros for exactly this, because a five-year clock reset on every job change kept active people indefinitely.

Which sourced records have no statutory floor?

Sourced applicant CVs, interview notes, and cold-outreach marketing contacts carry no statutory retention floor. For these, the purpose-limited clock is the only control, so once the original purpose ends and any claim window closes, the basis lapses and deletion is forced. Payroll, PAYE, NMW, and VAT records are different: they carry hard floors that override early deletion.

How often should the purge run?

Quarterly, with an annual review of the schedule itself. A written retention policy with no job that executes it is the most common cause of over-retention, and regulators now read schedules as maximums rather than minimums. A quarterly cadence keeps the expired-record backlog small and produces a regular audit artefact you can show on request.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next