What does a dirty student lead database actually look like?
A dirty database is one where a meaningful share of records can't be reached, aren't unique, or no longer reflect a real prospect. It rarely announces itself — the CRM shows a healthy contact count and campaigns still send, until open rates or call connection rates quietly slip.
The most common culprits are predictable once you go looking for them:
- Duplicate records from repeated form fills. A prospective student downloads a brochure in September under one email, requests a callback in October under a different one, then fills in an open-day form in November using a parent's phone number. Three records, one person.
- Mistyped or fake phone numbers. A rushed mobile-first submission drops a digit, transposes two, or gets "07000000000" typed in just to skip a required field.
- Expired or role-based email addresses. A school-leaver's college email stops working the summer after they leave; a parent enters
info@theirfamily.comor another shared inbox nobody reads. - Incomplete records with no course interest or intake year. These sit as a name and a stale email, unusable for any segmented campaign.
None of this is unusual — it's the normal residue of collecting prospect data across six or seven channels over a cycle. The problem isn't that it exists; it's sending a campaign into it without checking first.
Why clean before a campaign, not after?
Cleaning before a send protects deliverability, spend and staff time in ways cleaning after can't undo. Once a campaign has gone out to a dirty list, the damage — bounces, wasted ad impressions, a distorted read on what worked — is already done.
Four consequences compound quickly once a campaign fires into an unclean list:
- Deliverability takes the hit for the whole campaign. Mailbox providers score your sending domain on bounce and complaint rate across the entire send — Google's guidelines for bulk senders note that a high bounce rate on one campaign can push future emails from the same domain into spam.
- Ad spend is wasted retargeting contacts who were never reachable. Upload a lead list to a Meta or Google custom audience and every duplicate or dead email is still counted, and billed, as a targeted contact.
- Admissions staff burn hours chasing numbers that don't connect. A counsellor working a 200-number call list where 30 are wrong or disconnected gets nothing back for that share of the time.
- Campaign analytics stop meaning anything. If 15% of your list is duplicated, open rate and cost per lead are both calculated against an inflated denominator, skewing every decision made from that data.
Weigh that against what a UK admissions team pays to acquire a student: the average cost to acquire an enrolled student in the UK sits at £2,400-3,200 (Source: Skolbot estimates based on EAIE, StudyPortals and Campus France sector data, 2025-2026). A duplicate or dead contact in your CRM isn't a neutral inefficiency — it's spend you've already made to acquire that contact once, now wasted a second time because you can't reach them. Cleaning the list before another pound goes into a campaign is the cheapest fix available, covered alongside other cost drivers in our guide to calculating the true cost of student acquisition.
How do you actually deduplicate a student lead list?
Deduplication works by matching records on a combination of fields, deciding which matches are true duplicates rather than coincidences, then merging survivors by a fixed rule for which value wins. Skip any one of those steps and you either miss real duplicates or merge records that shouldn't be combined.
Matching rules that hold up in practice
Email alone is too narrow — the same prospect often uses two addresses across their journey — and name alone is too loose, given how many "James Smith" or "Sarah Patel" records a mid-size university receives in a single cycle. A workable rule combines several weaker signals:
- Exact email match is the strongest single signal and should always trigger a duplicate check.
- Normalised phone number match — strip spaces, dashes and the leading
+44/0before comparing, since07911 123456and+447911123456are the same number written two ways. - Fuzzy name match plus a secondary field — a close name match (allowing for typos and nickname variants) paired with the same postcode, date of birth, or intake year is a reliable third signal.
Merge or delete, and which field wins
Merge, don't delete, whenever the two records hold different useful information — one has a phone number, the other a course preference. Delete outright only for pure noise: a form submitted twice in the same minute, or an obvious test entry.
The rule that works best is "most recent non-empty value wins," field by field, rather than keeping the oldest or newest record wholesale. A phone number entered eight months ago still beats no phone number; an intake year entered yesterday should override one from a different cycle. Most CRMs used in UK admissions support this kind of field-level merge natively — worth checking in our comparison of CRM platforms for higher education.
Format validation vs real verification: what's worth automating?
Format validation checks that an email or phone number is structurally plausible; real verification checks that it actually reaches someone. A mid-size admissions team doesn't need both at full strength on every record — HubSpot's research on marketing data quality points the same way: automate the cheap, high-volume checks and reserve human review for records where getting it wrong actually costs something.
Email: format and MX-record checks (does the domain exist and accept mail) are cheap to automate on every record at entry, catching typos like gmial.com before they reach the CRM. Disposable-email detection — flagging temporary-inbox addresses, almost always abandoned within days — is worth automating too. Full verification (a confirmation link, waiting for a click) is best reserved for high-value segments, such as applicants close to a deposit deadline, rather than run database-wide.
Phone: format validation (correct UK mobile or landline structure, right digit count) can run automatically on every submission. An SMS ping — a short text confirming the number is live — costs more per contact and suits selective use, before a call campaign to a segment untouched in months rather than on every new enquiry.
| Check | What it catches | Automate or manual review? |
|---|---|---|
| Email format + MX check | Typos, non-existent domains | Automate on every record |
| Disposable email detection | Temp-inbox addresses, unlikely to be checked twice | Automate on every record |
| Email confirmation link | Addresses that exist but aren't actually monitored | Manual trigger, high-value segments only |
| Phone format validation | Wrong digit count, obvious fakes (e.g. repeated digits) | Automate on every record |
| SMS verification ping | Disconnected or reassigned numbers | Manual trigger, before call campaigns |
| Fuzzy name + secondary field match | Genuine duplicates that email/phone alone miss | Automate, review flagged pairs manually |
The pre-campaign cleaning checklist
Run this sequence in order — each step narrows the list further, so working out of order means redoing earlier steps once later checks surface more duplicates.
| Step | Action | Typical effect on list size |
|---|---|---|
| 1 | Remove exact duplicate emails and normalised duplicate phone numbers | -5 to -12% |
| 2 | Run fuzzy name + secondary field match, review flagged pairs manually | -2 to -5% additional |
| 3 | Strip role-based and disposable email addresses | -1 to -3% |
| 4 | Validate phone number format; flag or remove obvious fakes | -2 to -6% |
| 5 | Flag records inactive beyond your retention window for review or suppression | Varies by database age |
| 6 | Re-run campaign segment counts against the cleaned list before sending | N/A — sanity check |
Step 5 is a compliance checkpoint, not just a hygiene one. The ICO's guidance on the storage limitation principle is explicit that personal data shouldn't be kept longer than necessary for its original purpose — a prospect who enquired three years ago and never engaged again is a data protection question as much as a deliverability one, and a pre-campaign clean is a natural moment to review and purge those records.
How does capturing leads through a chatbot prevent dirty data in the first place?
A chatbot that validates format at the point of entry stops a large share of dirty data from ever reaching the CRM, instead of leaving it to be caught and cleaned later. A web form only checks that a field isn't empty; a well-configured chatbot conversation checks that an email looks structurally valid and a phone number has the right shape for a UK mobile before moving on — and simply asks again, in plain language, if it doesn't. That matters because most dirty data isn't malicious, it's careless: a prospect typing quickly on a phone, skipping a validation error rather than lose their place, or entering a placeholder number just to clear a required field. A conversational flow that catches the problem in the moment — "That doesn't look like a full UK mobile number, mind double-checking it?" — fixes it at the exact point it would otherwise enter the database, instead of surfacing three months later as a bounce.
The effect shows up in aggregate numbers too. Across 18 schools tracked over 2024-2025, qualified leads per month rose from 120 to 195 (+62%) after introducing a validating chatbot at intake, and cost per lead fell from roughly £42-equivalent to £26-equivalent (-38%) (Source: Skolbot median results, 18 schools, 2024-2025). Those gains combine the chatbot's own effect with concurrent funnel optimisations run over the same period, so the chatbot alone doesn't explain the whole shift — but a cleaner intake funnel is a consistent part of what drives cost per lead down, the same logic our guide to the wider higher education marketing funnel applies to acquisition generally: prevention at the point of capture beats correction downstream.
A chatbot complements the admissions team's judgement here, it doesn't replace it. Flagging a suspicious number is a mechanical check; deciding whether an ambiguous record deserves a manual follow-up call still needs a person — its job is to free up counsellor time from chasing bad data, not to make admissions decisions.
FAQ
How often should we clean our student lead database?
Run the full checklist before every campaign send, and schedule a lighter pass — deduplication and format checks only — monthly regardless of whether a campaign is planned. Databases degrade continuously as new records arrive through multiple channels, so a single annual clean can't keep pace.
Can we automate deduplication entirely, or does it need manual review?
Automate the matching and flagging step, but keep manual review for anything the fuzzy-match logic surfaces as probable rather than certain. Exact email or normalised phone matches are safe to merge automatically; fuzzy name matches alone carry too much false-positive risk — two different prospects sharing a common name — for that.
What's the difference between deleting a record and suppressing it?
Deleting removes the record entirely, appropriate for genuine noise like a duplicate submission or a clear test entry. Suppressing keeps the record but excludes it from active campaigns — the better choice for a prospect who's gone unreachable but might still hold reporting value or resurface later.
Does list cleaning conflict with UK GDPR record-keeping requirements?
No — the two point the same direction. The ICO's storage limitation principle expects organisations to review and remove personal data once it's no longer needed for the purpose it was collected for, so a pre-campaign clean that flags and purges long-inactive records is a compliance step as much as a deliverability one.
Try Skolbot on your school in 30 seconds


