What a dirty student lead database actually looks like
A dirty database isn't an abstract problem — it's the same prospect sitting in your CRM three times under three slightly different spellings, a phone number that's really "555-555-5555" typed to get past a required field, and a Hotmail account nobody has opened since the open house eighteen months ago. None of these records look wrong at a glance; they only reveal themselves once a campaign hits send and the bounce report comes back.
Four patterns account for most of the damage. Duplicate contacts pile up when the same prospect fills out a viewbook request, an event registration, and a program inquiry form separately, each creating a fresh record instead of updating the existing one. Mistyped or fabricated phone numbers slip through when a prospect isn't ready for a call and types a placeholder just to clear a required field. Expired or role-based emails — a graduating high schooler's school-board address, or a parent's shared "family@" inbox — pass format checks but go nowhere. And incomplete records, missing a program of interest or intake term, are technically valid but useless beyond a generic blast.
None of this reflects poorly on your admissions team; it's what happens by default when prospect data arrives from a dozen entry points — a recruitment fair card, a Facebook lead ad, a static form, a provincial application portal transfer — with no single validation layer catching errors before they land in the CRM.
Why cleaning before a campaign matters more than cleaning after
Cleaning your database before a send protects deliverability, budget, and staff time in ways cleaning after the fact can't recover. Once a campaign has gone out to a dirty list, the damage — a dented sender reputation, wasted ad spend, a skewed read on what worked — is already done.
Contact records decay on their own too, as people change schools, switch numbers, or abandon an email used once for a form. HubSpot Research puts general B2B database decay at roughly 22.5% a year — a list clean at the start of an intake cycle can already be out of date by the next campaign.
Email deliverability shows the cost clearest: every hard bounce and spam complaint tells Gmail, Outlook, and other providers your sending domain is unreliable, and that reputation follows you into the next campaign. Google's guidance for bulk senders confirms a high bounce rate depresses inbox placement for the sender as a whole, so a dirty send in March can suppress open rates on a clean send in September.
Paid acquisition and staff time absorb the rest. Syncing a duplicate-heavy list to an ad audience means paying to serve the same household repeat impressions while real reach shrinks, and a counsellor who dials a duplicated contact isn't calling a real prospect in that window. Campaign analytics then report results against a denominator inflated with dead contacts, making a solid message look worse than it performed.
Deduplication logic that actually catches duplicates
Effective deduplication combines several matching signals instead of relying on an exact name match, because prospects rarely enter their own details identically twice. Name alone is the weakest signal — "Mohammed," "Mohamed," and "Md." routinely refer to the same applicant, and a common name plus a typo defeats simple string comparison. A more reliable approach layers signals: exact email match first, a normalized phone number second, then a fuzzy name-plus-postal-code match to catch prospects who used two different emails — a personal account for the first form fill, a school-issued address for the second.
The harder decision is what to keep once a duplicate is confirmed. A merge, not a delete, should be the default, since deleting a record throws away engagement history a merge preserves. When two records disagree, the more recent value should generally win for status fields like program of interest — but for consent, the more restrictive value wins, so a prospect who unsubscribed doesn't get re-added because an older duplicate still shows as opted in.
| Field type | Merge rule | Why |
|---|---|---|
| Email, phone | Keep most recently verified | Older contact info is more likely to be stale |
| Program/intake interest | Keep most recent entry | Reflects the prospect's current plan, not their first inquiry |
| Consent/opt-out status | Most restrictive value wins | Protects against re-subscribing someone who opted out |
| Source/first-touch channel | Keep earliest record | Preserves accurate attribution for reporting |
| Engagement history (opens, clicks, chats) | Combine, don't overwrite | Losing history undercounts real interest |
Phone and email validation: format check versus real verification
Format validation confirms a number or address is structurally plausible; real verification confirms it's actually reachable. The two aren't the same thing, and mid-size admissions teams rarely need to run the expensive kind at full volume.
Format validation is nearly free and should run on every record: does the phone number match a valid pattern, does the email contain a real domain rather than a typo like "gmial.com"? This layer eliminates the obvious junk — "555-555-5555" placeholders, missing digits — automatically, the moment a record is created.
Real verification goes further: pinging the mail server to confirm a mailbox exists, or running a carrier lookup to confirm a number is active rather than a landline. It's worth paying for on the active campaign segment right before a send, not the entire historical database, where cost adds up fast for limited return. Moz's guidance on data quality makes a similar case for prioritizing effort where it changes an outcome. The practical split: automate format checks on every record, batch-verify only the active segment before each send, and reserve manual review for ambiguous cases a tool flags but doesn't reject outright.
The pre-campaign cleanup checklist
Run through this sequence in the two weeks before a campaign send, not the day of.
| Step | Action | Target |
|---|---|---|
| 1. Deduplicate | Merge contacts matching on email, normalized phone, or fuzzy name + postal code | Zero duplicate active contacts in the send segment |
| 2. Format-validate | Run every phone and email through a syntax check | 100% of records pass or are flagged |
| 3. Real-verify the segment | Confirm mailbox and number are live for contacts in this campaign only | <5% invalid in the final send list |
| 4. Suppress stale records | Move contacts with no engagement in 12+ months to a separate list | Active list reflects genuinely reachable prospects |
| 5. Fill critical gaps | Flag records missing program of interest or intake term for a lighter, more generic touch | No blind personalization on missing fields |
| 6. Confirm consent | Cross-check opt-out status against your suppression list before syncing to any ad platform | Zero opted-out contacts in the audience |
| 7. Spot-check a sample | Manually review 20-30 random records for anything the automated passes missed | Catches edge cases before, not after, send |
How structured intake prevents dirty data at the source
The cheapest way to clean a database is to stop letting bad data into it in the first place, at the point of entry rather than in a cleanup project six months later. A static web form can't tell a prospect that "555-555-5555" isn't a real number — it accepts whatever satisfies the required-field rule and moves on.
A chatbot handling first contact can validate format conversationally, the way a person would: if a prospect types a number that doesn't match a plausible pattern, the bot can ask again right there, before the record is ever created. The same goes for email — a chatbot can catch an obviously malformed address in the moment, instead of letting a campaign discover the problem at send time.
Across Skolbot's European partner panel — 18 schools, median results, 2024-2025 — qualified leads climbed from 120 to 195 a month (+62%) while cost per lead dropped 38%. That gain reflects the chatbot's effect combined with funnel optimizations schools ran at the same time, not the chatbot alone, and it's a European benchmark cited here for comparative context rather than a Canadian figure. Still, the mechanism translates directly: fewer malformed records at intake means less time chasing dead contacts and more attention on prospects who can actually be reached, which is what drives a lower cost per qualified lead. Our guide to calculating true student acquisition cost covers how that's worked out for a Canadian institution.
This doesn't replace an admissions counsellor's judgment on which prospects to prioritize — it frees up their time for contacts actually worth calling. Our comparison of CRM platforms for higher education looks at which systems make validation and deduplication easiest to configure natively.
Data retention: what PIPEDA and Loi 25 require of a clean list
A clean database isn't just a marketing best practice in Canada — it's a compliance requirement spanning two overlapping regimes for any institution recruiting nationally. The federal Personal Information Protection and Electronic Documents Act (PIPEDA) sets the baseline: organizations must not retain personal information longer than necessary for the purpose it was collected for, and a record kept indefinitely "just in case" doesn't meet that bar.
Quebec's Loi 25, in force since September 22, 2023, goes further and applies to any institution recruiting prospects based in Quebec, regardless of where the institution operates. It requires clearer, more specific consent at collection and carries penalties well above PIPEDA's. A national cleanup policy needs to satisfy Loi 25's stricter rules across the whole database, not just the baseline other provinces fall back on — build the deletion schedule to the stricter standard from the start, and use each cleanup pass to flag records past their retention window. A stale record that fails a phone or email check is usually also overdue for removal on privacy grounds, not just a marketing one.
FAQ
How often should a school clean its student lead database?
Run a light validation pass monthly — format checks, opt-out cross-referencing — and a full deduplication pass before each major campaign, at minimum quarterly. Admissions cycles concentrate new form fills in short windows, so the database gets dirtiest right before the periods that matter most.
What percentage of a student database is typically duplicate or invalid?
There's no universal figure; it depends on how many intake points feed the CRM. Institutions running multiple disconnected forms — event registration, viewbook download, program inquiry — without a unified capture layer commonly find a meaningful share of their list is duplicated, outdated, or unreachable.
Should invalid contacts be deleted or just suppressed?
Suppress first, delete on a schedule. Moving unreachable or opted-out contacts to a suppression list keeps them out of active campaigns while preserving the record for reporting, then schedule deletion once they've passed your retention window under PIPEDA or Loi 25.
Can a chatbot really stop bad data from entering the CRM?
It can catch a meaningful share at the point of entry, though it isn't a full substitute for a periodic audit. Asking again when a number doesn't look plausible prevents much of the placeholder and mistyped entries a static form waves through, but records already in the CRM from other channels still need a separate cleanup pass.
Does list cleaning hurt campaign reach by shrinking the list?
It shrinks the number of contacts sent to, but grows the number who can respond. A smaller, genuinely reachable list outperforms a larger one padded with dead contacts on every metric that matters — open rate, click rate, applications started — and protects the sender reputation your next campaign depends on.
For the broader strategy this fits into, see our digital marketing guide for higher education. If your database has gone quiet rather than dirty, our piece on reactivating dormant student leads covers winning stalled prospects back.
Try Skolbot on your school in 30 seconds


