Data Classification: The Step Most DLP Projects Skip

Last updated: 09/08/2026
Cybersecurity

Companies buy data loss prevention software, switch it on, and watch it flag a few thousand files in the first week. Almost none of them matter. The tool isn't broken. It has no idea which of those files are worth protecting, because nobody ever told it. Six weeks later the rollout is parked and everybody blames the vendor.

Data classification sorts your company's data into sensitivity levels so security controls apply by rule instead of by guess. It belongs before DLP, not after, and it starts with finding what data you already have.

Only 37% of breached organizations encrypt sensitive data both at rest and in transit, according to IBM's 2026 Cost of a Data Breach report. Encryption, access rules, retention schedules, and any working data loss prevention program all depend on the same thing first. Knowing which data is which.

That's classification. It's also the least popular part of the job, because there's no dashboard for it and no vendor demo ends with a slide about the six weeks you'll spend asking department heads where the contracts actually live. So it gets skipped. The tool gets bought, the tool misbehaves, and the project dies quietly in a backlog.

Worth saying plainly before anything else. Classification isn't a labeling exercise. It's a scoping decision, and you do it to make the DLP policy smaller.

What is data classification, and why does DLP depend on it?

Data classification is the process of sorting data into sensitivity levels, like Public or Confidential, so that storage, sharing, and access rules apply to a whole category at once instead of file by file.

Two different jobs get called the same name, and mixing them up is where the first month gets lost. Discovery answers where your data is. Classification answers how sensitive it is once you've found it. They run in sequence. You can't label an inventory you don't have.

Sitting on top of all that is data loss prevention, the software that watches sensitive files as they move to email, USB drives, or a browser tab. It reads the label and enforces the rule. Without labels it falls back to pattern matching on everything it can see, which is how you end up with thousands of alerts and a security analyst who stopped opening them months ago. NIST put out a draft guide on exactly this problem, SP 1800-39, Data Classification Practices, in February 2026. Its framing is that sensitive information tends to sit scattered across systems, chat threads, data lakes, and file repositories. Which is a polite way of saying nobody knows where it all lives.

We covered what DLP actually does separately, if you want the mechanics of the enforcement layer. This post is the step before it.

Why classification gets skipped

Nobody owns it. That's the honest answer.

Security owns the tool. Legal owns the retention rules. The department heads own the actual files and have never been asked to think about them in tiers. Classification sits in the gap between all three, and work in a gap doesn't get scheduled.

Splunk's 2019 State of Dark Data survey of more than 1,300 business and IT leaders put 55% of organizational data in the dark, meaning nobody was sure what was in it. That survey is dated. The current numbers point the same direction, and IBM's 2026 report found that just 34% of organizations have visibility into their cryptographic assets. Unread, unlabeled, fully backed up, and fully exposed. Paid for, too.

Then there's the quieter reason. Every vendor implies the tool will handle it. Automated classification is real and it works, but it works on things that have a shape. A Social Security number has a shape. A credit card number has a shape. Microsoft ships prebuilt sensitive information types for hundreds of those patterns, and they're genuinely good. Your unreleased product roadmap has no shape. Neither does the pricing model your competitor would pay for. Pattern matching cannot find those, and those are usually the files that actually hurt.

You feel it as noise. Microsoft publishes an entire page on tuning classifier accuracy, complete with guidance on reviewing true and false positives by hand, which tells you how routinely the first configuration gets it wrong. An alert stream nobody trusts is worse than no alert stream. Worse, not equal. At least with no alerts you know you aren't covered.

Isometric grid of identical unlabeled blocks with one cluster pulled open, representing dark data nobody has classified

The four levels, and why three is usually enough

Four levels is the standard model. Public, Internal, Confidential, Restricted. Every policy template you'll find online uses some version of it, and the definitions are broadly consistent across all of them.

  • Public. Marketing pages, published spec sheets, job postings. Open to anyone, and a leak costs you nothing.
  • Internal. Org charts, meeting notes, most of SharePoint. Any employee can open it. A leak is embarrassing, occasionally a competitor advantage.
  • Confidential. Contracts, pricing models, payroll, customer lists. Named teams only. A leak means contract disputes, lost deals, an angry customer call.
  • Restricted. Regulated records, source code, controlled unclassified information, trade secrets. Named individuals only, logged and reviewed. A leak means regulatory action, breach notification, and litigation.

I'd push back on the standard advice. For a company running 20 to 1000 users, four tiers is usually one tier too many, and the fourth is where adoption quietly dies.

It fails like this. Somebody has to decide, in the moment, whether a signed vendor contract is Confidential or Restricted. It's a genuinely arguable call. So they guess, or they pick the higher tier to be safe, or they skip labeling altogether because the dropdown made them think for eight seconds and they had a meeting to get to. Now your Restricted tier is full of ordinary contracts, and the access controls you built for the crown jewels are annoying 40 people a day. Somebody files a ticket. The controls get relaxed. You're back where you started, but with a policy document that says otherwise.

Three levels forces a cleaner question. Is this public, is it ordinary internal work, or would it genuinely hurt if it got out? People can answer that one without a meeting. Add the fourth tier later, once you have a real population of files that needs it and a real owner asking for it.

What NIST actually says about classifying data

Borrow the federal model even if you'll never sell to a government agency. It asks a better question than "how secret is this?"

FIPS 199 categorizes information on three axes rather than one. Confidentiality, integrity, and availability. What happens if this gets out, what happens if it gets altered, and what happens if you can't reach it. Each gets a low, moderate, or high impact rating, and the highest rating drives the control set. CISA will even run the categorization for federal systems as a service.

Private-sector schemes usually measure only the first axis. That's a real gap. Your production bill of materials isn't especially secret, and a competitor who saw it would shrug. But if somebody silently changes a tolerance in it, you scrap a run. Integrity, not confidentiality. A single-axis policy files that document as Internal and applies nothing to it.

NIST's newer draft, SP 1800-39, is more practical and much closer to what a mid-market company needs. The NCCoE built a synthetic dataset and ran commercially available classification tools against it to show discovery, identification, and labeling end to end. It also makes a point that's easy to miss. Classification is now the prerequisite for the things everyone wants next, including zero trust architecture, quantum-safe cryptography migration, and feeding your own data to AI models on purpose rather than by accident. Already writing a data inventory for one of your compliance frameworks? Then you're deep into this work and probably don't realize it.

Check what your Microsoft licensing already covers

Before you price a classification tool, open your Microsoft 365 bill. A meaningful share of this capability is already in there, and which share depends entirely on your plan tier.

  • Manual sensitivity labels, applied by whoever made the file. Business Premium, E3, F1 and F3, and everything above.
  • Scanner-based discovery of on-premises file shares. Microsoft 365 E3.
  • Content Explorer aggregation, meaning the index quietly builds in the background. Microsoft 365 E3.
  • Content and Activity Explorer, the screens you actually look at. Microsoft 365 E5 or E5 Compliance.
  • Automatic and policy-based labeling. Microsoft 365 E5, or Information Protection and Governance.
  • Endpoint DLP, covering files sitting on laptops. Microsoft 365 E5 or the Purview Suite.

Two lines in the Microsoft Purview service description are worth reading twice. "Scanner-based discovery is supported with a Microsoft 365 E3 license." And this one, on classification analytics, which says the Content and Activity Explorer interfaces need E5, while "the underlying data aggregation continues for E3/A3/G3 tenants."

Read together, those two sentences say something useful. On E3 you can discover, you can label by hand, and the index is quietly accumulating whether you're looking at it or not. The E5 upgrade buys automation and a window to see through. It does not buy permission to start.

Which matters. Starting is the part everyone defers while they wait on budget.

About that automated tier, since it's what people assume will rescue them. Trainable classifiers let you teach Purview to recognize your own document types by feeding it examples, which is genuinely the answer for shapeless things like statements of work or design docs. Two limits nobody mentions in the demo. Custom classifiers only support English. And a published classifier can't be retrained, so if it turns out mediocre you delete it and start over with a bigger sample set. There's a footnote with teeth, too. Classifiers don't work on items that are already encrypted, which means the documents someone already protected by hand are invisible to the thing meant to find them. Read that one twice.

If you'd rather compare purpose-built platforms than stretch what Microsoft ships, we ranked the DLP platforms for 2026 against a scoring model that weights mid-market fit separately from raw capability.

How to classify data without stalling the business

Sequence matters more than tooling here, and it runs opposite to how these projects usually go.

  1. Discover before you write anything. Run a scan and look at the results before drafting a single policy line. Microsoft calls this zero change management in its classifiers overview, and the logic is sound. See what's actually there, then decide what the rules should be. Writing the policy first means writing it against a file estate you imagined.
  2. Pick your levels. Three, per the argument above. Name them in words your staff already use.
  3. Label the top tier only. Not everything. Find the genuinely sensitive material, tag that, and let the rest sit at your default. Companies that set out to label 100% of their data at once are often still labeling eleven months later.
  4. For the shapeless stuff, use document fingerprinting before reaching for a trainable classifier. It recognizes files built from a template, which covers a surprising amount of contract and form traffic without any training data at all.
  5. Turn DLP on in monitor mode. Watch for a few weeks. Tune. Only then start blocking anything.

Step five is where the pain gets avoided. We laid out the sequencing and timeline for rolling out DLP without breaking the business in more detail, including what to do when the first policy catches your CFO.

Four stages of a data classification rollout, each stage more structured than the one before

One current wrinkle. Purview's DLP documentation now names generative AI destinations directly, including ChatGPT, Google Gemini, DeepSeek, and Microsoft Copilot. Copilot is a DLP location in its own right. That changes the urgency of this work, because an assistant with access to your tenant will happily surface whatever a user is technically permitted to see, including the payroll file in the wrong SharePoint folder. Labels are how you draw that boundary, and a tenant with no labels gives Copilot no boundary to respect. It's the same problem we walked through in governing Copilot access.

What the policy has to say

A data classification policy that only defines the levels is half a document. The half that gets used is the handling rules underneath each one.

For every level, write down where it can be stored, how it can be sent, who can share it outside the company, how long it's kept, and how it gets destroyed. Five answers per level. Fifteen answers total if you took the three-tier advice. That's a short document. Short documents get read.

Name an owner. Not a committee. One person who settles the arguable cases and reviews the scheme once a year, because the file estate keeps moving and a policy nobody revisits becomes fiction inside two years.

When you don't need this yet

Some companies should close this tab and go fix something else.

If you're under about 20 users, everything sits in one Microsoft 365 tenant, and you don't handle regulated data or anybody's payment information, formal classification is premature. Your risk is almost certainly somewhere else. Multi-factor authentication on every account, a backup you have actually restored from in the last six months, and admin rights taken off daily-driver laptops will buy you more real protection than a labeling scheme nobody will maintain.

Same answer if you're mid-migration. Classifying a file estate you're about to move, restructure, or largely delete is work you'll do twice. Finish the move, then classify what survived.

It stops being optional the day a customer contract, an insurer, or a framework like CMMC or SOC 2 asks you to produce a data inventory. Now it's on somebody else's deadline. Much worse way to do it.

Where this leaves you

One concrete next step. Pull your Microsoft 365 licensing detail and run a discovery scan against what it already allows. That's about a week of effort, it costs nothing extra on E3, and it tells you whether the rest of this is a six-week project or a six-month one before you commit to either.

Consilien is a managed IT and cybersecurity provider working with companies of 20 to 1000 users across manufacturing, distribution, professional services, and real estate. We run data protection as a program rather than a product, which usually starts with reading your Microsoft licensing before recommending anything new. If you're staring at a DLP tool that won't stop shouting, or a compliance questionnaire asking for a data inventory you don't have, speak to a data protection expert and we'll start with what you already own.

Start With What You Already Own

Pull your Microsoft 365 licensing detail and run a discovery scan against what it already allows. That is about a week of effort, it costs nothing extra on E3, and it tells you whether the rest of this is a six-week project or a six-month one.

Consilien runs data protection as a program rather than a product, for companies of 20 to 1000 users nationwide. If you are staring at a DLP tool that will not stop shouting, or a compliance questionnaire asking for a data inventory you do not have, we will start with what you already have in place.

Common Questions About Data Classification

How many classification levels should we actually have?
Three, for almost every company under 1000 users. Public, Internal, and Confidential covers the decisions people can make in the moment without stopping to think. The four-level model adds a Restricted tier that sounds rigorous and, in practice, fills up with ordinary contracts because nobody's sure where the line is. Add it when you have a specific population of files that genuinely needs separate handling, like controlled unclassified information under a defense contract, and a named person asking for it. Not before.
Can we just let the tool classify everything automatically?
Partly, and the split is predictable. Anything with a recognizable pattern, so Social Security numbers, card numbers, bank details, gets caught well by prebuilt detection. Your roadmap, your pricing model, and your unreleased designs don't have a pattern, and that's usually the material that would actually hurt you. Trainable classifiers close some of that gap if you feed them enough examples, though the custom ones you build yourself are English-only and can't be retrained once published. Automation handles the regulated data. A person still has to point at the rest.
Do we need to classify data we're about to delete?
Backwards, actually. Delete first. Every file you remove is one you never have to inventory, label, protect, or explain to an auditor. Run the cleanup before the discovery scan rather than after it, and the scope you end up classifying is a good deal smaller.
Is data classification required for CMMC, SOC 2, or PCI DSS?
Not always by that name, but the requirement is there under different wording in all three. Each expects you to know what sensitive data you hold and where it lives, whether that's controlled unclassified information for CMMC, whatever you scoped in a SOC 2 report, or cardholder data for PCI DSS. Auditors ask for the inventory and the handling rules. Whether you call the artifact a classification policy is up to you, but you'll be producing one either way.
What does this cost if we're already paying for Microsoft 365?
Manual labeling and scanner-based discovery are already in E3, and manual labels go down as far as Business Premium. So the tooling cost for a first pass is frequently zero. What you're actually spending is time, mostly interviewing the people who own the files and arguing about tiers. The E5 or E5 Compliance upgrade buys automatic labeling, endpoint coverage, and the Content and Activity Explorer screens. Worth pricing, but not worth waiting for.
Realistically, how long does a first pass take?
Plan on six to ten weeks for a company in the 100 to 300 user range, assuming you scope it to the top tier and don't try to label everything. Discovery scanning is quick and largely runs itself. The slow part is human. Getting finance, legal, and operations to agree on what counts as confidential takes as long as it takes, and pushing that conversation faster than it wants to go is how you end up with a scheme nobody follows.

Related Articles

Stay ahead with expert tips, industry trends, and actionable strategies.