Data classification is the practice of sorting an organization's data into a few levels by how much harm its exposure, alteration or loss would cause, then tying each level to handling rules.
Three or four levels are enough for most businesses, typically Public, Internal, Confidential and Restricted. What protects data is the rule attached to each level: who may access it, where it may be stored, whether it must be encrypted, and how it is shared and destroyed.
- Classify by the harm a leak or loss would cause, not by department
- Three or four levels. More levels means more mislabeled files
- A label with no handling rule attached changes nothing
- Default everything unlabeled to Internal, not Public
- Start with the data you know is sensitive, not a full inventory
On this page
- What data classification is, and what it is for
- How many data classification levels you need
- An example classification scheme with handling rules
- A label with no rule attached changes nothing
- Types of data classification
- How to start a data classification program
- Where classification meets compliance and risk
- Finding the data before you can label it
- Where people go wrong
- Comparison
- FAQ
The ideaWhat data classification is, and what it is for
Data classification is the process of organizing data and information into categories based on its sensitivity and value, so that each category gets security controls that fit it. A price list on the website and a spreadsheet of employee Social Security numbers should not be stored, shared or backed up the same way.
Classification is how an organization writes that difference down, across every system where its data lives: file shares, email, cloud storage, business applications and backups.
The point is proportion. Protecting everything to the highest standard is expensive and makes work painful, so people route around it. Protecting everything to the lowest standard leaves sensitive information exposed. A data classification scheme lets security teams spend effort where the risk is, and lets everyone else work normally.
It also answers questions that come up anyway: which files may go to a personal email address, which data may be pasted into an AI tool, which folders need encryption, what must be reported if a laptop is stolen.
Without classification, each of those questions is decided case by case, usually by the person in a hurry. With it, data security and compliance decisions follow a rule the organization agreed on in advance.
LevelsHow many data classification levels you need
Three or four, and there are good reasons not to go further.
Every extra level is another judgment call for the person saving a file. With seven levels, people cannot remember the difference between level four and level five, so they pick one at random or pick the highest to be safe. Both outcomes defeat the purpose. With three or four, the choice is obvious.
The US federal model makes the same point. NIST's FIPS 199 defines three levels of potential impact from a loss of confidentiality, integrity or availability: low, moderate and high.
In its wording, the impact is low when the loss could be expected to have a limited adverse effect on organizations or individuals, moderate when the effect is serious, and high when it is severe or catastrophic.
For a business, a workable set of data classification levels is:
Public. Approved for anyone. Disclosure causes no harm.
Internal. Normal business data. Disclosure is embarrassing or mildly harmful, not damaging.
Confidential. Sensitive information whose exposure would harm the business, its customers or its staff.
Restricted. The most sensitive data, whose exposure would cause serious legal, financial or personal harm, often information a law or contract regulates.
A company with little regulated data can merge Confidential and Restricted into one level and run three.
Example schemeAn example classification scheme with handling rules
The table below is an example, not a standard. It shows the shape a data classification policy should take: every level paired with rules a person can follow. Adapt the names, examples and rules to your own data and obligations.
| Level | Data classification examples | Access | Storage and encryption | Sharing | Disposal |
|---|---|---|---|---|---|
| Public | Website text, published prices, press releases, job ads | Anyone | No requirement | Free to share | No requirement |
| Internal | Internal procedures, org chart, most email, meeting notes | All staff | Company systems only | Staff and trusted partners under agreement | Normal deletion |
| Confidential | Contracts, customer lists, financials, salaries, source code | Named groups on a need to know basis | Company systems, encrypted devices | Named external recipients only, through company tools | Secure deletion, shredding |
| Restricted | Card numbers, health records, SSNs, passwords, bank details | Named individuals, with MFA | Approved systems only, encrypted at rest and in transit | Only where a process requires it, logged | Secure deletion with a record |
Two design choices in that example matter more than the level names. Access is defined in terms your systems already enforce, such as groups and MFA. And each rule is something a person can check without asking IT.
Handling rulesA label with no rule attached changes nothing
The most common failure of data classification is a scheme that exists as a document and a set of labels only. Files get marked Confidential and still sit in a folder shared with everyone, still get emailed to personal addresses, still land on unencrypted USB sticks. The label described the data. Nothing enforced it.
Each level needs handling rules, and where possible, a technical control that applies that protection automatically.
Access. Tie levels to permission groups. Confidential folders are shared with named groups, never with everyone in the organization. Restricted data sits behind MFA and, where available, conditional access rules.
Storage. Say where each level may live. Restricted data only in named systems, not in personal OneDrive or SharePoint sites created ad hoc, not on local desktops.
Encryption. Confidential and Restricted data on laptops and removable media is encrypted, and Restricted data is encrypted in transit. The page on encryption algorithms covers what that means in practice.
Sharing. Who outside the company may receive each level, and through which tool.
Retention and disposal. How long each level is kept, and how it is destroyed.
Backup. Restricted and Confidential data are in the backups, and the backups are protected as carefully as the source. A backup copy is still Restricted data, which the 3-2-1 backup rule page touches on.
Tools can apply labels and rules automatically. Microsoft describes its Purview sensitivity labels as letting you classify and protect your organization's data, and a label can apply encryption and content markings such as a Confidential watermark.
Microsoft also says Copilot and agents recognize sensitivity labels, which makes labeling more relevant as AI tools read company files and cloud data.
TypesTypes of data classification
The types of data classification describe how a label gets decided, and most programs combine all three.
Content based. The data itself is inspected for patterns: card numbers, Social Security numbers, bank account formats, keywords. This is what data loss prevention tools and automatic labeling do. It is good at finding regulated identifiers and poor at judging a contract or a strategy memo.
Context based. The label follows from where the data lives or who created it. Everything in the payroll system is Restricted. Everything in the finance share is at least Confidential. This is cheap and covers a lot with one decision.
User based. The person creating or handling the data chooses the label. It captures judgment no tool has, and it depends on training and on a scheme simple enough to remember.
Sensitive data classification in a small business usually starts context based, because a few systems hold most of the regulated data. Content inspection finds the copies that escaped those systems, and user labeling covers the documents only a person can judge.
Getting startedHow to start a data classification program
A full data inventory before anything else is how classification projects stall. Start where the risk is, and let compliance obligations point to the first data sets.
1. Write the levels and handling rules first, on one page, like the example above. 2. List the most sensitive information you already know about. Payroll, HR records, customer payment data, health information if you handle it, passwords and keys.
Name the system each one lives in and its business owner. 3. Lock those systems down to the Restricted rules: access groups, MFA, encryption, logging. 4. Set a default. Anything unlabeled is Internal. This one rule protects more data than any other. 5.
Find the copies. Exports, spreadsheets and email attachments that left the source system. Content based discovery helps here. 6. Train people on the levels with examples from your own business, not generic ones. 7. Review yearly and whenever a new system or data type arrives.
Data classification fits the wider zero trust approach many organizations follow: access to data is granted by need, and the level of the data sets how much verification and protection that access requires.
ComplianceWhere classification meets compliance and risk
Most organizations start classifying because something outside them requires it, and the requirement usually names a category of data rather than a label.
Regulated categories come with their own rules. Cardholder data falls under PCI DSS, protected health information under HIPAA, and personal data of people in the EU under the GDPR.
Each regime sets its own handling and breach obligations, and a classification scheme earns its keep by making the affected data findable. The regime decides the rule. Your label decides which files the rule reaches.
Contracts and frameworks ask the same question. A customer security questionnaire, a SOC 2 audit or an ISO 27001 certification asks how information is classified and protected. An answer that points at a scheme in use, with handling rules and an owner, is short. An answer that points at a policy nobody applies takes far longer to defend.
Risk is what the levels actually rank. The point of the top level is not secrecy for its own sake. It is that unauthorized disclosure of that data would do serious damage: a regulatory penalty, a lost contract, a safety issue.
Classification is a way of writing down where the risk sits, so security controls and budget follow the data that matters instead of spreading evenly.
Governance is the part that keeps it alive. Someone owns the scheme, reviews it, and decides the label for new kinds of data. Without that, the scheme is accurate on the day it is written and drifts from then on.
DiscoveryFinding the data before you can label it
Classification assumes you know what you hold, which for most organizations is the hard part. Data sits in the file server, in cloud storage, in mailboxes, in the line of business application, in backups, in a spreadsheet on a laptop and in whatever a former employee saved to a personal drive.
Start with a data inventory. List the systems that hold data, what kind of data each holds, who owns it and where it is backed up. This is less work than it sounds when the list is systems rather than files, and it is the input every later step needs.
Let discovery tools do the searching. Cloud platforms can scan storage for patterns such as card numbers and national identifiers, and report where sensitive data is concentrated. Vendors market this as data security posture management, or DSPM. The value for a smaller organization is the first scan, which usually finds data somewhere nobody expected.
Label at the source, then let tools inherit it. A label applied in the platform where the data is created travels with the file, and data loss prevention rules can act on it. Labeling copies after the fact never finishes.
Deal with what you find. Discovery tends to turn up old exports and test copies of production data. Deleting those is the cheapest security control in this whole process.
PitfallsWhere people go wrong
Too many levels. Five or more levels produce inconsistent labels and a habit of marking everything at the top.
Labels without handling rules. Covered above, and still the most common failure.
Classifying by department. Finance data is not all Restricted, and marketing data is not all Public. The level follows the data, not the team.
Defaulting to Public. Unlabeled data should be treated as Internal until someone decides otherwise.
Forgetting copies. Exports, backups, test databases and email attachments carry the same level as the source.
Treating it as an IT project. IT and security teams enforce the rules, but the business owner of the data decides its level, because they understand the risk. Without owners, nobody can answer whether a given file is Confidential.
Never reviewing. New systems, new cloud services, new AI tools and new customer contracts change what counts as sensitive data, and what compliance requires.
ComparisonFour data classification levels and the controls each one gets
| Criterion | Public | Internal | Confidential | Restricted |
|---|---|---|---|---|
| Harm if disclosed | None | Minor | Significant | Serious legal, financial or personal |
| Who may access | Anyone | All staff | Named groups | Named individuals |
| Encryption required | No | No | On devices and media | At rest and in transit |
| MFA required | No | Standard staff MFA | Standard staff MFA | Yes, enforced for the system |
| External sharing | Free | Under agreement | Named recipients only | Only by defined process |
| Logging of access | No | No | Recommended | Required |
| Share of a typical company's data | Small | Most | Some | Smallest |
The last row is the reason the scheme works. Most data is Internal and needs nothing special. The expensive security controls, enforced MFA, logging, encryption in transit, apply only to the small Restricted set, which is why it pays to keep that set small and well defined.
FAQFrequently asked questions
What is data classification?
Data classification is the process of sorting data into levels by how much harm its exposure, alteration or loss would cause, then applying handling rules to each level. It lets an organization protect sensitive data strongly without burdening everyday work with the same controls.
What are the data classification levels?
Most schemes use three or four: Public, Internal, Confidential and Restricted. Governments use their own sets, and NIST FIPS 199 uses low, moderate and high impact. The names matter less than having few levels, each with clear handling rules.
What are the types of data classification?
Content based classification inspects the data for patterns such as card numbers. Context based classification uses where the data lives or who created it. User based classification lets the person handling the data choose the label. Most programs combine all three.
What are some data classification examples?
Public: website content and press releases. Internal: procedures and most email. Confidential: contracts, customer lists and financial reports. Restricted: payment card numbers, health records, Social Security numbers and passwords. The right level depends on the harm exposure would cause.
What should a data classification policy include?
The levels and their definitions, examples of data at each level, who owns classification decisions, handling rules for access, storage, encryption, sharing, retention and disposal, the default for unlabeled data, and how often the policy is reviewed.
What is sensitive data classification?
It is the part of classification that identifies data whose exposure causes real harm, such as personal, financial, health and authentication data, so it can be placed in the Confidential or Restricted levels and protected with stricter controls.
How many classification levels should a small business use?
Three or four. Fewer levels produce more consistent labels. A business with little regulated data can run Public, Internal and Confidential, and add a Restricted level if it handles payment, health or identity data.
Who is responsible for classifying data?
The business owner of the data decides its level, because they understand its value and the harm of exposure. IT and security teams set up the systems that enforce the handling rules, and every employee applies labels to what they create.
What is the difference between data classification and data labeling?
Classification is the decision about which level data belongs to. Labeling is recording that decision on the file, email or record, often as metadata or a visible marking, so people and tools can apply the matching handling rules.
Can data classification be automated?
Partly. Tools can detect patterns such as card and Social Security numbers, and apply labels based on location. Contracts, plans and other documents that need judgment still depend on people choosing the right label.
What is the default classification for unlabeled data?
Internal is the safest practical default. It keeps unlabeled data inside the company without imposing the heavier controls of Confidential or Restricted, and it avoids the risk of treating unknown data as Public.
How often should data classification be reviewed?
At least once a year, and whenever the business adopts a new system, a new AI tool, a new type of customer data or a contract with data handling terms. Levels for individual data sets should be reviewed by their owners.
Is data classification the same as data privacy?
No. Privacy is about how personal information may be collected and used, and data classification is one of the security controls that supports it. Classifying which stores hold personal data is what makes a privacy obligation enforceable in practice.
Who owns data classification governance?
A named business owner sets the levels and approves changes, with security and IT supplying the tooling. Governance means somebody reviews the scheme, resolves disputes over a label and decides how new kinds of information are classified.
Keep readingRelated concepts
Read next · Identity and access Zero Trust Explained The access model that data classification feeds: verify every request, and scale checks to what is at stake. Open this next13 min- Cryptography · 11 min Encryption Algorithms, and the Two Families They Fall Into What encryption at rest and in transit actually uses, for the levels that require it.
- Platforms · 11 min OneDrive vs SharePoint, and What Happens When Someone Leaves Where company files should live in Microsoft 365, which decides how well a storage rule can be enforced.