In this guide
- What is data storage?
- Why data storage matters
- How data storage works
- Types of data storage
- Storage devices: HDD, SSD and tape
- DAS, NAS and SAN: how storage is attached
- Cloud storage
- File, block and object storage
- RAID basics: surviving a drive failure
- Storage performance: IOPS, throughput and latency
- Backup and recovery basics
- Data storage security
- How to choose a data storage solution
- Data storage solutions for a business
- Storage terms worth knowing
- FAQ
Data storage basics, from a single drive to shared storage
Data storage basics answer four questions about any piece of data: what it is stored on, how a computer reaches it, how it survives a failure, and how it comes back after a disaster. A laptop drive, a NAS in an office closet and a cloud bucket are all answers to the same four questions, at different scales.
This guide covers the storage fundamentals in the order they build on each other: what data storage is, the types of storage, the devices, direct-attached versus network storage, file, block and object storage, RAID, performance, backup, and how to choose. Each section links to a full page in this library.
DefinitionWhat is data storage?
Data storage is the recording of digital information on a medium so that it can be read back later. Everything is stored as bits, grouped into bytes, and capacity is counted in multiples of bytes: a gigabyte is a billion bytes, a terabyte is a thousand gigabytes.
Drive makers count in powers of ten and operating systems often count in powers of two, which is why a 1 TB drive shows up as about 931 GB.
The first distinction is between memory and storage:
- Primary storage is memory the processor works in directly: RAM and cache. It is very fast and volatile, which means it loses its contents when the power goes off.
- Secondary storage holds data permanently: hard drives, solid-state drives and tape. It is slower, far larger and non-volatile.
When people say storage, they almost always mean secondary storage, and that is what the rest of this guide covers.
StakesWhy data storage matters
An organization's data is usually worth more than the devices it is stored on. Customer files, accounting records, email and designs cannot be bought again, so the decisions about where to store information, who can access it and how many copies exist are business decisions as much as technical ones.
Good storage keeps data available to the people who need it, keeps it from everyone else, and keeps the cost of both under control.
MechanicsHow data storage works
Every storage device does the same two things: it writes data when an application saves, and it reads data back when an application asks. What differs is how the device finds the right location, and how long that takes.
- The application asks the operating system to store or open a file.
- The file system translates the file name into the locations on the device where the data lives.
- The controller sends the request over an interface: SATA or SAS for most drives, NVMe for fast flash, or a network for shared storage.
- The device reads or writes the data. A hard disk moves a head across a spinning platter to reach it, which takes milliseconds; flash storage addresses it electronically, which takes microseconds.
That difference in access time is the reason the types of storage devices behave so differently under load, and the reason the same data is often kept on more than one type: fast storage for what is in use, cheaper storage for what is kept.
OverviewTypes of data storage
Storage is classified in three independent ways, and most of the confusion comes from mixing them up. A single system has an answer to each.
| Question | Options | Decides |
|---|---|---|
| What medium? | Hard disk, solid-state, tape, optical | Speed, cost per terabyte, lifespan |
| How is it attached? | Direct-attached (DAS), network-attached (NAS), storage area network (SAN), cloud | Who can reach it, and how it is shared |
| How is data addressed? | File, block, object | What kind of application it suits |
MediaStorage devices: HDD, SSD and tape
The choice of medium is a trade between speed and cost per terabyte.
| Device | How it stores data | Strength | Best for |
|---|---|---|---|
| Hard disk drive (HDD) | Magnetic platters that spin under a moving head | Lowest cost per terabyte of any online medium | Bulk capacity, backup targets, archives |
| Solid-state drive (SSD), SATA | Flash memory, no moving parts | Far faster than a hard disk, silent, shock resistant | Operating systems and everyday workloads |
| NVMe SSD | Flash memory attached over PCIe instead of SATA | The fastest option, with very low latency | Databases, virtual machines, anything that waits on disk |
| Tape (LTO) | Magnetic tape in removable cartridges | Cheap, durable, and offline by nature | Long-term archive and ransomware-proof backup copies |
Smaller storage devices fill the gaps: external USB drives for a simple local backup, USB flash drives and SD cards for moving files between devices, and optical discs, now rare, for distribution and long-term archives. They are convenient and easy to lose, so anything sensitive stored on a portable device should be encrypted.
The SSD vs HDD page compares the two in detail, including how each one fails. How a drive is divided and booted is a separate question, covered in MBR vs GPT: MBR is limited to 2 TB and four primary partitions, which is why modern systems use GPT.
On top of the partitions sits a file system, which tracks where each file lives: NTFS on Windows, ext4 and XFS on Linux, APFS on a Mac, and ZFS where data integrity matters most.
AttachmentDAS, NAS and SAN: how storage is attached
How storage connects to the computers that use it decides who can share it.
| Type | What it is | Shared? | Typical use |
|---|---|---|---|
| DAS, direct-attached storage | Drives inside or plugged directly into one computer | No, one machine only | A workstation, a single server, an external backup drive |
| NAS, network-attached storage | A storage appliance that shares folders over the ordinary network | Yes, as files | Office file shares, home media, backup targets |
| SAN, storage area network | A dedicated network that presents raw disks to servers | Yes, as blocks | Virtualization clusters and databases |
| Cloud storage | Capacity rented from a provider and reached over the internet | Yes, from anywhere | Off-site backup, archives, application data |
A NAS speaks file protocols, SMB for Windows and NFS for Linux, so a user sees a shared folder. A SAN speaks block protocols, iSCSI or Fibre Channel, so a server sees what looks like a local disk. NAS vs SAN explains when each is the right answer, and SMB port 445 covers the protocol most office file sharing runs on.
CloudCloud storage
Cloud storage is capacity in a provider's data centers, rented by the gigabyte and reached over the internet. Users access files from any device, the provider replaces failed hardware, and capacity grows without a purchase order. It comes in three deployment models:
- Public cloud: shared infrastructure from a provider such as AWS, Microsoft Azure or Google Cloud. The lowest cost to start and the easiest to scale.
- Private cloud: the same technology on infrastructure dedicated to one organization, for data that must stay under its control.
- Hybrid cloud: a mix, typically keeping active files on local storage and sending backups and archives to the cloud.
Two things surprise people. The first is cost: storing data in the cloud is cheap, but reading it back out can carry a retrieval or egress charge, and the cheapest archive tiers take hours to return a file.
The second is responsibility: the provider secures the platform, while access, permissions, sharing links and a second copy remain the customer's job. File sync services such as OneDrive are storage, not backup; OneDrive vs SharePoint explains what each one is for.
AccessFile, block and object storage
The third classification is how data is addressed, and it determines which applications a storage system suits.
- File storage organizes data as files in folders, with a path to each one. It is what people expect, and it is what a NAS provides. It struggles when there are billions of files.
- Block storage splits data into fixed-size blocks with no structure of their own; the server's file system gives them meaning. It is the fastest form, so it backs databases and virtual machine disks, and it is what a SAN and cloud volumes provide.
- Object storage keeps each item as an object with its data, its metadata and a unique ID in one flat pool, reached over HTTP. It scales almost without limit and is cheap, so it holds backups, media and archives. Amazon S3 is the familiar example.
RedundancyRAID basics: surviving a drive failure
RAID combines several drives so that the failure of one does not lose data or stop the system. The common levels trade capacity against protection and speed:
| Level | How it works | Minimum drives | Survives | Usable capacity |
|---|---|---|---|---|
| RAID 0 | Striping, data split across drives | 2 | Nothing; one failure loses everything | All of it |
| RAID 1 | Mirroring, every drive holds a full copy | 2 | One drive | Half |
| RAID 5 | Striping with one drive's worth of parity | 3 | One drive | All but one drive |
| RAID 6 | Striping with two drives' worth of parity | 4 | Two drives | All but two drives |
| RAID 10 | Mirrored pairs, striped together | 4 | One drive in each pair | Half |
RAID 5 vs RAID 10 covers the choice most people actually face, including why rebuild times on large drives have pushed many toward RAID 6 or RAID 10.
The rule that matters most: RAID is not a backup. It protects against a failed drive. It does nothing about a deleted file, a corrupted database, ransomware, a fire or a mistake, because every one of those is faithfully copied to all the drives at once.
PerformanceStorage performance: IOPS, throughput and latency
Three numbers describe how fast storage is, and they measure different things:
- IOPS: input/output operations per second, the count of small reads and writes a device can handle. It matters for databases and virtual machines.
- Throughput: the volume of data moved per second, in megabytes. It matters for backups, video and large file copies.
- Latency: the delay before a single operation completes. It is what a user feels as responsiveness.
A hard disk has respectable throughput and poor IOPS, because the head must physically move for each random read. An SSD improves IOPS and latency by orders of magnitude, which is why moving a slow server to flash often fixes it outright.
ProtectionBackup and recovery basics
A backup is a separate copy of data, kept somewhere the original's problems cannot reach, from which the data can be restored. The standard to aim for is the 3-2-1 backup rule: three copies of the data, on two different media, with one copy off-site.
- Full, incremental and differential: a full backup copies everything; an incremental copies what changed since the last backup of any kind; a differential copies what changed since the last full.
- Snapshots are not backups when they live on the same system as the data. They are a fast way to undo a change, and they die with the array.
- RPO and RTO: the recovery point objective is how much data you can afford to lose; the recovery time objective is how long you can afford to be down. Those two numbers decide the backup design and its cost.
- Immutability: ransomware looks for backups first. A copy that cannot be altered or deleted for a set period, described on the immutable backups page, is the defense.
- Test the restore. A backup that has never been restored is an assumption.
The wider planning, including failover and the order in which systems come back, is covered in backup and disaster recovery.
SecurityData storage security
Stored data needs protecting from the people who should not read it as much as from hardware failure.
- Encryption at rest makes a stolen drive or a lost laptop unreadable. Most NAS devices, cloud services and operating systems offer it; it only has to be turned on.
- Access control limits each share and folder to the groups that need it, with permissions given to groups and not to individuals.
- Secure disposal matters at the end of a drive's life. Deleting files does not remove the data; a drive should be securely erased or physically destroyed before it leaves the building.
- Audit logs record who accessed or deleted what, which is often a compliance requirement for regulated information.
DecisionsHow to choose a data storage solution
Work through these in order and the answer usually falls out:
- How much data, and how fast is it growing? Size for three years, not for today.
- Who needs to reach it? One machine means DAS; an office means a NAS; a virtualization cluster means a SAN or shared block storage.
- How fast must it be? Databases and virtual machines need flash; archives and backups do not.
- What does losing it cost? That sets the RAID level, the backup schedule and whether an off-site copy is optional. It is not.
- How sensitive is it? Sensitive data needs encryption at rest, such as BitLocker or FileVault on endpoints, and access control on shares.
- Who will run it? Cloud storage moves the hardware work to a provider in exchange for a monthly bill and a dependency on the internet link.
In practiceData storage solutions for a business
Three setups cover most small and midsize organizations, and each is a combination of the types of storage described above:
- A small office: files in a cloud storage service or on a small NAS with mirrored drives, backed up nightly to a second device and to the cloud. No servers to look after.
- A growing office: a NAS or file server with RAID 6 for shared data, flash storage for the applications that need speed, and backup software that keeps a local copy for fast restores and an off-site copy for disasters.
- A virtualized environment: shared block storage, a SAN or a hyperconverged cluster, so that virtual machines can move between hosts, with snapshots for quick rollback and an immutable backup copy outside the cluster. The virtualization library covers the compute side.
The right data storage solution is the simplest one that meets the capacity, speed and recovery targets, because every extra device is something else to patch, monitor and secure.
GlossaryStorage terms worth knowing
The data storage basics above lean on a small vocabulary. These are the terms that come up in every storage conversation:
- Volume: a usable storage area with a file system on it, which may span part of a drive or many drives.
- LUN: a logical unit number, the name for a block device a SAN presents to a server.
- Thin provisioning: promising more capacity than physically exists and allocating it only as data is written.
- Deduplication: storing identical blocks once, which shrinks backups considerably.
- Hot, cool and archive tiers: cloud storage classes priced by how often the data is read; the colder the tier, the cheaper to keep and the dearer to retrieve.
- Hot spare: an idle drive in an array that takes over automatically when another fails.
- Bit rot: silent corruption of stored data over time, which checksumming file systems such as ZFS detect and repair.
