Updated Aug 12, 2026: added ZFS-vs-traditional-RAID comparison, a rebuild-risk worked example, and a monitoring checklist.
“RAID is not a backup” gets repeated so often it’s lost meaning. Here’s the version that actually helps you pick a redundancy level instead of just feeling vaguely guilty about your array.
What redundancy actually protects against
Redundancy (RAID-Z1, RAID-Z2, mirrors, whatever flavor) protects against drive failure, full stop. It does not protect against:
- Accidental deletion (redundancy dutifully replicates your mistake)
- Ransomware or file corruption that writes bad data (same problem)
- Fire, theft, flooding, or anything that takes out the whole box at once
- Controller or power supply failure that kills multiple drives simultaneously
That’s what backups are for, and it’s a separate problem from redundancy. This post is specifically about the “how many drives can I lose without losing data” question. If you only take one thing from this post, take the 3-2-1 rule for the data you actually care about: 3 copies, on 2 different types of media, with 1 copy off-site. Redundancy is not one of the three copies - it just keeps your primary copy available while a drive is being replaced.
ZFS vs. traditional RAID
Most of this guide is written in ZFS terms (RAID-Z1/Z2/Z3) because it’s the default for serious home data-hoarding setups now, but the same risk logic applies to hardware/software RAID (RAID 5, RAID 6, RAID 10) too. The practical differences worth knowing:
- ZFS checksums every block and can detect (and, with redundancy, correct) silent data corruption that traditional RAID has no way to notice - a failing drive can quietly return bad data that traditional RAID happily writes as “correct” because it has no checksum layer to catch it.
- ZFS scrubs proactively read every block and verify checksums against redundant copies, catching problems before you’d otherwise discover them (usually when it’s too late).
- Traditional RAID 5/6 maps roughly to RAID-Z1/Z2 in redundancy terms, but without the checksumming - if data integrity matters as much as uptime, that’s a real gap.
- Hardware RAID controllers add a new single point of failure of their own (a dead controller can make an array unreadable on different hardware) - one more reason ZFS’s software-defined approach has become the homelab default.
The real risk: rebuild time
The reason “just run RAID-Z1 with one parity drive” advice has gotten worse over time isn’t the math - it’s drive size. Rebuilding a failed 4TB drive from parity used to take hours. Rebuilding a failed 20TB drive can take the better part of a day or more depending on array load and drive speed, and during that entire window, your array has zero redundancy left. If a second drive fails - and rebuild stress is exactly when a second aging drive tends to fail, since a rebuild hammers every remaining disk with sustained read load - you lose the pool.
A worked example: picture a 6-drive RAID-Z1 array of 16TB drives. A single failed drive means the other 5 have to be read in full to reconstruct it - that’s a sustained, hours-to-a-day-plus operation depending on your array’s real-world throughput, all while the array sits at zero spare redundancy. Compare that to RAID-Z2 on the same array: a single drive failure during that same rebuild window still leaves you with one more parity drive of protection. That gap is the entire argument for Z2 on larger, higher-capacity arrays.
This is why RAID-Z1 (single parity) has fallen out of favor for anything beyond small arrays of smaller drives, and why RAID-Z2 (dual parity) has become the practical default for larger pools.
A sizing framework, not a rule of thumb
Instead of “always run Z2,” work through this:
- How many drives are in the vdev? More drives means more chances for a second failure during a rebuild. A 4-drive vdev is a different risk profile than a 12-drive vdev.
- How big are the drives? Bigger drives mean longer rebuilds mean longer exposure windows. Scale your parity up as your drive size goes up.
- How replaceable is this data? A media library you could re-rip is a different risk tolerance than family photos or financial records that exist in exactly one place.
- Do you actually have a backup for the important subset? If yes, you can run leaner redundancy on the pool since a rebuild failure isn’t catastrophic - it’s an inconvenience you recover from with the backup.
Practical starting points
- Small pool (4-6 drives, ≤8TB each), backed-up critical data: RAID-Z1 is defensible. Watch your drive ages and replace proactively.
- Larger pool (8+ drives) or drives above 12TB: RAID-Z2. The rebuild window is long enough that single parity is gambling.
- Anything truly irreplaceable: mirrors (RAID-10 equivalent) or RAID-Z2 and an off-site backup. Redundancy buys you uptime during a drive failure; it does not buy you disaster recovery.
- Very large pools (16+ drives): consider RAID-Z3 (triple parity) or splitting into multiple smaller vdevs rather than one giant one - a single huge vdev means a single huge rebuild.
Monitoring: the habit that matters more than any RAID level
Monitor drive health and replace drives showing early warning signs before they fail outright, not after. A proactive swap on a healthy array is a non-event. A failure during a rebuild is how pools actually get lost. A minimal checklist:
- Run regular ZFS scrubs (monthly is a common baseline) to catch silent corruption and confirm redundancy is actually intact, not just assumed to be.
- Watch SMART data on every drive - reallocated sectors, pending sectors, and rising temperatures are all early warnings worth acting on rather than dismissing.
- Check
zpool statusfor a clean pool state as a matter of routine, not only when something already seems wrong. - Track drive age and replace proactively on older drives showing any warning signs, especially ahead of a known-risky period like a heat wave or a planned hardware move.
Redundancy buys you a safety margin - monitoring is what tells you when you’re about to need it, and it’s the difference between a routine drive swap and a 2am scramble.