Updated Aug 12, 2026: added ZFS-vs-traditional-RAID comparison, a rebuild-risk worked example, and a monitoring checklist.

“RAID is not a backup” gets repeated so often it’s lost meaning. Here’s the version that actually helps you pick a redundancy level instead of just feeling vaguely guilty about your array.

What redundancy actually protects against

Redundancy (RAID-Z1, RAID-Z2, mirrors, whatever flavor) protects against drive failure, full stop. It does not protect against:

  • Accidental deletion (redundancy dutifully replicates your mistake)
  • Ransomware or file corruption that writes bad data (same problem)
  • Fire, theft, flooding, or anything that takes out the whole box at once
  • Controller or power supply failure that kills multiple drives simultaneously

That’s what backups are for, and it’s a separate problem from redundancy. This post is specifically about the “how many drives can I lose without losing data” question. If you only take one thing from this post, take the 3-2-1 rule for the data you actually care about: 3 copies, on 2 different types of media, with 1 copy off-site. Redundancy is not one of the three copies - it just keeps your primary copy available while a drive is being replaced.

ZFS vs. traditional RAID

Most of this guide is written in ZFS terms (RAID-Z1/Z2/Z3) because it’s the default for serious home data-hoarding setups now, but the same risk logic applies to hardware/software RAID (RAID 5, RAID 6, RAID 10) too. The practical differences worth knowing:

  • ZFS checksums every block and can detect (and, with redundancy, correct) silent data corruption that traditional RAID has no way to notice - a failing drive can quietly return bad data that traditional RAID happily writes as “correct” because it has no checksum layer to catch it.
  • ZFS scrubs proactively read every block and verify checksums against redundant copies, catching problems before you’d otherwise discover them (usually when it’s too late).
  • Traditional RAID 5/6 maps roughly to RAID-Z1/Z2 in redundancy terms, but without the checksumming - if data integrity matters as much as uptime, that’s a real gap.
  • Hardware RAID controllers add a new single point of failure of their own (a dead controller can make an array unreadable on different hardware) - one more reason ZFS’s software-defined approach has become the homelab default.

The real risk: rebuild time

The reason “just run RAID-Z1 with one parity drive” advice has gotten worse over time isn’t the math - it’s drive size. Rebuilding a failed 4TB drive from parity used to take hours. Rebuilding a failed 20TB drive can take the better part of a day or more depending on array load and drive speed, and during that entire window, your array has zero redundancy left. If a second drive fails - and rebuild stress is exactly when a second aging drive tends to fail, since a rebuild hammers every remaining disk with sustained read load - you lose the pool.

A worked example: picture a 6-drive RAID-Z1 array of 16TB drives. A single failed drive means the other 5 have to be read in full to reconstruct it - that’s a sustained, hours-to-a-day-plus operation depending on your array’s real-world throughput, all while the array sits at zero spare redundancy. Compare that to RAID-Z2 on the same array: a single drive failure during that same rebuild window still leaves you with one more parity drive of protection. That gap is the entire argument for Z2 on larger, higher-capacity arrays.

This is why RAID-Z1 (single parity) has fallen out of favor for anything beyond small arrays of smaller drives, and why RAID-Z2 (dual parity) has become the practical default for larger pools.

A sizing framework, not a rule of thumb

Instead of “always run Z2,” work through this:

  1. How many drives are in the vdev? More drives means more chances for a second failure during a rebuild. A 4-drive vdev is a different risk profile than a 12-drive vdev.
  2. How big are the drives? Bigger drives mean longer rebuilds mean longer exposure windows. Scale your parity up as your drive size goes up.
  3. How replaceable is this data? A media library you could re-rip is a different risk tolerance than family photos or financial records that exist in exactly one place.
  4. Do you actually have a backup for the important subset? If yes, you can run leaner redundancy on the pool since a rebuild failure isn’t catastrophic - it’s an inconvenience you recover from with the backup.

Practical starting points

  • Small pool (4-6 drives, ≤8TB each), backed-up critical data: RAID-Z1 is defensible. Watch your drive ages and replace proactively.
  • Larger pool (8+ drives) or drives above 12TB: RAID-Z2. The rebuild window is long enough that single parity is gambling.
  • Anything truly irreplaceable: mirrors (RAID-10 equivalent) or RAID-Z2 and an off-site backup. Redundancy buys you uptime during a drive failure; it does not buy you disaster recovery.
  • Very large pools (16+ drives): consider RAID-Z3 (triple parity) or splitting into multiple smaller vdevs rather than one giant one - a single huge vdev means a single huge rebuild.

Monitoring: the habit that matters more than any RAID level

Monitor drive health and replace drives showing early warning signs before they fail outright, not after. A proactive swap on a healthy array is a non-event. A failure during a rebuild is how pools actually get lost. A minimal checklist:

  • Run regular ZFS scrubs (monthly is a common baseline) to catch silent corruption and confirm redundancy is actually intact, not just assumed to be.
  • Watch SMART data on every drive - reallocated sectors, pending sectors, and rising temperatures are all early warnings worth acting on rather than dismissing.
  • Check zpool status for a clean pool state as a matter of routine, not only when something already seems wrong.
  • Track drive age and replace proactively on older drives showing any warning signs, especially ahead of a known-risky period like a heat wave or a planned hardware move.

Redundancy buys you a safety margin - monitoring is what tells you when you’re about to need it, and it’s the difference between a routine drive swap and a 2am scramble.