Burn-in testing catches a bad drive before it ever holds data. Bit rot scrubbing catches silent corruption on drives that are otherwise working fine. Neither one watches a healthy drive for the slow, specific warning signs that show up months or years into its working life, the gradually climbing error counts that mean a drive is dying, not corrupted or defective out of the box. That’s a separate job: ongoing SMART monitoring with an actual alert attached to it, not a dashboard you remember to check every few months.
Why “it still mounts fine” isn’t good enough
A failing drive doesn’t usually go from perfect to dead in one step. It degrades, and most of that degradation is visible in SMART attributes long before the drive throws a read error your filesystem can’t recover from. The problem is that nothing in a typical setup surfaces those attributes unless you go looking. A drive can sit in an array quietly accumulating reallocated sectors for months, still mounting, still serving reads, still passing a casual glance, right up until the day it doesn’t. By the time a drive fails outright during a resilver, you’ve lost the only window where you could have replaced it on your own schedule instead of under pressure with redundancy already gone.
This is exactly the gap RAID and ZFS redundancy don’t close on their own. Redundancy means an array survives one drive failing. It says nothing about which drive is going to fail next, or when, and finding out after the fact that a second drive in the same rebuild is also marginal is how single failures turn into data loss.
The SMART attributes that actually predict failure
SMART exposes dozens of attributes, and most of them are noise for this purpose. A handful are the ones backed by real failure-rate data (Backblaze’s published drive-stats work is the most widely cited source here) and worth building alerts around:
- Reallocated Sector Count (attribute 5). The drive has already remapped bad sectors to spare ones. A low, flat count that’s been stable for a long time is often tolerable. Any upward movement is the single strongest signal something is actively getting worse.
- Current Pending Sector Count (attribute 197). Sectors waiting to be reallocated, not yet confirmed bad. A nonzero value that clears back to zero after a rewrite is a gray area some people tolerate. One that climbs, or that comes back nonzero across multiple checks, is not.
- Offline Uncorrectable Sector Count (attribute 198). Sectors the drive couldn’t read or write even during its own internal recovery attempts. This correlates strongly with near-term failure and deserves an alert on any nonzero reading, not just a climbing one.
- UDMA CRC Error Count (attribute 199). This one is usually a cabling or connector problem, not the drive itself, loose SATA cable, bad backplane connection, marginal power. Worth checking before assuming the drive is at fault, but still worth fixing since a flaky interface link can itself cause data errors.
- Power-On Hours and Power Cycle Count. Not failure predictors by themselves, but useful context. A drive with low hours showing reallocated sectors is a very different situation than one with five years on the clock showing the same thing.
Raw SMART status (smartctl -H, the simple PASSED/FAILED line) is not enough on its own. Drives routinely report PASSED right up until the moment they fail completely, because that field only trips on a small set of catastrophic thresholds. The attributes above, tracked as trends rather than single snapshots, are where the actual early warning lives.
Setting up smartd for scheduled checks and alerts
smartmontools ships with smartd, a daemon built specifically for this, and it’s already installed alongside smartctl on most distros. The config lives at /etc/smartd.conf, and a reasonable baseline entry per drive looks like:
/dev/sda -a -o on -S on -s (S/../.././02|L/../../6/03) -m [email protected] -M exec /usr/share/smartmontools/smartd-runner
Breaking that down: -a enables monitoring of all attributes, -o on and -S on turn on automatic offline testing and attribute autosave, -s schedules a short self-test every day at 2am and a long self-test every Saturday at 3am, and -m/-M exec route alerts through your system’s mail setup. The schedule syntax is dense, but it’s worth getting right once rather than relying on whatever default smartd ships with, since the long test is the one that actually exercises the full surface.
For /dev/sdX device names specifically: they aren’t stable across reboots on a system with more than a handful of drives, so reference drives by their /dev/disk/by-id/ path in the actual config instead, and only use /dev/sdX for quick manual smartctl checks. This is the same path-stability problem that shows up in HBA card setups and anywhere else drives get addressed directly.
Email alerting only works if mail actually leaves the box. On a typical homelab setup that means either a working local MTA relaying through an external SMTP provider, or swapping -M exec for a webhook script that posts to something you’ll actually see, Discord, ntfy, ping Home Assistant, whatever notification path you’re already watching. An alert that lands in a mailbox nobody opens is functionally the same as no alert.
Scrutiny: a dashboard instead of log-grepping
smartd alerts on thresholds, but it won’t show you a trend, and parsing raw smartctl output across a dozen drives by hand gets old fast. Scrutiny fills that gap: it’s a self-hosted web dashboard, usually run as two Docker containers (a collector that runs on each host with physical drive access, and a web UI/database that can live anywhere), that stores SMART history over time and renders it as an actual graph per attribute per drive.
The practical value is seeing a reallocated sector count that went from zero to four over three months, which is a trend worth acting on, versus one that’s sat at two for two years, which usually isn’t. Scrutiny also ships with its own failure-prediction model built on the same kind of large-fleet failure data that informs which attributes matter in the first place, and it can push its own notifications (webhook, Discord, email) on top of what smartd already gives you, so the two aren’t competing, smartd for the immediate threshold alert, Scrutiny for the long-term trend view and a second opinion.
Running the collector requires the container to have direct access to the physical drives, which on a TrueNAS or Unraid box usually means passing through the actual block devices rather than relying on a VM’s virtualized disk layer, since SMART data doesn’t reliably pass through virtualized storage controllers.
Telling real wear from a real warning
Not every nonzero SMART reading means pull the drive. Normal, expected behavior on an aging-but-healthy drive includes a stable, non-climbing reallocated sector count, power-on hours climbing steadily with no corresponding error growth, and temperature that tracks ambient and chassis airflow rather than running hot in isolation. What separates that from a real warning isn’t the presence of a nonzero number, it’s whether the trend is flat or climbing, which is the entire reason a dashboard that tracks history is worth more here than a single smartctl -a snapshot.
When in doubt, the asymmetry favors caution. A drive you replace a few months early because of a climbing pending-sector count costs you one drive. A drive you leave in place because “it still mounts fine” costs you a rebuild window at best and the array at worst.
Where this fits with everything else
SMART monitoring isn’t a replacement for anything else in a solid data-hoarding setup, it’s the piece that watches the long middle of a drive’s life that burn-in testing and bit rot scrubbing don’t cover. Burn-in catches manufacturing defects before deployment. Scrubs catch silent corruption on drives that report healthy. SMART monitoring catches the drive itself telling you, slowly and quietly, that it’s wearing out, as long as something is actually listening. Set smartd up with a real alert destination, point a Scrutiny dashboard at the fleet if you’re running more than a couple of drives, and check the trend, not just the latest reading, before deciding whether a warning is real.