Every drive you own is going to fail eventually. The part that catches people off guard is when. Hard drive failure rates don’t sit flat over a drive’s life, they follow what’s usually called a bathtub curve: a spike of early failures right out of the gate, a long flat stretch of low failure rate through the drive’s useful life, and a rising tail of wear-related failures years later. Burn-in testing exists to deal with that first spike, on your terms, before the drive is holding data you care about.

Why infant mortality is a real thing, not just folklore

Manufacturing a hard drive means assembling a sealed unit with platters spinning at thousands of RPM, read/write heads flying nanometers off the platter surface, and firmware that has to get spin-up, seek, and thermal behavior right from the first power-on. Most manufacturing defects that are going to cause a problem show up fast, within the first days to weeks of real use, not months later. A drive with a marginal head, a slightly out-of-spec bearing, or firmware that mishandles a specific power state tends to fail early or not at all.

This isn’t unique to hard drives, it’s the same reliability curve that shows up across most electromechanical hardware, and it’s exactly why enterprise buyers and cloud storage operators don’t just unbox a drive and put it straight into production. They run it through a deliberate stress period first, because catching a bad drive during a controlled test costs a return shipping label. Catching the same bad drive three weeks into a rebuild, after it’s already resilvering into your array, can cost you the rebuild window and put a second drive’s SMR vs CMR read load right when you can least afford a second failure.

What burn-in actually catches

Burn-in won’t tell you anything about a drive’s long-term wear characteristics, that’s a different question entirely and not something you can compress into a few days of testing. What it’s good at is surfacing defects that are already present but haven’t announced themselves yet:

  • Bad sectors present from the factory. Every drive ships with some reallocated sectors already mapped out at the factory, that’s normal. What you’re checking for is sectors that are marginal right now and likely to fail under real write load.
  • Marginal heads or platters that produce read errors under sustained access patterns but might pass a quick spot-check.
  • Firmware and power-state bugs that only show up after enough spin-up/spin-down cycles or enough time at operating temperature.
  • Thermal problems specific to your enclosure, not the drive in isolation. A drive that runs fine on a bench can run hot enough to matter once it’s packed into a dense multi-bay chassis with six neighbors also generating heat, which is one of the only ways you’ll catch an airflow problem before it’s baked into your build.
  • Vibration sensitivity in the same dense-chassis scenario, since a batch of consumer drives crammed into a case they weren’t rated for can throw read errors under vibration that a single drive on a bench never will.

It won’t catch a drive that’s going to die from ordinary wear at year four. That’s what SMART monitoring, scrubs, and checking for bit rot over the drive’s actual service life are for. Burn-in is specifically about the front edge of the bathtub curve.

The actual sequence

This is the order that catches the most for the least wasted time, run in full before a drive goes into any array, new or shucked.

1. Check SMART before you do anything else. Run smartctl -a /dev/sdX and look at the raw attributes before you’ve touched the drive at all: reallocated sector count, current pending sector count, offline uncorrectable count, and power-on hours (should read close to zero on a genuinely new drive, a nonzero number here is worth investigating before you go further). This is your baseline, and it’s also your first chance to catch a drive that already shipped in bad shape.

2. Run a full destructive write test with badblocks. On Linux, badblocks -wsv /dev/sdX writes and reads back multiple patterns across the entire drive, which is a genuinely thorough way to exercise every sector at full capacity. This is destructive, it will erase anything already on the drive, so this step only makes sense on a drive with nothing on it yet. It’s also slow: expect somewhere in the range of a day or more per pass on a large modern drive, and running the full four-pattern pass rather than a single quick pass is worth the extra time on a drive about to go into production. If you can’t spare the time for a full write test, a read-only pass (badblocks -sv, no -w) or the drive’s own extended SMART self-test is a meaningfully weaker but still nonzero substitute.

3. Run the drive’s extended SMART self-test. smartctl -t long /dev/sdX kicks off an internal self-test that the drive runs itself in the background, using its own firmware to check the full surface. It’s non-destructive, so it’s safe to run even on a drive you’ve already written data to, but it’s not a replacement for badblocks since it’s testing different things (the drive validating its own mechanism versus you validating actual read/write behavior at the block layer). Check progress and result with smartctl -a /dev/sdX once it finishes, it logs a pass/fail directly into the SMART self-test log.

4. Recheck SMART attributes after both tests. Look specifically at whether reallocated sector count, pending sector count, or uncorrectable sector count moved from your baseline in step 1. Any of those climbing during burn-in is exactly the signal you’re running this whole process to catch.

5. Let it run under real load for several days before final commit. Once the drive passes steps 1 through 4 clean, don’t rush straight into a rebuild. A few days of real read/write activity, whether that’s a non-critical rsync job, a scratch dataset, or just leaving it in the array as a hot spare briefly before promoting it, gives thermal and vibration issues in your specific enclosure a chance to show up under conditions closer to how the drive will actually live.

What counts as a fail

Any of the following during burn-in is grounds to RMA or return the drive rather than trust it, full stop, no matter how good the rest of the results look:

  • Reallocated sector count, current pending sector count, or offline uncorrectable count above zero and climbing (a small nonzero pending-sector count that later clears back to zero after a rewrite is a gray area some people tolerate, a climbing one is not).
  • Any read or write error surfaced during the badblocks pass.
  • The SMART extended self-test itself reporting a failure in its result code.
  • Temperatures that run meaningfully hotter than the drive’s rated operating range under your actual chassis airflow, even if every other test passes clean. A drive that’s mechanically fine but is going to cook itself in your specific case is still a problem you want to know about now.

A drive that fails any of these during burn-in, while it’s empty and easy to replace, is a much better outcome than the same failure six months from now during a resilver, when a second drive in the same array is under the exact kind of sustained read load that tends to surface the failures you didn’t catch the first time.

How long is actually enough

There’s no single correct number here, and the honest answer is that it’s a tradeoff between thoroughness and how long you’re willing to have a drive sit idle before it’s earning its keep. A full badblocks four-pass write test plus an extended SMART self-test, run back to back, is a reasonable minimum for any drive going into a production array, and typically lands somewhere between one and several days depending on drive capacity. If you’re buying drives in bulk for a new build, staggering them through this process a few at a time rather than powering up an entire chassis of new drives simultaneously also has a practical side benefit: it spreads out the inrush current and spin-up load instead of hitting your PSU with a dozen drives spinning up at once.

For drives that are replacing a failed member in an already-degraded array, where every hour of delay is an hour spent without redundancy, it’s reasonable to compress this down to the SMART baseline check and the extended self-test, skip the multi-day write pass, and accept a slightly higher residual risk in exchange for getting back to full redundancy faster. That’s a real tradeoff, not a shortcut to pretend doesn’t exist, and it’s worth deciding on deliberately rather than by default.

Bottom line

Burn-in testing is cheap insurance against the specific, well-documented failure spike that happens right after a drive is manufactured. It costs you some idle time and a bit of electricity, and in exchange it moves the discovery of a bad drive from “during a live rebuild” to “before it ever held a byte of your data.” Run the SMART baseline, run a full badblocks pass, run the extended self-test, recheck SMART, and give it a few days of real activity before you trust it completely. Skipping straight from unboxing to production is the single easiest way to find out about a bad drive at the worst possible time.