HardDiskIndex

Choosing drives for a NAS: what actually differs from a desktop

By Harry Saarinen ·

The mechanics are the same. A NAS drive and a desktop drive of the same generation often share platters, heads and firmware base. What differs is a short list of settings and tolerances, and only one of them will actually break your array: how long the drive is willing to sit there retrying a bad sector before it answers.

Everything else - vibration tolerance, workload rating, warranty - shifts probabilities. Error recovery time changes behaviour. A desktop drive in an array is not slightly worse; it fails differently, and the failure looks like a healthy drive being ejected.

The second uncomfortable thing goes near the top too, because most buying advice depends on it being false. The ladder that matters runs inside a family, not between the two vendors. Western Digital’s January 2025 WD Red Pro datasheet rates the whole family, 2 TB through 26 TB, at 550 TB per year and calls it suitable for “RAID-optimized NAS systems with unlimited # of bays”. Seagate’s current IronWolf Pro datasheets rate its equivalent identically - 550 TB per year, “Drive Bays Supported: Unlimited” - so the Pro tier has converged. Any article still contrasting the two on those figures is quoting a Seagate datasheet that has since been revised upward.

What has not converged is the ladder inside a single family. WD Red Plus is 180 TB/year against Red Pro’s 550. Two WD Red Pro part numbers of the same 14 TB capacity differ by a factor of ten on unrecoverable read error rate. Two WD Red Plus part numbers of the same 12 TB capacity differ by more than two to one on idle power, because one is helium-filled and one is not. The part number is the specification. The family name is a shelf.

Error recovery time is the one specification that changes behaviour

A drive that cannot read a sector does not give up. It repositions the head, re-reads, applies progressively heavier correction, tries deliberate off-track offsets in both directions, and re-reads again. On a desktop that is exactly right: the drive is the only copy of the file, and a sector recovered on the fortieth attempt is a photograph you did not lose.

The Linux kernel RAID wiki puts the desktop-drive retry ceiling at over two minutes. No vendor publishes the retry algorithm, the number of passes or a hard limit, so treat two minutes as an observed order of magnitude, not a specification.

Now put that drive in an array. The controller asked for a sector, has had no answer, concludes the drive is dead, and drops it. You are now rebuilding - a full-surface read of every surviving drive, for days at current capacities - over one slow sector that was about to be delivered.

The timeout mismatch, exactly

Layer Timeout Where it is set
Hardware RAID controllers Commonly around 8 seconds Controller firmware, often undocumented
Linux SCSI/libata command timeout 30 seconds /sys/block/sdX/device/timeout
ZFS on Linux Inherits the 30-second kernel timeout As above
Drive with ERC enabled 7 seconds typical factory default SCT Error Recovery Control
Drive with ERC disabled Over 120 seconds Not settable
Drive-managed SMR, error path Over 10 minutes Not settable, not bounded

The kernel RAID wiki names the consequence in one sentence: when the kernel gives up first, “the raid code assumes the drive is dead and kicks it from the array”. The drive was working. The timeout killed it.

The fix is to make the drive give up before the layer above it does. The ATA feature is SCT Error Recovery Control; vendors rename it - Western Digital TLER (time-limited error recovery), Seagate ERC (error recovery control), Samsung and the old HGST lines CCTL (command completion time limit). They are the same SCT command set underneath.

$ sudo smartctl -l scterc /dev/sda
SCT Error Recovery Control:
           Read:     70 (7.0 seconds)
          Write:     70 (7.0 seconds)

$ sudo smartctl -l scterc,70,70 /dev/sdb    # read limit, write limit

The units are deciseconds, so 70 is seven seconds - deliberately just under the eight-second controller timeout. The smartctl manual page adds the floor most guides omit: 0 disables the feature entirely, and “other values less than 65 are probably not supported”. Read: Disabled means the drive has the feature but is sitting in desktop behaviour; SCT Error Recovery Control command not supported means it lacks it.

The persistence advice in most guides is now out of date

The standing folklore is that the setting never survives a power cycle. That was true, and it stopped being true. From smartmontools 7.3, the manual page documents a persistent form:

# set, persistent across power-on reset
$ sudo smartctl -l scterc,70,70,p /dev/sdb

# read back the persistent values
$ sudo smartctl -l scterc,p /dev/sdb

# restore the manufacturer's defaults
$ sudo smartctl -l scterc,reset /dev/sdb

The manual is explicit that “the ,p and ,reset options require the device to support ATA ACS-4 or higher”, so an older drive will refuse them and you are back to re-applying at boot. Try the persistent form first and verify it survives a cold power cycle, not a reboot - a warm reboot may not drop the drive’s power rail, which makes a non-persistent setting look persistent.

Where ,p is unavailable, the correct mechanism is a udev rule or a systemd unit that runs on every boot, and the correct verification is reading the value back afterwards rather than assuming the write took.

ZFS wants a far shorter limit than RAID does, and says so

The seven-second convention comes from hardware RAID controllers, and applying it under ZFS is inherited habit rather than reasoning. The OpenZFS hardware guidance recommends something an order of magnitude tighter:

it is advisable to write a script to set the error recovery time to a low value, such as 0.1 seconds until ZFS is modified to control it. This must be done on every boot.

That is smartctl -l scterc,1,1. The logic is that ZFS does not need the drive’s recovery attempt at all: it has a checksum, it knows the block is wrong, and it can reconstruct from parity or a mirror immediately. Every second the drive spends retrying is a second ZFS spends waiting for an answer it is about to discard. The same document notes that drives with the feature “typically default to 7 seconds”, so the factory value is tuned for somebody else’s array.

Expect most drives to refuse it. The smartctl floor quoted above applies here too: values below 65 are probably not supported. So set it, read it straight back with smartctl -l scterc, and fall back to 70,70 if the drive did not take the value. OpenZFS’s recommendation is what you would want rather than what most SATA drives will give you, and a setting you have not read back is a setting you do not have.

Stack Recommended ERC Why
Hardware RAID controller 70,70 (7.0 s) Just under the controller’s own ~8 s timeout
Linux mdadm 70,70 (7.0 s) Named as typical in the smartctl manual page
OpenZFS 1,1 (0.1 s) if the drive accepts it, else 70,70 Parity reconstruction is faster than drive retry
SAS, any stack Mode page 0x01 Ships at array-appropriate values

When the drive cannot do it at all

Move the other timeout. The kernel RAID wiki gives the fallback directly:

# raise the kernel's per-device command timeout past the drive's retry ceiling
for d in /sys/block/sd*/device/timeout; do echo 180 > "$d"; done

This must also be re-applied at boot. It is the worse fix - a genuinely dead drive now hangs the array for three minutes instead of thirty seconds - but it beats losing a member. The asymmetry that matters: software RAID’s timeout is a kernel setting you own, and a hardware controller’s is not. On a hardware controller you cannot raise the timeout, so if the drive cannot lower its own, there is no fix at all. That, rather than anything about the drives, is why hardware RAID and desktop drives pair so badly.

OpenZFS records which way the market moved, and names a vendor: “Since ZFS’ introduction, error recovery control has been removed from low-end drives from certain manufacturers, most notably Western Digital.” The historical pattern is worth knowing because it repeats. WD’s desktop models of the era - Caviar SE, SE16, GP, Raptor - shipped with TLER read and write configured as 0, disabled; the RAID-class models, Caviar RE2 and RE2-GP, shipped at 7 seconds on both. WD later withdrew the ability to change it on desktop drives and warned that the old WDTLER.EXE utility could brick newer firmware. Firmware giveth and firmware taketh away, and the marketing name is not the specification.

Which current drives allow it is a per-model fact, not a per-brand one. The NAS and enterprise lines - WD Red Plus, Red Pro, Gold and Ultrastar; Seagate IronWolf, IronWolf Pro and Exos; Toshiba N300 and MG - ship with it enabled or settable, and Seagate’s IronWolf Pro datasheet lists “Time-Limited Error Recovery” explicitly in its AgileArray bundle alongside dual-plane balancing. Consumer lines vary. Check the drive in front of you, not the line it is in.

SAS puts this in a mode page instead

A SAS drive sidesteps the whole argument. Its recovery limits live in the SCSI Read-Write Error Recovery mode page (0x01), which exposes a Recovery Time Limit field, and SAS drives generally ship at array-appropriate values because there is no desktop market to tune them for.

sudo sdparm --get=RTL /dev/sdX          # read the recovery time limit
sudo sdparm --set=RTL=3000 /dev/sdX     # set it, value in milliseconds

Seagate additionally ships openSeaChest for its own drives, which reaches settings sdparm does not. This is one of several reasons a SAS pull is an easier array member than a shucked SATA drive - see enterprise drive pulls for the rest.

SMR is not a slow sector, it is an unbounded one

The same kernel wiki that gives the two-minute figure for a desktop CMR drive gives a worse one for shingled media: an SMR drive can stall for over ten minutes during a read error. That is beyond any achievable kernel timeout and beyond the 180-second workaround. The mechanism is in the CMR/SMR guide; the conclusion here is that ERC tuning does not rescue a shingled drive, because the number you would have to tune to is not bounded.

Rotational vibration sensors measure less than their marketing implies

Vendors used to rate NAS drives by bay count, and the tiering has come apart. Western Digital’s current Red Pro datasheet claims “unlimited # of bays” and Seagate’s current IronWolf Pro datasheets print “Drive Bays Supported: Unlimited”; only WD Red Plus still carries a number, “up to 8 bays”. Both vendors have dropped the ceiling on their Pro lines, and neither has published the test behind either the old number or the new one. A limit that can go from 24 bays to unlimited without a matching change in the mechanical specification was a marketing tier rather than a measurement. Understand the mechanism instead.

The error budget, derived from geometry

Track pitch is not published per product, so it has to be derived. These numbers are mine, not a vendor’s, and the assumptions are all visible so you can disagree with any of them:

assume a 24 TB drive: 10 platters, 20 recording surfaces (helium-class)
assume usable annulus: outer radius 46 mm, inner radius 23 mm
assume bit aspect ratio BPI : TPI = 5 : 1

per-surface capacity   24e12 x 8 / 20        = 9.60e12 bits
annulus area           pi x (46^2 - 23^2) mm^2 = 4986 mm^2 = 7.73 in^2
areal density          9.60e12 / 7.73        = 1.24e12 bits/in^2

TPI                    sqrt(1.24e12 / 5)     = ~498,000 tracks per inch
track pitch            25.4e6 nm / 498,000   = ~51 nanometres
off-track write budget, roughly 10% of pitch = ~5.1 nanometres

The head has about five nanometres of position error budget while writing. The same arithmetic at a 32 TB, 22-surface drive gives about 4.6 nm. The exact figure is uncertain - platter geometry, surface count and aspect ratio are all assumptions, and vendors publish none of them per product - but the order of magnitude is single-digit nanometres. That is what the servo defends, and it is why a chassis that hums loses throughput.

Why it is rotational vibration specifically

Actuator arms are mass-balanced about their pivot, so linear vibration produces no net torque there and the mechanical design already rejects it. Rotational vibration is angular acceleration about an axis parallel to the spindle, measured in rad/s², and it acts directly on the arm’s moment of inertia. No balancing trick cancels it.

It comes from Newton’s third law. When a drive seeks, its actuator accelerates one way and its baseplate is pushed the other, and that reaction torque goes into the chassis and on to every drive bolted to the same metal. Eight drives seeking independently are eight uncorrelated torque sources exciting one structure. Add spindle imbalance at the rotation frequency - 120 Hz at 7200 rpm, 90 Hz at 5400 - and drives whose actual speeds differ by fractions of a per cent beat against each other at a few hertz, exactly where a sheet metal chassis damps worst.

What the sensor pair actually is

Two accelerometers on the printed circuit board assembly, and what matters is what is done with their outputs rather than what they are made of. Each one measures linear acceleration in the plane of the board. Subtract one from the other and the common-mode component - the whole drive being shaken bodily in that plane - cancels, leaving a difference proportional to rotational acceleration about the spindle axis. That difference is the quantity the servo wants, and no single accelerometer placed anywhere on the board could have produced it.

A note on a distinction that is not one, because it recurs in every discussion of this feature. “Piezoelectric” names the transduction element; “linear” names the axis measured. A piezoelectric accelerometer is a linear accelerometer, and the ASHRAE paper quoted further down calls the same sensors piezoelectric while describing exactly the arrangement above. Neither word tells you the thing you want to know, which is the separation distance, the bandwidth of the feedforward path and the rejection actually achieved - and no vendor publishes any of the three per product.

The difference signal is digitised and fed forward into the servo - added to the normal feedback effort, not replacing it - so the correcting current is applied before the position error signal has registered the disturbance. That is the entire value: feedback reacts after the head has moved, feedforward reacts before.

A drive without the pair finds out from its position error signal alone, by which point it must inhibit the write and wait a full revolution.

The missed revolution, and what it actually costs

ASHRAE Technical Committee 9.9, in a 2019 paper whose authors include Seagate, Dell, IBM and Intel engineers, describes the mechanism in the vendors’ own words:

the servo system suspends the write process until the head returns to an acceptable range to the track center, and the system must wait for at least one revolution of the disk to attempt again to write. This time delay increases latency and results in system performance degradation.

7200 rpm -> 60 / 7200 = 8.33 ms per revolution
5400 rpm -> 60 / 5400 = 11.11 ms per revolution

theoretical average access, 7200 rpm:
  rotational latency 4.17 ms + seek 4 to 9 ms   =  8 to 13 ms
empirical average access, 7200 rpm (OpenZFS)    = 13 to 16 ms
rated average access, 15k rpm (OpenZFS)         = 3.4 ms read, 3.9 ms write

This guide previously said a missed revolution roughly doubles the cost of an operation. Against the theoretical figure that is nearly true; against the empirical one it is not. OpenZFS’s hardware documentation gives 13 to 16 ms as the real measured average access time for 7200 rpm drives. A lost 8.33 ms revolution on that is a 52 to 64 per cent increase, not a doubling.

And nothing reports it. No SMART attribute increments, no error is logged, no counter moves. The array is simply slower than it was on the bench, with no artefact to point at. That is the most important property of this failure mode: it does not look like a failure.

The sensors are structurally blind to most of what now matters

This is the part that almost no buying guide mentions, and ASHRAE states it plainly:

Successful rotational vibration (RV) feedforward systems measure the rigid body motion of the HDD using two piezoelectric sensors on the printed circuit board assembly (PCBA). These sensors only measure in-plane motion… They do not measure rotational motion around x-axis or y-axis or z-axis linear motion and cannot accurately measure nonrigid body motion.

Disturbance Sensed?
In-plane linear motion (x, y) Yes, and cancelled as common mode
Rotation about z (the spindle axis) Yes - this is the whole feature
Rotation about x or y (chassis rocking) No
Linear motion along z No
Non-rigid-body motion (panel flexing, resonance) Not accurately

A NAS enclosure whose side panel resonates, or whose drive cage rocks about a horizontal axis, produces disturbance the RV feedforward system cannot see and therefore cannot correct. An “RV sensor” tick on a datasheet buys rejection of one of five disturbance modes.

Worse, there is a hard bandwidth ceiling. ASHRAE: “The current sets of technologies and mechanics allow for error rejection below 2 kHz, but disturbances from AMDs can input excitations up to 20 kHz”, and “Present HDDs show sensitivity up to ~20 kHz”. AMD there means air-moving device: a fan or blower. The servo rejects below 2 kHz and the drive is sensitive to 20 kHz. That gap is not a tuning problem to be fixed in a firmware revision; it is a property of the mechanical loop.

Acoustics, not neighbouring drives, is the dominant disturbance now

The classic story is drive-to-drive seek coupling. In a dense modern chassis the fans have overtaken it, and the scaling is brutal:

fan sound power level    Lw1 - Lw2 = 50 x log10(RPM1 / RPM2)  dB
blower sound power level             ~70 x log10(RPM1 / RPM2) dB
structural vibration amplitude       proportional to RPM^2
blade pass frequency  fb = Nb x RPM / 60

so raising fan speed by 25%:
  fan Lw           +50 x log10(1.25)  = +4.8 dB
  blower Lw        +70 x log10(1.25)  = +6.8 dB
  vibration amp    1.25^2             = 1.56x

and raising it by 50%:
  fan Lw           +8.8 dB,  blower +12.3 dB,  vibration 2.25x

50 x log10(R) is a power-level relation, so acoustic power goes as R^5
and sound pressure amplitude as R^2.5

Acoustic power tracking the fifth power of fan speed is why a NAS that was fine in winter loses throughput in a warm room. The thermal control loop raises fan RPM by a modest fraction, the acoustic energy reaching the drives jumps by several decibels, and the drives start inhibiting writes. Nothing in the monitoring will say so.

Blade pass frequency is worth computing for your own chassis: a seven-blade fan at 1200 rpm excites at 140 Hz, the same fan at 3000 rpm at 350 Hz, a nine-blade high-static-pressure fan at 6000 rpm at 900 Hz. All below 2 kHz, which is the good news; the harmonics are not.

Helium drives are quieter inside and more fragile outside

ASHRAE again:

Although the transition to using helium inside HDDs results in an environment with very low airflow disturbance to the head, the designs become more susceptible to external vibration, especially vibration caused by coupling with acoustic sources.

The mechanism is counter-intuitive. Air inside a sealed drive causes windage - turbulent flow that buffets the arm and the platters. Helium, at about one seventh the density, nearly removes it. With that internal disturbance gone, external excitation is no longer masked by it, and the same absolute disturbance becomes a larger fraction of the error budget. Helium designs also carry more platters, which means thinner platters and thinner arms, and lower mechanical stiffness.

So the highest-capacity drive on your shortlist is probably helium, probably more sensitive to chassis acoustics than the air drive it replaces, and the datasheet will not say so.

The datasheet figure is a design target, not a measurement

Seagate’s IronWolf Pro datasheet publishes rotational vibration as 12.5 at 10 to 1500 Hz - not 0 to 1500 Hz, which matters because the sub-10 Hz region where chassis rocking lives is excluded from the test. It prints the unit as “rad/s”, although the physical quantity is angular acceleration in rad/s².

The revealing detail is that the value is identical - 12.5 - at 4, 6, 8, 10, 12, 14, 16, 18 and 20 TB. A number that does not move across five generations of areal density and two fill gases is a design target the family is qualified against, not a per-model measurement. Treat it as “this family passed the internal RV spec” and nothing more precise.

And then ASHRAE closes the door on the number you actually want:

Because there are a myriad of AMDs (and resulting frequency characteristics) and chassis designs, there is no specific reference speed for an AMD that may be used to determine the HDD performance in an enclosure.

There is no published figure that predicts your chassis’s throughput loss, and there cannot be, because the answer is a property of the enclosure, the fan curve, the mounting hardware and the drive together. The paper also notes that measuring throughput alone will not identify the mechanism - you see the loss and cannot attribute it.

What you can do is bound it empirically. Measure sustained sequential throughput with the array’s fans pinned to minimum, then again pinned to maximum, on otherwise identical work. If the numbers differ materially, acoustic coupling is costing you something, and rubber-grommet mounting, a slower larger fan, or damping the panel is worth trying. In a four-bay plastic enclosure the effect is usually small enough to ignore. By eight bays in a metal chassis it is real, and it arrives as slowness rather than as an error.

Workload rating is a warranty condition quoted at an operating point you are not at

Workload rating is published in terabytes per year, and it is commonly described as covering host traffic plus the drive’s own background media scans. That is wrong, and the vendors are specific. Western Digital’s Red Plus datasheet footnote defines workload rate as “the amount of user data transferred to or from the hard drive”, and the Red Pro footnote gives the annualisation formula outright:

Annualized Workload Rate = TB transferred x (8760 / recorded power-on hours)

Host user data only. The drive’s own background media scan, its idle reorganisation, its read-after-write verification - none of that appears in the vendor definition. Whether it counts against the physical wear the rating is a proxy for is a separate question, and an unanswerable one from public documents.

The tier table, and where the step actually falls

Tier Workload rating Rated power-on hours Source
Desktop ~55 TB/year ~2400 hours/year (light-duty rating) Consumer datasheets, where they state it at all
NAS 180 TB/year 8760 hours/year WD Red Plus; Toshiba N300
NAS Pro 550 TB/year, both vendors 8760 hours/year WD Red Pro and Seagate IronWolf Pro datasheets
Enterprise nearline 550 TB/year 8760 hours/year Seagate Exos X24 product manual; Toshiba MG11 datasheet

The bottom two rows are the correction, and they are the same number. The NAS Pro tier and the enterprise nearline tier have the same published workload rating on both vendors’ current datasheets, so the workload figure is not what you are paying the enterprise premium for - see enterprise drives for what is.

The step that is real is the one from 180 to 550, between Red Plus and Red Pro inside a single vendor’s range. That is a factor of three on the warranty condition for a difference in shelf position most buyers read as a badge.

What exceeding it does is published - just not as a curve

“Nobody publishes what happens when you exceed it” is the usual claim, and it is too strong. WD publishes the mechanism in its Red Pro MTBF footnote: MTBF is estimated at a workload of 220 TB/year and 40 °C, and “Derating of MTBF will occur above these parameters, up to 550TB writes per year”.

Read that twice, because it reframes both numbers on the front of the datasheet.

WD Red Pro
  headline MTBF     2,500,000 hours (14-26 TB) / 2,000,000 hours (2-12 TB)
  quoted at         220 TB/year, 40 °C
  workload rating   550 TB/year
  between them      MTBF derates, curve not published

WD Red Plus
  headline MTBF     1,000,000 hours
  quoted at         90 TB/year, 40 °C
  workload rating   180 TB/year
  between them      MTBF derates, up to 65 °C drive temperature

The headline reliability figure and the headline workload figure are quoted at different operating points, and the reliability one sits at 40 to 50 per cent of the workload one - 220 against 550 on Red Pro, 90 against 180 on Red Plus. A drive run at its full rated workload does not have the MTBF on the front of its datasheet. What it does have is not published; only the fact that it is lower. That gives you a target: design for 220 TB/year on Red Pro and 90 TB/year on Red Plus if you want the advertised MTBF, not for the ceiling.

It is still a warranty condition, not a wear counter

An SSD’s TBW maps to a physical mechanism - oxide degradation per program/erase cycle, one countable event per block - which is why SSD endurance arithmetic works. A hard drive has no equivalent counter. It wears through head-disk interface contact, lubricant depletion, head-media spacing degradation and bearing wear, and those track seeking, flying hours and thermal cycling far more than bytes moved. The workload rating is a proxy the vendor chose because it is measurable from the host side, not because bytes cause the wear.

The hours line binds sooner for most people anyway: a desktop drive rated for 2400 hours a year, run 24x7, spends its annual allowance in 100 days.

As a continuous rate, these ratings are small

 55 TB/year / 31,536,000 s =   1.7 MB/s sustained, continuously
180 TB/year / 31,536,000 s =   5.7 MB/s sustained, continuously
220 TB/year / 31,536,000 s =   7.0 MB/s sustained, continuously
550 TB/year / 31,536,000 s =  17.4 MB/s sustained, continuously

saturated 1 GbE   (125 MB/s)  =  3,942 TB/year
saturated 2.5 GbE (312 MB/s)  =  9,839 TB/year
saturated 10 GbE  (1250 MB/s) = 39,420 TB/year

A saturated gigabit link is about twenty-two times a 180 TB/year rating and seven times a 550 TB/year one. Nobody saturates gigabit continuously, which is why these ratings rarely bind on household traffic: a family writing two terabytes of photos and backups a year sits at roughly one per cent of a NAS drive’s budget. Even a 10 GbE link only matters if you keep it busy, and the thing that keeps a home NAS busy is not the humans.

Scrubbing is what blows through it, and mdadm is far worse than ZFS

This is where the arithmetic changes materially depending on your filesystem, by a factor of 1.6 on the worked example below.

Six 12 TB drives in RAID6, 48 TB usable, 30 TB in use. ZFS scrubs allocated blocks only, so each drive reads roughly its own share of what is stored:

ZFS, pool 62% full: ~7.5 TB read per drive per scrub
  monthly   7.5 x 12 =  90 TB/year per drive
  weekly    7.5 x 52 = 390 TB/year per drive

mdadm’s check is allocation-blind. It reads every sector of every member regardless of whether anything is stored there, because it is comparing raw parity against raw data and has no idea what the filesystem above it considers used:

mdadm, same array:  12 TB read per drive per check
  monthly    12 x 12 = 144 TB/year per drive
  weekly     12 x 52 = 624 TB/year per drive

A weekly mdadm check on 12 TB drives is 624 TB/year of scrub traffic alone - more than a 550 TB/year enterprise rating, before the array has served a single file. Monthly is 144 TB/year, which already consumes 80 per cent of a 180 TB/year WD Red Plus allowance and 65 per cent of the 220 TB/year point where WD quotes Red Pro’s MTBF.

And the schedule is set for you either way, which is why this section exists. Debian ships mdadm’s checkarray on a cron entry that fires on the first Sunday of the month at 00:57. Red Hat and Fedora do not ship checkarray at all. They ship raid-check, driven by raid-check.timer, whose unit file reads OnCalendar=Sun *-*-* 01:00:00 and describes itself as a “Weekly RAID setup health check”. TrueNAS schedules a monthly ZFS scrub per pool by default.

Debian    checkarray   first Sunday of the month, 00:57   -> 144 TB/yr per 12 TB member
Red Hat   raid-check   every Sunday, 01:00                -> 624 TB/yr per 12 TB member
TrueNAS   ZFS scrub    monthly per pool                   ->  90 TB/yr per member at 62% full

On a Red Hat or Fedora mdadm array the shipped default is weekly, and on 12 TB members that is the 624 TB/year figure above - more than an enterprise workload rating, from a timer nobody chose. It is the single most expensive default in this guide, and the one most likely to be running right now on a machine whose owner has never read the unit file.

The trade cuts both ways: scrubbing is how you find latent sector errors while redundancy still exists to repair them from, and the evidence for that is the strongest single finding in the field literature - see the URE section below. So do not turn it off. Monthly is defensible on any of the three. Weekly on large mdadm members is not, unless you have deliberately decided the workload rating is a number you will exceed - and if you are on Red Hat you have decided it by default. systemctl edit raid-check.timer and an OnCalendar=Sun *-*-01..07 01:00:00 is the monthly equivalent.

Reliability figures, and what each one is actually measuring

Three numbers get quoted as if they were the same kind of thing. They are not, and the differences decide which ones you can use.

Figure What it is What it is not
MTBF / MTTF A model output at a stated operating point A life expectancy, or a field rate
AFR (datasheet) The same model, expressed per year A measurement
AFR (field) A measurement, of somebody else’s fleet Transferable to your drives
URE rate A specification ceiling A mean, or a distribution

MTBF is a number with fine print attached

Current published figures, from the vendors’ own documents:

Drive family MTBF/MTTF Published AFR Quoted at
WD Red Pro 14-26 TB 2,500,000 h not published 220 TB/yr, 40 °C
WD Red Pro 2-12 TB 2,000,000 h not published 220 TB/yr, 40 °C
WD Red Plus 2-12 TB 1,000,000 h not published 90 TB/yr, 40 °C
Seagate IronWolf Pro 12-32 TB 2,500,000 h not published not stated on the datasheet
Seagate IronWolf Pro 6-10 TB 2,000,000 h not published not stated on the datasheet
Toshiba MG11 (to 24 TB) 2,500,000 h 0.35% not stated

Five of those six rows say “not published” in the AFR column, and that is the finding rather than an omission. An AFR obtained by dividing 8,760 by the MTBF is not a published AFR; it is the same model figure in different units, and putting it in a column headed “published” is how a derived number acquires an authority it never had. Toshiba prints both, and they agree to the digit - 8,760 / 2,500,000 = 0.35 per cent - which tells you what a published AFR usually is: the MTBF, restated.

A 2.5 million hour MTBF is 285 years, which nobody claims is a life expectancy. It is the inverse of a modelled constant hazard rate during the flat part of the bathtub curve: “if you ran a large population for a year, this fraction would fail”, not “this drive will last that long”.

The field measurements disagree with the model, consistently

Schroeder and Gibson’s FAST 2007 study analysed about 100,000 disks across large production installations and found field replacement rates of 2 to 4 per cent typical, against datasheet MTTFs of 1.0 to 1.5 million hours - a nominal AFR of at most 0.88 per cent. Some populations ran at 13 per cent. Two structural findings from that paper matter more than the headline gap:

  • Time between replacements fits a Weibull distribution with a decreasing hazard rate, not the constant hazard rate that MTBF arithmetic assumes. The model underlying every MTBF figure is the wrong model.
  • Replacements showed autocorrelation significant at lags up to 30 weeks. Failures arrive in clusters, months apart. Independence is not a safe assumption, and every rebuild-risk calculation in the next section assumes it.

Backblaze’s Q1 2026 Drive Stats, the largest continuously published fleet data there is, gives the current picture, and its 2025 annual report the last full year:

drives in service       341,263
drive-days in quarter    30,203,180
failures in quarter           1,030
quarterly AFR                 1.24%
lifetime AFR                  1.39%   (529,968,464 drive-days, 20,212 failures)
20 TB+ cohort AFR             0.85%   across the quarter
2025 full year                1.36%   (down from 1.55% in 2024)
worst single model            4.90%   (1,251 drives of one Seagate 14 TB model)

Two things to take from it. First, the fleet-wide figure and the worst-model figure differ by a factor of four, so a fleet average is not a prediction about a specific model. Second, the 20 TB+ cohort at 0.85 per cent is the best number in the table, contradicting the intuition that bigger drives must be less reliable - though those drives are also the newest in the fleet, and furthest from the wear-out end of the bathtub.

The blind spot in every published failure statistic

Backblaze names its own methodology limit, and it is exactly the limit that matters to a buyer:

Because it defines a failure based on a drive’s presence in the pool the day before, that means we can’t define day one failures using only the Drive Stats program.

They put it more bluntly in the same post: “Have we been completely under-reporting day one failures? Short answer: yes, probably.”

Day-one failures are invisible in the statistics. A drive that arrives dead, or that fails in its first week, never enters the dataset as a failure - it was never present to go missing. That is precisely the class of failure burn-in testing exists to catch, and it is precisely the class the published AFR figures cannot tell you about. Do not read a 0.85 per cent AFR as your probability of a bad drive out of the box; read it as your probability of losing a drive that survived its first week in a datacentre.

SMART is a strong signal and a poor predictor

Google’s FAST 2007 disk failure study is still the reference, and its two useful findings point in opposite directions:

  • A drive was 39 times more likely to fail within 60 days of its first scan error. When SMART says something, listen.
  • A large fraction of failed drives showed no SMART errors at all. A clean SMART report is not evidence of health.

The same study found something that contradicts the standard advice: over most of the observed temperature range, lower temperatures correlated with higher failure rates, reversing only at the extreme top end. The correlation is probably confounded - cooler drives in that fleet were in different roles with different workloads - but the plain reading is that chasing the last few degrees is not where your effort belongs. The heat section below gives the ceiling that does matter.

The 10^14 argument, done with the correct probability

Consumer datasheets specify an unrecoverable read error rate of less than 1 sector per 10^14 bits read. NAS and enterprise drives are usually 10^15, and some SAS drives 10^16.

The first correction is that the tiers are not as clean as the marketing. The entire WD Red Plus family, 2 TB through 12 TB, is specified at less than 1 in 10^14 - the same as a desktop drive, on a NAS-branded product. And within WD Red Pro, the 14 TB WD141KFGX is specified at 10^14 while WD142KFGX, also 14 TB, is specified at 10^15. Two part numbers, the same capacity, the same family, a factor of ten apart on the specification that the entire RAID 5 argument turns on. Read the datasheet for the part number, not the family.

The common arithmetic is wrong in a nameable way

The version of this argument you will meet online divides bits read by the URE rate and calls the result a probability:

16 TB read x 8 = 1.28e14 bits
1.28e14 x 1e-14 = 1.28  ->  "128% chance of a URE"

A probability of 128 per cent should have ended the argument on the spot. The division gives the expected number of UREs, not the probability of at least one. The correct form treats each read as an independent Bernoulli trial, which in the limit is Poisson:

P(at least one URE) = 1 - e^(-N x p)

N = bits read, p = per-bit error rate

16 TB:   1 - e^(-1.28e14 x 1e-14) = 1 - e^(-1.28) = 72.2%

72.2 per cent, not 128 per cent. The distinction gets larger as the naive number grows: at 96 TB read the naive answer is 768 per cent and the correct one is 99.95 per cent.

The sector model agrees, which answers the usual objection

The standard attack on the bit model is that bits do not fail independently - errors happen at sector granularity, so the framing is wrong. Run it at sector granularity and the answer does not move:

p per 4096-byte physical sector = 1e-14 x 4096 x 8  = 3.28e-10
sectors in 16 TB                = 16e12 / 4096      = 3.906e9

P(at least one) = 1 - (1 - 3.28e-10)^3.906e9 = 72.196%
Poisson form    = 1 - e^(-3.28e-10 x 3.906e9) = 72.196%

Identical to three significant figures. The granularity objection is correct in principle and changes nothing numerically, because the expected count is the same either way and both distributions have the same mean.

The table, extended to current capacities and to 10^16

Five-drive RAID5, rebuild reads four drives in full. These are my calculations, from the published specification ceiling, not measurements:

Drive size Bytes read At 10^-14 At 10^-15 At 10^-16
2 TB 8 TB 47.3% 6.2% 0.6%
4 TB 16 TB 72.2% 12.0% 1.3%
8 TB 32 TB 92.3% 22.6% 2.5%
12 TB 48 TB 97.9% 31.9% 3.8%
16 TB 64 TB 99.4% 40.1% 5.0%
20 TB 80 TB 99.8% 47.3% 6.2%
24 TB 96 TB 99.95% 53.6% 7.4%
30 TB 120 TB 99.99% 61.7% 9.2%

Read down the columns rather than across the rows. The specification tier dominates the capacity. A 24 TB drive at 10^16 is safer on this metric than a 2 TB drive at 10^14 by a factor of six, while reading twelve times as much data to get there. Hold the capacity still and move only the tier - 2 TB at 10^14 against 2 TB at 10^16 - and the factor is roughly seventy-five. Capacity growth is not what broke RAID 5; it is what made the difference between tiers unignorable. A useful way to hold it:

bytes read per expected URE
  1 in 10^14  =    12.5 TB
  1 in 10^15  =   125   TB
  1 in 10^16  = 1,250   TB

Four reasons the percentages overstate the risk

The honest position is that those percentages are too pessimistic while the trend they describe is not in dispute.

  • The spec is a ceiling, not a mean. Datasheets say “less than 1 in 10^14”. Field rates appear better, but no vendor publishes the distribution and no independent measurement exists at scale. You are computing with an upper bound and reporting the answer as a point estimate.
  • The empirical test fails. If 72 per cent per rebuild were real, five-drive 4 TB RAID5 arrays would almost never complete one. They routinely do.
  • A URE is not automatically an array loss. This is the biggest error in the whole argument, and it gets its own subsection below.
  • Errors are not independent. Bairavasundaram and colleagues at NetApp studied latent sector errors across 1.53 million drives in over 50,000 arrays over 32 months. The quoted aggregate is 3.45 per cent, but the split is what a NAS buyer needs: 8.5 per cent of consumer-class disks against 1.9 per cent of enterprise-class, a factor of 4.5; in the first year alone, 3.15 per cent against 1.46 per cent. And 0.2 per cent of the drives that had any error at all had more than 1000 errors - roughly seven drives in every hundred thousand, not two in every thousand. That is where the clustering shows: errors concentrate in a few bad drives rather than spreading evenly. The same study found a drive with one error is much more likely to develop a second, that the second is likely to be near the first, and that the affected fraction rises with capacity and age. Clustering breaks the Poisson model in both directions: most drives are cleaner than the model says, and the bad ones are far worse.

That last study also supplies the strongest empirical argument for scrubbing at all: over 60 per cent of latent sector errors were found by scrubbing. They were sitting there, invisible to any status page, waiting for a rebuild to find them.

The consequence fallacy, which is the real error

The model computes P(URE during rebuild) and then silently equates it with P(array loss). Those are different events, and how different depends entirely on what is managing the array:

Stack What a URE during rebuild does
Older hardware RAID controller Aborts the rebuild and drops the array
Linux mdadm Records the block in its bad block list and continues
OpenZFS Names the affected file in zpool status -v and continues

The kernel documents the mdadm mechanism directly: md maintains a per-device bad_blocks file exposing “the list of all known bad blocks in the form of start address and length (in sectors respectively)”, plus unacknowledged_bad_blocks for entries known but not yet written to disk. That list is precisely what lets a degraded array survive an unreadable sector rather than aborting.

Those three outcomes differ by orders of magnitude in cost. One loses the array; one loses nothing; one loses a named file you can restore from backup. The “RAID 5 is dead” argument computes a number that is roughly right for the first row and applies it to all three.

Where the argument came from, stated accurately

Robin Harris made the original argument on ZDNet in July 2007, in “Why RAID 5 stops working in 2009”. Adam Leventhal analysed RAID 6 late in 2009, in “Triple-Parity RAID and Beyond” for ACM Queue, and Harris’s StorageMojo follow-up of February 2010 drew on that article. It put a seven-drive RAID5 of 2 TB SATA drives at roughly a 62 per cent chance of data loss on rebuild at 1 in 10^14, assuming about 2.5 hours per TB minimum at 115 MB/s with real times two to five times longer. The 2019 line is Harris’s summary of Leventhal in that post, not a quotation from him: “RAID 6 protection levels will be as good as RAID 5 was until 2019”.

The prediction was directionally right and the industry responded: RAID 6 and raidz2 became the default, and the 10^15 and 10^16 tiers became the things people check. The date has passed and arrays have not stopped working - which is what you would expect from an argument right about direction and wrong about magnitude.

The probability that actually governs array loss

On a modern software RAID, the event that loses the array is not a URE. It is a second whole-drive failure during the rebuild window. Here is that number, derived from Backblaze’s published AFR for its 20 TB+ cohort, assuming independence:

5-drive array, one member lost, 4 surviving members plus the replacement

20 TB at 100 MB/s  ->  55.6 h rebuild window
AFR 0.85%          ->  per-drive hazard over 55.6 h = 1 - (1-0.0085)^(55.6/8760)
P(at least one of 5 fails in the window)            = 0.027%
Scenario Rebuild window P(second failure)
12 TB, 100 MB/s, AFR 0.85% 33.3 h 0.016%
20 TB, 100 MB/s, AFR 0.85% 55.6 h 0.027%
24 TB, 100 MB/s, AFR 0.85% 66.7 h 0.032%
20 TB, 50 MB/s, AFR 0.85% 111 h 0.054%
20 TB, 100 MB/s, AFR 1.39% (fleet lifetime) 55.6 h 0.044%
20 TB, 50 MB/s, AFR 3% (Schroeder field range) 111 h 0.19%

Under independence these are small numbers, which is exactly why the correlation evidence matters more than the URE arithmetic. Schroeder and Gibson found autocorrelation out to 30 weeks; array members share a model, a firmware revision, a manufacturing batch, a thermal history, a vibration environment and a workload. The independence assumption in that table is the weakest thing in this guide, and it is the assumption every published rebuild-risk figure makes.

The exact probability cannot be derived from a published upper bound and a violated independence assumption, and this site will not pretend otherwise. But the data that must be read without error scales linearly with capacity while nothing protecting it has improved, so RAID5’s margin gets thinner with every capacity generation - a claim about direction, which does not need the percentage to be right.

Rebuild and resilver arithmetic at current capacities

Rebuild time is capacity divided by sustained throughput. Current 3.5-inch 7200 rpm drives do roughly 250 to 287 MB/s on their outermost tracks and 110 to 130 MB/s on their innermost, because outer tracks are longer and pass under the head faster at the same angular velocity. A whole-surface average near 200 MB/s is what a rebuild gets on an idle array. The published outer-track figures set the top of that range:

Drive Max internal transfer rate Source
WD Red Pro 24-26 TB 287 MB/s WD Red Pro datasheet
WD Red Pro 20 TB (WD202KFGX) 285 MB/s WD Red Pro datasheet
Seagate IronWolf Pro 20 TB 285 MB/s IronWolf Pro datasheet
WD Red Pro 12 TB (WD121KFBX) 240 MB/s WD Red Pro datasheet
WD Red Plus 2-12 TB 180-260 MB/s WD Red Plus datasheet
WD Red Pro 2 TB 164 MB/s WD Red Pro datasheet

Within one current family the sequential rate varies by 75 per cent, so a rebuild-time table that does not say which class it assumes is not telling you much.

Drive At 200 MB/s (idle) At 100 MB/s (in service) At 50 MB/s (busy)
4 TB 5.6 h 11.1 h 22.2 h
8 TB 11.1 h 22.2 h 44.4 h
12 TB 16.7 h 33.3 h 66.7 h
16 TB 22.2 h 44.4 h 88.9 h
20 TB 27.8 h 55.6 h 4.6 days
24 TB 33.3 h 66.7 h 5.6 days
30 TB 41.7 h 83.3 h 6.9 days

It is worse than it was because capacity has grown far faster than the rate at which a head can read a track - areal density improves in two dimensions, sequential throughput in only one:

2007:  1 TB  /  80 MB/s  =  3.5 hours
today: 24 TB / 200 MB/s  = 33.3 hours
       24x the capacity, 2.5x the throughput, ~9.6x the rebuild

What ZFS does differently, and where it does worse

ZFS and btrfs rebuild only allocated blocks - TrueNAS puts it as “ZFS only copies blocks in use, reducing the time it takes to rebuild the vdev” - so a half-full pool resilvers in roughly half the time. That is one of the better arguments for not filling an array, and it is a genuine advantage over mdadm, which copies every block regardless.

Against that, a raidz resilver walks the pool in block-pointer order rather than LBA order, turning a sequential read into a near-random one. The saving from reading less can be eaten by the cost of reading it out of order, and on a fragmented pool it routinely is. OpenZFS’s sequential resilver covers mirrors and dRAID, not raidz.

Three operational facts that surprise people mid-resilver:

  • A resilver pre-empts a scrub. “If a resilver is in progress, ZFS does not allow a scrub to be started until the resilver completes.” On older Solaris ZFS a running scrub was suspended and restarted from the beginning afterwards.
  • The time estimate is not reliable, by design. The zpool-scrub manual page states that “scrubs may progress beyond 100% completion” because the pool changes underneath the scan, and gives that as the reason time estimates become inaccurate.
  • Progress is two numbers, not one. The manual’s own example reads 403M / 405M scanned at 100M/s, 68.4M / 405M issued at 10.0M/s. Scanned is metadata traversal; issued is actual I/O. They differ by an order of magnitude routinely, and reading the first as progress will mislead you every time.

Oracle’s ZFS administration guide gives the scheduling behaviour: a scrub “proceeds as fast as the devices allow, though the priority of any I/O remains below that of normal operations”. It will use everything the array is not using, and stop being polite only when nothing else is asking.

Drive count against drive size, for a fixed usable capacity

“Fewer bigger drives” is the standard advice and it is half right. Here is the whole trade, at a fixed 60 TB usable on raidz2 or RAID6, with resilver times at 100 MB/s. These are my calculations, not measurements:

Layout Raw Parity tax Per-drive resilver Data read during resilver
8 x 10 TB 80 TB 25.0% 27.8 h 70 TB
7 x 12 TB 84 TB 28.6% 33.3 h 72 TB
6 x 15 TB 90 TB 33.3% 41.7 h 75 TB
5 x 20 TB 100 TB 40.0% 55.6 h 80 TB
4 x 30 TB 120 TB 50.0% 83.3 h 90 TB

Fewer bigger drives cost more parity, not less. At a fixed parity count, a narrower vdev spends a larger fraction of its raw capacity on redundancy: two drives of eight is 25 per cent, two drives of four is 50 per cent. The resilver is three times longer and the array reads 29 per cent more data to complete it.

Against that, drive count multiplies the annual chance of touching a rebuild at all. At Backblaze’s 20 TB+ AFR of 0.85 per cent, and at their lifetime 1.39 per cent:

Drives P(at least one fails in a year) at 0.85% at 1.39%
4 3.36% 5.45%
5 4.18% 6.76%
6 4.99% 8.06%
7 5.80% 9.33%
8 6.60% 10.59%
12 9.74% 15.46%

The decision trades how often you rebuild against how long each rebuild takes. Eight drives rebuild twice as often as four and each rebuild is a third as long, so total exposure time is roughly a wash and the eight-drive array wins on parity efficiency. Fewer bigger drives win instead on the things that scale with spindle count: bays, which are the scarce resource; noise, power and spin-up surge; and the cost of chassis, controller and cabling. Both answers are defensible. Only one of them is usually stated.

RAID levels, and what each one actually survives

With 12 TB drives:

Layout 4 drives 6 drives 8 drives Survives
RAID5 / raidz1 36 TB 60 TB 84 TB any 1
RAID6 / raidz2 24 TB 48 TB 72 TB any 2
raidz3 - 36 TB 60 TB any 3
RAID10 / mirrors 24 TB 36 TB 48 TB 1 per pair

At four drives, RAID6 and RAID10 cost exactly the same capacity. RAID10 rebuilds by copying one surviving drive to one replacement - no parity computation, no reading the rest of the array, often a third of the time - but it survives only the right second failure, where RAID6 survives any. At four bays that is the entire decision. By six drives RAID6 leads on capacity by a third, and by eight by half.

Mirrors have one advantage the capacity table hides: a mirror vdev’s read IOPS scale with the number of members, while a raidz vdev’s write IOPS are those of a single disk. For a NAS serving large sequential files, raidz is right. For one running virtual machines or a database, mirrors usually are, and the capacity loss is the price of the IOPS.

RAID5 at modern sizes is defensible only narrowly: small drives, a small array, and a tested backup. The URE table above is why, with the caveats about the URE table above attached.

RAID is not a backup, and that is not a technicality. It protects against one failure mode - a drive dying - and not against deletion, ransomware, filesystem corruption, a controller writing garbage to every member at once, a power event, fire or theft. Worse, it replicates faithfully: a rebuild reconstructs corrupted data exactly as corrupt as it was. A backup is three copies, on two media, one of them elsewhere; an array is one copy however many drives are in it.

ZFS: the decisions that are permanent, and the ones that are not

ZFS is the default recommendation for a NAS, and the reason is checksums - it can tell you that a block is wrong, which no traditional RAID layer can. The cost is a set of decisions made at creation time that cannot be undone.

ashift is permanent, and getting it wrong is permanent too

ashift is the base-2 logarithm of the vdev’s sector size. It is set when the vdev is created and cannot be changed afterwards - not by a property, not by a rewrite, only by destroying and recreating the vdev.

ashift  9 =    512 bytes
ashift 12 =  4,096 bytes    <- current recommendation for hard drives
ashift 13 =  8,192 bytes
ashift 14 = 16,384 bytes
valid range 9 to 16; 0 means autodetect

The failure mode is asymmetric. Set ashift too high and you waste some space on small blocks. Set it too low - typically by letting autodetect believe a 512e drive’s emulated 512-byte sector - and every partial-sector write becomes a read-modify-write at the drive for the life of the vdev. Not every drive reports its physical sector size honestly.

Always set ashift=12 explicitly on a hard-drive pool. Autodetect is the value that can be wrong, and it is wrong in the direction that cannot be fixed.

recordsize affects new files only

recordsize defaults to 128 KiB and takes powers of two from 512 bytes to 1 MiB, or to 16 MiB with the large_blocks feature. The OpenZFS tuning documentation is explicit that “changing the recordsize on a dataset will only take effect for new files”, so setting it after you have written the data does nothing to the data.

Workload Recommended recordsize Reason
Media, backups, large sequential files 1M Fewer, larger I/Os; better compression ratio
General file share 128K (default) No strong reason to move
MySQL / InnoDB 16K Match the engine’s page size
PostgreSQL 32K Match the engine’s page size
SQLite 64K Match the engine’s page size

For a media NAS, recordsize=1M on the bulk datasets is the documented recommendation and it is worth setting before the first file lands.

RAIDZ space efficiency is not the parity fraction

This is the ZFS surprise that costs people real capacity. The naive model says a 3-disk raidz1 gives you 2/3 of raw. What you actually get depends on recordsize against ashift:

3-disk raidz1, ashift=12 (4 KiB sectors), 2 data + 1 parity per stripe

4 KiB record:
  1 data sector, and a stripe cannot have less than 1 parity sector
  4 KiB stored as 8 KiB  ->  50% efficient

128 KiB record:
  32 data sectors, in 16 stripes of 2  ->  16 parity sectors
  48 sectors = 192 KiB of media to store 128 KiB
  128 / 192  ->  66.7% efficient

A small-recordsize dataset on a narrow raidz vdev loses far more capacity than the parity fraction suggests - here, a third more than you budgeted for. The effect shrinks as the vdev widens and as the recordsize grows, which is another reason media datasets want 1M.

vdev width, and why more narrow vdevs beat fewer wide ones

OpenZFS recommends 3 to 9 disks per raidz vdev, with the minimum being the parity count plus one. TrueNAS adds a hard ceiling: “We do not recommend using more than 12 disks per vdev.”

The performance reason is blunt. A raidz vdev’s write IOPS are approximately those of a single disk - every write touches every member, so the vdev goes as fast as its slowest member. Two 6-wide raidz2 vdevs have roughly twice the write IOPS of one 12-wide raidz2 vdev of the same drives, at the cost of two more drives of parity. For almost every workload that is the right trade.

OpenZFS 2.4 added a related mitigation: raidz sit-out, where a disk that lags behind its peers is temporarily bypassed and its data reconstructed from parity while writes continue, with a default threshold of 600 seconds. That is the ZFS-native analogue of the ERC problem, handled in the pool rather than at the drive - and it does not remove the reason to set ERC, because 600 seconds is a very long time to be slow.

raidz expansion exists now, with two caveats

OpenZFS 2.3 added raidz expansion: zpool attach a single disk to an existing raidz vdev, online, with the pool readable throughout. Two things it does not do:

  • Fault tolerance does not change. A raidz2 stays a raidz2, now spread across more disks. Widening does not buy you more parity.
  • Blocks written before the expansion keep their original data-to-parity ratio until they are rewritten. A block written on a 4-wide raidz1 still carries one parity sector per three data sectors after the vdev grows to six. The capacity you expected appears gradually, as data turns over, or immediately only for new writes.

special vdevs: real gains, total risk

A special allocation class vdev holds metadata, the indirect blocks of user data, and any dedup tables, on faster media. With special_small_blocks set, it additionally absorbs files below a size threshold - valid at zero or any power of two from 512 bytes to 1 MiB.

The thing to understand before adding one: losing a special vdev destroys the pool. Not the metadata, not some files - the pool. TrueNAS enforces the consequence at creation time by refusing to build a pool where the metadata vdev’s redundancy is lower than the data vdevs’. A three-way mirror for the special vdev under a raidz2 data vdev is not paranoia; it is the matching level.

The gain is genuine on a pool with many small files, because directory traversal and metadata reads stop hitting spinning media. On a media pool with 40,000 large files it is close to pointless.

SLOG is for synchronous writes and almost nothing else

A separate log vdev accelerates only synchronous writes - the ones an application requested with fsync() or O_SYNC, typically NFS exports, iSCSI targets and database commits. Asynchronous writes, which is nearly everything an SMB file share does, never touch it.

size needed:  ~4 GB of flash is sufficient
upper bound:  no benefit beyond max ARC size
              = half of system RAM on Linux, three quarters on illumos
TrueNAS:      ZFS currently uses 16 GiB of SLOG space
topology:     single device or mirror; raidz is not supported for a log vdev

If your NAS serves SMB to humans, an SLOG will do nothing measurable, and the device is better spent as a special vdev or left out.

L2ARC costs RAM, and on a small system it makes things worse

Every block held in L2ARC needs a header in ARC - roughly 70 bytes each - so the RAM cost scales with the record count, not the device size:

512 GiB L2ARC device, ARC header cost at 70 bytes per record

recordsize 1M    ->    524,288 records  ->   ~35 MiB of ARC
recordsize 128K  ->  4,194,304 records  ->  ~280 MiB of ARC
recordsize 16K   -> 33,554,432 records  ->  ~2.2 GiB of ARC

On a 16 GiB system with a database dataset at 16K records, a 512 GiB cache device eats an eighth of the RAM that was doing the caching. TrueNAS’s rule is direct: do not add L2ARC below 32 GiB of RAM, and keep the device under ten times RAM.

It also wears the SSD. At the default l2arc_write_max of 8 MiB/s, a consumer SSD rated for 500 TiB of writes reaches its endurance in roughly two years; raise the feed rate to 64 MiB/s and it is about three months. That is a cache device being consumed to hold data that RAM would have held better.

Klara Systems’ summary is the one to remember: ARC is roughly an order of magnitude faster than L2ARC, so the answer to “should I add L2ARC” is “not until RAM is maxed out”. The minimum useful hit ratio before the device earns its keep is around 25 per cent, and since OpenZFS 2.0 the cache does at least survive reboots.

Two settings that are not permanent and are worth changing

Hardware RAID controllers should not be under ZFS at all, and the reason is correctness rather than performance: “Hardware RAID will limit opportunities for ZFS to perform self healing on checksum failures.” Presenting each drive as a single-drive RAID0 volume to simulate an HBA is explicitly discouraged. Use an HBA - the enterprise guide covers IT mode and which cards do it.

And watch the fill level. TrueNAS documents a real allocator change, not a rule of thumb: “At 90% capacity, ZFS switches from performance- to space-based optimization, which has massive performance implications.” Add capacity before 80 per cent. On a 6 x 12 TB raidz2 pool that is 38.4 TB of the 48 TB usable, and it arrives sooner than you think.

Scrub scheduling

TrueNAS schedules a monthly scrub per pool by default. The widely used split is monthly for consumer-grade drives and quarterly for datacentre-grade, with two to four weeks for business VM pools. Weigh it against the workload arithmetic above: on ZFS a monthly scrub of a 62-per-cent-full six-drive array is about 90 TB/year per drive, which fits comfortably inside a 180 TB/year rating and inside the 220 TB/year point where WD quotes Red Pro’s MTBF. Monthly is the defensible default on ZFS. It is a much more expensive default on mdadm.

mdadm and btrfs: what each one gives up

mdadm is the most predictable of the three, and the least informed

Linux md exposes its scrub machinery through sync_action, which takes six values: resync, recover, idle, frozen, check and repair. frozen is the one worth knowing and the one most guides omit: it stops the current action and prevents a new one from starting, where idle only stops the current one and leaves the array free to begin another. During check and repair, the kernel documentation says, “md will count the number of errors that are found. The count in mismatch_cnt is the number of sectors that were re-written, or (for check) would have been re-written.”

# start a read-only consistency check
echo check > /sys/block/md0/md/sync_action

# watch it
cat /proc/mdstat
cat /sys/block/md0/md/mismatch_cnt

# stop it
echo idle > /sys/block/md0/md/sync_action

# stop it and keep it stopped
echo frozen > /sys/block/md0/md/sync_action

A non-zero mismatch_cnt on RAID5/6 means data and parity disagree. It does not tell you which one is wrong. Without checksums md knows there is an inconsistency and cannot identify the correct version; repair picks parity and recomputes, which is a guess. That is the strongest argument for ZFS or btrfs over mdadm, and it is a correctness argument rather than a performance one.

What md does have is the bad block list described above, which lets a degraded array survive a URE, and one genuinely useful tuning knob:

/sys/block/md0/md/stripe_cache_size
  default 256 entries, settable from 17 to 32768

Raising stripe_cache_size materially changes RAID5/6 write throughput on spinning disks, at a RAM cost proportional to the value times the member count. It is one of the few md defaults chosen for smallness rather than for your workload.

btrfs RAID5/6 is still classified unstable, and this is not gossip

The btrfs upstream status page classifies RAID56 stability as “unstable”, and the same page defines that word: “do not use for other then testing purposes, known severe problems, missing implementation of some core parts”. RAID1, RAID10 and scrub are all marked “OK”. The page is maintained against current kernels.

That is upstream’s own assessment in upstream’s own table. btrfs RAID1 and RAID10 on a NAS are fine. btrfs RAID5/6 is not a thing to put your array on.

A configuration detail most guides omit: even where btrfs RAID5/6 data is used, the documentation discourages RAID5/6 metadata and recommends raid1 or raid1c3 for that profile. A btrfs “RAID6” array is therefore normally a mixed-profile array, and anyone quoting you a usable-capacity figure from the parity fraction alone has not accounted for it.

btrfs does have what mdadm lacks: per-block checksums, so its scrub identifies which copy is correct rather than guessing. On RAID1 that is the whole value proposition, and it is a good one.

Synology SHR and Unraid parity are not RAID levels

Both get described as “RAID variants”. Neither is, and the differences decide whether they suit you.

SHR is layered RAID groups, not a new algorithm

Synology Hybrid RAID slices each drive at the size boundaries of the drive set and runs an independent RAID group inside each layer. Everything below the smallest drive’s capacity becomes one parity-protected group across every disk; the region between the smallest and second-smallest becomes its own group on the drives that reach that height; and so on up. SHR-1 layers behave like RAID5, or RAID1 where only two drives reach that layer; SHR-2 layers tolerate two failures.

Worked, with 2 TB + 3 TB + 4 TB + 4 TB:

layer 1: 0-2 TB on 4 drives, RAID5   ->  3 x 2 TB  =  6 TB usable
layer 2: 2-3 TB on 3 drives, RAID5   ->  2 x 1 TB  =  2 TB usable
layer 3: 3-4 TB on 2 drives, RAID1   ->  1 x 1 TB  =  1 TB usable
                                          total    =  9 TB = 8.18 TiB

plain RAID5 across the same drives, limited to the smallest member:
  4 x 2 TB, one parity            ->  3 x 2 TB    =  6 TB = 5.46 TiB

SHR recovers 50 per cent more usable capacity from a mismatched set, and that is its entire reason to exist. It costs nothing in redundancy per layer and costs you understanding: a “one drive failure” means a different thing in each layer, and the recovery path is a multi-group rebuild. With four identical drives SHR and RAID5 are the same thing, so the feature only pays when your drives are mismatched.

Unraid is not striping at all

Each data disk in an Unraid array carries its own ordinary filesystem, complete and independently mountable. The parity disk holds the XOR of every corresponding sector position across the data disks. There is no stripe, no chunk size, and no relationship between a file and any disk but the one it is on.

Three consequences follow, and they are the whole argument:

  • Only the disk holding a file must spin up. On a media library that is the dominant power saving, and no parity array can match it.
  • A file survives whole on its own disk if you lose more disks than parity covers. On a RAID6 array, losing three members loses everything. On Unraid with single parity, losing two disks loses the contents of those two disks and leaves the rest readable. That is a fundamentally different failure shape.
  • The parity disk must be at least as large as the largest data disk, because it must cover every sector position that exists.

The cost is write throughput. Unraid’s default write mode is read-modify-write:

default (read-modify-write):  read old data, read old parity,
                              write new data, write new parity
                              = 4 I/Os, 2 drives spinning

turbo / reconstruct write:    read all other data disks, compute parity,
                              write new data, write new parity
                              = 2 writes + one read per other data disk,
                                every drive spinning

That is the trade stated exactly, and note which way the I/O count goes: the default costs four I/Os on two drives; turbo write costs two writes plus a read of every other data disk, so on an eight-disk array it is eight I/Os on eight drives. Turbo write is faster because nothing waits on a read-modify cycle, not because it does less work - it does more, on more spindles, and gives up the power advantage entirely. Dual parity (P and Q) survives two failures, as RAID6 does.

Unraid suits a media library that is written rarely and read one file at a time. It does not suit anything that wants sustained write throughput or parallel reads across many files.

Mixing capacities, vendors and batches

Array members share every variable that drives failure: model, firmware revision, manufacturing batch, power-on hours, temperature, vibration environment and workload. That is about as far from independent samples as a population gets, and every probability in this guide assumes independence.

Schroeder and Gibson’s finding that time between replacements shows significant autocorrelation out to 30-week lags is the empirical form of the problem. Failures arrive in clusters. The constant-hazard model underlying MTBF arithmetic fit their data poorly.

So buy deliberately mixed: two vendors, or two retailers, or the same model bought months apart, and check the date codes on arrival rather than assuming. A firmware defect is the strongest case for it, being the one failure mode that can take a whole array inside an hour and the one that no amount of parity survives.

But be honest about the limit. Nobody has published a study quantifying how much batch diversity actually buys. It is a cheap hedge against an unquantified risk, not a measured improvement, and anyone who gives you a percentage has made it up.

Mixing capacities is a separate question with a per-platform answer:

Platform Mixed capacities
ZFS raidz Every drive is used to the size of the smallest. The excess is wasted.
mdadm Same - the array is built on the smallest member’s size.
Synology SHR Layered groups recover most of the excess, as above.
Unraid Fully supported; parity drive must be the largest.

On ZFS, adding one larger drive to a raidz vdev buys nothing until every member is replaced. That is the single most common mistake in home ZFS builds, and it is why the expansion plan below matters more than the drive you buy today.

SMR in an array is a different failure, not a slower one

Not in an array, and the reason is the best-documented failure in recent consumer storage. In 2020 it emerged that drive-managed SMR had been shipped without disclosure across more than one vendor’s consumer lines, and inside a NAS-branded one: the WD Red. Owners found out when they replaced a failed drive and the array could not rebuild.

The mechanism belongs to the CMR/SMR guide; the array-level consequence is specific. A drive-managed SMR drive absorbs writes into a persistent media cache and reorganises them into shingled bands later, while idle. A rebuild gives it hours of sustained writes and no idle time. The cache fills, and every subsequent write forces a read-modify-write of a whole shingled band. Throughput falls to single-digit megabytes per second and individual commands start taking seconds - which lands you back in the first section of this guide, with the kernel wiki’s figure of over ten minutes for a stall during a read error. The array does not see a slow drive, it sees a drive that stopped answering, and drops it. Adding a replacement can cost you a second member.

Nor can you test for it easily. /sys/block/sdX/queue/zoned flags only host-managed and host-aware drives; drive-managed SMR reports none, identically to a CMR drive, and no SMART attribute records it.

So filter to CMR in the hard drive listings, and treat an unlabelled listing as unknown rather than as CMR. A model number does not tell you whether a drive is shingled. The same marketing name has covered both technologies across generations and regions, sellers transcribe part numbers wrongly, and this site will not infer what it cannot establish.

Shucked drives, the power-disable pin, and the firmware you cannot change

Shucking an external drive is often the cheapest route to NAS capacity, and it carries two specific problems that are not the drive’s fault.

The 3.3 V pin is a specification change, not a defect

SATA 3.3, published on 2 February 2016, incorporated SATA-IO technical proposal TPR056 - authored by Frank Chu of HGST with James Hatfield and Alvin Cox of Seagate - which reassigned connector pin P3. Take the order carefully, because it is usually given backwards. P3 originally carried 3.3 V alongside P1 and P2, as part of the supply. SATA 3.2, in August 2013, reassigned it to DEVSLP. SATA 3.3 then gave the same pin a third job: the Power Disable control, driving it high at 2.1 to 3.6 V to cut power to the drive circuitry. The proposal notes that Power Disable and Device Sleep are asserted by the same voltage on P3 and are therefore mutually exclusive - one pin, three successive meanings, the last two of which cannot coexist. The interfaces guide has the full revision history. Power Disable exists so a datacentre can power-cycle one drive in a backplane without touching its neighbours.

Western Digital’s own technical brief states the symptom plainly:

if you put a new SATA HDD with this feature into a legacy chassis or enclosure, the drive may not spin up! The HDD is not defective. Some legacy power supplies provide 3.3V power on P3 (Pin 3), and this forces the HDD to get stuck in a hard reset condition.

Any power supply whose SATA connector actually delivers 3.3 V on P3 - which older supplies do, because that was the specification - holds a PWDIS drive in permanent reset. The drive is fine. It is being told to stay off.

Two field fixes, neither of which harms the drive, because both simply prevent 3.3 V reaching P3:

  • Kapton tape over pin 3 of the 15-pin SATA power connector. Pin 1 is the end nearest the L-shaped notch; count three. Kapton rather than electrical tape because it does not soften at drive-bay temperatures.
  • A Molex-to-SATA power adapter, which carries only 12 V and 5 V and has no 3.3 V rail at all.

The detail most guides miss is in WD’s own brief: the feature requires a distinct PCBA, so it is a per-part-number property rather than a firmware setting you could turn off. WD publishes separate part numbers with and without it

  • the DC HC510 8 TB is 0F27610 without and 0F27455 with - and the SATA models of the DC HC530 do not offer it at all. WD’s own recommendation is to buy drives without the feature unless you specifically need it. You cannot tell from the outside of a sealed external enclosure which one is in there, though on a bare Ultrastar the model number settles it before you bid - the interfaces guide decodes the suffix.

The other shucking risk is recording technology

An external enclosure’s marketing does not state recording technology, and the 2020 episode establishes that one product line can contain either. A shucked drive is a drive whose CMR status you cannot establish until it is out of the case, and often not then - see buying used drives on eBay for the rest of what a listing withholds.

Helium, and what the sensor can and cannot tell you

Helium fill is not a tier, it is a per-model manufacturing choice, and it does not track capacity cleanly:

Family Helium Air
WD Red Plus WD120EFBX (12 TB) only Everything else, 2-12 TB
WD Red Pro 14 TB and above, plus WD121KFBX (12 TB) WD122KFBX (12 TB), and 10 TB and below
Seagate IronWolf Pro 12-20 TB, and ST10000NE0008 8 TB and below, and ST10000NE0004

Every row of that table has a capacity sitting on both sides of the line. The two 10 TB IronWolf Pro part numbers, ST10000NE0004 and ST10000NE0008, are one digit apart and one fill gas apart. So are the two 12 TB WD Red Plus part numbers, WD120EFGX air and WD120EFBX helium. So are the two 12 TB WD Red Pro part numbers, WD122KFBX air and WD121KFBX helium - which is why the Red Pro row cannot be stated as “helium from 14 TB” however much cleaner that would read. Same family, same capacity, different fill gas, different power figures, different acoustic behaviour. This is a part-number fact, not a capacity fact, and the digit that carries it is not the one you would look at.

SMART 22 is a warning light with no gauge behind it

Helium-sealed drives expose SMART attribute 22, a pre-fail attribute with a published normalised threshold of 25. It starts at 100 and counts down. The raw value has no interpretable physical meaning to an owner, because HGST explicitly declined to publish either the amount of helium in a drive or the trip points; the attribute “trips once the drive detects that the internal environment is out of specification”.

So: any movement below 100 is a signal, and you cannot tell how much margin is left. In Backblaze’s fleet only one HGST drive ever read below 100, at values between 94 and 99 - a rare event, and an unambiguous one. Only HGST helium drives reported the attribute at all in that analysis; Seagate’s did not, so its absence on a helium drive is not evidence of anything.

Helium does not appear to change failure rates

Backblaze’s normalised comparison put helium drives at 1.06 per cent AFR against 1.61 per cent for comparable air drives, with the confound that the helium drives were newer. The effect you can verify from the datasheets is roughly 20 per cent lower spin power, and for a home NAS that matters more than the reliability difference does.

A drive that lies about its spindle speed

Filed here because it is the same class of trap. WD Red Plus footnote 7, for the WD120EFGX:

Actual spindle motor rotational speed for this model is 7200 RPM; although ID Device may report 5400 to reflect previous Performance Class designation.

The drive reports 5400 rpm over ATA IDENTIFY and physically spins at 7200. Any tool reading RPM from the drive - smartctl, hdparm -I, every NAS dashboard - reports the wrong number, and a decision made on quietness or power draw from that figure is wrong too. It is why filtering by RPM tells you what the listing claims rather than what the spindle does, and it generalises: the drive’s self-report is a field in firmware, not a measurement.

Burn-in, with the commands and the ceiling nobody mentions

Day-one failures are invisible in every published statistic, as Backblaze’s own methodology note concedes. Burn-in is the only way you find them, and the window in which finding them is useful is the return window.

badblocks has a hard capacity ceiling

badblocks uses a 32-bit block counter, which sets a maximum testable size per block size. This is why an 18 TB drive returns “Value too large for defined data type”, and why a great many people conclude their drive is broken:

Block size Maximum testable
-b 512 (default) 2.20 TB (2.00 TiB)
-b 1024 4.40 TB (4.00 TiB)
-b 4096 17.59 TB (16.00 TiB)
-b 8192 35.18 TB (32.00 TiB)

Use -b 4096 up to 16 TiB and -b 8192 above it. The default of 512 has been useless for over a decade.

The sequence

# 1. short self-test, about 5 minutes - catches the obviously dead
sudo smartctl -t short /dev/sdX

# 2. conveyance test, about 2 minutes - looks for shipping damage specifically
sudo smartctl -t conveyance /dev/sdX

# 3. long self-test - full surface read, hours
sudo smartctl -t long /dev/sdX

# 4. record the attributes BEFORE the destructive test
sudo smartctl -A /dev/sdX > /root/sdX.before

# 5. destructive four-pattern write/read - DESTROYS ALL DATA
sudo badblocks -b 4096 -wsv /dev/sdX          # up to 16 TiB
sudo badblocks -b 8192 -wsv /dev/sdX          # 16 to 32 TiB

# 6. long self-test again
sudo smartctl -t long /dev/sdX

# 7. compare
sudo smartctl -A /dev/sdX > /root/sdX.after
diff /root/sdX.before /root/sdX.after

Step 7 is the actual test. A single SMART reading tells you the drive’s history; the difference across a full-surface write and verify tells you whether the drive is degrading right now. Three raw values must be zero on a drive you are keeping:

  5  Reallocated_Sector_Ct    sectors already remapped
197  Current_Pending_Sector   unreadable, not yet remapped
198  Offline_Uncorrectable    unreadable during offline scan

Attributes 197 and 198 are the ones that become a failed rebuild. A pending sector is one the drive could not read and has not yet reallocated - precisely the URE you did not want to meet while degraded. A non-zero raw value on any of the three, on a drive you have owned a week, is a return rather than a risk to price in.

It takes days, and it costs a chunk of the workload rating

badblocks -w writes and reads four patterns, so a full run moves eight times the drive’s capacity. Reported field times: a 20 TB Seagate at 190 to 195 hours, about eight days; a 2 TB WD Red at just over 24 hours.

naive estimate, 20 TB at 150 MB/s average:
  20e12 x 8 / 150e6 / 3600 = 296 hours

reported: ~190 hours
implied average rate: 160e12 / (190 x 3600) = 234 MB/s

The implied rate is near the drive’s outer-track figure, so treat 190 hours as a best case rather than a typical one. Budget eight to ten days per large drive - not a week, which is less than the best case this paragraph has just rejected as optimistic - and run them in parallel if the chassis and the power supply allow.

Connect it back to the workload rating, because nobody does. The middle column is the 220 TB/year point at which WD quotes Red Pro’s MTBF, which is the number the drive’s advertised reliability figure actually depends on:

Drive Traffic per full run vs 180 TB/yr vs 220 TB/yr vs 550 TB/yr
8 TB 64 TB 36% 29% 12%
12 TB 96 TB 53% 44% 17%
20 TB 160 TB 89% 73% 29%
24 TB 192 TB 107% 87% 35%

A single burn-in of a 24 TB drive at the 180 TB/year tier exceeds its entire annual workload allowance. That is not a reason to skip burn-in - finding a bad drive inside the return window is worth more than a warranty condition - but it is a reason to do it once rather than routinely, and to plan the year’s scrub schedule knowing the budget is already partly spent.

SSD caching in a NAS, and when it does nothing

The ZFS analysis above applies: RAM first, L2ARC only when RAM is maxed, SLOG only for synchronous writes. Synology’s DSM caching is a different implementation, and its published arithmetic is unusually honest. From the DSM 7.1 SSD Cache white paper:

RAM cost:        ~400 KB of system memory per 1 GB of SSD cache
RAM ceiling:     DSM uses at most 25% of installed RAM for the mapping table
Model ceiling:   930 GB total cache on Alpine-CPU models
Read-write cache: 2 to 12 SSDs, in RAID1, RAID5 or RAID6
800 GB cache  ->  800 x 400 KB = 320,000 KB of mapping table
                               = 320 MB decimal
                               = 312.5 MiB, if Synology's KB is 1024 bytes
                  needs at least ~1.3 GB installed to stay under the 25% cap
                  (either reading; 320 x 4 = 1.28 GB, 327.7 x 4 = 1.31 GB)

Synology does not say which kind of kilobyte it means, and this is a guide that spends a paragraph on decimal against binary elsewhere, so both readings are above. They do not change the answer.

Their own advice is worth quoting, because it argues against selling you more SSD:

The recommended SSD cache size should be just enough to cover the size of frequently accessed data… increasing the SSD cache size to 500 GB for 100 GB of hot data will not result in significant performance increase. Excess cache space will only be used to store cold data.

Their in-house tests state the headline as 30 times the random read IOPS with cache against without, and the measurements behind it range from roughly 15 to 40 times depending on the test: their NVMe configuration goes from 6,581 to 263,667 IOPS on a 100 per cent random read, which is 40 times, while a random write test in the same paper is nearer 15. They publish no latency figure at all, in any unit, anywhere in the paper. If you meet “93 per cent lower latency” attributed to this document, check the arithmetic before you repeat it: 1 - 1/15 = 93.3 per cent. It is an understated IOPS ratio inverted and relabelled as a latency measurement, which is a derived number wearing a vendor’s authority.

Read the workload before the number, too: that is an OLTP-style small random pattern. A media NAS streaming large sequential files gets approximately nothing from an SSD cache, because there is no hot set to cache and the drives were never the sequential bottleneck.

Where it earns its place: many small files, virtual machine images, a photo library’s thumbnails and metadata, an iSCSI target. Where it does not: a film library, a backup target, an archive. Choosing between a cache SSD and more RAM, add RAM.

Noise, heat and power, budgeted

Noise adds incoherently, and that is good news

Uncorrelated sources add at +3 dB per doubling, so total sound pressure is L + 10 x log10(N):

Per-drive seek figure 4 drives 8 drives 12 drives
28 dBA 34.0 dBA 37.0 dBA 38.8 dBA
32 dBA 38.0 dBA 41.0 dBA 42.8 dBA
36 dBA 42.0 dBA 45.0 dBA 46.8 dBA

Eight drives are only about 9 dB louder than one, so a drive 4 dBA quieter buys more than halving the drive count. Datasheet figures across current NAS families run 20 to 34 dBA idle and 26 to 39 dBA seeking - WD Red Pro 20-34 and 31-39, Red Plus 20-34 and 26-39, IronWolf Pro 28 and 32. The spread within a single family is larger than the effect of drive count, and inside Red Pro the loud end is the air-filled part numbers: WD122KFBX, WD103KFBX and WD102KFBX are all 34 dBA idle against 20 for the quietest helium members.

Then remember the ASHRAE result from earlier: the fans are probably louder than the drives, and their noise is the thing degrading throughput.

Spin-up sizes the 12 V rail, not running load

Datasheet peak 12 V current is 1.2 to 2.1 A per drive - WD Red Pro 1.7 to 2.08 A, Red Plus 1.2 to 1.9 A, IronWolf Pro around 2.0 A typical at startup:

at 2.0 A per drive on the 12 V rail:
   4 drives  ->  96 W surge
   6 drives  -> 144 W surge
   8 drives  -> 192 W surge
  12 drives  -> 288 W surge

against idle, at 5.5 W per drive:
   8 drives  ->  44 W

Eight drives spinning up simultaneously draw roughly four times their idle power, all on the 12 V rail and all in the first few seconds. That is why NAS chassis and HBAs implement staggered spin-up, and why a power supply sized on running load browns out at boot. If your array fails to come up after a power cut but works when you start drives by hand, this is why.

Idle power sets the bill, and the spread is 2.5x

A NAS spends nearly all its life idle, so the idle figure is the one that multiplies by 8,760 hours:

Drive Idle Read/write Standby
WD Red Pro 26 TB 3.6 W 6.0 W 0.3-1.6 W
WD Red Plus 2 TB 2.4 W 4.0 W 0.3-1.6 W
WD Red Plus 12 TB (WD120EFBX, helium) 2.9 W 6.3 W 0.3-1.6 W
WD Red Plus 12 TB (WD120EFGX, air) 6.1 W 8.8 W 0.3-1.6 W
Seagate IronWolf Pro 20 TB 5.5 W 7.7 W 0.3-1.6 W

Look at the two 12 TB WD Red Plus rows. Same family, same capacity, 2.9 W against 6.1 W idle - the helium one uses less than half the power of the air one. Over six drives running continuously:

6 x 2.9 W = 17.4 W  ->  152 kWh/year
6 x 6.1 W = 36.6 W  ->  321 kWh/year
difference             169 kWh/year, from a part-number suffix

The 26 TB Red Pro at 3.6 W idle draws less than the 12 TB air-filled Red Plus at 6.1 W. Bigger drives do not necessarily cost more to run, and buying the newer generation is an argument on power grounds alone.

Temperature

Seagate’s IronWolf Pro datasheet gives 65 °C as the maximum drive-reported operating temperature for 6 to 24 TB, and 60 °C on the 28 to 32 TB models, and then advises against using it: “Seagate does not recommend operating at sustained drive temperatures above 60C.” Set against Google’s finding that lower temperatures correlated with higher failure rates across most of the observed range, the rule is: keep drives below 60 °C and stop there. Chasing 30 °C costs fan speed, fan speed costs acoustic energy to the fifth power, and acoustic energy costs throughput you cannot measure.

Capacity planning, and the cost of expanding later

What you see is not what is on the box, for two reasons that compound:

6 x "12 TB" in RAID6
  raw       72 x 10^12 bytes
  usable    48 x 10^12 bytes         (4 data drives)
  reported  48e12 / 2^40 = 43.7 TiB  (binary units)
  at 80%    35.0 TiB of actual room

Decimal terabytes against binary tebibytes is the capacity guide’s subject; it costs about 9 per cent before the array takes its share. As for that 80 per cent, do not plan to run past it, for three reasons that stack:

  • Zoned bit recording. On a 3.5-inch platter the outer data radius is around 46 mm and the inner around 23 mm, and that 2:1 geometric ratio is close to the 2:1 ratio between a drive’s best and worst sustained rates. Whatever lands in the last fifth of the pool lives on the slow half of the platter.
  • Allocator behaviour. Filesystems need contiguous free runs, and copy-on-write filesystems need them even to modify an existing file, never writing in place. ZFS’s documented change at 90 per cent - from performance- based to space-based optimisation - is the hard version of this.
  • Headroom for operations. Snapshots, a reorganisation, and the growth during a multi-day resilver all need space that is not there.

Which makes expansion concrete. Say you need 40 TB now and 80 TB in three years, in an eight-bay chassis:

Plan Now Later Bays left
6 x 12 TB RAID6 48 TB add 2 x 12 TB -> 72 TB 0
4 x 20 TB RAID6 40 TB add 2 -> 80 TB, add 2 -> 120 TB 2
6 x 20 TB RAID6 80 TB nothing needed 2

Bays are the scarce resource, not terabytes. The first plan spends six bays on cheap capacity today and the last two on the same cheap capacity later, and then leaves exactly one move: replace all eight. Growing by replacement means a full rebuild per drive, one at a time:

8 drives x 20 TB at ~100 MB/s  = 55.6 h each
                               = 444.8 h
                               = 18.5 days of degraded or resilvering operation

Eighteen days with no redundancy margin to spare, on drives three years old, with the scrub workload arithmetic above saying each of those resilvers is also a substantial fraction of an annual allowance. That is the price of “buy small now and upgrade later”, paid in risk rather than money. OpenZFS 2.3’s raidz expansion softens it a drive at a time, with the two caveats named earlier, and neither that nor mdadm’s reshape removes the bay problem.

Arguing the other way is price per terabyte, which is rarely flat across the capacity range and moves with the market - a live number rather than a fact about drives, so the 3.5-inch listings answer it better than a guide can.

What to do with this on a listing page

  1. Filter to CMR first, and treat “not stated” as SMR. Start at the hard drive listings and filter to CMR, or exclude the known shingled ones with ?smr=no. This is the one filter where a wrong guess costs you an array rather than an afternoon, because a shingled drive can stall for over ten minutes during a rebuild and no timeout setting rescues it.
  2. Narrow to the class you actually need, not the one with the best badge. NAS-class and enterprise-class listings are separate filters here. Then read the datasheet for the exact part number on the URE rate, the workload figure and the fill gas - all three vary inside a single family, and two part numbers of the same capacity have differed by a factor of ten on URE.
  3. Check the price step between capacities before deciding drive count. Sort by price per terabyte, then work the table above: fewer bigger drives cost more parity and a three-times-longer resilver, more smaller drives cost bays, power and spin-up surge. There is no universally right answer, only the one your chassis and your electricity price pick.
  4. Run smartctl -l scterc on every drive the day it arrives, used or new, before it joins anything. If the answer is Disabled, set it - 70,70 under mdadm or a hardware controller, 1,1 under ZFS if the drive accepts it and 70,70 if it does not - and try the persistent form scterc,70,70,p before writing a boot script. Read the value back every time; below 65 is probably not supported and the write can fail quietly. If the answer is “not supported”, either raise the kernel timeout to 180 seconds or put that drive somewhere else.
  5. Burn it in with the right block size, and diff the SMART output. badblocks -b 4096 -wsv up to 16 TiB, -b 8192 above. Attributes 5, 197 and 198 must be zero before and after. Budget eight to ten days per large drive, and count the eight-times-capacity of traffic against this year’s workload allowance.
  6. For used drives, do all of the above inside the return window. The used-drive guide covers what a listing withholds and how to read the hours; the enterprise pulls guide covers the HBA, the sector-size trap and why a SAS drive sidesteps the error-recovery argument entirely.
  7. Buy deliberately mixed, and accept that you cannot price the benefit. Two vendors, two retailers, or the same model months apart, date codes checked on arrival. Nobody has quantified what batch diversity buys. It is still the cheapest insurance against the one failure mode parity does not survive.
Affiliate disclosure:We are a member of the eBay Partner Network and earn a commission from qualifying purchases made through links to eBay on this site. Prices and availability are captured periodically and may have changed - the live price is always the one shown on eBay.