Choosing drives for a NAS: what actually differs from a desktop
By Harry Saarinen ·
The mechanics are the same. A NAS drive and a desktop drive of the same generation often share platters, heads and firmware base. What differs is a short list of settings and tolerances, and only one of them will actually break your array: how long the drive is willing to sit there retrying a bad sector before it answers.
Everything else - vibration tolerance, workload rating, warranty - shifts probabilities. Error recovery time changes behaviour. A desktop drive in an array is not slightly worse; it fails differently, and the failure looks like a healthy drive being ejected.
The second uncomfortable thing goes near the top too, because most buying advice depends on it being false. The ladder that matters runs inside a family, not between the two vendors. Western Digital’s January 2025 WD Red Pro datasheet rates the whole family, 2 TB through 26 TB, at 550 TB per year and calls it suitable for “RAID-optimized NAS systems with unlimited # of bays”. Seagate’s current IronWolf Pro datasheets rate its equivalent identically - 550 TB per year, “Drive Bays Supported: Unlimited” - so the Pro tier has converged. Any article still contrasting the two on those figures is quoting a Seagate datasheet that has since been revised upward.
What has not converged is the ladder inside a single family. WD Red Plus is 180 TB/year against Red Pro’s 550. Two WD Red Pro part numbers of the same 14 TB capacity differ by a factor of ten on unrecoverable read error rate. Two WD Red Plus part numbers of the same 12 TB capacity differ by more than two to one on idle power, because one is helium-filled and one is not. The part number is the specification. The family name is a shelf.
Error recovery time is the one specification that changes behaviour
A drive that cannot read a sector does not give up. It repositions the head, re-reads, applies progressively heavier correction, tries deliberate off-track offsets in both directions, and re-reads again. On a desktop that is exactly right: the drive is the only copy of the file, and a sector recovered on the fortieth attempt is a photograph you did not lose.
The Linux kernel RAID wiki puts the desktop-drive retry ceiling at over two minutes. No vendor publishes the retry algorithm, the number of passes or a hard limit, so treat two minutes as an observed order of magnitude, not a specification.
Now put that drive in an array. The controller asked for a sector, has had no answer, concludes the drive is dead, and drops it. You are now rebuilding - a full-surface read of every surviving drive, for days at current capacities - over one slow sector that was about to be delivered.
The timeout mismatch, exactly
| Layer | Timeout | Where it is set |
|---|---|---|
| Hardware RAID controllers | Commonly around 8 seconds | Controller firmware, often undocumented |
| Linux SCSI/libata command timeout | 30 seconds | /sys/block/sdX/device/timeout |
| ZFS on Linux | Inherits the 30-second kernel timeout | As above |
| Drive with ERC enabled | 7 seconds typical factory default | SCT Error Recovery Control |
| Drive with ERC disabled | Over 120 seconds | Not settable |
| Drive-managed SMR, error path | Over 10 minutes | Not settable, not bounded |
The kernel RAID wiki names the consequence in one sentence: when the kernel gives up first, “the raid code assumes the drive is dead and kicks it from the array”. The drive was working. The timeout killed it.
The fix is to make the drive give up before the layer above it does. The ATA feature is SCT Error Recovery Control; vendors rename it - Western Digital TLER (time-limited error recovery), Seagate ERC (error recovery control), Samsung and the old HGST lines CCTL (command completion time limit). They are the same SCT command set underneath.
$ sudo smartctl -l scterc /dev/sda
SCT Error Recovery Control:
Read: 70 (7.0 seconds)
Write: 70 (7.0 seconds)
$ sudo smartctl -l scterc,70,70 /dev/sdb # read limit, write limit
The units are deciseconds, so 70 is seven seconds - deliberately just under
the eight-second controller timeout. The smartctl manual page adds the floor most
guides omit: 0 disables the feature entirely, and “other values less than 65 are
probably not supported”. Read: Disabled means the drive has the feature but is
sitting in desktop behaviour; SCT Error Recovery Control command not supported
means it lacks it.
The persistence advice in most guides is now out of date
The standing folklore is that the setting never survives a power cycle. That was true, and it stopped being true. From smartmontools 7.3, the manual page documents a persistent form:
# set, persistent across power-on reset
$ sudo smartctl -l scterc,70,70,p /dev/sdb
# read back the persistent values
$ sudo smartctl -l scterc,p /dev/sdb
# restore the manufacturer's defaults
$ sudo smartctl -l scterc,reset /dev/sdb
The manual is explicit that “the ,p and ,reset options require the device to
support ATA ACS-4 or higher”, so an older drive will refuse them and you are back
to re-applying at boot. Try the persistent form first and verify it survives a
cold power cycle, not a reboot - a warm reboot may not drop the drive’s power
rail, which makes a non-persistent setting look persistent.
Where ,p is unavailable, the correct mechanism is a udev rule or a systemd unit
that runs on every boot, and the correct verification is reading the value back
afterwards rather than assuming the write took.
ZFS wants a far shorter limit than RAID does, and says so
The seven-second convention comes from hardware RAID controllers, and applying it under ZFS is inherited habit rather than reasoning. The OpenZFS hardware guidance recommends something an order of magnitude tighter:
it is advisable to write a script to set the error recovery time to a low value, such as 0.1 seconds until ZFS is modified to control it. This must be done on every boot.
That is smartctl -l scterc,1,1. The logic is that ZFS does not need the drive’s
recovery attempt at all: it has a checksum, it knows the block is wrong, and it
can reconstruct from parity or a mirror immediately. Every second the drive
spends retrying is a second ZFS spends waiting for an answer it is about to
discard. The same document notes that drives with the feature “typically default
to 7 seconds”, so the factory value is tuned for somebody else’s array.
Expect most drives to refuse it. The smartctl floor quoted above applies here
too: values below 65 are probably not supported. So set it, read it straight back
with smartctl -l scterc, and fall back to 70,70 if the drive did not take the
value. OpenZFS’s recommendation is what you would want rather than what most SATA
drives will give you, and a setting you have not read back is a setting you do not
have.
| Stack | Recommended ERC | Why |
|---|---|---|
| Hardware RAID controller | 70,70 (7.0 s) | Just under the controller’s own ~8 s timeout |
Linux mdadm |
70,70 (7.0 s) | Named as typical in the smartctl manual page |
| OpenZFS | 1,1 (0.1 s) if the drive accepts it, else 70,70 | Parity reconstruction is faster than drive retry |
| SAS, any stack | Mode page 0x01 | Ships at array-appropriate values |
When the drive cannot do it at all
Move the other timeout. The kernel RAID wiki gives the fallback directly:
# raise the kernel's per-device command timeout past the drive's retry ceiling
for d in /sys/block/sd*/device/timeout; do echo 180 > "$d"; done
This must also be re-applied at boot. It is the worse fix - a genuinely dead drive now hangs the array for three minutes instead of thirty seconds - but it beats losing a member. The asymmetry that matters: software RAID’s timeout is a kernel setting you own, and a hardware controller’s is not. On a hardware controller you cannot raise the timeout, so if the drive cannot lower its own, there is no fix at all. That, rather than anything about the drives, is why hardware RAID and desktop drives pair so badly.
OpenZFS records which way the market moved, and names a vendor: “Since ZFS’
introduction, error recovery control has been removed from low-end drives from
certain manufacturers, most notably Western Digital.” The historical pattern is
worth knowing because it repeats. WD’s desktop models of the era - Caviar SE,
SE16, GP, Raptor - shipped with TLER read and write configured as 0, disabled;
the RAID-class models, Caviar RE2 and RE2-GP, shipped at 7 seconds on both. WD
later withdrew the ability to change it on desktop drives and warned that the old
WDTLER.EXE utility could brick newer firmware. Firmware giveth and firmware
taketh away, and the marketing name is not the specification.
Which current drives allow it is a per-model fact, not a per-brand one. The NAS and enterprise lines - WD Red Plus, Red Pro, Gold and Ultrastar; Seagate IronWolf, IronWolf Pro and Exos; Toshiba N300 and MG - ship with it enabled or settable, and Seagate’s IronWolf Pro datasheet lists “Time-Limited Error Recovery” explicitly in its AgileArray bundle alongside dual-plane balancing. Consumer lines vary. Check the drive in front of you, not the line it is in.
SAS puts this in a mode page instead
A SAS drive sidesteps the whole argument. Its recovery limits live in the SCSI Read-Write Error Recovery mode page (0x01), which exposes a Recovery Time Limit field, and SAS drives generally ship at array-appropriate values because there is no desktop market to tune them for.
sudo sdparm --get=RTL /dev/sdX # read the recovery time limit
sudo sdparm --set=RTL=3000 /dev/sdX # set it, value in milliseconds
Seagate additionally ships openSeaChest for its own drives, which reaches
settings sdparm does not. This is one of several reasons a SAS pull is an
easier array member than a shucked SATA drive - see
enterprise drive pulls for the rest.
SMR is not a slow sector, it is an unbounded one
The same kernel wiki that gives the two-minute figure for a desktop CMR drive gives a worse one for shingled media: an SMR drive can stall for over ten minutes during a read error. That is beyond any achievable kernel timeout and beyond the 180-second workaround. The mechanism is in the CMR/SMR guide; the conclusion here is that ERC tuning does not rescue a shingled drive, because the number you would have to tune to is not bounded.
Rotational vibration sensors measure less than their marketing implies
Vendors used to rate NAS drives by bay count, and the tiering has come apart. Western Digital’s current Red Pro datasheet claims “unlimited # of bays” and Seagate’s current IronWolf Pro datasheets print “Drive Bays Supported: Unlimited”; only WD Red Plus still carries a number, “up to 8 bays”. Both vendors have dropped the ceiling on their Pro lines, and neither has published the test behind either the old number or the new one. A limit that can go from 24 bays to unlimited without a matching change in the mechanical specification was a marketing tier rather than a measurement. Understand the mechanism instead.
The error budget, derived from geometry
Track pitch is not published per product, so it has to be derived. These numbers are mine, not a vendor’s, and the assumptions are all visible so you can disagree with any of them:
assume a 24 TB drive: 10 platters, 20 recording surfaces (helium-class)
assume usable annulus: outer radius 46 mm, inner radius 23 mm
assume bit aspect ratio BPI : TPI = 5 : 1
per-surface capacity 24e12 x 8 / 20 = 9.60e12 bits
annulus area pi x (46^2 - 23^2) mm^2 = 4986 mm^2 = 7.73 in^2
areal density 9.60e12 / 7.73 = 1.24e12 bits/in^2
TPI sqrt(1.24e12 / 5) = ~498,000 tracks per inch
track pitch 25.4e6 nm / 498,000 = ~51 nanometres
off-track write budget, roughly 10% of pitch = ~5.1 nanometres
The head has about five nanometres of position error budget while writing. The same arithmetic at a 32 TB, 22-surface drive gives about 4.6 nm. The exact figure is uncertain - platter geometry, surface count and aspect ratio are all assumptions, and vendors publish none of them per product - but the order of magnitude is single-digit nanometres. That is what the servo defends, and it is why a chassis that hums loses throughput.
Why it is rotational vibration specifically
Actuator arms are mass-balanced about their pivot, so linear vibration produces no net torque there and the mechanical design already rejects it. Rotational vibration is angular acceleration about an axis parallel to the spindle, measured in rad/s², and it acts directly on the arm’s moment of inertia. No balancing trick cancels it.
It comes from Newton’s third law. When a drive seeks, its actuator accelerates one way and its baseplate is pushed the other, and that reaction torque goes into the chassis and on to every drive bolted to the same metal. Eight drives seeking independently are eight uncorrelated torque sources exciting one structure. Add spindle imbalance at the rotation frequency - 120 Hz at 7200 rpm, 90 Hz at 5400 - and drives whose actual speeds differ by fractions of a per cent beat against each other at a few hertz, exactly where a sheet metal chassis damps worst.
What the sensor pair actually is
Two accelerometers on the printed circuit board assembly, and what matters is what is done with their outputs rather than what they are made of. Each one measures linear acceleration in the plane of the board. Subtract one from the other and the common-mode component - the whole drive being shaken bodily in that plane - cancels, leaving a difference proportional to rotational acceleration about the spindle axis. That difference is the quantity the servo wants, and no single accelerometer placed anywhere on the board could have produced it.
A note on a distinction that is not one, because it recurs in every discussion of this feature. “Piezoelectric” names the transduction element; “linear” names the axis measured. A piezoelectric accelerometer is a linear accelerometer, and the ASHRAE paper quoted further down calls the same sensors piezoelectric while describing exactly the arrangement above. Neither word tells you the thing you want to know, which is the separation distance, the bandwidth of the feedforward path and the rejection actually achieved - and no vendor publishes any of the three per product.
The difference signal is digitised and fed forward into the servo - added to the normal feedback effort, not replacing it - so the correcting current is applied before the position error signal has registered the disturbance. That is the entire value: feedback reacts after the head has moved, feedforward reacts before.
A drive without the pair finds out from its position error signal alone, by which point it must inhibit the write and wait a full revolution.
The missed revolution, and what it actually costs
ASHRAE Technical Committee 9.9, in a 2019 paper whose authors include Seagate, Dell, IBM and Intel engineers, describes the mechanism in the vendors’ own words:
the servo system suspends the write process until the head returns to an acceptable range to the track center, and the system must wait for at least one revolution of the disk to attempt again to write. This time delay increases latency and results in system performance degradation.
7200 rpm -> 60 / 7200 = 8.33 ms per revolution
5400 rpm -> 60 / 5400 = 11.11 ms per revolution
theoretical average access, 7200 rpm:
rotational latency 4.17 ms + seek 4 to 9 ms = 8 to 13 ms
empirical average access, 7200 rpm (OpenZFS) = 13 to 16 ms
rated average access, 15k rpm (OpenZFS) = 3.4 ms read, 3.9 ms write
This guide previously said a missed revolution roughly doubles the cost of an operation. Against the theoretical figure that is nearly true; against the empirical one it is not. OpenZFS’s hardware documentation gives 13 to 16 ms as the real measured average access time for 7200 rpm drives. A lost 8.33 ms revolution on that is a 52 to 64 per cent increase, not a doubling.
And nothing reports it. No SMART attribute increments, no error is logged, no counter moves. The array is simply slower than it was on the bench, with no artefact to point at. That is the most important property of this failure mode: it does not look like a failure.
The sensors are structurally blind to most of what now matters
This is the part that almost no buying guide mentions, and ASHRAE states it plainly:
Successful rotational vibration (RV) feedforward systems measure the rigid body motion of the HDD using two piezoelectric sensors on the printed circuit board assembly (PCBA). These sensors only measure in-plane motion… They do not measure rotational motion around x-axis or y-axis or z-axis linear motion and cannot accurately measure nonrigid body motion.
| Disturbance | Sensed? |
|---|---|
| In-plane linear motion (x, y) | Yes, and cancelled as common mode |
| Rotation about z (the spindle axis) | Yes - this is the whole feature |
| Rotation about x or y (chassis rocking) | No |
| Linear motion along z | No |
| Non-rigid-body motion (panel flexing, resonance) | Not accurately |
A NAS enclosure whose side panel resonates, or whose drive cage rocks about a horizontal axis, produces disturbance the RV feedforward system cannot see and therefore cannot correct. An “RV sensor” tick on a datasheet buys rejection of one of five disturbance modes.
Worse, there is a hard bandwidth ceiling. ASHRAE: “The current sets of technologies and mechanics allow for error rejection below 2 kHz, but disturbances from AMDs can input excitations up to 20 kHz”, and “Present HDDs show sensitivity up to ~20 kHz”. AMD there means air-moving device: a fan or blower. The servo rejects below 2 kHz and the drive is sensitive to 20 kHz. That gap is not a tuning problem to be fixed in a firmware revision; it is a property of the mechanical loop.
Acoustics, not neighbouring drives, is the dominant disturbance now
The classic story is drive-to-drive seek coupling. In a dense modern chassis the fans have overtaken it, and the scaling is brutal:
fan sound power level Lw1 - Lw2 = 50 x log10(RPM1 / RPM2) dB
blower sound power level ~70 x log10(RPM1 / RPM2) dB
structural vibration amplitude proportional to RPM^2
blade pass frequency fb = Nb x RPM / 60
so raising fan speed by 25%:
fan Lw +50 x log10(1.25) = +4.8 dB
blower Lw +70 x log10(1.25) = +6.8 dB
vibration amp 1.25^2 = 1.56x
and raising it by 50%:
fan Lw +8.8 dB, blower +12.3 dB, vibration 2.25x
50 x log10(R) is a power-level relation, so acoustic power goes as R^5
and sound pressure amplitude as R^2.5
Acoustic power tracking the fifth power of fan speed is why a NAS that was fine in winter loses throughput in a warm room. The thermal control loop raises fan RPM by a modest fraction, the acoustic energy reaching the drives jumps by several decibels, and the drives start inhibiting writes. Nothing in the monitoring will say so.
Blade pass frequency is worth computing for your own chassis: a seven-blade fan at 1200 rpm excites at 140 Hz, the same fan at 3000 rpm at 350 Hz, a nine-blade high-static-pressure fan at 6000 rpm at 900 Hz. All below 2 kHz, which is the good news; the harmonics are not.
Helium drives are quieter inside and more fragile outside
ASHRAE again:
Although the transition to using helium inside HDDs results in an environment with very low airflow disturbance to the head, the designs become more susceptible to external vibration, especially vibration caused by coupling with acoustic sources.
The mechanism is counter-intuitive. Air inside a sealed drive causes windage - turbulent flow that buffets the arm and the platters. Helium, at about one seventh the density, nearly removes it. With that internal disturbance gone, external excitation is no longer masked by it, and the same absolute disturbance becomes a larger fraction of the error budget. Helium designs also carry more platters, which means thinner platters and thinner arms, and lower mechanical stiffness.
So the highest-capacity drive on your shortlist is probably helium, probably more sensitive to chassis acoustics than the air drive it replaces, and the datasheet will not say so.
The datasheet figure is a design target, not a measurement
Seagate’s IronWolf Pro datasheet publishes rotational vibration as 12.5 at 10 to 1500 Hz - not 0 to 1500 Hz, which matters because the sub-10 Hz region where chassis rocking lives is excluded from the test. It prints the unit as “rad/s”, although the physical quantity is angular acceleration in rad/s².
The revealing detail is that the value is identical - 12.5 - at 4, 6, 8, 10, 12, 14, 16, 18 and 20 TB. A number that does not move across five generations of areal density and two fill gases is a design target the family is qualified against, not a per-model measurement. Treat it as “this family passed the internal RV spec” and nothing more precise.
And then ASHRAE closes the door on the number you actually want:
Because there are a myriad of AMDs (and resulting frequency characteristics) and chassis designs, there is no specific reference speed for an AMD that may be used to determine the HDD performance in an enclosure.
There is no published figure that predicts your chassis’s throughput loss, and there cannot be, because the answer is a property of the enclosure, the fan curve, the mounting hardware and the drive together. The paper also notes that measuring throughput alone will not identify the mechanism - you see the loss and cannot attribute it.
What you can do is bound it empirically. Measure sustained sequential throughput with the array’s fans pinned to minimum, then again pinned to maximum, on otherwise identical work. If the numbers differ materially, acoustic coupling is costing you something, and rubber-grommet mounting, a slower larger fan, or damping the panel is worth trying. In a four-bay plastic enclosure the effect is usually small enough to ignore. By eight bays in a metal chassis it is real, and it arrives as slowness rather than as an error.
Workload rating is a warranty condition quoted at an operating point you are not at
Workload rating is published in terabytes per year, and it is commonly described as covering host traffic plus the drive’s own background media scans. That is wrong, and the vendors are specific. Western Digital’s Red Plus datasheet footnote defines workload rate as “the amount of user data transferred to or from the hard drive”, and the Red Pro footnote gives the annualisation formula outright:
Annualized Workload Rate = TB transferred x (8760 / recorded power-on hours)
Host user data only. The drive’s own background media scan, its idle reorganisation, its read-after-write verification - none of that appears in the vendor definition. Whether it counts against the physical wear the rating is a proxy for is a separate question, and an unanswerable one from public documents.
The tier table, and where the step actually falls
| Tier | Workload rating | Rated power-on hours | Source |
|---|---|---|---|
| Desktop | ~55 TB/year | ~2400 hours/year (light-duty rating) | Consumer datasheets, where they state it at all |
| NAS | 180 TB/year | 8760 hours/year | WD Red Plus; Toshiba N300 |
| NAS Pro | 550 TB/year, both vendors | 8760 hours/year | WD Red Pro and Seagate IronWolf Pro datasheets |
| Enterprise nearline | 550 TB/year | 8760 hours/year | Seagate Exos X24 product manual; Toshiba MG11 datasheet |
The bottom two rows are the correction, and they are the same number. The NAS Pro tier and the enterprise nearline tier have the same published workload rating on both vendors’ current datasheets, so the workload figure is not what you are paying the enterprise premium for - see enterprise drives for what is.
The step that is real is the one from 180 to 550, between Red Plus and Red Pro inside a single vendor’s range. That is a factor of three on the warranty condition for a difference in shelf position most buyers read as a badge.
What exceeding it does is published - just not as a curve
“Nobody publishes what happens when you exceed it” is the usual claim, and it is too strong. WD publishes the mechanism in its Red Pro MTBF footnote: MTBF is estimated at a workload of 220 TB/year and 40 °C, and “Derating of MTBF will occur above these parameters, up to 550TB writes per year”.
Read that twice, because it reframes both numbers on the front of the datasheet.
WD Red Pro
headline MTBF 2,500,000 hours (14-26 TB) / 2,000,000 hours (2-12 TB)
quoted at 220 TB/year, 40 °C
workload rating 550 TB/year
between them MTBF derates, curve not published
WD Red Plus
headline MTBF 1,000,000 hours
quoted at 90 TB/year, 40 °C
workload rating 180 TB/year
between them MTBF derates, up to 65 °C drive temperature
The headline reliability figure and the headline workload figure are quoted at different operating points, and the reliability one sits at 40 to 50 per cent of the workload one - 220 against 550 on Red Pro, 90 against 180 on Red Plus. A drive run at its full rated workload does not have the MTBF on the front of its datasheet. What it does have is not published; only the fact that it is lower. That gives you a target: design for 220 TB/year on Red Pro and 90 TB/year on Red Plus if you want the advertised MTBF, not for the ceiling.
It is still a warranty condition, not a wear counter
An SSD’s TBW maps to a physical mechanism - oxide degradation per program/erase cycle, one countable event per block - which is why SSD endurance arithmetic works. A hard drive has no equivalent counter. It wears through head-disk interface contact, lubricant depletion, head-media spacing degradation and bearing wear, and those track seeking, flying hours and thermal cycling far more than bytes moved. The workload rating is a proxy the vendor chose because it is measurable from the host side, not because bytes cause the wear.
The hours line binds sooner for most people anyway: a desktop drive rated for 2400 hours a year, run 24x7, spends its annual allowance in 100 days.
As a continuous rate, these ratings are small
55 TB/year / 31,536,000 s = 1.7 MB/s sustained, continuously
180 TB/year / 31,536,000 s = 5.7 MB/s sustained, continuously
220 TB/year / 31,536,000 s = 7.0 MB/s sustained, continuously
550 TB/year / 31,536,000 s = 17.4 MB/s sustained, continuously
saturated 1 GbE (125 MB/s) = 3,942 TB/year
saturated 2.5 GbE (312 MB/s) = 9,839 TB/year
saturated 10 GbE (1250 MB/s) = 39,420 TB/year
A saturated gigabit link is about twenty-two times a 180 TB/year rating and seven times a 550 TB/year one. Nobody saturates gigabit continuously, which is why these ratings rarely bind on household traffic: a family writing two terabytes of photos and backups a year sits at roughly one per cent of a NAS drive’s budget. Even a 10 GbE link only matters if you keep it busy, and the thing that keeps a home NAS busy is not the humans.
Scrubbing is what blows through it, and mdadm is far worse than ZFS
This is where the arithmetic changes materially depending on your filesystem, by a factor of 1.6 on the worked example below.
Six 12 TB drives in RAID6, 48 TB usable, 30 TB in use. ZFS scrubs allocated blocks only, so each drive reads roughly its own share of what is stored:
ZFS, pool 62% full: ~7.5 TB read per drive per scrub
monthly 7.5 x 12 = 90 TB/year per drive
weekly 7.5 x 52 = 390 TB/year per drive
mdadm’s check is allocation-blind. It reads every sector of every member
regardless of whether anything is stored there, because it is comparing raw parity
against raw data and has no idea what the filesystem above it considers used:
mdadm, same array: 12 TB read per drive per check
monthly 12 x 12 = 144 TB/year per drive
weekly 12 x 52 = 624 TB/year per drive
A weekly mdadm check on 12 TB drives is 624 TB/year of scrub traffic alone -
more than a 550 TB/year enterprise rating, before the array has served a single
file. Monthly is 144 TB/year, which already consumes 80 per cent of a 180
TB/year WD Red Plus allowance and 65 per cent of the 220 TB/year point where WD
quotes Red Pro’s MTBF.
And the schedule is set for you either way, which is why this section exists.
Debian ships mdadm’s checkarray on a cron entry that fires on the first
Sunday of the month at 00:57. Red Hat and Fedora do not ship checkarray at
all. They ship raid-check, driven by raid-check.timer, whose unit
file reads OnCalendar=Sun *-*-* 01:00:00 and describes itself as a “Weekly RAID
setup health check”. TrueNAS schedules a monthly ZFS scrub per pool by default.
Debian checkarray first Sunday of the month, 00:57 -> 144 TB/yr per 12 TB member
Red Hat raid-check every Sunday, 01:00 -> 624 TB/yr per 12 TB member
TrueNAS ZFS scrub monthly per pool -> 90 TB/yr per member at 62% full
On a Red Hat or Fedora mdadm array the shipped default is weekly, and on 12 TB
members that is the 624 TB/year figure above - more than an enterprise workload
rating, from a timer nobody chose. It is the single most expensive default in this
guide, and the one most likely to be running right now on a machine whose owner
has never read the unit file.
The trade cuts both ways: scrubbing is how you find latent sector errors while
redundancy still exists to repair them from, and the evidence for that is the
strongest single finding in the field literature - see the URE section below. So
do not turn it off. Monthly is defensible on any of the three. Weekly on large
mdadm members is not, unless you have deliberately decided the workload rating
is a number you will exceed - and if you are on Red Hat you have decided it by
default. systemctl edit raid-check.timer and an OnCalendar=Sun *-*-01..07 01:00:00 is the monthly equivalent.
Reliability figures, and what each one is actually measuring
Three numbers get quoted as if they were the same kind of thing. They are not, and the differences decide which ones you can use.
| Figure | What it is | What it is not |
|---|---|---|
| MTBF / MTTF | A model output at a stated operating point | A life expectancy, or a field rate |
| AFR (datasheet) | The same model, expressed per year | A measurement |
| AFR (field) | A measurement, of somebody else’s fleet | Transferable to your drives |
| URE rate | A specification ceiling | A mean, or a distribution |
MTBF is a number with fine print attached
Current published figures, from the vendors’ own documents:
| Drive family | MTBF/MTTF | Published AFR | Quoted at |
|---|---|---|---|
| WD Red Pro 14-26 TB | 2,500,000 h | not published | 220 TB/yr, 40 °C |
| WD Red Pro 2-12 TB | 2,000,000 h | not published | 220 TB/yr, 40 °C |
| WD Red Plus 2-12 TB | 1,000,000 h | not published | 90 TB/yr, 40 °C |
| Seagate IronWolf Pro 12-32 TB | 2,500,000 h | not published | not stated on the datasheet |
| Seagate IronWolf Pro 6-10 TB | 2,000,000 h | not published | not stated on the datasheet |
| Toshiba MG11 (to 24 TB) | 2,500,000 h | 0.35% | not stated |
Five of those six rows say “not published” in the AFR column, and that is the finding rather than an omission. An AFR obtained by dividing 8,760 by the MTBF is not a published AFR; it is the same model figure in different units, and putting it in a column headed “published” is how a derived number acquires an authority it never had. Toshiba prints both, and they agree to the digit - 8,760 / 2,500,000 = 0.35 per cent - which tells you what a published AFR usually is: the MTBF, restated.
A 2.5 million hour MTBF is 285 years, which nobody claims is a life expectancy. It is the inverse of a modelled constant hazard rate during the flat part of the bathtub curve: “if you ran a large population for a year, this fraction would fail”, not “this drive will last that long”.
The field measurements disagree with the model, consistently
Schroeder and Gibson’s FAST 2007 study analysed about 100,000 disks across large production installations and found field replacement rates of 2 to 4 per cent typical, against datasheet MTTFs of 1.0 to 1.5 million hours - a nominal AFR of at most 0.88 per cent. Some populations ran at 13 per cent. Two structural findings from that paper matter more than the headline gap:
- Time between replacements fits a Weibull distribution with a decreasing hazard rate, not the constant hazard rate that MTBF arithmetic assumes. The model underlying every MTBF figure is the wrong model.
- Replacements showed autocorrelation significant at lags up to 30 weeks. Failures arrive in clusters, months apart. Independence is not a safe assumption, and every rebuild-risk calculation in the next section assumes it.
Backblaze’s Q1 2026 Drive Stats, the largest continuously published fleet data there is, gives the current picture, and its 2025 annual report the last full year:
drives in service 341,263
drive-days in quarter 30,203,180
failures in quarter 1,030
quarterly AFR 1.24%
lifetime AFR 1.39% (529,968,464 drive-days, 20,212 failures)
20 TB+ cohort AFR 0.85% across the quarter
2025 full year 1.36% (down from 1.55% in 2024)
worst single model 4.90% (1,251 drives of one Seagate 14 TB model)
Two things to take from it. First, the fleet-wide figure and the worst-model figure differ by a factor of four, so a fleet average is not a prediction about a specific model. Second, the 20 TB+ cohort at 0.85 per cent is the best number in the table, contradicting the intuition that bigger drives must be less reliable - though those drives are also the newest in the fleet, and furthest from the wear-out end of the bathtub.
The blind spot in every published failure statistic
Backblaze names its own methodology limit, and it is exactly the limit that matters to a buyer:
Because it defines a failure based on a drive’s presence in the pool the day before, that means we can’t define day one failures using only the Drive Stats program.
They put it more bluntly in the same post: “Have we been completely under-reporting day one failures? Short answer: yes, probably.”
Day-one failures are invisible in the statistics. A drive that arrives dead, or that fails in its first week, never enters the dataset as a failure - it was never present to go missing. That is precisely the class of failure burn-in testing exists to catch, and it is precisely the class the published AFR figures cannot tell you about. Do not read a 0.85 per cent AFR as your probability of a bad drive out of the box; read it as your probability of losing a drive that survived its first week in a datacentre.
SMART is a strong signal and a poor predictor
Google’s FAST 2007 disk failure study is still the reference, and its two useful findings point in opposite directions:
- A drive was 39 times more likely to fail within 60 days of its first scan error. When SMART says something, listen.
- A large fraction of failed drives showed no SMART errors at all. A clean SMART report is not evidence of health.
The same study found something that contradicts the standard advice: over most of the observed temperature range, lower temperatures correlated with higher failure rates, reversing only at the extreme top end. The correlation is probably confounded - cooler drives in that fleet were in different roles with different workloads - but the plain reading is that chasing the last few degrees is not where your effort belongs. The heat section below gives the ceiling that does matter.
The 10^14 argument, done with the correct probability
Consumer datasheets specify an unrecoverable read error rate of less than 1 sector per 10^14 bits read. NAS and enterprise drives are usually 10^15, and some SAS drives 10^16.
The first correction is that the tiers are not as clean as the marketing. The entire WD Red Plus family, 2 TB through 12 TB, is specified at less than 1 in 10^14 - the same as a desktop drive, on a NAS-branded product. And within WD Red Pro, the 14 TB WD141KFGX is specified at 10^14 while WD142KFGX, also 14 TB, is specified at 10^15. Two part numbers, the same capacity, the same family, a factor of ten apart on the specification that the entire RAID 5 argument turns on. Read the datasheet for the part number, not the family.
The common arithmetic is wrong in a nameable way
The version of this argument you will meet online divides bits read by the URE rate and calls the result a probability:
16 TB read x 8 = 1.28e14 bits
1.28e14 x 1e-14 = 1.28 -> "128% chance of a URE"
A probability of 128 per cent should have ended the argument on the spot. The division gives the expected number of UREs, not the probability of at least one. The correct form treats each read as an independent Bernoulli trial, which in the limit is Poisson:
P(at least one URE) = 1 - e^(-N x p)
N = bits read, p = per-bit error rate
16 TB: 1 - e^(-1.28e14 x 1e-14) = 1 - e^(-1.28) = 72.2%
72.2 per cent, not 128 per cent. The distinction gets larger as the naive number grows: at 96 TB read the naive answer is 768 per cent and the correct one is 99.95 per cent.
The sector model agrees, which answers the usual objection
The standard attack on the bit model is that bits do not fail independently - errors happen at sector granularity, so the framing is wrong. Run it at sector granularity and the answer does not move:
p per 4096-byte physical sector = 1e-14 x 4096 x 8 = 3.28e-10
sectors in 16 TB = 16e12 / 4096 = 3.906e9
P(at least one) = 1 - (1 - 3.28e-10)^3.906e9 = 72.196%
Poisson form = 1 - e^(-3.28e-10 x 3.906e9) = 72.196%
Identical to three significant figures. The granularity objection is correct in principle and changes nothing numerically, because the expected count is the same either way and both distributions have the same mean.
The table, extended to current capacities and to 10^16
Five-drive RAID5, rebuild reads four drives in full. These are my calculations, from the published specification ceiling, not measurements:
| Drive size | Bytes read | At 10^-14 | At 10^-15 | At 10^-16 |
|---|---|---|---|---|
| 2 TB | 8 TB | 47.3% | 6.2% | 0.6% |
| 4 TB | 16 TB | 72.2% | 12.0% | 1.3% |
| 8 TB | 32 TB | 92.3% | 22.6% | 2.5% |
| 12 TB | 48 TB | 97.9% | 31.9% | 3.8% |
| 16 TB | 64 TB | 99.4% | 40.1% | 5.0% |
| 20 TB | 80 TB | 99.8% | 47.3% | 6.2% |
| 24 TB | 96 TB | 99.95% | 53.6% | 7.4% |
| 30 TB | 120 TB | 99.99% | 61.7% | 9.2% |
Read down the columns rather than across the rows. The specification tier dominates the capacity. A 24 TB drive at 10^16 is safer on this metric than a 2 TB drive at 10^14 by a factor of six, while reading twelve times as much data to get there. Hold the capacity still and move only the tier - 2 TB at 10^14 against 2 TB at 10^16 - and the factor is roughly seventy-five. Capacity growth is not what broke RAID 5; it is what made the difference between tiers unignorable. A useful way to hold it:
bytes read per expected URE
1 in 10^14 = 12.5 TB
1 in 10^15 = 125 TB
1 in 10^16 = 1,250 TB
Four reasons the percentages overstate the risk
The honest position is that those percentages are too pessimistic while the trend they describe is not in dispute.
- The spec is a ceiling, not a mean. Datasheets say “less than 1 in 10^14”. Field rates appear better, but no vendor publishes the distribution and no independent measurement exists at scale. You are computing with an upper bound and reporting the answer as a point estimate.
- The empirical test fails. If 72 per cent per rebuild were real, five-drive 4 TB RAID5 arrays would almost never complete one. They routinely do.
- A URE is not automatically an array loss. This is the biggest error in the whole argument, and it gets its own subsection below.
- Errors are not independent. Bairavasundaram and colleagues at NetApp studied latent sector errors across 1.53 million drives in over 50,000 arrays over 32 months. The quoted aggregate is 3.45 per cent, but the split is what a NAS buyer needs: 8.5 per cent of consumer-class disks against 1.9 per cent of enterprise-class, a factor of 4.5; in the first year alone, 3.15 per cent against 1.46 per cent. And 0.2 per cent of the drives that had any error at all had more than 1000 errors - roughly seven drives in every hundred thousand, not two in every thousand. That is where the clustering shows: errors concentrate in a few bad drives rather than spreading evenly. The same study found a drive with one error is much more likely to develop a second, that the second is likely to be near the first, and that the affected fraction rises with capacity and age. Clustering breaks the Poisson model in both directions: most drives are cleaner than the model says, and the bad ones are far worse.
That last study also supplies the strongest empirical argument for scrubbing at all: over 60 per cent of latent sector errors were found by scrubbing. They were sitting there, invisible to any status page, waiting for a rebuild to find them.
The consequence fallacy, which is the real error
The model computes P(URE during rebuild) and then silently equates it with P(array loss). Those are different events, and how different depends entirely on what is managing the array:
| Stack | What a URE during rebuild does |
|---|---|
| Older hardware RAID controller | Aborts the rebuild and drops the array |
Linux mdadm |
Records the block in its bad block list and continues |
| OpenZFS | Names the affected file in zpool status -v and continues |
The kernel documents the mdadm mechanism directly: md maintains a per-device
bad_blocks file exposing “the list of all known bad blocks in the form of start
address and length (in sectors respectively)”, plus unacknowledged_bad_blocks
for entries known but not yet written to disk. That list is precisely what lets a
degraded array survive an unreadable sector rather than aborting.
Those three outcomes differ by orders of magnitude in cost. One loses the array; one loses nothing; one loses a named file you can restore from backup. The “RAID 5 is dead” argument computes a number that is roughly right for the first row and applies it to all three.
Where the argument came from, stated accurately
Robin Harris made the original argument on ZDNet in July 2007, in “Why RAID 5 stops working in 2009”. Adam Leventhal analysed RAID 6 late in 2009, in “Triple-Parity RAID and Beyond” for ACM Queue, and Harris’s StorageMojo follow-up of February 2010 drew on that article. It put a seven-drive RAID5 of 2 TB SATA drives at roughly a 62 per cent chance of data loss on rebuild at 1 in 10^14, assuming about 2.5 hours per TB minimum at 115 MB/s with real times two to five times longer. The 2019 line is Harris’s summary of Leventhal in that post, not a quotation from him: “RAID 6 protection levels will be as good as RAID 5 was until 2019”.
The prediction was directionally right and the industry responded: RAID 6 and raidz2 became the default, and the 10^15 and 10^16 tiers became the things people check. The date has passed and arrays have not stopped working - which is what you would expect from an argument right about direction and wrong about magnitude.
The probability that actually governs array loss
On a modern software RAID, the event that loses the array is not a URE. It is a second whole-drive failure during the rebuild window. Here is that number, derived from Backblaze’s published AFR for its 20 TB+ cohort, assuming independence:
5-drive array, one member lost, 4 surviving members plus the replacement
20 TB at 100 MB/s -> 55.6 h rebuild window
AFR 0.85% -> per-drive hazard over 55.6 h = 1 - (1-0.0085)^(55.6/8760)
P(at least one of 5 fails in the window) = 0.027%
| Scenario | Rebuild window | P(second failure) |
|---|---|---|
| 12 TB, 100 MB/s, AFR 0.85% | 33.3 h | 0.016% |
| 20 TB, 100 MB/s, AFR 0.85% | 55.6 h | 0.027% |
| 24 TB, 100 MB/s, AFR 0.85% | 66.7 h | 0.032% |
| 20 TB, 50 MB/s, AFR 0.85% | 111 h | 0.054% |
| 20 TB, 100 MB/s, AFR 1.39% (fleet lifetime) | 55.6 h | 0.044% |
| 20 TB, 50 MB/s, AFR 3% (Schroeder field range) | 111 h | 0.19% |
Under independence these are small numbers, which is exactly why the correlation evidence matters more than the URE arithmetic. Schroeder and Gibson found autocorrelation out to 30 weeks; array members share a model, a firmware revision, a manufacturing batch, a thermal history, a vibration environment and a workload. The independence assumption in that table is the weakest thing in this guide, and it is the assumption every published rebuild-risk figure makes.
The exact probability cannot be derived from a published upper bound and a violated independence assumption, and this site will not pretend otherwise. But the data that must be read without error scales linearly with capacity while nothing protecting it has improved, so RAID5’s margin gets thinner with every capacity generation - a claim about direction, which does not need the percentage to be right.
Rebuild and resilver arithmetic at current capacities
Rebuild time is capacity divided by sustained throughput. Current 3.5-inch 7200 rpm drives do roughly 250 to 287 MB/s on their outermost tracks and 110 to 130 MB/s on their innermost, because outer tracks are longer and pass under the head faster at the same angular velocity. A whole-surface average near 200 MB/s is what a rebuild gets on an idle array. The published outer-track figures set the top of that range:
| Drive | Max internal transfer rate | Source |
|---|---|---|
| WD Red Pro 24-26 TB | 287 MB/s | WD Red Pro datasheet |
| WD Red Pro 20 TB (WD202KFGX) | 285 MB/s | WD Red Pro datasheet |
| Seagate IronWolf Pro 20 TB | 285 MB/s | IronWolf Pro datasheet |
| WD Red Pro 12 TB (WD121KFBX) | 240 MB/s | WD Red Pro datasheet |
| WD Red Plus 2-12 TB | 180-260 MB/s | WD Red Plus datasheet |
| WD Red Pro 2 TB | 164 MB/s | WD Red Pro datasheet |
Within one current family the sequential rate varies by 75 per cent, so a rebuild-time table that does not say which class it assumes is not telling you much.
| Drive | At 200 MB/s (idle) | At 100 MB/s (in service) | At 50 MB/s (busy) |
|---|---|---|---|
| 4 TB | 5.6 h | 11.1 h | 22.2 h |
| 8 TB | 11.1 h | 22.2 h | 44.4 h |
| 12 TB | 16.7 h | 33.3 h | 66.7 h |
| 16 TB | 22.2 h | 44.4 h | 88.9 h |
| 20 TB | 27.8 h | 55.6 h | 4.6 days |
| 24 TB | 33.3 h | 66.7 h | 5.6 days |
| 30 TB | 41.7 h | 83.3 h | 6.9 days |
It is worse than it was because capacity has grown far faster than the rate at which a head can read a track - areal density improves in two dimensions, sequential throughput in only one:
2007: 1 TB / 80 MB/s = 3.5 hours
today: 24 TB / 200 MB/s = 33.3 hours
24x the capacity, 2.5x the throughput, ~9.6x the rebuild
What ZFS does differently, and where it does worse
ZFS and btrfs rebuild only allocated blocks - TrueNAS puts it as “ZFS only copies
blocks in use, reducing the time it takes to rebuild the vdev” - so a half-full
pool resilvers in roughly half the time. That is one of the better arguments for
not filling an array, and it is a genuine advantage over mdadm, which copies
every block regardless.
Against that, a raidz resilver walks the pool in block-pointer order rather than LBA order, turning a sequential read into a near-random one. The saving from reading less can be eaten by the cost of reading it out of order, and on a fragmented pool it routinely is. OpenZFS’s sequential resilver covers mirrors and dRAID, not raidz.
Three operational facts that surprise people mid-resilver:
- A resilver pre-empts a scrub. “If a resilver is in progress, ZFS does not allow a scrub to be started until the resilver completes.” On older Solaris ZFS a running scrub was suspended and restarted from the beginning afterwards.
- The time estimate is not reliable, by design. The
zpool-scrubmanual page states that “scrubs may progress beyond 100% completion” because the pool changes underneath the scan, and gives that as the reason time estimates become inaccurate. - Progress is two numbers, not one. The manual’s own example reads
403M / 405M scanned at 100M/s, 68.4M / 405M issued at 10.0M/s. Scanned is metadata traversal; issued is actual I/O. They differ by an order of magnitude routinely, and reading the first as progress will mislead you every time.
Oracle’s ZFS administration guide gives the scheduling behaviour: a scrub “proceeds as fast as the devices allow, though the priority of any I/O remains below that of normal operations”. It will use everything the array is not using, and stop being polite only when nothing else is asking.
Drive count against drive size, for a fixed usable capacity
“Fewer bigger drives” is the standard advice and it is half right. Here is the whole trade, at a fixed 60 TB usable on raidz2 or RAID6, with resilver times at 100 MB/s. These are my calculations, not measurements:
| Layout | Raw | Parity tax | Per-drive resilver | Data read during resilver |
|---|---|---|---|---|
| 8 x 10 TB | 80 TB | 25.0% | 27.8 h | 70 TB |
| 7 x 12 TB | 84 TB | 28.6% | 33.3 h | 72 TB |
| 6 x 15 TB | 90 TB | 33.3% | 41.7 h | 75 TB |
| 5 x 20 TB | 100 TB | 40.0% | 55.6 h | 80 TB |
| 4 x 30 TB | 120 TB | 50.0% | 83.3 h | 90 TB |
Fewer bigger drives cost more parity, not less. At a fixed parity count, a narrower vdev spends a larger fraction of its raw capacity on redundancy: two drives of eight is 25 per cent, two drives of four is 50 per cent. The resilver is three times longer and the array reads 29 per cent more data to complete it.
Against that, drive count multiplies the annual chance of touching a rebuild at all. At Backblaze’s 20 TB+ AFR of 0.85 per cent, and at their lifetime 1.39 per cent:
| Drives | P(at least one fails in a year) at 0.85% | at 1.39% |
|---|---|---|
| 4 | 3.36% | 5.45% |
| 5 | 4.18% | 6.76% |
| 6 | 4.99% | 8.06% |
| 7 | 5.80% | 9.33% |
| 8 | 6.60% | 10.59% |
| 12 | 9.74% | 15.46% |
The decision trades how often you rebuild against how long each rebuild takes. Eight drives rebuild twice as often as four and each rebuild is a third as long, so total exposure time is roughly a wash and the eight-drive array wins on parity efficiency. Fewer bigger drives win instead on the things that scale with spindle count: bays, which are the scarce resource; noise, power and spin-up surge; and the cost of chassis, controller and cabling. Both answers are defensible. Only one of them is usually stated.
RAID levels, and what each one actually survives
With 12 TB drives:
| Layout | 4 drives | 6 drives | 8 drives | Survives |
|---|---|---|---|---|
| RAID5 / raidz1 | 36 TB | 60 TB | 84 TB | any 1 |
| RAID6 / raidz2 | 24 TB | 48 TB | 72 TB | any 2 |
| raidz3 | - | 36 TB | 60 TB | any 3 |
| RAID10 / mirrors | 24 TB | 36 TB | 48 TB | 1 per pair |
At four drives, RAID6 and RAID10 cost exactly the same capacity. RAID10 rebuilds by copying one surviving drive to one replacement - no parity computation, no reading the rest of the array, often a third of the time - but it survives only the right second failure, where RAID6 survives any. At four bays that is the entire decision. By six drives RAID6 leads on capacity by a third, and by eight by half.
Mirrors have one advantage the capacity table hides: a mirror vdev’s read IOPS scale with the number of members, while a raidz vdev’s write IOPS are those of a single disk. For a NAS serving large sequential files, raidz is right. For one running virtual machines or a database, mirrors usually are, and the capacity loss is the price of the IOPS.
RAID5 at modern sizes is defensible only narrowly: small drives, a small array, and a tested backup. The URE table above is why, with the caveats about the URE table above attached.
RAID is not a backup, and that is not a technicality. It protects against one failure mode - a drive dying - and not against deletion, ransomware, filesystem corruption, a controller writing garbage to every member at once, a power event, fire or theft. Worse, it replicates faithfully: a rebuild reconstructs corrupted data exactly as corrupt as it was. A backup is three copies, on two media, one of them elsewhere; an array is one copy however many drives are in it.
ZFS: the decisions that are permanent, and the ones that are not
ZFS is the default recommendation for a NAS, and the reason is checksums - it can tell you that a block is wrong, which no traditional RAID layer can. The cost is a set of decisions made at creation time that cannot be undone.
ashift is permanent, and getting it wrong is permanent too
ashift is the base-2 logarithm of the vdev’s sector size. It is set when the vdev
is created and cannot be changed afterwards - not by a property, not by a
rewrite, only by destroying and recreating the vdev.
ashift 9 = 512 bytes
ashift 12 = 4,096 bytes <- current recommendation for hard drives
ashift 13 = 8,192 bytes
ashift 14 = 16,384 bytes
valid range 9 to 16; 0 means autodetect
The failure mode is asymmetric. Set ashift too high and you waste some space on
small blocks. Set it too low - typically by letting autodetect believe a 512e
drive’s emulated 512-byte sector - and every partial-sector write becomes a
read-modify-write at the drive for the life of the vdev. Not every drive reports
its physical sector size honestly.
Always set ashift=12 explicitly on a hard-drive pool. Autodetect is the
value that can be wrong, and it is wrong in the direction that cannot be fixed.
recordsize affects new files only
recordsize defaults to 128 KiB and takes powers of two from 512 bytes to 1
MiB, or to 16 MiB with the large_blocks feature. The OpenZFS tuning
documentation is explicit that “changing the recordsize on a dataset will only
take effect for new files”, so setting it after you have written the data does
nothing to the data.
| Workload | Recommended recordsize | Reason |
|---|---|---|
| Media, backups, large sequential files | 1M | Fewer, larger I/Os; better compression ratio |
| General file share | 128K (default) | No strong reason to move |
| MySQL / InnoDB | 16K | Match the engine’s page size |
| PostgreSQL | 32K | Match the engine’s page size |
| SQLite | 64K | Match the engine’s page size |
For a media NAS, recordsize=1M on the bulk datasets is the documented
recommendation and it is worth setting before the first file lands.
RAIDZ space efficiency is not the parity fraction
This is the ZFS surprise that costs people real capacity. The naive model says a
3-disk raidz1 gives you 2/3 of raw. What you actually get depends on recordsize
against ashift:
3-disk raidz1, ashift=12 (4 KiB sectors), 2 data + 1 parity per stripe
4 KiB record:
1 data sector, and a stripe cannot have less than 1 parity sector
4 KiB stored as 8 KiB -> 50% efficient
128 KiB record:
32 data sectors, in 16 stripes of 2 -> 16 parity sectors
48 sectors = 192 KiB of media to store 128 KiB
128 / 192 -> 66.7% efficient
A small-recordsize dataset on a narrow raidz vdev loses far more capacity than the parity fraction suggests - here, a third more than you budgeted for. The effect shrinks as the vdev widens and as the recordsize grows, which is another reason media datasets want 1M.
vdev width, and why more narrow vdevs beat fewer wide ones
OpenZFS recommends 3 to 9 disks per raidz vdev, with the minimum being the parity count plus one. TrueNAS adds a hard ceiling: “We do not recommend using more than 12 disks per vdev.”
The performance reason is blunt. A raidz vdev’s write IOPS are approximately those of a single disk - every write touches every member, so the vdev goes as fast as its slowest member. Two 6-wide raidz2 vdevs have roughly twice the write IOPS of one 12-wide raidz2 vdev of the same drives, at the cost of two more drives of parity. For almost every workload that is the right trade.
OpenZFS 2.4 added a related mitigation: raidz sit-out, where a disk that lags behind its peers is temporarily bypassed and its data reconstructed from parity while writes continue, with a default threshold of 600 seconds. That is the ZFS-native analogue of the ERC problem, handled in the pool rather than at the drive - and it does not remove the reason to set ERC, because 600 seconds is a very long time to be slow.
raidz expansion exists now, with two caveats
OpenZFS 2.3 added raidz expansion: zpool attach a single disk to an existing
raidz vdev, online, with the pool readable throughout. Two things it does not do:
- Fault tolerance does not change. A raidz2 stays a raidz2, now spread across more disks. Widening does not buy you more parity.
- Blocks written before the expansion keep their original data-to-parity ratio until they are rewritten. A block written on a 4-wide raidz1 still carries one parity sector per three data sectors after the vdev grows to six. The capacity you expected appears gradually, as data turns over, or immediately only for new writes.
special vdevs: real gains, total risk
A special allocation class vdev holds metadata, the indirect blocks of user
data, and any dedup tables, on faster media. With special_small_blocks set, it
additionally absorbs files below a size threshold - valid at zero or any power of
two from 512 bytes to 1 MiB.
The thing to understand before adding one: losing a special vdev destroys the pool. Not the metadata, not some files - the pool. TrueNAS enforces the consequence at creation time by refusing to build a pool where the metadata vdev’s redundancy is lower than the data vdevs’. A three-way mirror for the special vdev under a raidz2 data vdev is not paranoia; it is the matching level.
The gain is genuine on a pool with many small files, because directory traversal and metadata reads stop hitting spinning media. On a media pool with 40,000 large files it is close to pointless.
SLOG is for synchronous writes and almost nothing else
A separate log vdev accelerates only synchronous writes - the ones an
application requested with fsync() or O_SYNC, typically NFS exports, iSCSI
targets and database commits. Asynchronous writes, which is nearly everything an
SMB file share does, never touch it.
size needed: ~4 GB of flash is sufficient
upper bound: no benefit beyond max ARC size
= half of system RAM on Linux, three quarters on illumos
TrueNAS: ZFS currently uses 16 GiB of SLOG space
topology: single device or mirror; raidz is not supported for a log vdev
If your NAS serves SMB to humans, an SLOG will do nothing measurable, and the device is better spent as a special vdev or left out.
L2ARC costs RAM, and on a small system it makes things worse
Every block held in L2ARC needs a header in ARC - roughly 70 bytes each - so the RAM cost scales with the record count, not the device size:
512 GiB L2ARC device, ARC header cost at 70 bytes per record
recordsize 1M -> 524,288 records -> ~35 MiB of ARC
recordsize 128K -> 4,194,304 records -> ~280 MiB of ARC
recordsize 16K -> 33,554,432 records -> ~2.2 GiB of ARC
On a 16 GiB system with a database dataset at 16K records, a 512 GiB cache device eats an eighth of the RAM that was doing the caching. TrueNAS’s rule is direct: do not add L2ARC below 32 GiB of RAM, and keep the device under ten times RAM.
It also wears the SSD. At the default l2arc_write_max of 8 MiB/s, a consumer SSD
rated for 500 TiB of writes reaches its endurance in roughly two years; raise
the feed rate to 64 MiB/s and it is about three months. That is a cache device
being consumed to hold data that RAM would have held better.
Klara Systems’ summary is the one to remember: ARC is roughly an order of magnitude faster than L2ARC, so the answer to “should I add L2ARC” is “not until RAM is maxed out”. The minimum useful hit ratio before the device earns its keep is around 25 per cent, and since OpenZFS 2.0 the cache does at least survive reboots.
Two settings that are not permanent and are worth changing
Hardware RAID controllers should not be under ZFS at all, and the reason is correctness rather than performance: “Hardware RAID will limit opportunities for ZFS to perform self healing on checksum failures.” Presenting each drive as a single-drive RAID0 volume to simulate an HBA is explicitly discouraged. Use an HBA - the enterprise guide covers IT mode and which cards do it.
And watch the fill level. TrueNAS documents a real allocator change, not a rule of thumb: “At 90% capacity, ZFS switches from performance- to space-based optimization, which has massive performance implications.” Add capacity before 80 per cent. On a 6 x 12 TB raidz2 pool that is 38.4 TB of the 48 TB usable, and it arrives sooner than you think.
Scrub scheduling
TrueNAS schedules a monthly scrub per pool by default. The widely used split is
monthly for consumer-grade drives and quarterly for datacentre-grade, with two to
four weeks for business VM pools. Weigh it against the workload arithmetic above:
on ZFS a monthly scrub of a 62-per-cent-full six-drive array is about 90 TB/year
per drive, which fits comfortably inside a 180 TB/year rating and inside the 220
TB/year point where WD quotes Red Pro’s MTBF. Monthly is the defensible default
on ZFS. It is a much more expensive default on mdadm.
mdadm and btrfs: what each one gives up
mdadm is the most predictable of the three, and the least informed
Linux md exposes its scrub machinery through sync_action, which takes six
values: resync, recover, idle, frozen, check and repair. frozen is
the one worth knowing and the one most guides omit: it stops the current action
and prevents a new one from starting, where idle only stops the current one
and leaves the array free to begin another. During check and
repair, the kernel documentation says, “md will count the number of errors that
are found. The count in mismatch_cnt is the number of sectors that were
re-written, or (for check) would have been re-written.”
# start a read-only consistency check
echo check > /sys/block/md0/md/sync_action
# watch it
cat /proc/mdstat
cat /sys/block/md0/md/mismatch_cnt
# stop it
echo idle > /sys/block/md0/md/sync_action
# stop it and keep it stopped
echo frozen > /sys/block/md0/md/sync_action
A non-zero mismatch_cnt on RAID5/6 means data and parity disagree. It does not
tell you which one is wrong. Without checksums md knows there is an
inconsistency and cannot identify the correct version; repair picks parity and
recomputes, which is a guess. That is the strongest argument for ZFS or btrfs over
mdadm, and it is a correctness argument rather than a performance one.
What md does have is the bad block list described above, which lets a degraded array survive a URE, and one genuinely useful tuning knob:
/sys/block/md0/md/stripe_cache_size
default 256 entries, settable from 17 to 32768
Raising stripe_cache_size materially changes RAID5/6 write throughput on
spinning disks, at a RAM cost proportional to the value times the member count. It
is one of the few md defaults chosen for smallness rather than for your
workload.
btrfs RAID5/6 is still classified unstable, and this is not gossip
The btrfs upstream status page classifies RAID56 stability as “unstable”, and the same page defines that word: “do not use for other then testing purposes, known severe problems, missing implementation of some core parts”. RAID1, RAID10 and scrub are all marked “OK”. The page is maintained against current kernels.
That is upstream’s own assessment in upstream’s own table. btrfs RAID1 and RAID10 on a NAS are fine. btrfs RAID5/6 is not a thing to put your array on.
A configuration detail most guides omit: even where btrfs RAID5/6 data is used,
the documentation discourages RAID5/6 metadata and recommends raid1 or
raid1c3 for that profile. A btrfs “RAID6” array is therefore normally a
mixed-profile array, and anyone quoting you a usable-capacity figure from the
parity fraction alone has not accounted for it.
btrfs does have what mdadm lacks: per-block checksums, so its scrub identifies
which copy is correct rather than guessing. On RAID1 that is the whole value
proposition, and it is a good one.
Synology SHR and Unraid parity are not RAID levels
Both get described as “RAID variants”. Neither is, and the differences decide whether they suit you.
SHR is layered RAID groups, not a new algorithm
Synology Hybrid RAID slices each drive at the size boundaries of the drive set and runs an independent RAID group inside each layer. Everything below the smallest drive’s capacity becomes one parity-protected group across every disk; the region between the smallest and second-smallest becomes its own group on the drives that reach that height; and so on up. SHR-1 layers behave like RAID5, or RAID1 where only two drives reach that layer; SHR-2 layers tolerate two failures.
Worked, with 2 TB + 3 TB + 4 TB + 4 TB:
layer 1: 0-2 TB on 4 drives, RAID5 -> 3 x 2 TB = 6 TB usable
layer 2: 2-3 TB on 3 drives, RAID5 -> 2 x 1 TB = 2 TB usable
layer 3: 3-4 TB on 2 drives, RAID1 -> 1 x 1 TB = 1 TB usable
total = 9 TB = 8.18 TiB
plain RAID5 across the same drives, limited to the smallest member:
4 x 2 TB, one parity -> 3 x 2 TB = 6 TB = 5.46 TiB
SHR recovers 50 per cent more usable capacity from a mismatched set, and that is its entire reason to exist. It costs nothing in redundancy per layer and costs you understanding: a “one drive failure” means a different thing in each layer, and the recovery path is a multi-group rebuild. With four identical drives SHR and RAID5 are the same thing, so the feature only pays when your drives are mismatched.
Unraid is not striping at all
Each data disk in an Unraid array carries its own ordinary filesystem, complete and independently mountable. The parity disk holds the XOR of every corresponding sector position across the data disks. There is no stripe, no chunk size, and no relationship between a file and any disk but the one it is on.
Three consequences follow, and they are the whole argument:
- Only the disk holding a file must spin up. On a media library that is the dominant power saving, and no parity array can match it.
- A file survives whole on its own disk if you lose more disks than parity covers. On a RAID6 array, losing three members loses everything. On Unraid with single parity, losing two disks loses the contents of those two disks and leaves the rest readable. That is a fundamentally different failure shape.
- The parity disk must be at least as large as the largest data disk, because it must cover every sector position that exists.
The cost is write throughput. Unraid’s default write mode is read-modify-write:
default (read-modify-write): read old data, read old parity,
write new data, write new parity
= 4 I/Os, 2 drives spinning
turbo / reconstruct write: read all other data disks, compute parity,
write new data, write new parity
= 2 writes + one read per other data disk,
every drive spinning
That is the trade stated exactly, and note which way the I/O count goes: the default costs four I/Os on two drives; turbo write costs two writes plus a read of every other data disk, so on an eight-disk array it is eight I/Os on eight drives. Turbo write is faster because nothing waits on a read-modify cycle, not because it does less work - it does more, on more spindles, and gives up the power advantage entirely. Dual parity (P and Q) survives two failures, as RAID6 does.
Unraid suits a media library that is written rarely and read one file at a time. It does not suit anything that wants sustained write throughput or parallel reads across many files.
Mixing capacities, vendors and batches
Array members share every variable that drives failure: model, firmware revision, manufacturing batch, power-on hours, temperature, vibration environment and workload. That is about as far from independent samples as a population gets, and every probability in this guide assumes independence.
Schroeder and Gibson’s finding that time between replacements shows significant autocorrelation out to 30-week lags is the empirical form of the problem. Failures arrive in clusters. The constant-hazard model underlying MTBF arithmetic fit their data poorly.
So buy deliberately mixed: two vendors, or two retailers, or the same model bought months apart, and check the date codes on arrival rather than assuming. A firmware defect is the strongest case for it, being the one failure mode that can take a whole array inside an hour and the one that no amount of parity survives.
But be honest about the limit. Nobody has published a study quantifying how much batch diversity actually buys. It is a cheap hedge against an unquantified risk, not a measured improvement, and anyone who gives you a percentage has made it up.
Mixing capacities is a separate question with a per-platform answer:
| Platform | Mixed capacities |
|---|---|
| ZFS raidz | Every drive is used to the size of the smallest. The excess is wasted. |
mdadm |
Same - the array is built on the smallest member’s size. |
| Synology SHR | Layered groups recover most of the excess, as above. |
| Unraid | Fully supported; parity drive must be the largest. |
On ZFS, adding one larger drive to a raidz vdev buys nothing until every member is replaced. That is the single most common mistake in home ZFS builds, and it is why the expansion plan below matters more than the drive you buy today.
SMR in an array is a different failure, not a slower one
Not in an array, and the reason is the best-documented failure in recent consumer storage. In 2020 it emerged that drive-managed SMR had been shipped without disclosure across more than one vendor’s consumer lines, and inside a NAS-branded one: the WD Red. Owners found out when they replaced a failed drive and the array could not rebuild.
The mechanism belongs to the CMR/SMR guide; the array-level consequence is specific. A drive-managed SMR drive absorbs writes into a persistent media cache and reorganises them into shingled bands later, while idle. A rebuild gives it hours of sustained writes and no idle time. The cache fills, and every subsequent write forces a read-modify-write of a whole shingled band. Throughput falls to single-digit megabytes per second and individual commands start taking seconds - which lands you back in the first section of this guide, with the kernel wiki’s figure of over ten minutes for a stall during a read error. The array does not see a slow drive, it sees a drive that stopped answering, and drops it. Adding a replacement can cost you a second member.
Nor can you test for it easily. /sys/block/sdX/queue/zoned flags only
host-managed and host-aware drives; drive-managed SMR reports none,
identically to a CMR drive, and no SMART attribute records it.
So filter to CMR in the hard drive listings, and treat an unlabelled listing as unknown rather than as CMR. A model number does not tell you whether a drive is shingled. The same marketing name has covered both technologies across generations and regions, sellers transcribe part numbers wrongly, and this site will not infer what it cannot establish.
Shucked drives, the power-disable pin, and the firmware you cannot change
Shucking an external drive is often the cheapest route to NAS capacity, and it carries two specific problems that are not the drive’s fault.
The 3.3 V pin is a specification change, not a defect
SATA 3.3, published on 2 February 2016, incorporated SATA-IO technical proposal TPR056 - authored by Frank Chu of HGST with James Hatfield and Alvin Cox of Seagate - which reassigned connector pin P3. Take the order carefully, because it is usually given backwards. P3 originally carried 3.3 V alongside P1 and P2, as part of the supply. SATA 3.2, in August 2013, reassigned it to DEVSLP. SATA 3.3 then gave the same pin a third job: the Power Disable control, driving it high at 2.1 to 3.6 V to cut power to the drive circuitry. The proposal notes that Power Disable and Device Sleep are asserted by the same voltage on P3 and are therefore mutually exclusive - one pin, three successive meanings, the last two of which cannot coexist. The interfaces guide has the full revision history. Power Disable exists so a datacentre can power-cycle one drive in a backplane without touching its neighbours.
Western Digital’s own technical brief states the symptom plainly:
if you put a new SATA HDD with this feature into a legacy chassis or enclosure, the drive may not spin up! The HDD is not defective. Some legacy power supplies provide 3.3V power on P3 (Pin 3), and this forces the HDD to get stuck in a hard reset condition.
Any power supply whose SATA connector actually delivers 3.3 V on P3 - which older supplies do, because that was the specification - holds a PWDIS drive in permanent reset. The drive is fine. It is being told to stay off.
Two field fixes, neither of which harms the drive, because both simply prevent 3.3 V reaching P3:
- Kapton tape over pin 3 of the 15-pin SATA power connector. Pin 1 is the end nearest the L-shaped notch; count three. Kapton rather than electrical tape because it does not soften at drive-bay temperatures.
- A Molex-to-SATA power adapter, which carries only 12 V and 5 V and has no 3.3 V rail at all.
The detail most guides miss is in WD’s own brief: the feature requires a distinct PCBA, so it is a per-part-number property rather than a firmware setting you could turn off. WD publishes separate part numbers with and without it
- the DC HC510 8 TB is 0F27610 without and 0F27455 with - and the SATA models of the DC HC530 do not offer it at all. WD’s own recommendation is to buy drives without the feature unless you specifically need it. You cannot tell from the outside of a sealed external enclosure which one is in there, though on a bare Ultrastar the model number settles it before you bid - the interfaces guide decodes the suffix.
The other shucking risk is recording technology
An external enclosure’s marketing does not state recording technology, and the 2020 episode establishes that one product line can contain either. A shucked drive is a drive whose CMR status you cannot establish until it is out of the case, and often not then - see buying used drives on eBay for the rest of what a listing withholds.
Helium, and what the sensor can and cannot tell you
Helium fill is not a tier, it is a per-model manufacturing choice, and it does not track capacity cleanly:
| Family | Helium | Air |
|---|---|---|
| WD Red Plus | WD120EFBX (12 TB) only | Everything else, 2-12 TB |
| WD Red Pro | 14 TB and above, plus WD121KFBX (12 TB) | WD122KFBX (12 TB), and 10 TB and below |
| Seagate IronWolf Pro | 12-20 TB, and ST10000NE0008 | 8 TB and below, and ST10000NE0004 |
Every row of that table has a capacity sitting on both sides of the line. The two 10 TB IronWolf Pro part numbers, ST10000NE0004 and ST10000NE0008, are one digit apart and one fill gas apart. So are the two 12 TB WD Red Plus part numbers, WD120EFGX air and WD120EFBX helium. So are the two 12 TB WD Red Pro part numbers, WD122KFBX air and WD121KFBX helium - which is why the Red Pro row cannot be stated as “helium from 14 TB” however much cleaner that would read. Same family, same capacity, different fill gas, different power figures, different acoustic behaviour. This is a part-number fact, not a capacity fact, and the digit that carries it is not the one you would look at.
SMART 22 is a warning light with no gauge behind it
Helium-sealed drives expose SMART attribute 22, a pre-fail attribute with a published normalised threshold of 25. It starts at 100 and counts down. The raw value has no interpretable physical meaning to an owner, because HGST explicitly declined to publish either the amount of helium in a drive or the trip points; the attribute “trips once the drive detects that the internal environment is out of specification”.
So: any movement below 100 is a signal, and you cannot tell how much margin is left. In Backblaze’s fleet only one HGST drive ever read below 100, at values between 94 and 99 - a rare event, and an unambiguous one. Only HGST helium drives reported the attribute at all in that analysis; Seagate’s did not, so its absence on a helium drive is not evidence of anything.
Helium does not appear to change failure rates
Backblaze’s normalised comparison put helium drives at 1.06 per cent AFR against 1.61 per cent for comparable air drives, with the confound that the helium drives were newer. The effect you can verify from the datasheets is roughly 20 per cent lower spin power, and for a home NAS that matters more than the reliability difference does.
A drive that lies about its spindle speed
Filed here because it is the same class of trap. WD Red Plus footnote 7, for the WD120EFGX:
Actual spindle motor rotational speed for this model is 7200 RPM; although ID Device may report 5400 to reflect previous Performance Class designation.
The drive reports 5400 rpm over ATA IDENTIFY and physically spins at 7200. Any
tool reading RPM from the drive - smartctl, hdparm -I, every NAS dashboard -
reports the wrong number, and a decision made on quietness or power draw from that
figure is wrong too. It is why filtering by RPM tells you
what the listing claims rather than what the spindle does, and it generalises:
the drive’s self-report is a field in firmware, not a measurement.
Burn-in, with the commands and the ceiling nobody mentions
Day-one failures are invisible in every published statistic, as Backblaze’s own methodology note concedes. Burn-in is the only way you find them, and the window in which finding them is useful is the return window.
badblocks has a hard capacity ceiling
badblocks uses a 32-bit block counter, which sets a maximum testable size per
block size. This is why an 18 TB drive returns “Value too large for defined data
type”, and why a great many people conclude their drive is broken:
| Block size | Maximum testable |
|---|---|
-b 512 (default) |
2.20 TB (2.00 TiB) |
-b 1024 |
4.40 TB (4.00 TiB) |
-b 4096 |
17.59 TB (16.00 TiB) |
-b 8192 |
35.18 TB (32.00 TiB) |
Use -b 4096 up to 16 TiB and -b 8192 above it. The default of 512 has been
useless for over a decade.
The sequence
# 1. short self-test, about 5 minutes - catches the obviously dead
sudo smartctl -t short /dev/sdX
# 2. conveyance test, about 2 minutes - looks for shipping damage specifically
sudo smartctl -t conveyance /dev/sdX
# 3. long self-test - full surface read, hours
sudo smartctl -t long /dev/sdX
# 4. record the attributes BEFORE the destructive test
sudo smartctl -A /dev/sdX > /root/sdX.before
# 5. destructive four-pattern write/read - DESTROYS ALL DATA
sudo badblocks -b 4096 -wsv /dev/sdX # up to 16 TiB
sudo badblocks -b 8192 -wsv /dev/sdX # 16 to 32 TiB
# 6. long self-test again
sudo smartctl -t long /dev/sdX
# 7. compare
sudo smartctl -A /dev/sdX > /root/sdX.after
diff /root/sdX.before /root/sdX.after
Step 7 is the actual test. A single SMART reading tells you the drive’s history; the difference across a full-surface write and verify tells you whether the drive is degrading right now. Three raw values must be zero on a drive you are keeping:
5 Reallocated_Sector_Ct sectors already remapped
197 Current_Pending_Sector unreadable, not yet remapped
198 Offline_Uncorrectable unreadable during offline scan
Attributes 197 and 198 are the ones that become a failed rebuild. A pending sector is one the drive could not read and has not yet reallocated - precisely the URE you did not want to meet while degraded. A non-zero raw value on any of the three, on a drive you have owned a week, is a return rather than a risk to price in.
It takes days, and it costs a chunk of the workload rating
badblocks -w writes and reads four patterns, so a full run moves eight times
the drive’s capacity. Reported field times: a 20 TB Seagate at 190 to 195
hours, about eight days; a 2 TB WD Red at just over 24 hours.
naive estimate, 20 TB at 150 MB/s average:
20e12 x 8 / 150e6 / 3600 = 296 hours
reported: ~190 hours
implied average rate: 160e12 / (190 x 3600) = 234 MB/s
The implied rate is near the drive’s outer-track figure, so treat 190 hours as a best case rather than a typical one. Budget eight to ten days per large drive - not a week, which is less than the best case this paragraph has just rejected as optimistic - and run them in parallel if the chassis and the power supply allow.
Connect it back to the workload rating, because nobody does. The middle column is the 220 TB/year point at which WD quotes Red Pro’s MTBF, which is the number the drive’s advertised reliability figure actually depends on:
| Drive | Traffic per full run | vs 180 TB/yr | vs 220 TB/yr | vs 550 TB/yr |
|---|---|---|---|---|
| 8 TB | 64 TB | 36% | 29% | 12% |
| 12 TB | 96 TB | 53% | 44% | 17% |
| 20 TB | 160 TB | 89% | 73% | 29% |
| 24 TB | 192 TB | 107% | 87% | 35% |
A single burn-in of a 24 TB drive at the 180 TB/year tier exceeds its entire annual workload allowance. That is not a reason to skip burn-in - finding a bad drive inside the return window is worth more than a warranty condition - but it is a reason to do it once rather than routinely, and to plan the year’s scrub schedule knowing the budget is already partly spent.
SSD caching in a NAS, and when it does nothing
The ZFS analysis above applies: RAM first, L2ARC only when RAM is maxed, SLOG only for synchronous writes. Synology’s DSM caching is a different implementation, and its published arithmetic is unusually honest. From the DSM 7.1 SSD Cache white paper:
RAM cost: ~400 KB of system memory per 1 GB of SSD cache
RAM ceiling: DSM uses at most 25% of installed RAM for the mapping table
Model ceiling: 930 GB total cache on Alpine-CPU models
Read-write cache: 2 to 12 SSDs, in RAID1, RAID5 or RAID6
800 GB cache -> 800 x 400 KB = 320,000 KB of mapping table
= 320 MB decimal
= 312.5 MiB, if Synology's KB is 1024 bytes
needs at least ~1.3 GB installed to stay under the 25% cap
(either reading; 320 x 4 = 1.28 GB, 327.7 x 4 = 1.31 GB)
Synology does not say which kind of kilobyte it means, and this is a guide that spends a paragraph on decimal against binary elsewhere, so both readings are above. They do not change the answer.
Their own advice is worth quoting, because it argues against selling you more SSD:
The recommended SSD cache size should be just enough to cover the size of frequently accessed data… increasing the SSD cache size to 500 GB for 100 GB of hot data will not result in significant performance increase. Excess cache space will only be used to store cold data.
Their in-house tests state the headline as 30 times the random read IOPS with cache against without, and the measurements behind it range from roughly 15 to 40 times depending on the test: their NVMe configuration goes from 6,581 to 263,667 IOPS on a 100 per cent random read, which is 40 times, while a random write test in the same paper is nearer 15. They publish no latency figure at all, in any unit, anywhere in the paper. If you meet “93 per cent lower latency” attributed to this document, check the arithmetic before you repeat it: 1 - 1/15 = 93.3 per cent. It is an understated IOPS ratio inverted and relabelled as a latency measurement, which is a derived number wearing a vendor’s authority.
Read the workload before the number, too: that is an OLTP-style small random pattern. A media NAS streaming large sequential files gets approximately nothing from an SSD cache, because there is no hot set to cache and the drives were never the sequential bottleneck.
Where it earns its place: many small files, virtual machine images, a photo library’s thumbnails and metadata, an iSCSI target. Where it does not: a film library, a backup target, an archive. Choosing between a cache SSD and more RAM, add RAM.
Noise, heat and power, budgeted
Noise adds incoherently, and that is good news
Uncorrelated sources add at +3 dB per doubling, so total sound pressure is
L + 10 x log10(N):
| Per-drive seek figure | 4 drives | 8 drives | 12 drives |
|---|---|---|---|
| 28 dBA | 34.0 dBA | 37.0 dBA | 38.8 dBA |
| 32 dBA | 38.0 dBA | 41.0 dBA | 42.8 dBA |
| 36 dBA | 42.0 dBA | 45.0 dBA | 46.8 dBA |
Eight drives are only about 9 dB louder than one, so a drive 4 dBA quieter buys more than halving the drive count. Datasheet figures across current NAS families run 20 to 34 dBA idle and 26 to 39 dBA seeking - WD Red Pro 20-34 and 31-39, Red Plus 20-34 and 26-39, IronWolf Pro 28 and 32. The spread within a single family is larger than the effect of drive count, and inside Red Pro the loud end is the air-filled part numbers: WD122KFBX, WD103KFBX and WD102KFBX are all 34 dBA idle against 20 for the quietest helium members.
Then remember the ASHRAE result from earlier: the fans are probably louder than the drives, and their noise is the thing degrading throughput.
Spin-up sizes the 12 V rail, not running load
Datasheet peak 12 V current is 1.2 to 2.1 A per drive - WD Red Pro 1.7 to 2.08 A, Red Plus 1.2 to 1.9 A, IronWolf Pro around 2.0 A typical at startup:
at 2.0 A per drive on the 12 V rail:
4 drives -> 96 W surge
6 drives -> 144 W surge
8 drives -> 192 W surge
12 drives -> 288 W surge
against idle, at 5.5 W per drive:
8 drives -> 44 W
Eight drives spinning up simultaneously draw roughly four times their idle power, all on the 12 V rail and all in the first few seconds. That is why NAS chassis and HBAs implement staggered spin-up, and why a power supply sized on running load browns out at boot. If your array fails to come up after a power cut but works when you start drives by hand, this is why.
Idle power sets the bill, and the spread is 2.5x
A NAS spends nearly all its life idle, so the idle figure is the one that multiplies by 8,760 hours:
| Drive | Idle | Read/write | Standby |
|---|---|---|---|
| WD Red Pro 26 TB | 3.6 W | 6.0 W | 0.3-1.6 W |
| WD Red Plus 2 TB | 2.4 W | 4.0 W | 0.3-1.6 W |
| WD Red Plus 12 TB (WD120EFBX, helium) | 2.9 W | 6.3 W | 0.3-1.6 W |
| WD Red Plus 12 TB (WD120EFGX, air) | 6.1 W | 8.8 W | 0.3-1.6 W |
| Seagate IronWolf Pro 20 TB | 5.5 W | 7.7 W | 0.3-1.6 W |
Look at the two 12 TB WD Red Plus rows. Same family, same capacity, 2.9 W against 6.1 W idle - the helium one uses less than half the power of the air one. Over six drives running continuously:
6 x 2.9 W = 17.4 W -> 152 kWh/year
6 x 6.1 W = 36.6 W -> 321 kWh/year
difference 169 kWh/year, from a part-number suffix
The 26 TB Red Pro at 3.6 W idle draws less than the 12 TB air-filled Red Plus at 6.1 W. Bigger drives do not necessarily cost more to run, and buying the newer generation is an argument on power grounds alone.
Temperature
Seagate’s IronWolf Pro datasheet gives 65 °C as the maximum drive-reported operating temperature for 6 to 24 TB, and 60 °C on the 28 to 32 TB models, and then advises against using it: “Seagate does not recommend operating at sustained drive temperatures above 60C.” Set against Google’s finding that lower temperatures correlated with higher failure rates across most of the observed range, the rule is: keep drives below 60 °C and stop there. Chasing 30 °C costs fan speed, fan speed costs acoustic energy to the fifth power, and acoustic energy costs throughput you cannot measure.
Capacity planning, and the cost of expanding later
What you see is not what is on the box, for two reasons that compound:
6 x "12 TB" in RAID6
raw 72 x 10^12 bytes
usable 48 x 10^12 bytes (4 data drives)
reported 48e12 / 2^40 = 43.7 TiB (binary units)
at 80% 35.0 TiB of actual room
Decimal terabytes against binary tebibytes is the capacity guide’s subject; it costs about 9 per cent before the array takes its share. As for that 80 per cent, do not plan to run past it, for three reasons that stack:
- Zoned bit recording. On a 3.5-inch platter the outer data radius is around 46 mm and the inner around 23 mm, and that 2:1 geometric ratio is close to the 2:1 ratio between a drive’s best and worst sustained rates. Whatever lands in the last fifth of the pool lives on the slow half of the platter.
- Allocator behaviour. Filesystems need contiguous free runs, and copy-on-write filesystems need them even to modify an existing file, never writing in place. ZFS’s documented change at 90 per cent - from performance- based to space-based optimisation - is the hard version of this.
- Headroom for operations. Snapshots, a reorganisation, and the growth during a multi-day resilver all need space that is not there.
Which makes expansion concrete. Say you need 40 TB now and 80 TB in three years, in an eight-bay chassis:
| Plan | Now | Later | Bays left |
|---|---|---|---|
| 6 x 12 TB RAID6 | 48 TB | add 2 x 12 TB -> 72 TB | 0 |
| 4 x 20 TB RAID6 | 40 TB | add 2 -> 80 TB, add 2 -> 120 TB | 2 |
| 6 x 20 TB RAID6 | 80 TB | nothing needed | 2 |
Bays are the scarce resource, not terabytes. The first plan spends six bays on cheap capacity today and the last two on the same cheap capacity later, and then leaves exactly one move: replace all eight. Growing by replacement means a full rebuild per drive, one at a time:
8 drives x 20 TB at ~100 MB/s = 55.6 h each
= 444.8 h
= 18.5 days of degraded or resilvering operation
Eighteen days with no redundancy margin to spare, on drives three years old, with
the scrub workload arithmetic above saying each of those resilvers is also a
substantial fraction of an annual allowance. That is the price of “buy small now
and upgrade later”, paid in risk rather than money. OpenZFS 2.3’s raidz expansion
softens it a drive at a time, with the two caveats named earlier, and neither that
nor mdadm’s reshape removes the bay problem.
Arguing the other way is price per terabyte, which is rarely flat across the capacity range and moves with the market - a live number rather than a fact about drives, so the 3.5-inch listings answer it better than a guide can.
What to do with this on a listing page
- Filter to CMR first, and treat “not stated” as SMR. Start at
the hard drive listings and filter to CMR, or exclude the known
shingled ones with
?smr=no. This is the one filter where a wrong guess costs you an array rather than an afternoon, because a shingled drive can stall for over ten minutes during a rebuild and no timeout setting rescues it. - Narrow to the class you actually need, not the one with the best badge. NAS-class and enterprise-class listings are separate filters here. Then read the datasheet for the exact part number on the URE rate, the workload figure and the fill gas - all three vary inside a single family, and two part numbers of the same capacity have differed by a factor of ten on URE.
- Check the price step between capacities before deciding drive count. Sort by price per terabyte, then work the table above: fewer bigger drives cost more parity and a three-times-longer resilver, more smaller drives cost bays, power and spin-up surge. There is no universally right answer, only the one your chassis and your electricity price pick.
- Run
smartctl -l sctercon every drive the day it arrives, used or new, before it joins anything. If the answer isDisabled, set it -70,70undermdadmor a hardware controller,1,1under ZFS if the drive accepts it and70,70if it does not - and try the persistent formscterc,70,70,pbefore writing a boot script. Read the value back every time; below 65 is probably not supported and the write can fail quietly. If the answer is “not supported”, either raise the kernel timeout to 180 seconds or put that drive somewhere else. - Burn it in with the right block size, and diff the SMART output.
badblocks -b 4096 -wsvup to 16 TiB,-b 8192above. Attributes 5, 197 and 198 must be zero before and after. Budget eight to ten days per large drive, and count the eight-times-capacity of traffic against this year’s workload allowance. - For used drives, do all of the above inside the return window. The used-drive guide covers what a listing withholds and how to read the hours; the enterprise pulls guide covers the HBA, the sector-size trap and why a SAS drive sidesteps the error-recovery argument entirely.
- Buy deliberately mixed, and accept that you cannot price the benefit. Two vendors, two retailers, or the same model months apart, date codes checked on arrival. Nobody has quantified what batch diversity buys. It is still the cheapest insurance against the one failure mode parity does not survive.