HDD or SSD: where each still wins, and why the gap persists
By Harry Saarinen ·
Flash for anything a person waits on. Platters for anything measured in terabytes that a person does not wait on. That split has been stable for a decade, the reasons are structural rather than fashionable, and nothing in the last two years has moved it.
What moved, violently, is the price of both. The comfortable assumption that flash swings wildly while disk drifts steadily downwards inverted in 2025 and 2026: on TrendForce’s figures NAND contract prices roughly quadrupled across three quarters, and the most popular hard drives in German shops cost 134 per cent more in September 2026 than a year earlier, with nearline lead times stretching past a year. Both media got much more expensive at once, for opposite reasons, and the ratio between them did not converge - it widened.
So the useful version of the question is not “which is better”. It is why the price per terabyte has stayed roughly an order of magnitude apart while both halves got enormously cheaper over fifteen years, and then both got enormously more expensive in eighteen months. It is not inertia and it is not a cartel. The two technologies have different cost structures, those structures improve through different mechanisms at different rates, and they respond to a demand shock in different directions.
This guide works that through from published specifications. Where a figure comes from a vendor document, the vendor and the document are named. Where a figure is derived here, it says so.
The cost argument is about where the money sits
Build a hard drive and your bill of materials looks roughly like this:
cost = fixed assembly + (platters × per-platter) + (heads × per-head)
fixed assembly = base casting, spindle motor, voice-coil actuator,
preamp, controller PCB, helium seal, test time
Capacity is platters × 2 surfaces × areal density × usable area. The important
term is areal density, because raising it costs nothing in the bill of
materials at all. A drive with the same ten platters, the same motor and the
same actuator holds more because the bits are smaller. The fixed assembly cost,
which is most of the drive, is then amortised over more terabytes.
Build an SSD and the shape is different:
cost = controller + DRAM + PCB + (NAND dies × per-die)
capacity = NAND dies × per-die capacity
Divide through by capacity and as the die count rises the fixed part becomes irrelevant. A large SSD’s cost per terabyte converges on one number: the cost of a NAND die divided by the bits on it. The hard drive improves by spreading a large fixed cost over more bits; the SSD is already at its asymptote and can only improve by making silicon cheaper.
That is why the low end of the hard drive market disappeared rather than getting cheap. Nobody manufactures a new 250 GB 3.5-inch drive, because the motor, the casting and the test time cost the same as they do on a 30 TB one. Below roughly a terabyte the fixed cost dominates and flash simply wins, which is exactly what happened.
The amortisation is visible in two product manuals
This is usually asserted. It can be read directly off Seagate’s own documentation. The Exos X24 SATA product manual specifies the 24 TB model as 20 heads on 10 disks. The Exos M SATA product manual, Rev C of March 2025, specifies the 32 TB model (ST32000NM004K) as 20 heads on 10 disks. Same head count, same disk count, same 7200 rpm spindle, same helium-sealed enclosure:
Exos X24, 24 TB 20 heads / 10 disks = 2.4 TB per platter
Exos M, 32 TB 20 heads / 10 disks = 3.2 TB per platter
identical mechanism, +33.3% capacity, zero additional parts
The same manual shows the other end of the lever. The Exos X24 at 16 TB drops to 18 heads on 9 disks - a platter and two heads deleted to make a cheaper part, because that is the only variable cost there is to delete. Capacity within a nearline family is set by how many platters are fitted and how dense they are, and by nothing else. That is the fixed-cost argument, printed in a specification table.
The floor under an SSD’s price, in dollars
The NAND die is the floor, and for once there is a public number for it. In November 2025 Phison’s chief executive stated that the price of a 1 terabit TLC NAND die had gone from $4.80 to $10.70 in six months. A terabit die is a binary terabit, so:
1 Tb die = 2^40 bits = 1,099,511,627,776 bits
/ 8 = 137,438,953,472 bytes
= 0.1374 TB (decimal)
$4.80 per die -> 4.80 / 0.1374 = $34.9 per TB of bare NAND
$10.70 per die -> 10.70 / 0.1374 = $77.9 per TB of bare NAND
That is silicon only. It is before the controller, the DRAM, the PCB, the packaging, the test time, the firmware and anybody’s margin. No SSD can be sold below it for long, which is why the retail price of flash tracks the die price with a lag and a multiplier rather than tracking demand for SSDs.
The wafer side sets the same floor from the other direction, and here the arithmetic is mine rather than a vendor’s, because wafer prices are not published. Micron’s G9 generation, 276 layers, is quoted at a TLC bit density of 21.0 Gb per square millimetre using CMOS-under-array. A 300 mm wafer is 70,686 square millimetres gross; call 85 per cent of that usable after edge exclusion, scribe lanes and test structures:
gross wafer area π × 150² = 70,686 mm²
usable, at 85% = 60,083 mm²
bit density, G9 TLC 21.0 Gb/mm²
gross bits per wafer 60,083 × 21.0 = 1,261,743 Gb
÷ 8 = 157,718 GB
≈ 158 TB per wafer, pre-yield
Divide a wafer’s cost by that and you have the irreducible cost per terabyte of flash. The 85 per cent, the yield and the wafer price are all mine or missing, so treat the 158 TB as an order-of-magnitude figure and not a specification. The consistency check is worth one line: 158 TB of good die at the November 2025 price of $77.9 per terabyte is about $12,000 of die revenue per wafer, which is in the region people quote for NAND wafers. It does not prove the model; it proves the model is not nonsense.
The important structural point is that layer scaling raises the numerator as well as the denominator. Going from 276 layers to 400 and beyond means more deposition steps, more etch steps and a harder high-aspect-ratio etch, so wafer cost rises roughly with layer count while bits per wafer also rise. The gain is a difference between two growing numbers, which is exactly why it has slowed.
Now put the two structures side by side. A hard drive maker who improves areal density by 33 per cent sells 33 per cent more terabytes out of the same castings, motors and test slots. A NAND maker who improves bit density by 33 per cent has to build the capacity to deposit and etch a third more layers first. One of these is free and the other is a fab.
What happened to the price per terabyte in 2025-26
TrendForce’s contract-price figures put NAND up 20 to 25 per cent quarter on quarter in Q4 2025 (its December revision; the widely repeated 33 to 38 per cent was its first forecast for Q1 2026, not a Q4 figure), 85 to 90 per cent in Q1 2026, with 70 to 75 per cent forecast for Q2 2026. Compound them:
low end 1.20 × 1.85 × 1.70 = 3.77×
high end 1.25 × 1.90 × 1.75 = 4.16×
Roughly four times in three quarters, and every quarter since has been forecast higher still; the 2026 price guide traces each figure to its release. The cause is not mysterious: AI infrastructure demand outran fab capacity, and fab capacity arrives in lumps eighteen months to three years after somebody decides to build it.
The hard drive side rose for the opposite reason. There was no overhang to absorb the step. Three suppliers build almost exactly to demand, so when demand stepped up there was no inventory and no idle line. Reported consequences by early 2026: a basket of twelve popular drives in German shops up 46 per cent from mid-September 2025 to mid-January 2026 (ComputerBase; Tom’s Hardware found US listings following), lead times out to roughly twelve months, a 4 TB WD Blue moving from the $67 to $85 band to $99. Western Digital has said it is sold out through calendar 2026. Seagate’s chief executive has said calendar-2026 exabyte capacity is fully allocated, with backlog building into 2027 and customer conversations reaching into 2028. TrendForce reported nearline lead times going from a few weeks to more than 52 weeks over the course of 2025.
Where the volume went is on Seagate’s own results. FY2026: 789 exabytes shipped, of which 695 exabytes was nearline, up 39.8 per cent year over year; revenue up 34 per cent at 46 per cent gross margin; nearline now 87 per cent of hard drive sales, up from 83. Western Digital’s cloud business is 89 per cent of revenue and consumer retail about 5 per cent. The retail shopper is now buying the leftovers of a datacentre supply chain, and the price reflects that rather than reflecting anything about manufacturing cost.
Rough price levels at the time of writing, and these are the least reliable numbers in this guide - they come from price trackers and market reports rather than from anybody’s datasheet, and they are wrong by the time you read them:
| Category | Approximate price per terabyte |
|---|---|
| Budget SATA SSD, consumer | $55 - $70 |
| Gen4 NVMe, consumer | $76 - $92 |
| Gen5 NVMe, consumer | $90 - $150 |
| High-capacity CMR hard drive, new retail | $23 - $50 |
| Enterprise NVMe | $117 - $188 |
| Nearline enterprise HDD | $12.50 - $16.63 |
The enterprise pair is the clean comparison because both are bought by the same buyer for the same rack: roughly 10 to 1 on the midpoints, and anywhere from 7 to 15 to 1 depending which ends of the two bands you pair. At the extreme high end one index put a 30 TB enterprise SSD at a $22,600 reference price in August 2026 against $1,216 for an equivalent-capacity hard drive, which is $753 per terabyte against $41, a ratio of 18.6. The same class of 30 TB TLC enterprise SSD was quoted near $3,000 in early 2025.
Those two hard drive figures do not agree, and pretending otherwise would be worse than saying so. The nearline band in the table is $12.50 to $16.63 per terabyte; the reference price in the paragraph above works out at $41. They come from different sources measuring different slices of the same market, and this guide cannot reconcile them to better than a factor of two and a half. When two respectable sources for one category differ by that much, the category is being allocated rather than priced. Every ratio in this section inherits that uncertainty, which is why two of them appear rather than one.
The consumer rows have the same problem pointing the other way. A budget SATA SSD at $55 to $70 per terabyte and a Gen4 NVMe at $76 to $92 sit at or below the $77.9 per terabyte of bare TLC die derived above, and a finished drive cannot stay under the cost of the silicon inside it. That is the lag in “tracks the die price with a lag and a multiplier” rather than a refutation of it. The two candidate explanations are retail inventory bought at an older die price and QLC parts, which carry less die area per terabyte than a TLC figure assumes; this guide cannot tell you which dominates. A finished drive priced under current bare die is a stock position, not a cost position, and stock positions close upward.
This site still will not tell you where the ratio sits today, and the reason has changed. It is no longer that flash is volatile and disk is steady. It is that both are now priced by allocation rather than by cost, and an allocated market has no stable ratio at all. Sort the hard drive listings and the SSD listings by price per terabyte and read today’s ratio off the medians yourself.
The historical curve, and every forecast of a crossover
Backblaze publishes its own purchase records, which is the cleanest long-run hard drive cost series available because it is transaction data rather than survey data: $0.11 per gigabyte in 2009, about $0.03 in 2017, $0.014 in November 2022. That is a fall of 56.36 per cent over 2017 to 2022 by their arithmetic and 87.4 per cent since 2009. The 2017 figure is rounded here and theirs is not, so the first of those two percentages will not reproduce exactly from the numbers as printed; the second will, against $0.111. Compounded over the whole span, and this compounding is mine:
2009 $0.110/GB = $110/TB
2022 $0.014/GB = $14/TB
(0.014 / 0.110)^(1/13) = 0.853 -> -14.7% per year for thirteen years
Backblaze explicitly retracted its own 2017 prediction that the decline was bottoming out. Then the curve did something neither they nor anybody else forecast: $14 per terabyte in late 2022 became $23 to $50 per terabyte of new retail capacity by 2026. That is not a slowdown in the decline, it is a reversal, and it is the first one in the history of the product.
The flash forecasts fared worse. A widely cited 2021 analysis projected QLC flash and hard drives reaching price parity at $1.00 per terabyte “during 2026”, from $20.67 per terabyte for QLC and $3.58 for HDD in 2020, and argued from that crossover that hard drive vendors could not justify the investment in HAMR. The forecast date has arrived. Neither medium is within an order of magnitude of a dollar, the ratio between them is wider than when the forecast was made, and HAMR is shipping in volume. Between 2015 and 2023 enterprise SSD cost per terabyte fell roughly 80 per cent while hard drives fell roughly 40, and convergence still did not happen, because an 80 per cent fall from a much higher number does not catch a 40 per cent fall from a much lower one.
The crossover people quote for shipped bytes does not survive contact with the filings. It is usually placed around 2020, at a few hundred exabytes of flash against a similar figure for hard drives, and the hard drive half of it is too small by more than a factor of two. Seagate’s Form 10-K for fiscal 2020 states that the company shipped 442 exabytes of hard drive capacity in that year, and Seagate is one of three suppliers. One vendor alone shipped more exabytes than the figure usually attributed to the whole hard drive industry. Published industry totals for that year differ by more than a factor of two depending whose count you take, so no crossover number appears here as settled.
Where flash overtook disk is money, not capacity, which is a different claim and a weaker one for a buyer. A terabyte of flash carries many times the revenue of a terabyte of platter, so the revenue lines crossed long before the exabyte lines came close, and on the figures in this guide they still have not. Whatever the year, it did not stop hard drive exabytes growing 39.8 per cent year over year in Seagate’s fiscal 2026. The two media stopped competing for the same bytes. Flash took the bytes somebody waits on and disk kept the bytes nobody waits on, and both grew.
Areal density: the wall was called in 1994 and moved three times
The reason hard drives stayed cheap per terabyte is that areal density kept rising, and the reason that stopped being reliable is that it slowed badly. The decade averages, from industry compilations rather than a single primary source:
| Period | Areal density growth |
|---|---|
| 1980s | about 30 per cent a year |
| 1990s | about 60 per cent a year, peaking near 100 around 1997-2002 |
| 2000s | about 39 per cent a year |
| 2009-2018 | about 7.6 per cent a year |
Seven and a half per cent a year is the number that broke the business model. Capacity per drive kept climbing because vendors added platters and sealed the enclosure with helium so they could fit more of them, but adding platters is the expensive lever, and it runs out at ten or eleven.
The physics wall was called long before it arrived. Stan Charap at Carnegie Mellon warned in 1994 that normal product evolution would hit thermal-stability trouble at about 40 Gbit per square inch within a decade. The Exos M ships at 1841 Gbit per square inch. That is roughly 46 times past the wall, and the industry got there by changing the media, the head and the write mechanism repeatedly rather than by proving Charap wrong. He was right about the physics and wrong about the industry’s willingness to replace the physics.
The recording trilemma, stated properly
Seagate states the constraint in its own technology paper TP707 and it is worth following exactly, because the popular version of the HAMR explanation gets the causality backwards.
One. Signal-to-noise ratio scales with the logarithm of the number of grains per bit. Seagate writes it as SNR proportional to log10(N). To keep SNR usable as bits get smaller, you need to keep roughly the same number of grains under each bit, so the grains must shrink in proportion.
Two. A grain’s resistance to being flipped by thermal energy depends on the product of its magnetic anisotropy K and its volume V. Shrink V and the energy barrier falls. Below a certain point ambient thermal energy flips grains on its own and the bit rots at room temperature. That is superparamagnetism, and it is the wall Charap named.
Three. The obvious fix is to raise K to compensate for the smaller V. The problem is that a high-K grain needs a bigger field to write. Head field comes from the saturation magnetisation of the primary pole, and that hit a hard physical ceiling - the Slater-Pauling limit on the alloys available - around 2010, after sixty years of increases.
So: you must shrink the grain, shrinking it makes it unstable, stabilising it makes it unwritable, and the write field cannot be increased. Three constraints, no free variable. That is the trilemma, and heat is the only thing that resolves it.
Here is where the common explanation goes wrong. Heat does not make a grain more stable. It does the opposite, deliberately and temporarily. Coercivity falls sharply as a material approaches its Curie temperature, so a spot of iron-platinum superlattice media that is unwritable at room temperature becomes writable for the instant it is hot. The head writes it in that instant, the spot cools, and the high anisotropy that made it unwritable now makes it stable for years. Seagate’s paper states the whole heating, writing and cooling cycle takes under one nanosecond, and that the media overcoat had to be re-engineered to survive being heated past 400 °C while still maintaining a reliable head-media interface.
Why HAMR needs a plasmonic transducer and not a brighter laser
The hot spot has to be smaller than a track. Tracks on the Exos M are a little over 34 nanometres apart, from its published 738 kTPI. Focused light cannot do that. Far-field optics are diffraction-limited: Blu-ray, at the short end of consumer optical, has a spot around 238 nanometres, and Seagate’s paper puts the floor for near-field optical tricks at about 100 nanometres. Both are an order of magnitude too big.
The answer is a near-field transducer: a plasmonic antenna that takes light from the laser diode and converts it into a surface plasmon concentrated at a peg whose width, not the wavelength, sets the size of the hot spot. The optics stop mattering and lithography starts mattering, which is a problem the industry knows how to attack. Seagate states it had manufactured more than 25 million near-field transducers over the development programme, which is the useful measure of how hard the component was: the physics was demonstrated in the 2000s and the manufacturing took another fifteen years.
Seagate’s own stated ceilings, from the same paper: perpendicular recording “will eventually run out of steam just over 1 terabit per square inch” and HAMR “should take data density up to about 5” terabits per square inch. Public HAMR demonstrations had reached 2.0 terabits per square inch at the time the paper was written. Shipping product is at 1.841. The gap between the shipping figure and the stated ceiling is under a factor of three, which is the most honest thing anybody has published about how long this generation lasts.
What HAMR actually bought, measured from two product manuals
The Exos X24 is the perpendicular drive HAMR replaces and the Exos M is the HAMR drive that replaces it. Both product manuals publish track density, linear recording density and average areal density, which almost nothing else in consumer storage does. Put them side by side:
Exos X24 (PMR) Exos M (HAMR) change
capacity 24 TB 32 TB +33.3%
heads / disks 20 / 10 20 / 10 unchanged
per-platter capacity 2.4 TB 3.2 TB +33.3%
spindle 7200 rpm 7200 rpm unchanged
track density 512 kTPI 738 kTPI +44.1%
linear density 2552 KBPI 2470 KBPI -3.2%
areal density (average) 1260 Gb/in² 1841 Gb/in² +46.1%
cache buffer 512 MB 512 MB unchanged
internal transfer rate 2951 Mb/s 2852 Mb/s -3.4%
sustained transfer, OD 285 MB/s 285 MB/s unchanged
Read the middle block twice. HAMR’s gain is almost entirely track density. Linear density went backwards. The bits are not shorter along the track; they are the same length, very slightly longer in fact, and the tracks are packed 44 per cent more densely - a pitch of 34.4 nanometres against 49.6, which is about 31 per cent narrower. Density and pitch are reciprocals, so the two percentages are not interchangeable and the smaller one is the one describing the gap between tracks. That is the single most under-reported fact about the technology, and it falls straight out of two published tables.
It also explains the row that looks like a misprint. The internal data transfer rate fell 3.4 per cent, which is what you would expect if the bits under the head got marginally longer. Bytes per revolution are set by linear density and track circumference, not by how many tracks there are, so squeezing tracks together adds capacity without adding a single byte per revolution.
Two caveats on that table, in the interest of not overstating it.
The first is small. Multiplying the two published densities does not exactly reproduce the published average areal density on either drive - 738 × 2470 gives 1823 against a stated 1841, and 512 × 2552 gives 1307 against a stated 1260. The published averages are measured over the whole surface while track and linear densities are quoted at particular radii, so a disagreement of a few per cent is expected rather than alarming.
The second is larger, and this guide cannot resolve it. Both drives carry 20 heads on 10 disks, so the recording surface is identical, and on identical surfaces capacity is average areal density multiplied by that area and nothing else. A 46.1 per cent rise in average areal density on the same surfaces does not give a 33.3 per cent rise in capacity:
24 TB × 1.461 = 35.1 TB implied by the density row
published capacity = 32 TB
capacity actually gained = +33.3%
areal density consistent with +33.3% 1841 / 1.333 = 1381 Gb/in²
stated for the X24 = 1260 Gb/in²
Either the X24’s 1260 is measured on a different basis from the Exos M’s 1841, or one of the two figures is misquoted in the source this guide worked from. The +44.1 per cent track density and the +33.3 per cent capacity agree with each other and with the head and disk counts; the +46.1 per cent areal density is the row that does not. The 1381 is derived here to show the size of the disagreement and is not a specification anybody publishes. Read the track density row as the story and treat the areal density row as provisional.
Now the consequence, which is the part that decides the purchase:
bandwidth per terabyte Exos X24, 24 TB 285 / 24 = 11.9 MB/s per TB
Exos M, 32 TB 285 / 32 = 8.9 MB/s per TB
-25.0%
Sequential bandwidth per terabyte fell a quarter in one product generation. The Exos M’s published sustained outer-diameter rate scales with capacity point - 285 MB/s at 32 TB, 275 at 30, 270 at 28, 240 at 24 - so the effect holds across the family:
| Drive | Capacity | Sustained OD | MB/s per TB |
|---|---|---|---|
| Exos X24 | 24 TB | 285 MB/s | 11.88 |
| Exos M | 24 TB | 240 MB/s | 10.00 |
| Exos M | 28 TB | 270 MB/s | 9.64 |
| Exos M | 30 TB | 275 MB/s | 9.17 |
| Exos M | 32 TB | 285 MB/s | 8.91 |
Note the second row. At the same 24 TB capacity, the HAMR part is rated slower than the perpendicular part it replaces. Why is not stated in the manual - fewer platters fitted, a different format, or lower-density media on the smaller parts would all produce it - and the manual does not say which. Every step up the capacity ladder makes bandwidth per terabyte worse, and nothing in the specification offsets it: both drives publish the same 512 MB multisegmented cache, so there is no firmware compensation to point at even if you wanted one.
Access density is the number that actually broke
Areal density is the number the industry publishes. Access density is the number that determines whether a drive is usable, and nobody publishes it because it has been going the wrong way for fifteen years.
A hard drive’s random IOPS is set by one actuator and one spindle speed. Neither has changed since the 2000s. Capacity per drive has gone up by more than an order of magnitude in the same period. Divide one by the other:
2010 2 TB at ~80 random IOPS = 40.0 IOPS per TB
2026 32 TB at ~80 random IOPS = 2.5 IOPS per TB
capacity ×16, IOPS unchanged, IOPS per terabyte ÷16
The same collapse in sequential terms, from the numbers above: 11.9 MB/s per terabyte on the drive being retired, 8.9 on the drive replacing it. A modern nearline drive is a slower device per unit of stored data than anything sold in the last twenty years, and it gets slower every generation by design.
The operational consequence is the whole-drive pass. Every rebuild, every scrub, every migration and every full backup has to read the entire surface:
best case, outer tracks only, which no real workload achieves
32 TB ÷ 285 MB/s = 112,281 s = 31.2 hours
44 TB ÷ 300 MB/s = 146,667 s = 40.7 hours
realistic, inner tracks run at roughly half the outer rate
average ≈ (285 + 143) / 2 = 214 MB/s
32 TB ÷ 214 MB/s = 149,533 s = 41.5 hours
A single-drive full-surface read is now a forty-hour operation, before any contention from live traffic. That is the real constraint on RAID rebuild, and drives for a NAS works through what it does to array design. It is worth holding that figure next to the drive’s own workload rating, which the section below on annualised workload rate does.
This is also the best available evidence for the argument, because the vendors have started spending money on it. Seagate’s MACH.2 dual-actuator drives put two independent actuators in one enclosure, transferring concurrently, reaching up to 480 MB/s sustained - Seagate’s phrase is “the fastest ever from a single hard drive” - and roughly doubling per-drive IOPS; Microsoft reported near-doubled IOPS with the Exos 2X14. Western Digital announced Dual Pivot actuators for 2028, with up to twice the sequential I/O, and reduced platter spacing to fit eleven disks where Seagate fits ten.
Which means the flat assertion in the older version of this guide needs a qualifier. A hard drive has one actuator - every drive you are likely to buy, and every drive on this site’s listings, but no longer every drive that exists. And even with two:
dual-actuator 14 TB at ~160 IOPS = 11.4 IOPS per TB
Still a quarter of what a 2 TB drive delivered in 2010. Doubling the actuator count buys back one of the four doublings that were lost.
The roadmap to 2030, and where delivered product stops
Roadmaps are marketing until they ship, and the honest version of that rule requires drawing the line where it actually sits rather than where it was two years ago.
Delivered, shipping in volume. Seagate’s Mozaic 4+ platform: 4 TB or more per disk, ten disks, up to 44 TB per drive, qualified and deployed at two hyperscale cloud providers as of March 2026, with second-generation superlattice platinum-alloy media, a second-generation plasmonic writer with an integrated nanophotonic laser, an eighth-generation spintronic reader and a 7 nm integrated controller. Western Digital’s 30 TB conventional ePMR is available now. Toshiba’s MG11 at 24 TB conventional and MA11 at 28 TB shingled ship on ten-disk helium-sealed 7200 rpm platforms using flux-control microwave-assisted recording, in which a spin-torque oscillator on the write head emits a microwave field that assists writing to high-coercivity media. Three suppliers, three different write mechanisms, all in production.
In qualification. Western Digital’s 40 TB UltraSMR ePMR, at two hyperscalers, with volume production stated for the second half of 2026.
Announced with dates, not delivered. Western Digital’s 60 TB ePMR using HAMR-derived innovations; its own HAMR, with customer certification completing and mass production stated for 2027, scaling to 100 TB by 2029; its High Bandwidth Drive technology claiming up to twice current bandwidth with a path to eight times. Seagate’s stated path from 4 TB per disk to 10 TB per disk and therefore 100 TB drives. Toshiba’s 40 TB in 2027 and 55 TB by 2030.
Seagate’s January 2025 investor release is the clearest statement of what the platform is supposed to buy: 3.6 TB per platter shipping at the time, a ten-platter design reaching 36 TB, laboratory demonstrations above 6 TB per disk, a path to 10, and claims of 25 per cent lower cost per terabyte, 60 per cent lower power per terabyte and 300 per cent more capacity in the same data-centre footprint. Read the last three as vendor claims, because they are.
Shingling deserves its own line in this, because it is the cheap lever and vendors reach for it first. Seagate’s own account puts SMR’s gain at 25 per cent when introduced in 2014, achieved without new production capital because it reuses conventional reader and writer elements - it is a track layout, not a write technology. The current premium is roughly a third: Western Digital offers 30 TB conventional against 40 TB UltraSMR on the same platform, having offered 26 against 32 in the previous generation. That is what a buyer is offered in exchange for a drive that only writes sequentially.
The standards are worth naming because they are how you tell which kind you have: SCSI Zoned Block Commands, ANSI INCITS 536-2016, and the ATA Zoned Device Command Set, INCITS 537-2016. Three models exist - drive-managed, where firmware hides everything and the drive presents as a conventional replacement; host-aware, with sequential-write-preferred zones and backward compatibility; and host-managed, where writes to a zone must be sequential and the host has to obey. One detail that catches filesystem authors: the zone-append command defined for zoned-namespace SSDs is not defined for SMR drives and must be emulated in host software. CMR vs SMR covers the behaviour and, more usefully, how to find out which one a listing is actually selling.
NAND wafer economics, and why QLC did not collapse the price
NAND cost per bit falls through three levers, and all three are getting harder.
Layers. 3D NAND stacks cells vertically. In production as of the 2025-26 cycle: SK hynix at 321 layers, Samsung at 286, Micron at 276. Samsung’s next generation targets more than 400 layers for 2026, TechInsights projects products above 500 layers within two to three years, and 600 to 700 layer packages within five using hybrid bonding. Every one of those layers is deposition and etch, and the high-aspect-ratio etch through several hundred layers is the hardest process step in the industry.
Bits per cell. Each extra bit needs twice as many distinguishable voltage levels: SLC 2, MLC 4, TLC 8, QLC 16, PLC 32. Going from TLC to QLC buys 33 per cent more capacity in exchange for halving the margin between states, which costs program time, read time, ECC strength and endurance together. Read latency has roughly doubled with each additional bit per cell, because the controller has to distinguish more thresholds. The returns shrink as the costs climb, which is why PLC has been discussed for years and shipped almost nowhere.
Shrink. Lateral scaling continues, modestly. The era of easy planar shrinks ended before 3D NAND started.
QLC was supposed to be the lever that ended the argument, and it delivered the density it promised. Micron ships QLC die up to 2 terabits, the largest single NAND die in volume production. What it did not deliver is the price, and the reason is not technical. Producers steered wafer capacity towards higher-margin enterprise SSDs, and fab transitions to higher layer counts lengthened ramp cycles. A 30 TB QLC drive that cost about $2,450 in Q2 2025 rose more than 470 per cent within a year. QLC did not fail; its output was redirected. That is a commercial fact rather than a physical one, which means it can reverse, and which is exactly why this guide will not forecast it.
The vendor counter-case deserves a fair hearing, because QLC datacentre parts are genuinely good products. Solidigm’s D5-P5316 is specified at up to 7,000 MB/s sequential read, 800,000 4K random read IOPS - up 38 per cent on the previous generation’s 580,000 - a 99.999th-percentile read latency of 600 microseconds against 1,150 for the prior generation, and 0.41 drive writes per day, which is 22,930 TBW at 30.72 TB. The D5-P5336 reaches 122.88 TB on a PCIe 4.0 x4 link with up to 1,005,000 4K random read IOPS. One of those drives holds nearly three times what the largest hard drive shipping in volume holds - 122.88 TB against the 44 TB Mozaic 4+ listed above - in a form factor a fraction of the size. It also costs, at the reference prices above, something in the region of eighteen times as much per terabyte.
The latency argument, derived
A hard drive’s random access time is the sum of two mechanical waits.
Rotational latency is the time for the sector you want to arrive under the head.
On average that is half a revolution. One revolution takes 60000 / rpm
milliseconds, so average rotational latency is 30000 / rpm milliseconds and
nothing a firmware engineer does changes it:
| Spindle | One revolution | Average rotational latency |
|---|---|---|
| 5,400 rpm | 11.11 ms | 5.56 ms |
| 7,200 rpm | 8.33 ms | 4.17 ms |
| 10,000 rpm | 6.00 ms | 3.00 ms |
| 15,000 rpm | 4.00 ms | 2.00 ms |
The Exos M product manual publishes 4.16 ms average latency, which is
30000 / 7200 = 4.167 rounded down. It is the one performance figure on a hard
drive datasheet that is arithmetic rather than measurement.
Seek time is the actuator moving the head to the right cylinder, and here the documentation has quietly changed. The 32 TB Exos M product manual does not publish a seek time anywhere. The word does not appear in its specification tables at all; it occurs only in the list of ATA command opcodes. Consumer drives stopped publishing seek years ago and now the flagship enterprise manual has stopped too. Add the two waits and divide into one second for a queue-depth-one random IOPS figure, remembering that the seek column is now a range across vendors and generations rather than a specification:
| Class | Rotational | Seek (typical, not specified) | Access | Random IOPS at QD1 |
|---|---|---|---|---|
| 5,400 rpm 3.5” | 5.56 ms | unpublished | ~17 ms | ~60 |
| 7,200 rpm 3.5” | 4.17 ms | 8-9 ms | ~12.5 ms | ~80 |
| 10K 2.5” SAS | 3.00 ms | 3.5-4 ms | ~7 ms | ~140 |
| 15K 2.5” SAS | 2.00 ms | 2-3 ms | ~4.3 ms | ~230 |
The rotational column is exact. The seek column is a range, the drive you buy may sit outside it, and on a current nearline drive it is not a published quantity at all. Flash gets its own table below, because the numbers are not comparable in the way people assume.
The spindle stopped at 15,000 rpm for reasons that are also physics. At 15K the outer track of the reduced-diameter platter those drives used moves at about 49 metres per second. Drag force rises with the square of velocity and the power to overcome it with the cube:
7200 -> 15000 rpm at the same radius
velocity ratio 15000 / 7200 = 2.08×
drag force ratio 2.08² = 4.34×
windage power ratio 2.08³ = 9.04×
Nine times the windage power, all of it turning into heat and turbulence inside a sealed enclosure where the head flies a few nanometres above the surface. That is why 15K drives used smaller platters, and why nobody built a 20K.
Queue depth, decomposed
The gap widens dramatically once you queue work up, and this is where the hand-waving usually happens. It does not have to.
The SATA queue-depth ceiling of 32 is not folklore; the drive reports it. IDENTIFY
DEVICE word 75 reads 001Fh on both the Exos M and the Exos X24, which is 31,
and the field is maximum queue depth minus one, so 32 outstanding native command
queueing commands is the hard ceiling for anything on a SATA link. NVMe’s
equivalent is 65,536 queues of 65,536 commands each. That is not an incremental
improvement, it is a different order of object.
The per-command host cost is quantified too. NVM Express’s own comparison puts AHCI at four uncacheable register reads per command at roughly 2,000 CPU cycles each, and NVMe at zero. AHCI also requires a lock to issue a command; NVMe has a doorbell register per queue and no locking.
AHCI per-command host tax
4 uncacheable register reads × ~2,000 cycles = 8,000 cycles
at ~3.2 GHz ≈ 2.5 µs
against a 7,200 rpm access of 12,500 µs = 0.02% irrelevant
against an NVMe device read of 20-70 µs = 3.6-12.5% not irrelevant
at 90,000 IOPS: 90,000 × 8,000 = 7.2e8 cycles/s = 22% of one 3.2 GHz core
at 1,000,000 IOPS: 1,000,000 × 8,000 = 8.0e9 cycles/s = 2.5 cores, issuing only
The overhead that was invisible behind a spinning platter became as much as an eighth of the latency budget behind flash, and a fifth of a CPU core at SATA SSD rates. That is the whole reason NVMe exists, and it is measurable rather than rhetorical. NVM Express’s measured protocol overhead comparison makes the same point against SAS: 6.0 microseconds and 19,500 cycles per command for SCSI over SAS, against 2.8 microseconds and 9,100 cycles for NVMe. NVMe also fetches all the parameters for a 4 KB command in a single 64-byte DMA fetch, and has an I/O command set of six commands.
Now the QD1 and queued figures for flash, and these need splitting in a way the
older version of this guide did not do. Device latency and end-to-end latency
are different numbers, and quoting one while computing IOPS from the other
produces nonsense. Device-level random read latency is roughly 20 to 70
microseconds for NVMe and 100 to 200 for a SATA SSD. Dividing one second by 70
microseconds gives 14,000 IOPS, which is a real number for a single outstanding
read through a real kernel with a real filesystem, but it is not 1 / device latency, because the protocol and software budget above sits on top.
The queued ceilings are where the useful comparison lives:
SATA SSD at the interface ceiling
perfect 32-way concurrency at 100 µs device latency
32 / 100 µs = 320,000 IOPS theoretical
observed ceiling ≈ 90,000 IOPS
implied end-to-end latency 32 / 90,000 = 356 µs per command
7,200 rpm drive
80 IOPS × 4 KiB = 320 KiB/s of random 4K reads
SATA SSD at QD32
90,000 × 4 KiB = 352 MiB/s
Where the missing 250 microseconds go is not something a datasheet will tell you: link round trips, frame overhead, the serialisation forced by having exactly one queue, and the drive’s own internal scheduling. Three orders of magnitude, from the same interface, on the same cable. That is the entire reason 15K SAS died as a product category. It delivered roughly three times a desktop drive’s IOPS for many times the price per terabyte, and flash arrived offering a thousand times. The market did not shrink, it vanished, and the remaining SAS listings on the SAS pages are capacity drives rather than performance drives.
At the top of the flash range the ceiling has moved again. High-end consumer PCIe 5.0 parts are now rated above two million 4K random read IOPS, and a datacentre QLC drive like the D5-P5336 is rated at 1,005,000. Those are ratings at high queue depth across many threads and they say nothing at all about what one process doing one read at a time will see, which remains the 20 to 70 microsecond number.
Inside a flash read: a ladder whose bottom rung is a hard drive
Flash’s median latency is excellent and its tail is not, and the reason is published in the academic literature in a form that almost nobody translates for buyers. A read from NAND is not one operation with one cost. It is a ladder of error-correction attempts, and the controller climbs it until the data comes back clean. The rungs and their costs, from Cai, Ghose, Haratsch, Luo and Mutlu’s survey of flash errors and recovery:
| Rung | Cost | Codeword failure rate after |
|---|---|---|
| BCH hard decoding with read-retry | 70 µs per iteration | 10^-4 at the first attempt |
| LDPC hard decoding | 80 µs | improving |
| Neighbour-cell-assisted correction | 140 µs for two neighbour reads, plus 70 µs per additional adjacent value | improving |
| LDPC soft decoding | 80 µs per level, so about 400 µs at five levels | improving |
| Superpage-level parity recovery | about 10 ms | 10^-15 |
Look at the bottom row. Ten milliseconds is a hard drive. A 7,200 rpm access is about 12.5 milliseconds, so a worn or hot SSD that falls all the way through to superpage parity recovery serves that read in roughly the time the platter it replaced would have taken. The ratio from the top of the ladder to the bottom is about 125 to 1.
fast path 80 µs
worst path 10,000 µs
ratio 125×
7,200 rpm HDD random access ≈ 12,500 µs
This is why enterprise SSD datasheets quote percentile latency rather than average, and why the percentile they quote keeps getting deeper. Solidigm’s D5-P5316 publishes a 99.999th-percentile read latency of 600 microseconds - roughly six times a typical median, and an improvement from 1,150 microseconds in the previous generation. Flash wins decisively in the median and the gap closes sharply in the tail, and a consumer drive that publishes no percentile figure at all is telling you it would rather you did not look.
The practical reading is that the ladder gets climbed more often as a drive wears, heats up, or sits unpowered between reads, because all three widen the threshold voltage distributions the controller is trying to separate. A drive that used to be fast and is now intermittently slow is usually not broken. It is spending more time on rung three.
Program and erase are the expensive operations
Reading NAND is fast. Writing it is not, and erasing it is very much not. The representative device timings below come from generation summaries and conference papers rather than a current datasheet, because current NAND datasheets sit behind vendor registration walls and this site has not verified one. Treat them as the right order of magnitude for the generation and not as a specification for any part you can buy:
| Operation | Typical | Maximum |
|---|---|---|
| TLC page read, tR | 77-100 µs | - |
| TLC page program, tPROG | 1,300-1,630 µs | 2,500-5,000 µs |
| Block erase, tBERS | about 15 ms | about 45 ms |
| QLC page program (1 Tb device, published target) | 3 ms | - |
| QLC page read (same device) | 127 µs | - |
program ÷ read 1,500 / 90 ≈ 17×
erase ÷ read 15,000 / 90 ≈ 167×
erase, worst case = 45 ms, which is 3.6 hard drive accesses
Those are per-die numbers, and they are the reason an SSD is built the way it is. A single NAND die is not fast. Take a drive with eight channels and four dies per channel, which is a plausible mid-range arrangement and mine rather than a vendor’s:
32 dies operating concurrently at tR = 90 µs
32 / 90 µs = 355,000 page reads per second
The IOPS come from parallelism, not from speed. That has two consequences worth carrying. A small drive has fewer dies, so a 500 GB model of the same product is genuinely slower than the 2 TB model and not just artificially limited. And a workload that hits the same die repeatedly gets none of the parallelism, which is why a queue depth of one looks so much worse than the headline figure.
The program-time number is also where TLC and QLC stop being able to write in one pass. Neither can be programmed in one shot: they use two-step and foggy-fine programming, staging intermediate bit values into reserved single-level buffers on the die and then applying the smallest incremental step-pulse programming voltages to reach the final threshold voltages. That staging is the mechanism behind both the SLC cache and the program-time penalty, and it is the same mechanism, not two. An MLC-class flash block holds 256 to 1,024 pages, which is the granularity at which erase has to happen and the origin of every write-amplification argument in the next section but one.
Sequential throughput is much closer than people assume
Random access is where flash wins by a thousand. Sequential is where it wins by two or three over SATA and twenty or thirty over PCIe, which surprises people in both directions.
A current nearline drive is faster than the older version of this guide said. The Exos M is rated at 285 MB/s sustained on its outer diameter in its product manual. Western Digital quotes up to 302 MB/s on its 26 TB Ultrastar, with current nearline parts reported in the 250 to 275 MB/s band, and claims its High Bandwidth Drive work takes that above 500. The ceiling on a SATA SSD is set by the interface, not the flash:
| Link | Raw rate | Encoding | Usable ceiling |
|---|---|---|---|
| SATA 6 Gb/s | 6.0 Gb/s | 8b/10b | 600 MB/s, ~550 in practice |
| SAS-3 12 Gb/s | 12.0 Gb/s | 8b/10b | 1,200 MB/s |
| SAS-4 22.5 Gb/s | 22.5 Gb/s | 128b/150b | 2,400 MB/s |
| PCIe 3.0 x4 (2010) | 8 GT/s per lane | 128b/130b | 3.938 GB/s |
| PCIe 4.0 x4 (2017) | 16 GT/s per lane | 128b/130b | 7.877 GB/s |
| PCIe 5.0 x4 (2019) | 32 GT/s per lane | 128b/130b | 15.754 GB/s |
| PCIe 6.0 x4 (2022) | 64 GT/s per lane | PAM-4, FLIT with FEC | 30.25 GB/s |
| PCIe 7.0 x4 (2025) | 128 GT/s per lane | PAM-4, FLIT with FEC | 60.5 GB/s |
Two things in that table are worth stopping on. PCIe 6.0 abandoned 128b/130b entirely: it uses four-level pulse amplitude modulation with forward error correction and a flow-control-unit framing, which is why its usable figure is not simply twice the previous row’s raw rate divided by eight. SAS-4 made a smaller version of the same move, dropping 8b/10b for 128b/150b, so the encoding column stops generalising after SAS-3. Drive interfaces works through which of these a given slot actually gives you, which is rarely the one printed on the box.
So a good hard drive against a SATA SSD on a long sequential copy is a factor of two, not a hundred. Against a PCIe 4.0 NVMe drive it is nearer thirty. And there is a new observation hiding in the table: hard drives are about to run out of SATA. A dual-actuator drive already reaches 480 MB/s sustained against SATA’s practical 550, and a High Bandwidth Drive at 500-plus would saturate the link. The interface that has been over-provisioned for platters since 2009 stops being so within this roadmap.
The hard drive’s sequential figure is also not one number. Platters spin at constant angular velocity while linear bit density is held roughly constant, so bytes per revolution scale with the circumference of the track, which scales with radius. The data band on a 3.5-inch platter runs from an outer radius to an inner one roughly half of it:
throughput ∝ 2πr × linear density × revolutions per second
∝ r
outer tracks : inner tracks ≈ 2 : 1
A drive that opens a benchmark sweep at 285 MB/s will finish it near 145. That is geometry, not a fault, and it is why manufacturers quote “up to” and why the whole-surface pass arithmetic above uses an average rather than the headline. Confirm it on a drive you own with a buffered read at the start and near the end:
sudo hdparm -t --direct /dev/sda # buffered read from the outer tracks
sudo hdparm -t --direct --offset 14000 /dev/sda # 14,000 GiB in, near the end of a 16 TB drive
--offset counts in GiB, so pick a number that actually lands inside your drive.
The --direct flag reads with O_DIRECT, so the transfer skips the page cache
entirely; plain -t flushes that cache first and still reads the platter, just
with an extra copy in the path. It is -T that times your RAM.
Both slow down when full, and only flash falls off a cliff
The hard drive’s decline is the radius argument above. Files written late land on inner tracks, so a drive at 95 per cent full is writing at roughly half its headline rate. Predictable, gradual, and visible in a sweep.
The SSD’s decline is about free blocks. NAND is written in pages and erased in much larger blocks - 256 to 1,024 pages per block - so the controller can never overwrite in place. It writes elsewhere, updates a mapping table, and garbage-collects the stale pages later. That collection needs somewhere to put live data, and the pool of spare blocks is smaller than people think.
Every drive has a hidden reserve from the base-2 versus base-10 gap alone, and the convention that creates it is documented rather than accidental. JEDEC defines SSD capacity with the same formula IDEMA uses for hard drives, and Seagate’s Alvin Cox, chairing JEDEC’s JC-64.8 committee, states it was adopted deliberately at OEM request so that one formula would cover both media:
capacity in GB = (user-addressable LBA count - 21168) / 1953504
1,953,504 sectors × 512 bytes = 1,000,194,048 bytes ≈ 1.0002 GB
The small excess is deliberate, and it guarantees the drive delivers at least its stated decimal capacity. Drive capacity explained covers what that does to the number your operating system shows you. What it does to an SSD is create the spare pool:
Advertised 1 TB = 1,000,000,000,000 bytes
NAND actually fitted 1 TiB = 1,099,511,627,776 bytes
Hidden spare = 99,511,627,776 bytes
= 9.05% of the installed NAND
Enterprise drives add far more on top, deliberately, which is part of what you pay for - see enterprise drives. The label gives it away: a 400 / 800 / 1,600 / 3,200 GB ladder is 37.4 per cent over-provisioned and write-intensive; a 480 / 960 / 1,920 / 3,840 ladder is 14.5 per cent and read-intensive.
The second mechanism is the SLC cache, and its arithmetic is the one to understand. A consumer TLC drive writes incoming data one bit per cell into blocks that will eventually hold three, because single-level programming needs only one coarse threshold rather than the foggy-fine staging described above. Caching one gigabyte therefore occupies three gigabytes of eventual TLC capacity, and on a QLC drive one occupies four. Most consumer drives size this cache dynamically from whatever is free, so:
Empty 2 TB TLC drive: large dynamic cache, writes at interface speed
Drive at 90% full: ~200 GB free ÷ 3 ≈ 67 GB of cache left
Drive at 98% full: essentially none - writes go straight to TLC
Direct-to-TLC write speed on consumer parts is commonly well under 200 MB/s, and direct-to-QLC lower still. A nearly full QLC drive, past its cache, can write more slowly than a hard drive. That is not a defect; it is the product working as designed, and it is the single most common reason someone concludes their SSD has “gone bad”.
The practical rule is the same for both and easier to follow than to like: keep 15 to 20 per cent free. On flash it preserves the cache and keeps write amplification down; on platters it keeps you off the inner tracks.
Write amplification, and the JEDEC endurance formula
Write amplification is the ratio of bytes written to NAND to bytes written by the host, and the JEDEC presentation gives a worked case that makes the whole argument in four lines. Take a block of 64 pages holding 256 sectors, and have the host write 8 sectors into the middle of it by read-modify-write:
host write 8 sectors
NAND write 256 sectors (the whole block, rewritten)
write amplification 256 / 8 = 32×
same block written sequentially, in full
write amplification ≈ 1×
Small random writes consume flash. Total volume does not. Two workloads that write the same number of gigabytes can differ by a factor of thirty in how much NAND they burn, which is why a drive’s endurance rating is a statement about an assumed workload rather than about the drive.
JESD218’s endurance formula makes that explicit, and it shows where the margin goes:
TBW < (C × NAND cycling capability) / (2 × WAF)
C is capacity and the factor of two is a guard band for uneven wear levelling, the assumption being that the most heavily cycled block receives twice the average number of cycles. Work it for a 2 TB TLC drive at the 1,000 program/erase cycles per block that Cai and colleagues give for TLC, and this working is mine:
WAF = 1.0 (large sequential writes) 2 × 1000 / (2 × 1.0) = 1,000 TB
WAF = 2.5 (mixed desktop) 2 × 1000 / (2 × 2.5) = 400 TB
WAF = 32 (JEDEC's small-random case) 2 × 1000 / (2 × 32) = 31.25 TB
The same drive, the same silicon, a thirty-two-fold spread in rated life depending entirely on how it is written. Endurance per cell has been falling for a decade as lithography shrank and bits per cell rose - roughly 150,000 cycles for SLC, 10,000 for 5x-nanometre MLC, 3,000 for 1x-nanometre MLC, and about 1,000 for TLC - which makes the write amplification term the dominant one. SSD endurance works through what to do about it and how to read the odometer on a used drive.
Power, and why it decides a multi-bay build
The Exos M product manual publishes a full DC power table measured at 35 °C ambient, which makes this one of the few places in storage where you can size a power supply from primary documentation rather than from guessing.
| State | Exos M, 32 TB | Exos X24, 24 TB |
|---|---|---|
| Average idle | 6.763 W | 6.22 W |
| Idle_A | 6.696 W | - |
| Idle_B | 4.125 W | 3.89 W |
| Idle_C | 3.136 W | 2.91 W |
| Standby | 1.196 W | 1.09 W |
| Random read, 4K QD16 | 9.269 W | - |
| Random write, 4K QD16 | 8.327 W | - |
| Sequential read, 64K QD16 | 8.020 W | - |
| Sequential write, 64K QD16 | 8.929 W | - |
| Typical operating | - | 8.88 W |
| Maximum | - | 9.02 W |
Those running figures are unremarkable. The number that actually sizes an enclosure is the one nobody looks at: both manuals specify a typical startup current of 2.6 A peak on the +12 V rail, with a 2.0 A reduced-current profile selectable through SMART Command Transport. Spin-up draws roughly four and a half times what the drive draws at idle and three and a half times what it draws under load, and it does it on one rail:
per drive at spin-up 2.6 A × 12 V = 31.2 W
per drive running 6.763 W idle / 9.269 W under random read
twelve-bay chassis, all spindles starting together
12 × 2.6 A = 31.2 A on +12 V
12 × 31.2 W = 374.4 W
the same twelve drives, running
12 × 6.763 W = 81.2 W idle
12 × 9.269 W = 111.2 W under random read
ratio 374.4 / 111.2 = 3.4×
A twelve-bay enclosure needs three and a half times the 12 V capacity at the moment of power-on that it needs for the rest of its life. Sizing the supply for that is wasteful and sizing it for the running figure produces a chassis that browns out at boot and appears to have flaky drives.
The mechanism that resolves it is on both drives and is worth asking about when you buy a backplane: PUIS, Power-Up In Standby, an ATA feature bit that lets the drive come up without spinning until the host tells it to. Both the Exos M and the Exos X24 report the PUIS supported and enabled bits in their IDENTIFY data. A backplane that staggers spin-up two drives at a time turns a 31.2 A inrush into a 5.2 A one, and three at a time into 7.8 A. It costs only time. Budget that time: the Exos M specifies 30 seconds typical and 60 seconds maximum from power-on to ready, the X24 25 and 30, and both manuals note that an unexpected power loss or a cold or hot extreme can add five to twenty seconds beyond typical. A twelve-bay array staggered in pairs is therefore three minutes to ready in the good case and six in the bad one, which is worth knowing before you conclude a controller has hung.
Watts per terabyte is where the dense drive earns its keep, and it is the metric that genuinely favours platters at scale:
Exos M 6.763 W ÷ 32 TB = 0.211 W/TB (÷ 30 TB = 0.225 W/TB)
Exos X24 6.220 W ÷ 24 TB = 0.259 W/TB
Divide each drive’s measured idle power by its own capacity, which is the only pairing the power table supports: the Exos M column is headed 32 TB, so 32 is the divisor and the 30 TB line is there only because the 30 TB part exists.
Seagate claims 60 per cent lower power per terabyte for the HAMR platform as a whole, which is a platform claim across a generation rather than the idle-state comparison above; the idle figures alone, 32 TB against 24 TB, give about 18 per cent. Over a year of continuous idle, 0.211 W/TB is about 1.85 kWh per terabyte per year, which at typical domestic tariffs is small enough that a home user should ignore it and large enough that a rack operator cannot.
The comparison a buyer wants is against flash, and this guide will not print one, because enterprise SSD idle power is not published in a form comparable to a hard drive’s state table. The state machines are different, the measurement conditions are different, and the honest position is that the comparison is not available rather than that it favours one side.
Noise and vibration in a box with other drives
Acoustics are published in bels of A-weighted sound power per ISO 7779, which is not the same thing as the decibel figure a phone app measures, and both the HAMR and perpendicular nearline drives are identical:
| Mode | Typical | Maximum |
|---|---|---|
| Idle | 2.8 bels | 3.0 bels |
| Performance seek | 3.2 bels | 3.4 bels |
Seagate defines the seek-mode test rate as 0.4 / (average latency + average access time) seeks per second, which is a specific duty cycle rather than
continuous thrashing, so a drive doing genuinely constant random work is louder
than its published seek figure. One footnote in the manual is worth carrying
across: idle acoustics and idle power can both rise to operational levels when
the drive runs its own SMART offline activity. A drive that gets audibly busy
while nothing is using it is usually doing its background self-test, not
failing.
Vibration is the more useful specification because it is stated as a throughput guarantee rather than a survival guarantee. The Exos M manual says the drive “will exhibit greater than 90 per cent throughput for sequential and random write operations” while subjected to shaped random rotary vibration of 12.5 radians per second squared over 20 to 1500 Hz. Linear random operating vibration is specified separately at 0.70 g rms over 5 to 500 Hz.
Read that carefully. Losing up to a tenth of your throughput to vibration is the specified behaviour of a good enterprise drive in a vibrating chassis, not a fault condition. It is not a guarantee that you keep your performance; it is a guarantee about how much of it you lose.
The mechanism is that rotational vibration displaces the head off-track, and the servo system responds by rejecting the write and retrying on the next revolution - which costs 8.33 milliseconds at 7,200 rpm each time. Writes suffer more than reads because an off-track write damages the neighbouring track, so the firmware must be conservative. Seagate’s own NAS drive selection guidance puts a number on when to care: rotational vibration sensors are recommended for systems with more than five drives. The other two mitigations it names are dual-plane balance of the spindle assembly and a top-cover-attached motor, and the Exos M lists “Top Cover Attached motor for excellent vibration tolerance” among its features - the motor is bolted to the cover as well as the base, so the assembly is stiffer and resonates less.
The practical consequence for a home builder is that a six-bay enclosure is not a three-bay one twice, and a cheap chassis with rubber grommets and a flexible tray is a worse environment than a stiff one with hard mounts. Drives for a NAS goes further into this. Flash, of course, has no equivalent specification, because there is nothing to shake off-track, which is a real and rarely stated advantage of an all-flash box in a domestic setting.
The 10^15 comparison nobody makes
Failure rates get compared endlessly. Uncorrectable error rates almost never do, and they are the more directly comparable pair, because both industries specify them the same way.
| Device | Specification | Source |
|---|---|---|
| Exos M nearline HDD | 1 unrecoverable sector per 10^15 bits read | Seagate Exos M product manual |
| Client SSD | UBER 10^-15 | JESD218 application class table |
| Enterprise SSD | UBER 10^-16 | JESD218 application class table |
A client SSD’s guaranteed error floor is the same as an enterprise hard drive’s. That is not the result most people expect, and it is worth knowing before paying for flash on reliability grounds alone. Only the enterprise SSD class is specified an order of magnitude better.
Work out what the number means across one whole-drive read:
UBER 1 unrecoverable sector per 10^15 bits read
one full pass of 30 TB 30e12 bytes × 8 = 2.4e14 bits = 0.24 × 10^15
expected bad sectors 0.24 per whole-drive read
i.e. one every ~4.2 full passes, at the spec limit
rebuild of a 10-wide array with one drive lost
bits read 9 × 2.4e14 = 2.16e15
expected bad sectors 2.16
Two honest qualifications. The specification is an upper bound on the rate, not an expectation - real drives are typically much better than their UBER, or arrays would fail constantly. And the failure it describes is one sector, not one drive, so the consequence depends entirely on whether the layer above can reconstruct a sector. That is the argument drives for a NAS works through as the 10^14 problem, and a nearline drive at 10^15 is ten times better than the consumer drives that argument was originally written about.
The scrubbing consequence is the practical one. If a whole-drive read is expected to produce a quarter of an unreadable sector at the specification limit, then scrubbing an array monthly finds those sectors while redundancy still exists, and scrubbing it never finds them during a rebuild when it does not.
Vendor claims run the other way and should be read as advocacy rather than as measurement. Solidigm states that its QLC SSDs show uncorrectable error rates two orders of magnitude better than the enterprise hard drives it compared against, and cites annual replacement rates 7.5 to 28 times lower. That is a vendor comparing its own product against a competitor’s category using its own methodology, and it belongs in the argument as a claim rather than as a number.
The workload rate on the datasheet, and its flash equivalent
Every buyer of enterprise flash knows about drive writes per day. Almost nobody knows that nearline hard drives carry the same kind of rating, that it is stricter, and that it is printed in the manual.
The Exos M product manual specifies a maximum annualised workload rate of under 550 TB per year, with a formula:
Workload Rate = TB transferred × (8760 / recorded power-on hours)
and the stated consequence of exceeding it is that doing so “may degrade drive MTBF and impact product reliability”. Convert it:
550 TB/year ÷ 30 TB per drive = 18.3 full-drive passes per year
550 TB/year ÷ 365 = 1.507 TB per day
1.507 TB/day ÷ 30 TB = 0.050 "drive writes per day" equivalent
Solidigm D5-P5316, 30.72 TB 0.41 DWPD = 22,930 TBW over 5 years
22,930 ÷ 5 = 4,586 TB per year
4,586 ÷ 550 = 8.3× the hard drive's annual allowance
A 30 TB nearline hard drive is rated for about eighteen full-drive passes a year. That is the platter’s DWPD, it is on every enterprise datasheet, and it is a stricter rating than the read-intensive QLC flash part by a factor of more than eight.
Two things stop that from being a clean comparison, and both should be stated. The hard drive figure counts bytes transferred, which includes reads; the flash figure counts bytes written only, because reads do not wear NAND. So the hard drive’s number is stricter in coverage as well as in magnitude. And the two consequences are different in kind: exceeding TBW on flash exhausts a physical budget and the drive eventually goes read-only, while exceeding workload rate on a hard drive degrades a population statistic in a way no individual drive will ever demonstrate to you.
Now hold that figure next to the rebuild arithmetic from earlier. A rebuild of a ten-wide array reads nine surviving drives in full:
9 × 32 TB = 288 TB in one operation
288 / 550 = 52.4% of a drive's entire annual workload allowance
One rebuild spends half the year’s rated workload. Two rebuilds and a monthly scrub schedule will take a home array past its drives’ specified annual workload without anybody doing anything unusual, and that is before a single byte of actual use. This is not an argument against scrubbing, which remains correct. It is an argument that the workload rating on a nearline drive assumes a datacentre’s access pattern - mostly idle, written once, read rarely - and that a small array doing maintenance behaves less like that than its owner assumes.
Failure modes, contrasted honestly
The important difference is not which fails more often. It is what failure looks like from where you are standing.
| Hard drive | SSD | |
|---|---|---|
| Wear mechanism | Bearings, actuator, head-media spacing, lubricant | NAND oxide degradation per program/erase cycle |
| Typical warning | Reallocated and pending sectors, slow reads, noise | Often none |
| Typical end state | Degrades; parts of it still read | Read-only, or gone entirely |
| Media defects | Grow slowly, remapped from a spare pool | Blocks retire into spare area, tracked as spare remaining |
| Controller failure | Donor PCB swap is sometimes viable | The mapping table is the data; recovery is specialist |
| Sudden total loss | Head crash, seized spindle | Controller or firmware failure, power loss during write |
| Environmental sensitivity | Rotational vibration, shock while operating | Temperature, both for retention and for throttling |
| Slow-failure signature | Rising pending sector count | Rising time on the ECC ladder |
A hard drive usually tells you it is dying, and an SSD usually does not. The
platters carry data in a layout with a comprehensible relationship to logical
block addresses, so a drive that is failing mechanically can often be imaged
slowly with ddrescue and most of it recovered. When a growing defect list is
the failure mode you get weeks of warning:
$ sudo smartctl -a /dev/sda
ID# ATTRIBUTE_NAME FLAGS VALUE WORST THRESH RAW_VALUE
5 Reallocated_Sector_Ct PO--CK 100 100 010 0
9 Power_On_Hours -O--CK 079 079 000 18942
194 Temperature_Celsius -O---K 114 102 000 36
197 Current_Pending_Sector -O--C- 100 100 000 0
198 Offline_Uncorrectable ----CK 100 100 000 0
Attributes 197 and 198 moving off zero is the signal that matters: sectors the drive could not read and has not yet been able to reallocate. Attribute 5 climbing steadily means the spare pool is being consumed. Be aware that raw value encodings are vendor-specific and the normalised VALUE column is close to meaningless across manufacturers; watch the direction of travel on one drive rather than comparing two.
An SSD’s data has no such relationship to anything. The flash translation layer maps logical blocks to physical pages, that map is rebuilt and rewritten constantly, and it lives on the drive. Lose the controller and the NAND packages are an encrypted-looking jumble that requires chip-off desoldering and reverse engineering of a proprietary FTL to read. This is why SSD failure feels sudden and total: there is no partial-read path.
The one warning flash does give is standardised on NVMe, and worth reading every few months:
$ sudo smartctl -a /dev/nvme0
Percentage Used: 7%
Available Spare: 100%
Available Spare Threshold: 10%
Data Units Written: 20,000,000 [10.2 TB]
Unsafe Shutdowns: 14
Media and Data Integrity Errors: 0
Data Units Written counts 1,000 blocks of 512 bytes each, so multiply by
512,000 to get host bytes: 20,000,000 × 512,000 = 10.24 TB. Percentage Used
is the controller’s own estimate of consumed endurance and can exceed 100.
Available Spare falling towards its threshold is the closest thing to a pending
sector count that flash offers. On SATA SSDs the equivalent wear attribute has no
standard ID and varies by vendor, which is one more reason NVMe is easier to
manage. SSD endurance covers what the numbers are worth.
Two more asymmetries. Power loss during a write is a real risk on consumer SSDs and largely not one on enterprise drives, which carry capacitors to flush the mapping table; a hard drive losing power mid-write corrupts the sectors in flight and nothing else. And temperature acts on the two media in opposite directions: a hot hard drive wears its bearings and lubricant faster, while a hot SSD throttles to protect itself and, separately, loses retention margin.
What Backblaze’s numbers do and do not support
Backblaze publishes quarterly annualised failure rates for a fleet large enough to mean something, which makes it the only public dataset of its kind. The Q1 2026 report gives a quarterly AFR of 1.24 per cent, down from 1.42 in Q1 2025, and a lifetime AFR of 1.39 per cent across 529,968,464 drive-days. The quarterly figure is arithmetic you can check:
AFR = (failures / drive-days) × 365 × 100
1,030 failures, 30,203,180 drive-days
(1,030 / 30,203,180) × 36,500 = 1.2447%
Full-year 2025 came in at 1.36 per cent, down from 1.55 in 2024 and the best since 1.37 in 2022, across 341,664 drives and 30 models. The fleet was 25.13 per cent at 0 to 12 TB, 52.06 per cent at 14 to 16 TB and 22.81 per cent at 20 TB and above.
The most useful single finding for a buyer is the large-drive figure. Drives of 20 TB and above ran at 0.85 per cent AFR across more than 86,000 units - better than the fleet average, not worse, which contradicts the standing folk belief that bigger drives are less reliable. Backblaze deployed 9,404 drives over 20 TB out of 10,220 in that quarter, so they are voting with the purchase order.
The average hides a spread that matters far more than the average does. In the same quarter:
| Model | Capacity | AFR | Basis |
|---|---|---|---|
| Toshiba MG11ACA24TE | 24 TB | 0.22% | quarterly |
| WDC WUH722222ALE6L4 | 22 TB | 0.38%, across 45,000+ units | quarterly |
| Seagate ST10000NM0086 | 10 TB | 3.13% | lifetime |
| Seagate ST14000NM0138 | 14 TB | 5.78% | lifetime |
A twenty-six-fold spread between models inside one fleet, in one building, with one operator - and note the basis column, because two of those four are lifetime rates rather than quarterly ones, so it is a spread across unlike quantities rather than a clean ratio. Any sentence of the form “brand X is more reliable than brand Y” is doing violence to that table. Model and vintage dominate brand, and neither is knowable in advance.
On the SSD comparison, the honest version is more interesting than the version usually quoted. Unadjusted, Backblaze’s Q2 2021 figures gave SSDs 0.59 per cent AFR at an average age of 14.2 months against hard drives at 2.19 per cent and 52.4 months, which compares a young population against an old one and proves nothing. Age-adjusted - comparing the hard drive fleet at the same average age - the figures were 0.59 per cent for SSDs against 1.27 for hard drives, with 95 per cent confidence intervals of roughly plus or minus 0.5 on each, so the two intervals overlap and the difference is not significant. The point estimates themselves are 0.68 apart, which is wider than either interval, so the version of this claim that says each estimate falls inside the other’s interval is wrong even though the conclusion drawn from it survives. Backblaze’s own written conclusion is that the difference “was certainly not enough by itself to justify the extra cost.”
Three caveats that anyone citing this data owes the reader.
The SSD population is small and atypical. Lifetime SSD AFR of 0.90 per cent over Q4 2018 to Q2 2023 covers about 3,144 drives, and they are boot drives - a light, sequential, mostly-read workload that resembles nothing a buyer does with a data SSD. Backblaze’s own minimum threshold for a meaningful AFR is 100 drives and 10,000 drive-days, and the SSD fleet only just clears the sort of scale where percentage points are not noise.
Day-one failures are systematically missing. Backblaze states that Drive Stats “is a function of conditional logic” keyed on a drive’s daily presence, and writes that they have “probably been completely under-reporting day one failures”. Their mitigation is a pre-deployment qualification process. You do not have one. For a buyer, and especially a buyer of used drives, the infant-mortality window that this dataset cannot see is precisely the window that matters - see buying used drives on eBay.
A population statistic is not a prediction about your drive. Seagate says so itself in the Exos M manual: the annualised failure rate of 0.35 per cent, which corresponds to a stated MTBF of 2,500,000 hours, is “a population statistic not relevant to individual units”. And it comes with conditions almost nobody reads: 8,760 power-on hours per year, a drive-reported head-disk-assembly temperature of 10 to 30 °C, ambient wet bulb at or below 26 °C, typical workload, ISA S71.04-2013 G2 contamination levels and ISO 14644-1 Class 8 dust. A drive in a warm cupboard, powered intermittently, is not the drive that figure describes. The same section rates 600,000 load-unload cycles and a five-year service life.
Note the size of the gap between the specification and the measurement: Seagate specifies 0.35 per cent AFR and Backblaze measures 1.24 per cent across a real fleet. Roughly three and a half times worse than specification is what observation looks like, and that is in a professionally operated datacentre.
Unpowered retention, and the archive question
An SSD stores bits as trapped charge. Charge leaks. A hard drive stores bits as magnetic domain orientation, which does not.
JEDEC specifies the floor in JESD218, and the application class table is the whole of it:
| Class | Active use | Powered-off retention | FFR | UBER |
|---|---|---|---|---|
| Client | 40 °C, 8 hours a day | 1 year at 30 °C | 3% | 10^-15 |
| Enterprise | 55 °C, 24 hours a day | 3 months at 40 °C | 3% | 10^-16 |
Four things about that table are routinely misread.
It is a post-endurance requirement, and this is the detail everybody gets wrong. JESD218 defines the endurance rating - the TBW number on the box - as the maximum host writes such that the drive simultaneously maintains its capacity, maintains its class UBER, meets the functional failure requirement, and retains data with power off for the required time for its class. The one year is not what a new drive does. It is what the drive must still do after being written to its full rated endurance. A lightly used drive has far more intact tunnel oxide and retains charge for very much longer. How much longer is not published by anybody, for any consumer drive, and this site will not guess it.
The verification is stricter than the headline. JESD218 requires retention to be demonstrated for two mechanisms rather than one: a temperature-accelerated mechanism with an activation energy of 1.1 eV, and a non-temperature-accelerated mechanism, both at 60 per cent statistical confidence, and under the assumption that the endurance stressing took place over no longer than one year at the endurance-use temperature.
Temperature dominates, and the famous table is attributable. The JEDEC JC-64.8 presentation by Seagate’s Alvin Cox that spawned the “SSDs lose data in a week” scare contains a chart of retention in weeks indexed by both active-use temperature and power-off temperature, and states its own provenance: the numbers “are based on Intel’s published acceleration model for the detrapping retention mechanism (the official JEDEC model in JESD47 and JEP122 for this mechanism).” The commonly quoted halving of retention per 5 °C is read off that chart rather than stated in the text. Anchoring on the client specification and applying that rule gives about six months at 35 °C and about three at 40 °C - derived from a rule of thumb read off a slide image, not measured on your drive. What the chart actually describes is a worn-out drive in a hot room, which is a real condition and not the one most people are in.
Enterprise is worse, not better. The shorter figure is not a weakness, it is a different design point: enterprise drives are optimised assuming they are always powered, so more of the endurance budget goes to performance and write cycles rather than retention margin. A decommissioned enterprise SSD is the worst thing you could choose for a shelf, and the fact that it is cheap per terabyte on the used market makes that a trap rather than a bargain.
Platters have no equivalent decay mechanism on any timescale you care about. The energy barrier holding a grain’s orientation is engineered to be large compared with thermal energy at room temperature - the ratio Seagate’s HAMR paper writes as KuV over kT - but that ratio is discussed qualitatively in the literature and is not published per product, so no specific figure appears here. What actually kills a shelved hard drive is mechanical: lubricant migrating off the platter surface, bearing grease settling, and stiction on a drive that has not turned in years. The failure is “it does not spin up”, not “the data faded”.
Tape is the only medium in this comparison designed for the job, and its published specifications are worth the space because they explain why it still has a roadmap. LTO-10 ships in two media, 30 TB native and 40 TB native, both rated at 2.5 to 1 compression and 400 MB/s native, so 100 TB and 1,000 MB/s at the larger point with compression applied. The 40 TB cartridge is a 122 per cent capacity improvement over LTO-9’s 18 TB. Both media use strontium-doped barium ferrite; the 40 TB cartridge adds an aramid base film that allows a thinner, smoother, longer tape in the same shell.
Then the number that explains everything about tape’s headroom:
LTO-10 30 TB cartridge 12 Gb/in²
18 TB hard drive 1,022 Gb/in²
ratio 1 : 85
Tape stores comparable capacity at one eighty-fifth of the areal density, because it has square metres to work with instead of square inches. It has not had to touch the trilemma yet, which is why INSIC projects tape can continue increasing capacity at historical rates through approximately 2034 on that headroom alone, while hard drives needed a laser to keep going.
No archival shelf-life figure appears here for tape, because the figure that circulates is not one this site could trace to the LTO consortium’s own published material. Neither hard drive nor SSD vendors publish one either, and it is worth noticing that none of them wants to.
For cold storage, the honest answer is that the medium matters less than the count. Two copies on different drives in different places, checksummed, and verified on a schedule beats any single medium’s specification. If you are choosing one anyway, choose platters, power them up annually, and read the whole surface when you do - which, on a 30 TB drive, is the forty-hour operation computed earlier, so schedule it rather than starting it on a Friday.
There is no workload where a platter is cheaper per IOPS
This deserves saying flatly, because the folk version of the argument holds that hard drives win somewhere on performance economics. They do not. Take the August 2026 enterprise reference prices from earlier and divide three ways. The sequential and IOPS figures come from vendor datasheets and the prices from a market index, which is a splice of two different kinds of document, so read the ratios rather than the absolute values:
30 TB nearline HDD, $1,216
per terabyte 1,216 / 30 = $40.53 / TB
per MB/s 1,216 / 275 = $4.42 per MB/s
per random IOPS 1,216 / 80 = $15.20 per IOPS
30 TB enterprise SSD, $22,600
per terabyte 22,600 / 30 = $753.33 / TB
per MB/s 22,600 / 7,000 = $3.23 per MB/s
per random IOPS 22,600 / 800,000 = $0.028 per IOPS
ratios
dollars per terabyte HDD wins 18.6×
dollars per MB/s flash wins 1.37×
dollars per random IOPS flash wins 543×
Three findings fall out of nine lines of division.
The hard drive’s advantage is capacity and only capacity. 18.6 to 1 on dollars per terabyte, and nothing else.
On sequential bandwidth the two are within a factor of 1.4, even at an eighteen-fold price gap, because a hard drive’s sequential rate is respectable and a flash drive’s is bounded by its interface. That is the most surprising row in this guide and it is why bulk sequential workloads - backup targets, media libraries, archives - remain platter workloads on economics rather than on tradition.
On random IOPS the ratio is five hundred to one and no arrangement of hard drives fixes it, because buying IOPS from platters means buying spindles, and each spindle arrives with terabytes you did not want and watts you must pay for. A 2010-era array bought thirty 15K drives to get 7,000 IOPS. One consumer NVMe drive now exceeds that by two orders of magnitude for less than the price of one of those thirty.
The correct engineering response is not to argue about the ratio. It is to stop buying IOPS from platters at all, which is what the tiering section below is about.
One argument that does not transfer, and it gets quoted at home users constantly. Rack density genuinely favours flash on total cost of ownership regardless of price per terabyte: a 15 TB U.2 SSD occupies about 6.4 cubic inches against about 24 for a 3.5-inch drive of similar capacity, roughly four times less volume, and one analysis put a 10 PB deployment at 19 racks all-flash against 102 racks on disk. Those two ratios are not the same number and should not be: the SFF-8301 envelope printed in Seagate’s own Exos X24 datasheet is 26.1 by 101.85 by 147 millimetres, which is 23.8 cubic inches, while a rack figure also counts backplanes, power and cooling. For a buyer paying no rack rent, no cooling contract and no floor-loading charge, that argument is worth exactly nothing, which is why datacentre TCO conclusions should not be repeated at people choosing a drive for a desktop.
The recommendation by workload
| Workload | Choose | Why |
|---|---|---|
| Boot and applications | NVMe SSD, SATA SSD if no slot | Queue-depth-one random reads, where the gap is already 100× before any queue |
| Gaming library | SSD, and prioritise capacity over tier | Level loads are random reads; SATA to NVMe is a small further gain |
| Video editing scratch | NVMe, chosen on sustained write | The SLC cache cliff is exactly this workload |
| Video archive, single stream | Hard drive | Sequential playback needs tens of MB/s, not thousands |
| NAS bulk storage | Hard drives, CMR | Cost per terabyte, and rebuild behaviour |
| Backup target | Hard drive | Large, sequential, cheap, powered occasionally |
| Cold archive | Hard drives or tape, plural | Charge leaks; magnetism does not |
| Databases, VMs, many users | SSD | Uncorrelated random streams are the worst case for one actuator |
| Metadata-heavy scans and indexing | SSD, or flash metadata on a platter pool | Directory traversal is the worst pattern an actuator can be given |
| Write-once, read-almost-never bulk | Hard drive, and check the workload rate | 550 TB/year is eighteen passes; archives do not come close |
Three things to watch at the edges of that table.
For NAS bulk, shingled recording is the trap, because an SMR drive’s random-write and rebuild behaviour gives you the worst of both technologies at hard drive prices. No vendor publishes a reliable model-number-to-recording-mode mapping and this site will not invent one; CMR vs SMR covers how to find out for a specific drive and what to do when you cannot. Within the hard drive listings, the drives that state their recording technology are a shorter list than you want and an honest one.
For the many-users case, note that this is the CPU-scheduler argument applied to storage. Forty concurrent readers produce forty uncorrelated address streams, which is precisely the pattern that pushes a hard drive towards its full seek-plus-rotation cost on every access while an SSD answers them across thirty-two dies in parallel. The access-density arithmetic earlier says the same thing in units: 2.5 IOPS per terabyte is not a number that supports many readers, at any capacity.
For anything where latency variance matters more than latency, read the ECC ladder section again. A median of 80 microseconds with a tail at 10 milliseconds behaves differently from a uniform 12.5 milliseconds, and which of those is worse depends entirely on whether anything downstream has a timeout.
Hybrid arrangements, and when they earn their place
Putting flash in front of platters works, but only under one condition: the hot working set has to be much smaller than the capacity, and it has to be reused. A cache does nothing for a linear backup pass, a media scan, or a first-run import, because every block is cold exactly once.
Distinguish two things that get the same name. A cache holds a copy; losing the device costs you performance. A tier holds the only copy of what is on it; losing the device costs you data. Check which one you are configuring.
| Arrangement | What it accelerates | The catch |
|---|---|---|
| ZFS L2ARC | Reads that miss RAM but fit the cache | Its index headers live in ARC, so an oversized L2ARC on a small-RAM box makes things slower |
| ZFS SLOG | Synchronous writes only | Not a write cache. Async writes never touch it, and it is only ever read back after a crash |
| ZFS special vdev | Metadata and small blocks | Holds the only copy. Mirror it or lose the pool |
| bcache / LVM cache | General block-level reuse | Writeback mode puts a consumer SSD in the durability path |
| Windows Storage Spaces tiering | Whole-file promotion on a schedule | Moves data on a timer, not on demand |
| SSHD (hybrid drive) | Boot, historically | The onboard cache was a few gigabytes; largely a dead product |
OpenZFS’s own tuning documentation is unusually direct about two of those, and it
is worth quoting rather than paraphrasing. On the SLOG: it accelerates only
fsync and O_SYNC workloads, and about 4 GB of usable space suffices, because
“most systems do not write anything close to 4GB to ZIL between transaction group
commits”. People buy 500 GB devices for this job. On L2ARC: its entries consume
ARC, which is RAM, and the primarycache and secondarycache properties can be
set to metadata only - which the documentation notes is the right setting for
databases and virtual machines that maintain their own caches, because caching
the same block twice in two layers wastes both.
The documentation gives no sizing guidance at all for special vdevs, which is worth knowing before you go looking for a rule. There is not one.
The arrangement that most reliably justifies itself is the special vdev, and for
an unglamorous reason that the access-density arithmetic explains: directory
traversal, find, backup scans and snapshot operations are metadata-heavy random
reads, which is the single worst thing you can ask an actuator to do. At 2.5 IOPS
per terabyte, a metadata scan of a large pool is arithmetic you can watch. Moving
just the metadata to flash makes a platter array feel responsive without buying
flash for the bulk. It also means that if the special vdev dies the pool dies, so
it gets mirrored, always.
The arrangement that most often disappoints is a large L2ARC bought instead of RAM. Add RAM first; it is the same cache one level up and it costs nothing to index.
What is not published, and why that is the honest answer
A list, because the gaps are as informative as the figures and because a guide that only prints what is available gives a misleading impression of how much is known:
- Seek time is absent from the 32 TB Exos M product manual entirely. Consumer drives stopped publishing it a decade ago and the flagship enterprise manual has now followed.
- Writer and reader head widths are not published per product, which is why the SMR density arithmetic in CMR vs SMR has to be modelled rather than quoted.
- Drive-managed SMR band sizes are not disclosed by any vendor.
- Current NAND page read, program and erase timings sit behind vendor registration walls. The figures in this guide are generation-representative, and labelled as such where they appear.
- Wafer cost is not published, so the flash cost floor has to be approached through die prices instead.
- Consumer SSD retention beyond the JESD218 floor is not published by anybody, for any drive. The one-year figure is a worn-drive minimum and the real number for a lightly used drive is unknown.
- Enterprise SSD idle power in a form comparable to a hard drive’s state table does not exist, so watts per terabyte cannot be compared across media honestly.
- SLC cache sizing algorithms and FTL internals are proprietary on every consumer drive.
- Day-one drive failures are missing from the only large public reliability dataset, by its publisher’s own admission.
- Annualised failure rate for your drive is not a thing that exists. Seagate says so in the manual: it is a population statistic, not relevant to individual units.
Every one of those is a place where somebody will confidently tell you a number. Treat a confident number in a gap as a guess wearing a suit.
What to do with this on a listing page
- Read today’s ratio off the medians, not off this page. Sort the hard drive listings and the SSD listings by price per terabyte and divide. Every dollar figure in this guide is dated by construction, and in an allocated market it is dated faster than usual.
- Decide recording mode before capacity on the platter side. Start at the hard drive listings, check the stated recording mode, and treat “not stated” as SMR if you are building an array. If shingling is fine for what you are doing, ignore the recording mode and buy the cheaper terabytes.
- Filter flash by interface, not by form factor. Use the NVMe
listings or the
ifacefilter on the SSD page. An M.2 slot is not automatically NVMe - see drive interfaces before you buy the wrong key. - Pick a capacity point rather than a price band. The site has no capacity
range:
capselects one exact capacity, andminandmaxbound price, not size. Compare 1 TB drives against a larger point and watch the price per terabyte fall away. Below roughly a terabyte a hard drive is all motor and casting, and the price per terabyte reflects that rather than reflecting the storage. That is the floor the whole first section explains. - Check the spindle where latency matters. The
rpmfilter separates 7,200 rpm drives from 5,400 rpm ones, and the difference is an exact 1.39 ms of average rotational latency on every single access, forever. - For flash, read the sustained figure and not the peak. Every consumer drive quotes its SLC cache speed. Ask what it does after the cache, and if the listing quotes a 99.999th-percentile latency at all, that is a better drive than one that does not.
- Cost the power before you cost the drive, if you are filling bays. 0.21 W per terabyte idle is cheap; 2.6 A per drive at spin-up is what actually constrains the chassis. Check that the backplane staggers spin-up.
- Check the
partsandbidsfilters before trusting a price per terabyte. A for-parts listing and an auction with two days to run are not comparable with a fixed-price working drive, and the sort does not know that. - Check the power-on hours on arrival, on either medium. A used drive of either kind carries an odometer, and buying used drives on eBay is about reading it honestly.
The one thing neither technology will do is make a single copy safe. Flash fails suddenly and platters fail slowly, the tail of a flash read is a hard drive’s median, and the plan that survives both is the same plan.