CMR vs SMR: what shingled recording costs, and how to tell
By Harry Saarinen ·
Often you cannot, and anyone who tells you a model number settles it is guessing.
That is the uncomfortable part, so it goes first. There is one product line where a part number really does decode - Western Digital prints the key on its own Ultrastar datasheets - and it is the exception that proves how bad the general case is. For everything else, the honest position is that a drive-managed shingled drive is under no obligation to say so, almost never does, and the one field in the ATA standard where it could have said so has been withdrawn from the current revision of that standard.
The rest of this explains what shingled magnetic recording does to a drive, why the damage is invisible to any benchmark shorter than about three minutes, what eleven-year-old academic measurements still tell you that no vendor will, and what the real sources of truth are. It is long because the subject has three layers - the physics, the firmware, and the standards - and advice that only covers one of them is where most of the bad advice comes from.
Why the writer is wider than the reader, and why that is permanent
A disk head is two devices on one slider. The writer is an electromagnet that flips the magnetisation of grains in the medium beneath it; the reader is a magnetoresistive sensor that detects which way those grains point. They hit different physical limits, and that asymmetry is the entire origin of SMR.
The writer always lays down a wider stripe than the reader needs to read it back. A tunnelling magnetoresistive sensor can be shrunk by narrowing the sensor stack, because sensitivity is a materials problem and the signal it needs is whatever fringing field escapes the medium. A write field cannot be shrunk the same way. The field has to exceed the coercivity of the grains, the grains have to be high-coercivity or thermal energy will flip them on their own, and the field from a narrow pole spreads as it leaves the pole tip. Shrink the writer far enough and the field is too weak to switch anything; keep the field strong and it is wider than the track you wanted.
This is one corner of the recording trilemma. Grains must be small enough for a good signal-to-noise ratio at the track width you want, stable enough that ambient thermal energy does not erase them over a decade, and switchable by a field a head can actually produce. Any two are easy. All three at once is the whole history of the industry, and every named technology of the last fifteen years - perpendicular recording, energy assist, two-dimensional readback, shingling - is an attack on one corner of it.
Conventional magnetic recording (CMR, also called PMR after the perpendicular grain orientation it uses) accepts the waste. Each track is written at the writer’s full width, with a guard band so that head positioning error and adjacent track interference do not eat into the neighbours. The CMR track pitch is set by the writer, which is the wider of the two devices, and the reader’s extra resolution is simply unused.
Shingled recording refuses the waste. Tracks are written in order, each overlapping the last like courses of roof tiles. Track 1 goes down at full writer width. Track 2 is written a fraction of that width away, destroying most of track 1 and leaving a narrow exposed strip. Track 3 trims track 2 the same way. What survives of each track is only as wide as the reader needs, so the SMR track pitch is set by the reader - the narrower device. Nothing about the head, the medium or the servo system has changed. Only the order and the spacing of the writes.
The exact arithmetic of a shingled band
Shingling is not free at the edges. The last track of a run has to be written
somewhere that does not damage whatever comes next, so runs of shingled tracks
are separated by a guard region. The Skylight paper (Aghayev and Desnoyers,
FAST ’15) states the trade-off exactly. For a band of b tracks written with a
head k tracks wide - meaning the writer covers k shingle steps - the guard
region costs k - 1 steps, and the storage efficiency of the band is:
efficiency = b / (b + k - 1)
That single expression contains the whole design tension. Larger bands are more space-efficient and more expensive to modify; smaller bands are cheap to modify and waste more surface on guards.
Band size b |
Efficiency, k = 2 |
Efficiency, k = 3 |
|---|---|---|
| 8 tracks | 88.9 per cent | 80.0 per cent |
| 16 tracks | 94.1 per cent | 88.9 per cent |
| 32 tracks | 97.0 per cent | 94.1 per cent |
| 64 tracks | 98.5 per cent | 97.0 per cent |
| 128 tracks | 99.2 per cent | 98.5 per cent |
Now the raw density gain, with numbers that are mine, not a vendor’s - writer and reader widths are not published per product - but of the right order for a recent drive:
Ww = width the writer lays down 60 nm, guard = 10 nm -> Tc = 70 nm
Wr = width the reader needs to resolve 40 nm, margin = 10 nm -> Ts = 50 nm
(Wr is always less than Ww)
CMR track pitch Tc = Ww + guard
SMR shingle step Ts = Wr + positioning margin
raw gain Tc / Ts = 70 / 50 = 1.400
band of 48 tracks, k = 2 1.400 x 48 / (48 + 2 - 1) = 1.371
band of 48 tracks, k = 3 1.400 x 48 / (48 + 3 - 1) = 1.344
Thirty-seven per cent more tracks on the same platter, from the same head, on the same medium, with no new physics. Note how little of that gain the guard region eats: at 48 tracks per band the band overhead costs about two points. The density gain is dominated by the ratio of reader width to writer width, not by the band geometry. That matters later, because it means an SMR drive cannot buy back much capacity by using larger bands, and larger bands are exactly what makes it slow.
The density gain is smaller than the folklore, and it is shrinking
The model above says 37 per cent. Folklore says 20 to 25 per cent. Vendors, when they publish anything at all, say less than either. The earlier version of this guide told you to carry 20 to 25 per cent; that is now wrong, or at least badly undated, and the correction is worth the space.
| Source | Claimed SMR gain over CMR | Comparator |
|---|---|---|
| Seagate, at the September 2013 launch of drive-managed SMR | about 25 per cent | non-shingled storage, unspecified |
| Western Digital UltraSMR technology brief, July 2026, describing the previous SMR generation | 11 per cent, stated as 2 TB per drive | CMR of the same generation |
| Western Digital UltraSMR technology brief, July 2026, describing UltraSMR | 18 per cent, stated as 4 TB per drive | CMR of the same generation |
| Western Digital blog, 4 December 2025 | 23 per cent | a 32 TB UltraSMR drive against a 26 TB CMR drive on the same platform |
| Seagate public CMR/SMR list, read 13 September 2026 | 6.7 per cent | a 32 TB Exos SMR drive against a 30 TB Mozaic HAMR CMR drive |
Those are not contradictions. They are the same quantity measured against different comparators at different times, and the spread is the finding. Two observations make the vendor figures legible.
First, Western Digital’s own numbers are internally consistent and let you
recover the comparator. If 11 per cent is 2 TB per drive, the CMR baseline was an
18 TB drive. If 18 per cent is 4 TB per drive, the baseline was 22 TB. Those are
my arithmetic, not WD’s statements, but they are forced: 2 / 18 = 0.111 and
4 / 22 = 0.182. The gain is a percentage of a moving base, and quoting the
percentage without the base is how a 2 TB advantage and a 4 TB advantage end up
sounding like the same thing.
Second, the December 2025 figure of 23 per cent is a 32 TB SMR drive against a
26 TB CMR drive, and (32 - 26) / 26 = 0.231. Pick a different CMR comparator
from the same catalogue and the same SMR drive produces a different percentage.
The last two rows of the table are that effect in its purest form: both are a
32 TB shingled drive, and the only thing that changes is what it is measured
against.
against a 26 TB CMR drive (32 - 26) / 26 = 0.231 = 23.1 per cent
against a 30 TB CMR drive (32 - 30) / 30 = 0.067 = 6.7 per cent
This is why vendor SMR gain figures are not comparable with each other and why a number quoted in a forum post without its comparator is worthless.
The honest statement is that the shingling gain is generation-dependent and
comparator-dependent, has ranged from roughly 7 to roughly 25 per cent across
sources, and is smaller now than it was in 2013. The reason it shrinks is
mechanical: as areal density rises, the reader and writer widths converge, and
the ratio Tc / Ts that drives the whole gain converges with them. Shingling
harvests the gap between two devices, and the gap is closing.
What has not shrunk is the commercial importance. Western Digital states that SMR now accounts for roughly 50 per cent of the exabytes it ships into capacity-optimised data centre applications, and that SMR buys about two HDD generations of capacity lead over CMR where it used to buy one. SMR is not a niche or a consumer trick. It is half of the high-capacity market by capacity, sold almost entirely to buyers who write software to accommodate it.
Bands, cascades, and the cost of one four-kilobyte write
Overlapping tracks cannot be rewritten in place. Rewrite track 1 and the full-width writer spills into track 2, which is only a narrow surviving strip; track 2 is gone. Rewrite track 2 to repair it and track 3 goes. The damage cascades forward to the end of the shingle.
If the shingle ran the whole radius of the platter, changing one sector would mean rewriting everything outboard of it. So drives break the surface into bands - host-managed drives call them zones - separated by a guard wide enough that writing the last track of one band does not touch the first of the next. The cascade is bounded by the band, and the procedure for modifying anything inside one has a name: read-modify-write. Read every track from the target to the end of the band, hold it somewhere, apply the change, then rewrite the whole run in shingle order.
There are two very different band sizes in the world, and the arithmetic has to be done twice.
Host-managed drives. Shipping host-managed drives commonly report 256 MiB zones. A 20 TB host-managed drive documented in the field reported a zone size of 256.00 MiB and 74,508 zones, which multiplies out to exactly 20.0 TB - a useful consistency check on a third-party report. Cost a single 4 KiB write at that zone size, worst case:
zone size 256 MiB = 268,435,456 bytes
host write 4 KiB = 4,096 bytes
media traffic read 256 MiB + write 256 MiB = 536,870,912 bytes
write amplification 536,870,912 / 4,096 = 131,072x
media time at 200 MB/s 536,870,912 / 200e6 = 2.68 s
One four-kilobyte write, 2.7 seconds of media time. This is the number a host-managed drive exists to make impossible: it does not perform this operation at all, it rejects the write. The cost is not paid, it is refused.
Drive-managed drives. Skylight measured the internal band size directly on Seagate drive-managed drives and found 15 to 40 MiB - six to seventeen times smaller than a 256 MiB host-managed zone, and still three to four orders of magnitude larger than a sector. Both halves of that comparison matter, so do the arithmetic rather than assert either:
256 MiB / 40 MiB = 6.4 256 MiB / 15 MiB = 17.1
40 MiB / 4 KiB = 10,240 15 MiB / 4 KiB = 3,840
A drive-managed band is nearer a host-managed zone than it is to anything the filesystem thinks it is writing. Drive-managed bands are not disclosed by any vendor; these remain the published measurements. The same sum at the measured sizes, with one extra wrinkle covered in the next section - a statically mapped drive must stage the updated band in scratch space before overwriting the original, because a power loss midway through an in-place band rewrite would destroy the band, so cleaning costs a read plus two band writes rather than a read plus one:
band = 15 MiB read + 1 write = 31,457,280 B amp 7,680x 0.157 s
read + 2 writes = 47,185,920 B amp 11,520x 0.236 s
band = 40 MiB read + 1 write = 83,886,080 B amp 20,480x 0.419 s
read + 2 writes = 125,829,120 B amp 30,720x 0.629 s
(media time at a flat 200 MB/s; mine, not measured)
That last line is worth pausing on. Skylight measured actual cleaning duration at 0.6 to 1.6 seconds per modified band, and the derived floor for a statically mapped 40 MiB band at 200 MB/s is 0.63 seconds. The model lands on the bottom of the measured range, which is what you would expect if the measured range is the media time plus seek overhead plus the lower transfer rates near the inner diameter. The FAST ’17 ext4-lazy paper, from the same group, puts typical cleaning at 1 to 2 seconds per band and reports up to 45 seconds in extreme cases.
Both sums point the same way. A drive-managed SMR drive does not have a 2.7-second worst case per 4 KiB write, it has a 0.6 to 1.6 second one, and it has to do it hundreds of thousands of times to clear a full cache. The difference between 2.7 seconds and 0.6 seconds is not the difference between broken and fine. It is the difference between two kinds of broken.
The persistent cache is a journal on the platters, and it has been measured
Every drive-managed SMR drive reserves a region written conventionally - not shingled, individually rewritable. Vendors call it the media cache or the persistent cache. It is on the platters, not in DRAM and not in flash, so it survives power loss; that is what “persistent” means here. Incoming writes are appended to it sequentially, which is the one thing spinning media does well, and acknowledged. The drive tells the host it is done. The data is safe, and it is in the wrong place.
Later, when the drive judges itself idle, it works through the backlog: reads the cached writes, groups them by destination band, and does the read-modify-write to fold them into the shingled region. That is background reorganisation, and it is the thing to understand, because it means an SMR drive has homework, and the bill arrives later than the transaction.
No vendor publishes the cache size. Skylight measured it, by inference from timing, on two Seagate drives:
| Drive | Capacity | Measured persistent cache |
|---|---|---|
| ST5000AS0011 | 5 TB | 20 GiB |
| ST8000AS0011 | 8 TB | 25 GiB |
The FAST ’17 paper reports approximately 25 GiB on the ST8000AS0002 as well. Three caveats attach to those figures and all three matter.
They are eleven years old, and they describe 2013 to 2015 Seagate parts. They are inferred, not read out of the drive. And most importantly, Skylight found the effective persistent-cache size is not a constant: it varies with write size and with queue depth. There is no single number to publish even if a vendor wanted to publish one, which is a better reason for the silence than embarrassment.
Skylight also worked out the cache’s internal structure, and it is not what the literature at the time assumed. The persistent cache is written as journal entries with quantised sizes - an append-only log, not a mapped scratch area - and data that is not in the cache uses static mapping: a fixed, arithmetic LBA-to-physical assignment exactly like a conventional drive’s.
Three consequences follow from the architecture, and they are the practical core of the whole subject.
- Performance depends on history, not on the current request. The same 4 KiB write takes a millisecond or several seconds depending on the preceding hour, and nothing in the interface exposes which state the drive is in.
- Idle time is a resource. A drive that never goes idle never clears the backlog. A workload at 70 per cent duty cycle is not 70 per cent of a problem; it may be all of it.
- Free space is a resource too. A nearly full drive gives the reorganiser less room and worse choices. SMR drives get worse as they fill, in a way CMR drives essentially do not.
Static and dynamic mapping, and why two vendors’ drives fail differently
The shingle translation layer - the STL, the SMR analogue of an SSD’s flash translation layer - can map logical blocks to physical bands in two ways, and the choice determines both the cleaning cost and the shape of the failure.
Static mapping keeps a fixed LBA-to-band assignment. Band 17 always lives at the same radius. To fold cached writes into band 17 the STL must read band 17, apply the updates in memory, write the result to scratch space, and only then rewrite band 17 in place. The staging write is not optional: overwriting a band in place is destructive by construction, so a power loss partway through would leave the band unrecoverable. Cost: one band read plus two band writes.
Dynamic mapping lets the STL put the updated band anywhere. Read band 17, apply the updates, write the result to a free band somewhere else, update the mapping table. Cost: one band read plus one band write - and no staging, because the original band is intact until the mapping flips.
FAST ’17 describes the two major vendors as having made opposite choices, with opposite consequences:
| Seagate drive-managed | Western Digital drive-managed | |
|---|---|---|
| Mapping | Static | Dynamic |
| Cleaning policy | During idle periods | Continuously |
| Cost per band cleaned | Read plus two writes | Read plus one write |
| Before the cache fills | Fast, indistinguishable from CMR | Slower than CMR |
| After the cache fills | Collapses hard | Degraded, but less catastrophically |
That table is the single most useful thing in this guide for anyone comparing test reports. Two drive-managed SMR drives from two vendors do not have the same failure curve, so two reviews measuring two drives can both be right and disagree completely. A reviewer who runs a short test on a Seagate part finds CMR-class numbers; one who runs the same test on a Western Digital part finds it mysteriously slow and concludes the drive is bad. Run both for four hours and the ranking inverts.
There is a second-order point hiding in the dynamic-mapping row. A dynamically mapped drive that cleans continuously is spending media time on reorganisation while you are waiting for it, which is why it never looks as good as CMR even fresh out of the box. It is paying the bill as it goes. Continuous taxation or occasional bankruptcy is a genuine engineering choice with a real case on both sides, and neither vendor tells you which one you bought.
The drives are laid out in horizontal zones, not vertical cylinders
One more Skylight result deserves its own space, because it invalidates an assumption that almost every drive benchmark is built on.
A conventional drive is organised in cylinders. Consecutive logical blocks step through every recording surface at one radius before the actuator moves, so the sequential transfer rate declines smoothly and monotonically from the outer diameter to the inner, and a benchmark can sample “outer”, “middle” and “inner” by sampling low, middle and high LBAs.
The drive-managed SMR drives Skylight examined are organised in horizontal zones instead: wide radial strokes of thousands of tracks on a single surface. Those strokes were measured at 18 to 20 GiB near the outer diameter, falling to about 4 GiB at higher LBAs. The sequential read curve consequently shows eight distinct plateaus on each side of a symmetry axis rather than a smooth decline. Sixteen plateaus in the curve, eight per side, with the axis of symmetry falling partway through the read.
The natural reading of that shape - mine, not the paper’s - is that the drive sweeps one surface from one edge to the other, switches heads, and sweeps back, so each plateau is one surface and the symmetry axis is the turnaround. Whatever the exact mechanism, two practical consequences follow and neither depends on the interpretation being right:
- LBA position does not map to radius the way you assume. A short-stroke partition at the front of the drive is not the outer tracks of every platter, it is one surface. Benchmarks that report an “outer / inner” spread on such a drive are reporting something else.
- The plateaus are a detectable signature. A whole-device sequential read that produces a stepped curve with a symmetry axis, rather than a smooth decline, is behavioural evidence about the layout - and it is far cheaper to run than the four-hour write probe, because a sequential read does no damage and takes as long as the drive takes to stream.
That second point is the only non-destructive behavioural test in this guide. It is not proof of shingling, since layout and recording technology are separate decisions, but a stepped read curve on a drive whose vendor will not say is one more reason to assume the worst.
A steady-state model of the cliff, and why published numbers disagree
Published figures for drive-managed random-write throughput after cache exhaustion range from “low single-digit megabytes per second” to, in the FAST ’17 measurements, the sub-1 MiB/s region on a log axis that runs down to 0.01 MiB/s. That is a spread of two orders of magnitude in the published literature for what sounds like the same test. The spread is not sloppiness. It is the model below, run with different inputs.
The following is my model, not a vendor’s and not a measurement, built from published measured inputs. Its purpose is that you can re-run it with your own numbers.
In steady state, once the persistent cache is full, the drive can accept new host writes only as fast as it frees cache space, and it frees cache space only by cleaning bands. So the host-visible throughput is:
throughput = (host bytes folded per band cleaned) / (seconds per band cleaned)
The denominator is measured: 0.6 to 1.6 seconds. The numerator depends entirely
on how the workload is spread. If the cache holds C bytes of dirty data and
that data is distributed over D distinct bands, each cleaning pass retires
C / D bytes of host data. Spread the same writes over more bands and each
cleaning pass retires less.
Take an 8 TB drive, 30 MiB bands, a 25 GiB persistent cache and 4 KiB random
writes. The drive has 8e12 / 31,457,280 = 254,313 bands, and a full cache holds
26,843,545,600 / 4,096 = 6,553,600 outstanding writes. The number of distinct
bands those writes touch, over a span of n bands, is the balls-in-bins
occupancy expression n x (1 - e^(-m/n)) with m writes - the expected number
of bins hit at least once, not the coupon-collector result, which answers the
opposite question:
LBA span bands in span distinct dirtied bytes per band @0.6 s @1.6 s
50 GiB 1,707 1,707 15,728,640 26.2 MB/s 9.8 MB/s
200 GB 6,358 6,358 4,222,125 7.0 MB/s 2.6 MB/s
1 TB 31,789 31,789 844,425 1.41 MB/s 0.53 MB/s
2 TB 63,578 63,578 422,212 0.70 MB/s 0.26 MB/s
8 TB (all) 254,313 254,313 105,553 0.18 MB/s 0.07 MB/s
Read that table twice, because it dissolves most of the disagreement in the literature at a stroke.
A 4 KiB random write test confined to 50 GiB of a drive-managed 8 TB drive
predicts about 26 MB/s. The same test spread across the whole drive predicts
about 0.18 MB/s. Nothing about the drive changed. A published independent fio
recipe that uses --size=50g measured a Seagate ST8000DM004 at roughly 50 to
70 MiB/s against 120-plus MiB/s for two conventional drives - the same order as
this model’s 50 GiB row, once you allow that not every write lands on a
previously clean band. The FAST ’17 figures, which exercise a much larger
footprint, sit two orders of magnitude lower. Both are correct measurements of
different questions.
Three things fall out of the model that are worth more than the numbers:
- Block size matters because it changes bands dirtied per byte, not because large writes are inherently kinder. A 128 KiB random write dirties one band per 128 KiB; thirty-two scattered 4 KiB writes dirty up to thirty-two bands for the same 128 KiB of data. That is the whole reason 4 KiB numbers are so much worse than 128 KiB numbers, and it is why quoting a random-write figure without the block size is meaningless.
- Confining a workload to a small LBA range genuinely helps. This is not a trick; it is the numerator of the model. Partitioning a drive-managed SMR drive so that the churning data lives in a small region and the static data lives everywhere else is a real mitigation, and about the only one available from outside the firmware.
- The collapse depth is a property of the workload’s footprint, not of the drive. Anyone who quotes you “SMR drives do 2 MB/s” without saying over what span, at what block size, and on which vendor’s mapping scheme, is quoting a number that could be off by a factor of a hundred in either direction.
Your benchmark measured the cache and reported it as the drive
The write curve has three regions, and only the third is interesting.
| Workload | Behaviour | Roughly |
|---|---|---|
| Sequential writes, large, volatile cache enabled | Detected as sequential and streamed straight into bands, bypassing the persistent cache | Near CMR speed |
| Random writes, bursty, total under the cache size | Absorbed by the persistent cache, folded back during idle | CMR speed or better |
| Random writes, sustained past the cache with no idle time | Every write waits on a band read-modify-write | One to three orders of magnitude down |
The first row is why SMR exists; the second is why the problem is hard to find. The second row understates it, in fact: Skylight measured random I/O throughput on a drive-managed drive at fifteen times that of the equivalent conventional drive, with the volatile cache enabled or at high queue depth, for as long as the persistent cache had room. A drive-managed SMR drive is genuinely, measurably faster than CMR at random writes right up until it is catastrophically slower. That is the single most treacherous fact in the subject, and it is why short benchmarks do not merely miss the problem - they actively reward the drive.
Now the arithmetic of why almost no benchmark finds it. Take the measured 25 GiB persistent cache and a drive absorbing writes at 150 MB/s:
cache size 25 GiB = 26,843,545,600 bytes
time to fill the cache, sequential 26.84e9 / 150e6 = 179 s = 3.0 min
time to fill it at 50 MB/s effective 26.84e9 / 50e6 = 537 s = 8.9 min
CrystalDiskMark default test file 1 GiB = 4.0% of cache
CrystalDiskMark 1 GiB x 5 passes 5.4 GB = 20.0% of cache
ATTO default 256 MB = 0.95% of cache
A default benchmark run writes a fraction of the cache and stops. It never reaches the mechanism it is meant to measure, reports the cache’s speed as the drive’s speed, and is not wrong about what it measured. It measured the wrong thing.
The pauses make it worse. A suite that runs a write test, prints results and starts the next has handed the drive the exact idle window it needs to fold the backlog, so every test begins with a clean cache - and on a Seagate part, which cleans during idle by design, the gift is maximal. To see the cliff you must write past the cache in one unbroken run and then keep writing - a test measured in hours, over a large LBA span, at a small block size. Which is why a used SMR drive can arrive, pass every short test you throw at it, post better random-write numbers than the CMR drive next to it, and fail you three weeks later on the first real workload.
Reads are not free, and a disabled write cache makes writes worse
An earlier version of this guide said that shingling costs nothing on reads and that every problem SMR has is a write problem. The second half survives. The first half is too strong, in three separate ways, and the corrections are more interesting than the original claim.
Reads of data still sitting in the persistent cache are not where the map says they are. The cache is at one radius; the band the data logically belongs to is at another. A sequential read across a region that has been sparsely overwritten therefore has to keep breaking off to fetch the cached fragments from the cache region, which is a seek in each direction. Skylight measured exactly this effect. Reads of never-modified data really are free; reads of recently modified data are not, and the penalty is a function of how much folding the drive still owes you.
Sequential writes are only fast if the volatile write cache is on. This is the least-known result in the Skylight paper and the most likely to catch a careful administrator. With the volatile write cache disabled, sequential write throughput on the Seagate drive-managed drives tested was less than a third that of a conventional drive, because static mapping forces the drive to visit physical locations in an order the host’s sequential stream does not supply. Full sequential throughput required the volatile cache enabled. Anyone who turns write caching off for power-loss safety, or who buys an enterprise drive shipped in write-cache-disabled mode, is not getting the behaviour that the first row of the table above promises.
And an SMR drive can be sequentially slower than the CMR drive it replaces, on the same platform. Western Digital’s own datasheets make the comparison available, provided you read the units. Both datasheets state the sustained transfer rate twice, once in MB/s and once in MiB/s, and the two figures sit next to each other in the same row. The Ultrastar DC HC690, an 11-disk helium platform at 30 TB and 32 TB, states SMR as its recording technology and a sustained transfer rate of 269 MB/s at 32 TB. The Ultrastar DC HC590, the same 11-disk platform at 24 TB and 26 TB, states CMR and up to 302 MB/s. The shingled drive holds 23 per cent more data and streams it 11 per cent slower:
MB/s compared with MB/s (269 / 302) - 1 = -0.109
MiB/s compared with MiB/s (257 / 288) - 1 = -0.108
Comparing the SMR drive’s MB/s against the CMR drive’s MiB/s is the easy mistake here, and it halves the deficit to about 6.6 per cent. Quote one unit or the other and stay in it.
There is no mystery in that. Sustained transfer rate is bits per second under the head, which is linear density times velocity; shingling raises track density, not linear density, and the wider platter-count and format overheads of the SMR layout have to come from somewhere. But it does dispose of the idea that shingling is free as long as you write sequentially. It is cheap. It is not free.
UltraSMR is four technologies, and each one is a confession
Western Digital’s UltraSMR is worth reading closely, not as marketing but as a disclosure of what shingling actually requires. Its technology brief describes four components, and three of the four exist only because SMR forbids update-in-place. The constraint that makes SMR slow is the same constraint that makes these error-correction techniques possible.
Two-dimensional magnetic recording. A second read sensor on the slider, offset from the first, reads the adjacent track at the same time as the target track. Knowing the neighbour’s signal lets the channel subtract inter-track interference rather than merely tolerate it, which permits a narrower surviving strip, which permits a smaller shingle step. This is the one component that attacks the physics directly.
Distributed sectors (DSEC). Sixteen logical 4 KB sectors are interleaved across a 64 KB physical area rather than each being written as one contiguous run. A localised media defect - a scratch, a contaminant, a thermal asperity - then damages a small fraction of sixteen sectors instead of destroying one sector outright, and a small fraction of a sector is well within what per-sector ECC recovers. WD states DSEC “is ideal for SMR”, and the reason is the confession: aggregating sixteen sectors into one 64 KB block write costs nothing on a device that cannot update a single sector in place anyway. On a CMR drive, DSEC would turn every 4 KB update into a 64 KB read-modify-write. On an SMR drive the read-modify-write is already mandatory, so the interleave is free.
Soft-decoded track ECC (sTECC). Conventional track-level parity works in
erasure mode: with N parity sectors you can reconstruct at most N known-bad
data sectors, because you need to know which ones failed. sTECC uses
likelihood information from the read channel - how confident the detector is
about each bit, rather than a hard yes or no - so a single parity sector can
contribute to correcting several damaged data sectors. WD claims the efficiency
of the correction improves “by an order of magnitude or more”. That is a vendor
claim about a vendor technology and is not independently verified here, but the
mechanism is standard soft-decision coding and the direction of the claim is
unsurprising.
Track parity itself is the fourth confession. The parity sectors are generated by a bitwise XOR across all the data sectors on the track at the moment the zone is written, and they never need updating until the whole zone is rewritten. On a CMR drive, maintaining track-level parity would mean a read-modify-write of the parity sector on every single sector update, which is why CMR drives do not do it. SMR writes whole zones or nothing, so the parity is computed once, for free, in the write stream.
OptiNAND, an embedded iNAND flash device on the drive, is described as a required enabler rather than an optional extra. An SMR drive holds partially written data for multiple open zones at any moment, and that transient state must survive an emergency power loss before it is committed to the platter. Flash on the drive is where it goes. If you have wondered why very high capacity SMR drives carry an on-board NAND device, this is the answer: shingling creates in-flight state that a purely magnetic drive has nowhere safe to put.
The pattern across all four is the same. Shingling removes the ability to update in place, and then spends the savings on error correction techniques that only work because you cannot update in place. That is a coherent engineering position. It is also a complete explanation of why an SMR drive cannot be made to behave like a CMR drive by firmware effort alone.
Host-aware SMR is dead, and the standards have deleted it
Shingling can be managed by the drive, by the host, or by both. The three-way split is usually presented as a spectrum with a live middle. It is not, any more.
| DM-SMR | HA-SMR | HM-SMR | |
|---|---|---|---|
| Name | Drive-managed | Host-aware | Host-managed |
| Who does the read-modify-write | The firmware, invisibly | Either | The host, always |
| Reports its zones | Never - at most it sets the ZONED field to declare itself | Yes | Yes |
| Accepts a random write | Yes, silently | Yes, drive handles it | No - rejected with an error |
| Works as an ordinary disk | Yes | Yes | No |
| Standard | No zoned command set | Removed from ZBC and SBC | ZBC / ZAC |
| In the Linux kernel | Treated as a regular disk | Removed in 6.8 | Supported since 4.10 |
| Status | Ubiquitous in the retail channel | Dead | Cloud, hyperscale, sold as such |
Host-aware is not rare. It is gone. The model has been removed from the ZBC and SBC specifications, it was never implemented in NVMe, and Christoph Hellwig’s patch series removing it from Linux was posted on 17 December 2023 and landed in kernel 6.8. The core commit, “block: remove support for the host aware zone model”, touches 19 files for a net removal of about 140 lines, which tells you how little was left to remove. A handful of HDD prototypes shipped host-aware and none reached mass production.
The stated reason is worth carrying, because it is a real lesson about half-measures in interfaces. A host-aware drive reported a write pointer per zone but also accepted writes that ignored it, so the pointer could drift into an under-defined state. Host software could never rely on the write pointer to actually be useful for, say, recovery. An interface that reports state it does not enforce reports nothing. Either the drive rejects the misaligned write, in which case the pointer means something, or it does not, in which case the host gains a number it cannot trust and a complexity budget it has spent for nothing.
The consequence for a shopper is simple: there are two kinds of SMR drive you can buy today, one that tells you everything and refuses to misbehave, and one that tells you nothing.
DM-SMR is the one a consumer encounters, and it is under no obligation to say what it is. Any sector can be written at any time; the operating system, the filesystem and the SMART data see an ordinary disk. It is not quite true that the shingling is always a private detail of the firmware - there is a standard field for it, covered below, and a few drives set it - but it is true that almost none do. That invisibility was deliberate, and commercially rational: it let SMR into channels that could not have absorbed a new command set. It is also the direct cause of everything else in this guide.
ZBC and ZAC: what a zoned drive actually promises
Host-managed drives are not a trap at all, because they cannot pretend. They implement a published command set, they advertise themselves unambiguously, and they refuse the operations they cannot do cheaply. Knowing the shape of that interface is also the fastest way to understand what the drive-managed firmware is doing behind your back, because it is doing the same work with the same constraints and no vocabulary to describe it.
The standards are two parallel documents, one per transport, from two INCITS technical committees:
| Standard | Transport | Committee | Document |
|---|---|---|---|
| ZBC | SCSI / SAS | INCITS T10 | ANSI INCITS 536-2016, reaffirmed R2021 |
| ZAC | ATA / SATA | INCITS T13 | ANSI INCITS 537-2016, reaffirmed R2026 |
| ZBC-2 | SCSI / SAS | INCITS T10 | INCITS 550-2023 |
| ZAC-2 | ATA / SATA | INCITS T13 | INCITS 549-2022, also ISO/IEC 17760-302:2025 |
| ZBC-3 | SCSI / SAS | INCITS T10 | In development as INCITS 579 |
A host-managed device declares itself in the identification path before anyone
issues a zone command. On SCSI it reports PERIPHERAL DEVICE TYPE 14h, Host
Managed Zoned Block Device - a distinct device type, not a flag on an ordinary
disk. On SATA it presents the device signature 0xABCD, which is how a SATA
host distinguishes one at the link layer. The now-removed host-aware model did it
the other way around: device type 00h, an ordinary disk, with a HAW_ZBC bit set
to 1b. That difference is precisely why lsscsi showed host-aware drives as
plain disk and shows host-managed ones as zbc, and why a SAS HBA that does
not implement SAT properly and recognise device type 0x14 can make a host-managed
SATA drive invisible as a zoned device entirely.
ZBC defines three zone types and eight zone conditions. Linux mirrors them
one-for-one in enum blk_zone_type and enum blk_zone_cond:
| Zone type | Code | Meaning |
|---|---|---|
| Conventional | 1h | Ordinary rewritable region, no write pointer |
| Sequential write required | 2h | Writes must start at the write pointer or be rejected |
| Sequential write preferred | 3h | Host-aware only; no longer present |
| Zone condition | Code |
|---|---|
| NOT_WP (no write pointer) | 0h |
| EMPTY | 1h |
| IMPLICITLY OPENED | 2h |
| EXPLICITLY OPENED | 3h |
| CLOSED | 4h |
| READ-ONLY | Dh |
| FULL | Eh |
| OFFLINE | Fh |
What ZBC requires of a host-managed device is narrower than people assume. At least one sequential-write-required zone is mandatory. Conventional zones are optional. Sequential-write-preferred zones are not supported at all on host-managed devices. A host-managed drive may therefore have no conventional region whatsoever, which matters because some filesystem support paths depend on having one.
Five zone management commands exist, with EXT variants on ATA:
REPORT ZONES enumerate zones, their types, conditions, write pointers
RESET WRITE POINTER return a zone to EMPTY; the zoned equivalent of erase
OPEN ZONE explicitly open a zone, reserving device resources
CLOSE ZONE close an open zone, releasing those resources
FINISH ZONE mark a zone FULL without writing the remainder
A write to a sequential-write-required zone that does not begin at that zone’s write pointer returns UNALIGNED WRITE COMMAND. Writes must also be aligned to the device’s physical block size. This is the enforcement that makes the write pointer trustworthy, and the absence of it is what killed host-aware.
Two further practical notes. Linux supports only host-managed disks that have unrestricted reads enabled - the URSWRZ bit - meaning reads above the write pointer return defined data rather than an error. That covers all commercially available SMR drives, so it is not a purchasing consideration, but it is the reason the kernel can treat a zoned device as a block device at all. And zone sizes are fixed at manufacturing time; there is no user-accessible way to change them under ZBC, and the only standard mechanism at all is ZBC-2’s FORMAT WITH PRESET, which is destructive.
The zone state machine the host has to run
The reason host-managed SMR is a software project rather than a driver setting is that the host now owns a state machine per zone, with device-imposed limits on how many zones may be in which state. A zone moves EMPTY to OPENED as soon as it is written, OPENED to CLOSED when the host releases device resources but has not finished the zone, CLOSED back to OPENED on the next write, and OPENED or CLOSED to FULL when the write pointer reaches the end or the host issues FINISH ZONE. RESET WRITE POINTER takes any of them back to EMPTY, and that is the only way to reclaim space.
Two consequences follow that have nothing to do with speed:
Deletion is not a thing. There is no per-sector free operation. The only unit of reclamation is the zone, so any filesystem on a host-managed device must do its own garbage collection: relocate the live data out of a partly dead zone, then reset the whole zone. This is exactly what an SSD’s FTL does, moved up into software where you can see it, and it is the real reason the software stack is non-trivial.
Open zones are a rationed resource. A device advertises a maximum number of open zones, and exceeding it is an error rather than a slowdown. A filesystem that wants to keep many independent write streams has to budget them. An application that assumed it could append to arbitrarily many files at once has to be rewritten.
ZBC-2 turned zones into a resource you can reallocate
After host-aware died, the standards went in a more honest direction: instead of a drive that pretends random writes are fine, a drive that lets the host convert regions between conventional and sequential at runtime. ZBC-2 adds zone domains and zone realms, the ZONE ACTIVATE command, and the REPORT REALMS and REPORT ZONE DOMAINS commands to inspect them. Drives built on this are sold as DH-SMR, sometimes written XMR.
The trade is explicit and metered. A realm can be activated as conventional capacity or as sequential-write-required capacity, and conventional capacity costs more surface area because it is not shingled. You choose, at runtime, how much of the drive is fast-and-random and how much is dense-and-sequential, and the drive tells you the exchange rate. That is the thing host-aware failed to be: a middle ground where the host knows exactly what it is getting.
sg_rep_zones from sg3_utils is the tool for inspecting all of this, and it is
three commands rather than one. Its default is REPORT ZONES per ZBC; --domain
issues REPORT ZONE DOMAINS and --realm issues REPORT REALMS, both from ZBC-2.
Its report filter accepts 0 for all zones, 1 through 7 to filter by zone
condition, 0x10 for zones where a reset write pointer is recommended, 0x11 for
zones with non-sequential write resources active, 0x3e for all but gap zones and
0x3f for not-write-pointer zones. It is also worth correcting a common gloss:
sg_rep_zones is not SAS-only. It works on SATA ZAC drives through SCSI-to-ATA
translation, provided the HBA supports SAT.
The two bits that could have answered this question
Here is the correction that matters most, because the previous version of this guide got it wrong and the wrong version is the version everyone repeats.
It is commonly said that there is no recording-technology field in ATA IDENTIFY DEVICE or SCSI INQUIRY. That is false. There is one in each, and they are the same field bridged across transports.
On ATA: IDENTIFY DEVICE word 69, bits 1:0 is the ZONED field.
00b not reported
01b host aware
10b device managed
Linux reads it as ata_id_zoned_cap() in include/linux/ata.h, which is
literally id[69] & 0x3.
On SCSI: the Block Device Characteristics VPD page, page code B1h, byte 8,
bits 5:4 carries the same three values. The Linux SCSI disk driver reads it in
sd_read_block_characteristics() as sdkp->zoned = (vpd->data[8] >> 4) & 3.
And libata bridges them. For a SATA drive behind the SCSI layer, libata reads
ATA word 69 bits 1:0 and writes the value into the synthesised SCSI B1h page at
byte 8 bits 5:4 - rbuf[8] = (zoned << 4) in ata_scsiop_inq_b1(). It also
advertises VPD page B6h, zoned block device characteristics, only if the device
is genuinely zoned.
So a drive-managed SMR drive can declare itself, in a standard field, on either transport. And when one does, three separate pieces of software will tell you:
dmesg | grep -i SMR
The Linux SCSI disk driver prints exactly one of three messages on first scan.
For SCSI device type ZBC it prints Host-managed zoned block device. For a ZONED
value of 1 it prints Host-aware SMR disk used as regular disk. For a value of 2
it prints Drive-managed SMR disk, and that last string is a real detection
method for a drive-managed drive. It returns a positive on any drive that sets
the bit, and silence otherwise.
smartctl -i prints a Zoned Device: line reading either
Device managed zones or Host aware zones, decoded straight from ATA word 69
bits 1:0, and emits the line at all only if the field is non-zero. Both strings
are sentence case, which matters if you are grepping for them. The JSON output
carries it as zoned_device.capabilities with the value device_managed or
host_aware. So the older claim in this guide that smartctl -i gives you a
great deal and none of it is this was also wrong. smartctl will name a
drive-managed SMR drive if the drive lets it, and stays silent otherwise. The
sample output people circulate showing no recording-technology line is what you
see when the ZONED field reads zero - the common case, but not the capability.
sg_vpd -p bdc reads the raw B1h page, which is the SCSI original rather than a
decode, and is the right tool when you want to see the field itself rather than
somebody’s interpretation of it.
One tool that does not decode it is hdparm. Its word-69 feature-name table
in identify.c labels bits 1 and 0 as reserved 69[1] and reserved 69[0]. A
drive-managed SMR drive that honestly sets ZONED to 10b will appear in
hdparm -I as reserved 69[1] and nothing else. If you have been using
hdparm -I as your check, you have been looking at the right word with the wrong
decoder.
And now the twist that makes the whole section bleak. smartmontools’ own source records that the ATA ZONED field was added in ACS-4 and obsoleted in ACS-5. The one standard place in the ATA command set where a drive could have declared itself drive-managed SMR has been withdrawn from the current revision of the standard. It was optional while it existed, it was almost never set, and it is now formally on the way out.
So the accurate framing is: the field exists, it is optional, drive-managed drives almost always leave it at 00b, and the standards body has removed it. That is a much sharper indictment than “there is no such field”, and it is also checkable, which the older claim was not.
TRIM on a spinning disk is a hint, not a test
Data Set Management with the TRIM bit is an SSD feature. Its presence on a spinning hard disk is unusual enough that it circulates as an SMR detection heuristic, and the reasoning behind the heuristic is sound even though the heuristic is not.
Why a shingle translation layer wants TRIM: the STL’s expensive operation is folding a band, and folding a band means reading every live sector in it. Sectors the filesystem has discarded are not live. A drive that knows which logical blocks are dead can skip them during band cleaning, drop the corresponding journal entries from the persistent cache, and in the dynamic-mapping case retire whole bands without reading them at all. TRIM on a drive-managed SMR drive does the same job it does on an SSD - it reduces garbage-collection work - for the same structural reason, which is that both devices have a translation layer with a reclamation problem.
How to read it:
$ sudo smartctl -i /dev/sda
TRIM Command: Available, deterministic, zeroed
$ sudo hdparm -I /dev/sda | grep -i trim
* Data Set Management TRIM supported (limit 8 blocks)
smartctl decodes those three properties from IDENTIFY word 169 bit 0 for supported, word 69 bit 14 for deterministic and word 69 bit 5 for zeroed. hdparm reads word 169 bit 0 and the block limit from word 105. Note in passing that hdparm reads other bits of word 69 perfectly well; it simply has no name for bits 1 and 0.
Why it is not a test, in both directions. Not all SMR drives implement TRIM, and a drive-managed drive with a well-behaved idle-time cleaner has less need of it. And some conventional drives do: Western Digital has published a whitepaper on the general benefits of TRIM for hard disk drives, covering non-shingled lines including WD Purple, where the drive uses discard information for its own internal housekeeping. A positive on TRIM raises your prior that a drive is shingled. It does not settle anything, and a negative settles even less.
A detection procedure that terminates
Here is the whole set of checks, in order, with what each one proves. Read the warning first, because it applies to every step: almost every check in this list returns only positives. A positive result identifies a zoned or self-declaring drive. A negative result is consistent with a conventional drive and with a silent drive-managed SMR drive, and cannot distinguish them.
Step 1. The vendor datasheet or product manual. Some state the recording technology plainly, usually in the specification table of the datasheet PDF rather than the marketing page. Western Digital’s Ultrastar datasheets state “Recording Technology: SMR” or “CMR” outright. Toshiba now states it per model on its product pages. Seagate maintains a live public CMR/SMR list. This is the only source that is both authoritative and available before purchase. Its coverage is patchy and it is generation-specific.
Step 2. lsblk -z, which has seven zone columns and not one. The shortcut
-z is equivalent to --zoned, and the available columns are ZONED (zone
model), ZONE-SZ (zone size), ZONE-WGRAN (write granularity), ZONE-APP (zone
append maximum bytes), ZONE-NR (number of zones), ZONE-OMAX (maximum open zones)
and ZONE-AMAX (maximum active zones).
lsblk -z
lsblk -o NAME,ROTA,SIZE,MODEL,ZONED,ZONE-SZ,ZONE-NR,ZONE-OMAX
A ZONED value of host-managed is proof. none proves nothing.
Step 3. sysfs, which is more informative than the single zoned file.
cat /sys/block/sda/queue/zoned # host-managed | none
cat /sys/block/sda/queue/chunk_sectors # zone size in 512-byte sectors
cat /sys/block/sda/queue/nr_zones # number of zones
cat /sys/block/sda/queue/max_open_zones
cat /sys/block/sda/queue/max_active_zones
cat /sys/block/sda/queue/zone_write_granularity
Two corrections to the folklore here. On a current kernel zoned can only
return host-managed or none. queue_zoned_show() in block/blk-sysfs.c
has exactly two emit strings; the host-aware value no longer exists, having
gone with the rest of host-aware support in 6.8. And the kernel’s own stable
sysfs ABI documentation states the negative case explicitly, which is the best
citation available for the central point of this guide: since drive-managed zoned
block devices do not support zone commands, they will be treated as regular block
devices and zoned will report none.
zoned: none does not mean the drive is CMR. It means the drive is not
telling the host about zones, and a drive-managed SMR drive by definition does
not. A CMR drive and a DM-SMR drive produce byte-identical output here. The
negative result carries no information at all, and the kernel documentation says
so in as many words.
Step 4. dmesg | grep -i SMR. This catches the self-declaring
drive-managed drive that steps 2 and 3 will miss, because the message is printed
from the B1h ZONED field rather than from zone support. A hit on the
Drive-managed SMR disk string is proof. Silence is not.
Step 5. smartctl -i, looking for two specific lines. Zoned Device: if the
ZONED field is non-zero, and TRIM Command: as a weak prior. Use
smartctl --json and read zoned_device.capabilities if you are scripting it.
Step 6. sg_vpd -p bdc, for the raw field. The abbreviation bdc is the
Block Device Characteristics page, B1h, which carries the ZONED field; zbdc is
the Zoned Block Device Characteristics page, B6h, which only exists on genuinely
zoned devices. ZBC support arrived in sg3_utils 1.39, which introduced
sg_rep_zones and sg_reset_wp; 1.47 added REPORT ZONE DOMAINS and REPORT
REALMS support along with zone alignment mode and zone starting LBA granularity.
Step 7. lsscsi -g, for the device type. Host-managed ZBC and ZAC disks show
as device type zbc in the second column. This is the SCSI PERIPHERAL DEVICE
TYPE 14h surfacing. Host-aware drives showed as disk because their signature
was identical to a regular disk, which is one more reason the model did not
survive.
Step 8. Zone enumeration, if any of the above said yes.
sudo blkzone report /dev/sda
sudo sg_rep_zones /dev/sda
sudo sg_rep_zones --domain /dev/sda # ZBC-2 zone domains
sudo sg_rep_zones --realm /dev/sda # ZBC-2 zone realms
blkzone report prints columns start, len, cap, wptr, reset, non-seq, cond and
type, with zone conditions abbreviated as cl closed, nw not write pointer,
em empty, fu full, oe explicitly opened, oi implicitly opened, ol
offline and ro read only. For richer inspection, libzbc implements INCITS 550
ZBC-2 and ZAC-2 revision 15 and ships zbc_info, zbc_report_zones, the
open/close/finish/reset zone utilities, and the gzbc and gzviewer GUIs for
watching zone state and write pointers move in real time.
Step 9. The listing itself. Sometimes it says. Usually it does not. A search of listing text is at least fast, and it runs over the hard drive listings.
Step 10. Measuring it, after you have already bought it. The next section.
And now the closing statement of the whole procedure, which is the same as the opening one. Every step above returns a positive or silence. There is no sequence of commands that proves a drive is CMR. Western Digital’s own support page on determining CMR versus SMR offers no diagnostic tool at all: its advice is that legacy HGST and WD drives built before 2018 use CMR, so check the build date on the physical label, and that for everything else you should consult the product datasheet. The vendor does not claim a software method exists, because there is not one.
Measuring it yourself, and why the span of the test decides the answer
The only reliable test on a drive-managed drive is behavioural: write past the cache and watch for the collapse. The model earlier in this guide says the answer you get depends almost entirely on how wide a region you write, so the parameters below are the point of the exercise, not boilerplate.
# DESTRUCTIVE. This writes to the raw device and will destroy its contents.
sudo fio --name=smr-probe --filename=/dev/sdX --rw=randwrite \
--bs=4k --iodepth=32 --ioengine=libaio --direct=1 \
--size=100% --runtime=14400 --time_based \
--write_bw_log=smr --write_lat_log=smr --log_avg_msec=1000
Four parameters do the work:
--size=100%spans the whole device. This is the single most important choice. A published independent recipe uses--size=50g, which is a reasonable compromise for a quick comparison but measures the 26 MB/s row of the model rather than the 0.18 MB/s row. Both are real; only one is the worst case.--bs=4kmaximises bands dirtied per byte written. At 128 KiB you will measure something an order of magnitude kinder.--time_basedwith a long--runtimestops fio finishing before the cache does. Four hours is a reasonable floor and eight is better.--write_lat_logmatters as much as the bandwidth log, because the distribution shape is the signal. In the published recipe the SMR drive averaged roughly 50 to 70 MiB/s with a broad, smeared bandwidth distribution and a latency distribution about double and much wider than the conventional comparators, which held around 120 MiB/s or more with tight clustering. A wide distribution at an acceptable mean is the fingerprint, because it is the cache absorbing some writes at full speed while others wait behind a fold.
Plot the bandwidth log. A conventional drive gives a flat, boring line for the whole run. A drive-managed drive holds cache speed, falls off a step, and does not come back - and if it is a statically mapped Seagate part the step is a cliff, while a dynamically mapped Western Digital part starts lower and descends more gently, per the vendor comparison earlier.
Two honest limits on the test. It tells you nothing before the money is spent. And a run that shows no collapse is weak evidence rather than proof - you may simply not have written past the cache, or not over a wide enough span, or the drive may have found idle time you did not notice giving it.
Why it broke RAID rebuilds and ZFS resilvers
A RAID rebuild or a ZFS resilver is the worst workload you can hand an SMR drive, and a NAS drive is bought specifically to do it. That is why undisclosed SMR became a scandal rather than a footnote.
A rebuild writes to the replacement continuously for hours or days with no idle windows, so the cache fills early and never drains. It is not purely sequential either: a ZFS resilver walks the pool’s block tree in transaction-group order rather than disk order, so the target sees a substantially random stream spread over the whole device - the bottom row of the model’s table, not the top. Recent OpenZFS versions sort and batch that traffic, which helps, and does not convert it into a sequential stream.
The published field measurements:
| Source | Measurement |
|---|---|
| ServeTheHome, 28 May 2020: WD40EFAX (SMR) against WD40EFRX, Seagate ST4000VN008 and HGST 0F26902 (all CMR), four-drive FreeNAS RAIDZ about 60 per cent full, with concurrent 1 MB file copies and 2 TB of reads during the rebuild | SMR “can put data at risk 13-16x longer than CMR” |
| An independent ZFS resilver of a two-way mirror of 2.5-inch SMR SATA drives, May 2024 | 2.78 TB resilvered in 3 days, 14 hours, 33 minutes, 47 seconds, ZFS reporting 14.2 MiB/s scanned and 13.1 MiB/s issued, with zero errors |
That second one is the more useful data point precisely because it succeeded. The
overall average across the elapsed time is mine, not the author’s:
2.78e12 / 311,627 s = 8.9 MB/s. Nothing failed. It simply took most of a week
to restore redundancy on a pair of 2.5-inch drives, and for the whole of that
week the mirror had no redundancy to spare.
Run the general arithmetic yourself for a member of your own array:
| Sustained write rate at the target | Rebuild 4 TB | Rebuild 20 TB |
|---|---|---|
| 180 MB/s (healthy CMR) | 6.2 hours | 30.9 hours |
| 40 MB/s | 27.8 hours | 5.8 days |
| 13.7 MB/s (13.1 MiB/s, the measured resilver above) | 3.4 days | 16.9 days |
| 8 MB/s | 5.8 days | 28.9 days |
| 2 MB/s | 23.1 days | 115.7 days |
the resilver row, converted first because ZFS reports binary megabytes
13.1 MiB/s x 1,048,576 = 13.74e6 B/s = 13.7 MB/s
4e12 / 13.74e6 = 291,000 s = 3.4 days
20e12 / 13.74e6 = 1,456,000 s = 16.9 days
4e12 / 8e6 = 500,000 s = 5.8 days
20e12 / 8e6 = 2,500,000 s = 28.9 days
That conversion is not pedantry. ZFS reports scan and issue rates in MiB/s and drive datasheets quote MB/s, and carrying the bare number across the boundary inflates every duration you derive from it by 4.9 per cent.
The array has no redundancy margin for the whole of that window, and a second failure inside it is the exact event the array was built to survive. At 20 TB members the SMR rows are not degraded operation, they are a different mode of existence.
But the sharpest failure was not slowness. It was a hard error. iXsystems documents, in its notice on Western Digital Red SMR drive compatibility with ZFS, that WD Red DM-SMR drives can be driven into an unresponsive state by TRIM-queue overflow: when the TRIM commands overflow, the drive cannot handle normal I/O and returns IDNF responses. Its recommended recovery is to stop all I/O, leave the drive powered, and let it drain its TRIM queue - a process that iXsystems says “can take many hours or days”. It states that it cannot recommend these drives for FreeNAS or TrueNAS.
IDNF is ID Not Found. That is an addressing error, not a timeout. A timeout means the drive was slow; IDNF means the drive told the host it could not locate the sector it was asked for. A translation layer that has run out of room to track discards and answers the host with an addressing error is a firmware defect, not a performance characteristic, and it is qualitatively worse than everything else in this guide.
Note the mechanism, because it is not the one people repeat. The documented trigger is discard pressure, not write pressure: a host issuing TRIM commands faster than the drive can retire them. The section above explains why a shingle translation layer wants TRIM at all, and this is the bill for that want. The guide’s earlier editions attributed the IDNF failures to write load during resilvering; heavy write load is the circumstance in which the failures were reported, not the cause iXsystems documents, and the distinction matters because it means a quiet drive with a large pending discard queue is also exposed.
A concrete instance is on record. OpenZFS issue #10214, filed 16 April 2020,
documents a WD40EFAX resilvering at 101 MB/s scan and 96.7 MB/s issue until
27.61 per cent - about 947 GB - at which point IOPS fell to zero, write errors
climbed, and the kernel logged DID_SOFT_ERROR and I/O errors. The resilver
never completed. A rebuild meant to restore redundancy instead consumed it.
Note the shape of that failure against the model. The drive ran at full speed for 947 GB - far more than a 25 GiB cache - because a resilver’s initial phase is largely sequential and the STL streams sequential writes straight into bands. It collapsed when the stream stopped being sequential. This is why “it was fine for the first terabyte” is not evidence of anything.
A mixed array is no halfway house either, because one SMR member paces the rebuild for all of them. What a NAS actually wants is the subject of drives for a NAS; the short version is that recording technology matters more there than anywhere else.
The 2020 disclosure episode, vendor by vendor
In 2020 it emerged that drive-managed SMR had been shipping in NAS-marketed drives without being stated. The three vendors behaved differently, and the differences are still visible in how each one documents its products today. It is worth getting the details right, because the summary version - “all three published disclosure lists” - is not what happened.
Western Digital issued a statement on 20 April 2020 that read, verbatim: “WD Red capacities 2TB-6TB currently employ device-managed shingled magnetic recording (DMSMR) to maximize areal density and capacity. WD Red 8-14TB drives use conventional magnetic recording (CMR).” On 23 June 2020 it restructured the line and published the split explicitly: WD Red as DMSMR at 2, 3, 4 and 6 TB with a 180 TB/year workload rating; WD Red Plus as CMR; WD Red Pro as CMR. The capacity ranges quoted at the time were 1-14 TB for Red Plus and 2-18 TB for Red Pro, and both lines have extended since. WD explicitly directed ZFS users to Red Plus and Red Pro.
Toshiba published an actual list, in a press release dated 28 April 2020: P300 at 6 and 4 TB, DT02 at 6 and 4 TB, DT02-V at 6 and 4 TB, L200 at 2 and 1 TB, and MQ04 at 2 and 1 TB. Its wording was direct about the consequence - SMR “is recognized as having an impact on write-speeds in drives where this technology is used, especially in the case of continuous random writing” - and it pointed NAS buyers at the N300.
Seagate did not publish a list. It confirmed four specific models to the press - ST2000DM008 at 2 TB, ST4000DM004 at 4 TB, ST8000DM004 at 8 TB and ST5000DM000 at 5 TB - and answered the question about non-disclosure with: “We provide technical information consistent with the positioning and intended workload for each drive.” Seagate’s Exos documentation did describe SMR; the BarraCuda, Desktop HDD and Archive product manuals did not. The live public CMR/SMR list Seagate maintains today is a later and separate artefact, and it is a current-products page rather than a 2020 disclosure - which matters, because it no longer lists the 2 TB ST2000DM008 that was one of the four drives at the centre of the story.
The non-disclosure was older than the argument, and it is checkable. The Seagate Archive HDD Product Manual, document 100757960 Rev. A, July 2014, is the manual for Seagate’s first consumer SMR family. Extract its text and search it: it contains zero occurrences of “SMR”, “shingled” or “media cache”. The only “bands” documented in it are the sixteen self-encrypting-drive encryption bands, which are unrelated to shingling. That is a vendor’s own product manual for a shingled product, six years before anyone made a fuss, with no mention of the defining characteristic of the product.
The legal outcome was modest. The WD Red class action settled with a $2.7 million fund; claimants received $4 to $7 per drive, with a possible pro-rata adjustment up to 85 per cent of retail. The class covered US purchasers of SMR WD Red NAS drives between October 2018 and 21 July 2021, with a claims deadline in November 2021. Two non-monetary terms were worth more than the money: WD agreed to disclose SMR on product packaging for at least four years after final approval, and it conceded that the SMR WD Reds are not suitable for NAS and RAID use.
Six years on, the durable outcome is that Toshiba states recording technology per model on its product pages, Seagate maintains a public list, and Western Digital prints it on Ultrastar datasheets and on WD Red packaging. That is genuine progress, and it still leaves most of the used market - OEM variants, appliance pulls, drives whose packaging is long gone - exactly where it was.
What the host-managed software stack costs in 2026
Host-managed SMR at a striking price per terabyte is not a bargain unless you already run the software. Here is what “the software” means now, because it has changed substantially and most advice online predates the changes.
The Linux kernel’s zoned support arrived in layers:
| Kernel | What landed |
|---|---|
| 3.18 | SG passthrough for ZBC and ZAC commands |
| 4.10 | Block layer zoned block device support; f2fs zoned support |
| 4.13 | dm-linear, dm-flakey, dm-zoned |
| 4.16 | blk-mq and scsi-mq zoned support |
| 5.6 | zonefs |
| 5.8 | Zone append |
| 5.9 | NVMe ZNS |
| 5.12 | btrfs zoned mode |
| 6.8 | Host-aware support removed |
| 6.10 | Zone write plugging |
| 6.15 | Native XFS zoned support |
And the filesystem options, with their real restrictions:
f2fs has supported zoned devices since 4.10 and is the oldest option. It is a log-structured filesystem, so the fit is natural.
zonefs (5.6) is the most honest and the least convenient: it exposes each
zone as a file. Sequential zone files accept only direct I/O append writes -
buffered writes and writable shared mappings are prevented outright. File size
tracks the write pointer. Truncation is permitted only to zero, which resets the
zone, or to the zone capacity, which finishes it. It requires a block elevator
implementing ELEVATOR_F_ZBD_SEQ_WRITE. This is not a general-purpose
filesystem; it is a thin, safe interface for an application that already
understands zones.
btrfs zoned mode (5.12 for SMR HDDs, 5.16 for ZNS SSDs) is the most
attractive-sounding and the most constrained. All data block group writes use the
Zone Append operation; no SMR hard drive implements it natively, so on a hard
drive the block layer supplies it - zone append emulation, reworked as zone write
plugging in 6.10 - which is one reason SMR zoned writes are serialised. Read
“requires Zone Append” as a device requirement and you will wrongly conclude
btrfs zoned mode cannot run on an SMR disk at all; it can, at a queue depth the
kernel chooses for you. Only the single, RAID0, RAID1 and RAID10 profiles
exist, and the RAID profiles are experimental, requiring
CONFIG_BTRFS_EXPERIMENTAL=y. RAID5 and RAID6 are unsupported, as are NOCOW,
fallocate(2) and mixed data/metadata block groups. All devices in the volume
must share the same zone model. If your plan for cheap host-managed drives
involved btrfs parity RAID, the plan does not exist.
XFS gained native zoned device support in kernel 6.15, implemented through the XFS realtime device. For an SMR disk that has both conventional and sequential zones, an internal realtime device is used and the disk can simply be formatted and mounted, which is by some distance the least painful path available today. It keeps a dedicated open zone for garbage collection and is still flagged experimental in kernel messages. This option did not exist when most of the advice you will find online was written, and it is the single biggest change to the practical picture since 2020.
dm-zoned is the escape hatch: it presents a host-managed drive as an ordinary block device so anything can run on it. The costs are concrete. It forces a 4096-byte logical block size regardless of the device’s physical sector size. It consumes conventional zones for both write buffering and metadata. It triggers zone reclaim when fewer than 50 per cent of random zones remain free. Its memory footprint is small - for a 10 TB disk with 256 MB zones, at most 4.5 MB of RAM and as few as 5 zones for metadata and reclaim - but the reclaim behaviour means you have reinvented the drive-managed cliff in software, where at least you can see it.
OpenZFS has no zoned block device support. An OpenZFS issue filed on 13 January 2026 requesting documentation warnings states plainly that the limiting factor is drive-managed SMR firmware behaviour rather than anything OpenZFS could tune around, and asks for explicit warnings about RAIDZ and mirror membership and about consumer USB external drives. ext4 does not support zoned devices either, though the FAST ’17 ext4-lazy work is the most interesting footnote in the whole subject: relocating the ext4 journal - 80 modified lines plus about 600 new ones - produced a 1.7 to 5.4x improvement on a metadata-light file server benchmark and 2 to 13x on metadata-heavy benchmarks on drive-managed SMR disks. Filesystem layout, not just workload shape, determines how badly SMR hurts. The same drive, the same workload, a different journal location, and up to thirteen times the throughput.
Finally, a change from January 2026 that nobody has absorbed yet. A new sysfs
attribute /sys/block/sda/queue/zoned_qd1_writes was added, and for rotational
zoned block devices - which is to say SMR hard drives - the default is 1,
meaning writes are serialised through a single kernel thread at a maximum queue
depth of one. For ZNS SSDs and zoned UFS the default is 0. The reasoning is
structural: SMR drives have no standardised Zone Append command, so the ordering
guarantees that let a ZNS device accept concurrent writes into a zone do not
exist on an SMR HDD, and the kernel emulates them by refusing to have more than
one write in flight. The kernel now deliberately throttles SMR HDD write
concurrency by default, which is the correct decision and is also a cost you
should know about before you plan capacity around a queue depth you will not be
allowed to use.
Zoned namespaces are the same idea, admitted out loud
NVMe Zoned Namespaces is the SSD analogue of ZBC and ZAC, and comparing them is the clearest way to see what SMR’s interface got wrong. ZNS revision 1.1 was ratified on 3 June 2021 and revision 1.2 on 5 August 2024. It reuses the same zone state model - ZSE Empty, ZSIO Implicitly Opened, ZSEO Explicitly Opened, ZSC Closed, ZSF Full, ZSRO Read Only, ZSO Offline - and then differs in four ways that all favour the host.
ZNS defines exactly one zone type: sequential write required. There are no conventional zones and there is no drive-managed option. A ZNS device cannot pretend to be an ordinary namespace, so the entire category of silent, undisclosed zoning that this guide is about does not exist on ZNS. That is not an accident of the flash medium; it is a decision that SMR’s designers declined to make.
Zone Capacity is separated from Zone Size. ZNS lets a zone’s writable capacity be smaller than its address-space size, so the zone size can stay a power of two for cheap address arithmetic while the capacity matches the actual erase-block geometry. The original ZBC and ZAC have no such separation, which is why SMR zone sizes are what they are and why a device’s zone geometry is an awkward fixed fact rather than a tunable one.
Zone Append lets the device choose the write position and report it back. This is the command that allows queue depth greater than one per zone without unaligned-write errors: the host says “put this somewhere in zone 17”, the device says “I put it at offset X”. It is optional in the specification and mandatory in practice - Linux zoned support and btrfs zoned mode are both built on it. SMR has no standardised equivalent, so on a hard drive Linux emulates the operation in the block layer rather than issuing it to the device, which is exactly why it now defaults SMR HDDs to queue depth one.
Resources are declared, in three classes. Active Resources counts zones that are implicitly open, explicitly open or closed; Open Resources counts implicitly or explicitly open zones; ZRWA Resources counts zone random write areas. Maximum Open Resources is at most Maximum Active Resources. A host can plan against those numbers. SMR gives you maximum open zones and maximum active zones and nothing as structured.
And then the point that ties the whole guide together. ZNS defines the Zone Random Write Area, and the specification’s own description of it is that “a ZRWA may be thought of as being analogous to a type of non-volatile cache.” It is a sliding window of ZRWASZ logical blocks starting at the zone’s write pointer, inside which the host may write out of order. It is optional, advertised by the ZRWASUP bit. It is explicitly allocated per zone by the host. And it is released when the zone becomes Full, Empty, Read Only or Offline.
The ZRWA is the media cache, made honest. Bounded instead of unknown, per-zone instead of global, host-allocated instead of firmware-managed, and described in a published specification instead of inferred from timing by academics whose measurements are now eleven years old. Every property the SMR persistent cache lacks, the ZRWA has, and it does the same job. If you want one sentence for why drive-managed SMR is a bad interface rather than a bad technology, it is that the industry built exactly this mechanism a second time on flash, and wrote it down.
What SMR is genuinely good at
The 2020 argument was about undisclosed SMR in the wrong application, not about SMR being bad. Half the exabytes Western Digital ships into capacity-optimised data centres are shingled; the technology works, at enormous scale, for people who built for it. The shape it suits states easily: written once, sequentially, read many times.
- Archive and cold storage. Written once and left. The read-modify-write path is never exercised because nothing is ever modified.
- Media libraries. Large files, written whole, read sequentially, rarely changed.
- Backup targets written sequentially. Full images and append-only formats are ideal; incremental schemes that rewrite in place are not.
- Video surveillance. A continuous sequential stream per camera, which is why SMR survives comfortably in that market.
- Hyperscale object storage, host-managed, with software written to know about zones. This is where the density gain is actually harvested.
The documented field example is instructive about both the fit and the cost. A 20 TB host-managed SMR drive under btrfs zoned mode reported 256.00 MiB zones and 74,508 zones, and filling it took 29 hours at an average of about 190 MB/s with the writing application seeing 0.54 to 1.81 second latency spikes. Those spikes are the zone-boundary work, and they were acceptable only because the workload was genuinely write-once-read-many. Put an interactive workload behind those numbers and 1.81 seconds is a stall a user notices.
The mirror image, where SMR is simply the wrong purchase: any RAID or ZFS member; virtual machine images and databases, which write scattered and never stop; general-purpose system drives, where metadata churn never stops either; and any drive you intend to run near full.
Seagate’s own framing on its public CMR/SMR page is a warning rather than a pitch. It says that “deploying SMR typically requires teams with architectural expertise and control that includes the ability to adapt or manage zone-based storage behavior and tune the software stack accordingly”, and the page’s own section headings then list the prerequisites one by one - ownership of the storage software stack, and SMR-aware filesystems and storage engines, among them. The vendor selling the densest SMR drives in the world says out loud that most buyers should not buy them. That is the most reliable vendor statement in this entire guide, and it is the one people quote least.
That sentence is worth quoting exactly rather than tightening into something punchier. A paraphrase inside quotation marks is a fabrication, and this is the passage where one would do the most damage: the whole point of citing it is that Seagate said it, not that it sounds true.
HAMR and MAMR are not SMR
Two other acronyms are spreading across listings, and they answer a different question. Both address the superparamagnetic limit: as grains shrink they need higher coercivity to stay stable, until the write head cannot switch them at room temperature at all. Both attack the “switchable” corner of the trilemma, where shingling attacks none of the corners and simply rearranges the tracks.
HAMR, heat-assisted magnetic recording, puts a laser diode in the head. It heats a spot on a high-anisotropy medium for a fraction of a nanosecond, dropping its coercivity far enough to be written; the spot cools in nanoseconds and locks the bit in as the head moves on. The medium can then be made of grains that would be unwritable at room temperature, which is the whole point. Seagate ships this in its Mozaic platform.
MAMR, microwave-assisted, uses a spin-torque oscillator to apply a high-frequency field that drives the grain’s magnetisation into precession, lowering the field needed to flip it. What shipped first was less than the full idea: Western Digital’s ePMR applies a current to the write head to improve the field gradient, and WD has described it as a step towards MAMR rather than MAMR itself. Toshiba has shipped flux-control MAMR in some enterprise lines.
These are writing technologies. SMR is a track layout. They are orthogonal. The clearest possible confirmation comes from the vendors themselves. Seagate manufactures the same Mozaic HAMR platform in both CMR and SMR configurations - its public list shows a 32 TB part in each. Western Digital’s UltraSMR drives combine SMR with ePMR and OptiNAND on one platform. The two density gains multiply, so energy-assisted recording and shingling are routinely combined in the highest-capacity parts.
“HAMR” or “ePMR” in a listing title tells you something real about how the bits are written and nothing whatsoever about whether the tracks overlap. If anything, seeing an energy-assist acronym on a very high capacity drive should raise your estimate that it is shingled, because the densest parts in each generation are where both technologies are applied at once.
Which families are shingled in 2026
This is the fastest-decaying section in the guide and it is dated deliberately. The following reflects Seagate’s public CMR/SMR list as read on 13 September 2026, and Toshiba’s product pages as of the same date. Seagate’s page carries no last-updated marker, which is its own small problem.
Seagate, 3.5-inch. Exos is CMR from 4 to 24 TB; the Mozaic HAMR parts at 28, 30 and 32 TB are CMR; Exos SMR is listed at 32, 36 and 44 TB. BarraCuda 3.5-inch is SMR at 4 TB and 8 TB and CMR at 12, 16, 20 and 24 TB - note that the shingled BarraCudas are the small ones, which is the opposite of what most people guess. Archive 8 TB is SMR. IronWolf and IronWolf Pro are entirely CMR. SkyHawk and SkyHawk AI are entirely CMR. FireCuda 3.5-inch is CMR at 1 and 2 TB.
Seagate, 2.5-inch. BarraCuda is SMR at 2 and 4 TB. FireCuda is SMR at 1 and 2 TB. Exos E is CMR from 600 GB to 2.4 TB.
Toshiba states recording technology per model on its product pages, which is the clearest post-2020 vendor practice of the three: DT02 at 6, 4 and 2 TB is SMR; MQ04 at 2 and 1 TB is SMR; MD07ACA at 14 and 12 TB is CMR; MD04 from 6 down to 2 TB is CMR; MQ01ABF at 500 and 320 GB is CMR.
Western Digital states it on Ultrastar datasheets. The Ultrastar DC HC690 at 30 and 32 TB is SMR, on an 11-disk helium platform - the world’s first - at 7200 rpm with a 512 MB buffer, an areal density of 1385 Gb per square inch at 30 TB and 1480 at 32 TB, a 550 TB/year workload rating, 2.5 million hours MTBF, a 0.35 per cent annualised failure rate, SATA 6Gb/s and SAS 12Gb/s options and a five-year warranty. The Ultrastar DC HC590 at 24 and 26 TB is CMR on the same 11-disk platform.
And the observation that should change what you buy for a laptop:
Every 2.5-inch consumer drive Seagate currently lists is SMR. BarraCuda 2 TB and 4 TB, FireCuda 1 TB and 2 TB - all shingled. Toshiba’s MQ04 at 1 and 2 TB is SMR while the MQ01ABF at 320 and 500 GB is CMR. Western Digital’s WD10SPZX and WD20SPZX are SMR while the WD3200LPCX, WD5000LPCX and WD5000LPVX are CMR. Above roughly 500 GB in a 2.5-inch form factor there is effectively no conventional option in the consumer channel. If you are replacing a laptop drive with a spinning 1 or 2 TB unit, you are almost certainly buying SMR, and the honest recommendation is to buy an SSD instead - the trade-offs are in HDD or SSD.
Three cautions on all of the above. Vendor lists are snapshots of current products, so a discontinued drive drops off and takes its disclosure with it. They do not reliably extend to OEM variants or appliance pulls, which is most of what the used market sells. And a model designation can be reused across generations with different internals.
Model-number decoding works for exactly one product line
The previous version of this guide said flatly that no vendor publishes a model-number-to-recording-technology mapping. That is too strong, and a reader checking an Ultrastar part number would catch it.
Western Digital publishes exactly such a decode, printed on the datasheets
themselves under “How to Read the Ultrastar Model Number”. The HC690 datasheet
decodes its shingled part numbers with “S = Ultrastar SMR Technology”; the
HC590 datasheet decodes WUH722626ALxxyz with “U = Ultrastar”, its
conventional line - a 26 TB drive on a 26 TB platform, which is how the two
capacity fields read. The second character of a WD Ultrastar part number is
therefore a documented recording-technology field:
W S H 7 2 c c c c A L x x y z
| | | \_____/
| | | +-- platform capacity then drive capacity, in terabytes
| | +-- H = Helium, S = Standard (air)
| +---- S = Ultrastar SMR Technology, U = Ultrastar (conventional)
+------ Western Digital
Only the second character is load-bearing for this question, which is why the capacity digits are left as placeholders above rather than printed from memory - quoting a specific part number you have not read off a datasheet is exactly the error this section exists to warn about. Check it against the datasheet for the model in front of you.
That is a real, vendor-published, checkable decode. It also has a strictly bounded scope, and the boundary is the important part:
- It applies to the WD Ultrastar data centre line only.
- It does not apply to WD Red, WD Blue, WD Black, WD Purple or WD Elements.
- It does not apply to Seagate or Toshiba part numbers at all.
- No vendor publishes a general mapping across its whole catalogue.
Everything outside that one line is folklore. Suffix rules circulate on forums, are occasionally right for one generation of one line, and are then invalidated by the next generation, an OEM variant, or a product change that reuses the designation. The WD40EFAX and WD40EFRX pair at the centre of the 2020 story is the textbook case: one character apart, opposite recording technologies, and no published key. Treating a character in a part number as a recording technology is, outside the Ultrastar line, a guess wearing the clothes of a specification.
Why this site says “Not stated”
Where a listing does not state the recording technology, this site records it as Not stated and filters accordingly, rather than inferring it. The reasoning is asymmetry of harm.
A guess that is right saves a reader thirty seconds. A guess that is wrong sends someone who specifically needs CMR - because they are building an array, which is the main reason anyone asks - to buy a drive that will fail them during a rebuild weeks later, in the one situation where their redundancy is already gone. The two errors are not the same size.
So the filter for drives stated as CMR shows drives where the recording technology is stated, not drives believed to be conventional. It is shorter than you want, and that is what the evidence supports - the methodology page says the same about every other field where the source data is silent.
The rule that follows is simple: if you need CMR, treat “not stated” as SMR. That rule is not pessimism, it is the base rate. Half the capacity Western Digital ships into capacity-optimised data centres is shingled, every 2.5-inch consumer drive above 500 GB in the retail channel is shingled, and the detection procedure above returns positives only. There is no reading of the evidence in which “not stated” should be resolved in the drive’s favour.
What to do with this on a listing page
- Start from drives stated as CMR if the drive is going into an array. Not “drives that are probably CMR” - drives where a source states it. The list is shorter than the full catalogue, and shortness is the feature. For anything that will be a RAID or ZFS member, close the unfiltered listing and work from the stated-CMR one.
- If the workload is write-once-read-many, drop the filter and sort on price per terabyte instead. Archive, media libraries, sequential backup targets and surveillance are the cases where shingling costs you nothing and saves you money. Browse all hard drives and use power-on hours as the screen that actually matters there.
- Search listing text before you trust silence. Two quick passes over
the hard drive listings, one for SMR and one for Ultrastar:
the first finds sellers who disclosed, the second finds the one product line
whose part number decodes. On an Ultrastar, read the second character -
WSHis shingled,WUHis conventional. - Treat a 2.5-inch spinning drive above 500 GB as SMR unless a datasheet says otherwise. On 3.5-inch drives the question is open; at 2.5 inches in the consumer channel it effectively is not. If the target is a laptop, price an SSD against it before deciding - see HDD or SSD.
- Do not read HAMR, ePMR or MAMR in a title as evidence of CMR. They are writing technologies and they combine with shingling routinely. On the densest parts in any generation, an energy-assist acronym should raise your estimate that the drive is shingled, not lower it.
- Check SAS listings separately for host-managed parts. A
host-managed drive at a striking price per terabyte is cheap because almost
nobody can use it. If
lsscsi -gwould showzbc, you need XFS on kernel 6.15 or later, btrfs zoned mode, f2fs, zonefs or dm-zoned - and not ZFS, not ext4, and not btrfs parity RAID. Enterprise drive pulls covers the rest of what those drives need. - Ask the seller in writing, through eBay messages, and keep the answer. A written “this is CMR” that turns out to be false is a not-as-described claim rather than an argument. A verbal one is nothing. This is the only step that converts an unknowable specification into a recoverable loss.
- Run the four-hour fio probe on arrival, inside the returns window. Full
device span, 4 KiB blocks,
--time_based, bandwidth and latency logs. It is the only test that can produce a negative, it takes an afternoon, and the 30-day clock is the reason to do it now rather than when the array needs rebuilding.