SSD endurance: TBW, DWPD, and how worried to actually be
By Harry Saarinen ·
Less than the internet thinks, and the reason is arithmetic rather than optimism. For an ordinary desktop, the endurance rating is not what ends the drive’s life - the capacity is. You will replace a 1 TB SSD because 1 TB stopped being enough, decades before you write the 600 terabytes its warranty covers.
That is the honest answer for most buyers and emphatically not the answer for all of them, and the distance between the two cases is a factor of several hundred. It is also not the whole question, because the rating on the box is not a wear counter. It is a four-way conjunction of guarantees, verified in most cases by extrapolation rather than by writing a drive to death, and reported back through a firmware field the specification itself calls a vendor-specific estimate.
What follows is the arithmetic for working out which case you are in, the physics underneath the arithmetic, and the exact log pages and attribute identifiers that let you read the wear off a second-hand drive before you pay for it. It is long because the subject has four layers - the oxide, the controller, the standards, and the counters - and the advice that goes wrong is almost always advice that covers one layer and assumes the others.
One housekeeping note, because it matters for how you read the rest. The command output shown below is the shape of the output, carrying one consistent set of figures through the article, and is not a capture from one particular drive. Where a number is a vendor’s, the vendor and the document are named in the text. Where a number is a derivation of mine, it says so.
TBW and DWPD are the same number twice
Two ratings, one fact, and every drive is sold with whichever one flatters it.
TBW - terabytes written - is the total volume of host writes the warranty covers. DWPD - drive writes per day - is how many times you could overwrite the whole drive daily for the length of that warranty and still land inside the total. Consumer drives are sold on TBW because the number is large; enterprise drives on DWPD because it describes the duty cycle the buyer is choosing between.
The conversion is division:
DWPD = TBW ÷ (capacity_TB × 365 × warranty_years)
TBW = DWPD × capacity_TB × 365 × warranty_years
Worked on a typical 1 TB consumer TLC NVMe drive, rated 600 TBW with a five-year warranty:
DWPD = 600 ÷ (1 × 365 × 5)
= 600 ÷ 1825
= 0.33
A third of the drive, every day, for five years.
The conversion is not a convention this site invented, and it can be confirmed on a shipping datasheet where a vendor publishes both halves. Solidigm’s D7-P5810 product brief gives an 800 GB SLC drive at 50 DWPD over five years and 73 PBW:
50 DWPD × 0.8 TB × 1825 days = 73,000 TB = 73 PBW exactly
When a vendor publishes both numbers they agree to the digit, which means the one you are missing is always recoverable. Kioxia’s datacentre product briefs are the case where you need that: the CD8 series brief publishes DWPD and a five-year warranty and gives no TBW figure at all, no UBER and no retention specification, alongside an MTTF of 2,500,000 hours carrying the vendor’s own caveat that it “is not a guarantee or estimate of product life”.
The consumer rule scales linearly and the enterprise rule does not
Run the DWPD sum on the 2 TB model in the same consumer family, rated 1200 TBW, and you get 1200 ÷ 3650 = 0.33 again. That is not a coincidence: consumer vendors scale TBW linearly with capacity, so DWPD is a property of the product line rather than of the drive you picked out of it.
That rule does not carry across to enterprise families, and assuming it does is how people mis-price a capacity they have not looked up. Within one family the DWPD rises with capacity, because the bigger part has more parallel dies to spread the same duty cycle across and a controller that is no longer the bottleneck. Solidigm’s D5-P5336 product brief, one family, one 192-layer QLC die:
| Capacity | DWPD over five years | Petabytes written |
|---|---|---|
| 7.68 TB | 0.42 | 5.9 |
| 15.36 TB | 0.53 | 14.7 |
| 30.72 TB | 0.56 | 32.1 |
| 61.44 TB | 0.58 | 65.2 |
| 122.88 TB | 0.60 | 134.3 |
Forty-three per cent more writes per day of rated duty cycle, from the smallest part to the largest, inside one product family. Back-converting the PBW column reproduces the DWPD column to the second decimal place in three rows of five and misses by 0.01 in the other two, which is rounding in a three-significant-figure PBW column rather than an inconsistency worth chasing.
Now the ratings enterprise drives actually carry, converted both ways:
| Drive | Rating | Arithmetic | The other number |
|---|---|---|---|
| 3.84 TB read-intensive, 5 yr | 1 DWPD | 1 × 3.84 × 1825 | 7,008 TBW (7.0 PBW) |
| 3.2 TB mixed-use, 5 yr | 3 DWPD | 3 × 3.2 × 1825 | 17,520 TBW (17.5 PBW) |
| 800 GB SLC, 5 yr | 50 DWPD | 50 × 0.8 × 1825 | 73,000 TBW (73 PBW) |
| 1 TB consumer TLC, 5 yr | 600 TBW | 600 ÷ 1825 | 0.33 DWPD |
| 2 TB consumer QLC, 5 yr | 400 TBW | 400 ÷ 3650 | 0.11 DWPD |
The spread between the cheapest consumer QLC part and a genuinely write-intensive one is a factor of four hundred and fifty in DWPD. That spread is the whole subject.
One correction to an older version of this article, and to a lot of advice still in circulation. The tier once described as “write-intensive at 10 DWPD, on the 400 / 800 / 1.6 / 3.2 TB ladder” has largely moved. The ladder still means what it meant - the arithmetic is below - but in 2026 it is commonly a 3 DWPD mixed-use part. Kioxia’s CD8-V ships 800 / 1,600 / 3,200 / 6,400 / 12,800 GB at 3 DWPD on the same 112-layer BiCS TLC as the CD8-R at 960 / 1,920 / 3,840 / 7,680 / 15,360 GB and 1 DWPD. Same flash, same warranty, different amount hidden. The genuinely write-intensive tier today is SLC or XL-Flash, not high-DWPD TLC, and it exists at capacities like 800 GB rather than at 3.2 TB.
Three things the conversion quietly assumes. Capacity is the marketed decimal figure, so a “1.92 TB” drive is 1.92 in the sum, not 1.75 TiB - see drive capacity explained for why those are different numbers. Both ratings count host writes, not what the controller actually put on the NAND, which is larger and the subject of the write amplification section below. And both are measured against a defined workload, which is the next section.
What the rating is a guarantee of, and what it is not
This is the part almost every article gets wrong, including the earlier version of this one, and it comes straight out of the standard.
TBW is not a wear counter. It is the write volume at which four separate guarantees are still simultaneously true. Seagate’s technical paper TP618, which reproduces JESD218 Table 1 with JEDEC’s permission, states them: at the rated write volume the drive must still present its advertised user capacity, still meet the uncorrectable bit error rate for its class, still be inside the functional failure requirement, and still retain data with the power off for the required time at the required temperature. Fail any one and the rating is not met, however healthy the flash looks.
That has a consequence people find counter-intuitive. A drive can be past its TBW and still be writing happily, because the guarantee that broke first was probably retention, and retention is invisible until you unplug the drive and come back later.
It also has a contractual consequence, and at least one vendor puts it in writing. Crucial’s warranty terms make the firmware counter the arbiter. They say a claim must be made “within three (3) years or within five (5) years, depending on applicable warranty period for the SSD product purchased, from the original date of purchase or before writing the maximum total bytes written (TBW) as published in the product datasheet and as measured in the product’s SMART data, whichever comes first”. The term itself varies by product; the flat five years usually quoted for it is not what the document says. The second clause is the one that matters here. The SMART field is not advisory. On that drive it is the legal instrument, which is a good reason to know exactly what it counts.
The warranty expires before the flash does, in all but the heaviest rows
Take the 600 TBW consumer drive and divide.
| Daily host writes | What that is | Days to 600 TBW | Years |
|---|---|---|---|
| 10 GB | Browsing, mail, office documents | 60,000 | 164 |
| 30 GB | Ordinary desktop use | 20,000 | 55 |
| 50 GB | Frequent game installs, large downloads | 12,000 | 33 |
| 100 GB | Developer workstation, VMs, big builds | 6,000 | 16 |
| 500 GB | Video scratch disk, sustained | 1,200 | 3.3 |
| 2 TB | Database, capture, cache tier | 300 | 0.8 |
The five-year warranty expires first in every row above the fifth. In the top four rows the drive is obsolete, dead of something unrelated, or thrown out for being too small long before the flash is tired, and the correct amount of time to spend comparing TBW figures for that machine is none.
Do not trust the row you think you are in. Measure it: read the host-write counter, wait a week, read it again, divide. People consistently guess high, because a 60 GB game install is memorable and three hundred days of 4 GB background writes are not.
Two qualifications, in opposite directions. TBW is a warranty limit, not a failure point - the destruction runs described later took consumer drives to between three and nine times their rating, which is a handful of samples from 2015 rather than a promise. And crossing it voids the warranty whether or not the drive works. Western Digital’s endurance white paper is blunt about the other side: “once the NAND inside the SSD has exhausted its rated endurance there is no longer any expectation that the data will be retained for a specific length of time”. Retention past the rating is not shortened. It is unspecified, which is a different and worse thing.
The wear unit is a physics event, not a counter
A flash cell is a transistor whose threshold voltage you can move by putting charge somewhere it cannot easily leave: a floating gate in the older technology, an insulating nitride trap layer in the newer one. Either way the charge has to get in and out through a thin insulating oxide, and the only way through an insulator is to make the field across it high enough that electrons tunnel.
That tunnelling is the wear. Fowler-Nordheim tunnelling drives a current through a layer that is not supposed to conduct, and a high-field current through an oxide generates defects in it - traps, where charge can be caught. Every program and every erase does it again. Nothing else in normal operation does it at all: Western Digital’s endurance paper states that reads “do not meaningfully affect the isolation elements”, because a read only senses the threshold voltage and neither adds nor removes charge from the storage node.
So the unit of wear is the program/erase cycle, and it is a count of times you have driven charge through a specific piece of oxide. Not hours. Not power cycles. Not bytes read.
The two ways charge escapes, and why retention is the same problem
Traps do not simply accumulate until something snaps. They have two distinct effects on the stored charge, and both are documented in the Cai et al. NAND reliability survey.
Trap-assisted tunnelling. Charge trapped in the oxide forms an electrical stepping stone through it, which exacerbates stress-induced leakage current. The electrons on the storage node now have a much easier path out than they did on a fresh cell, so the threshold voltage drifts down faster with time.
Charge detrapping. Charge previously caught in the oxide escapes spontaneously. Because the escaping carrier can be an electron or a hole, this one can move the threshold voltage in either direction, which is why worn flash does not simply read low.
At high cycle counts the trapped charge is dense enough to form percolation paths through the gate dielectric, which is the point at which the cell stops being a capacitor with a leak and starts being a resistor.
Endurance and retention are therefore not two specifications. They are one specification quoted at two points on the same curve. As the cycle count rises, the time the cell will hold its state falls. That is why you cannot state a P/E rating without stating the retention it was measured against, and why the same die is legitimately rated differently for a client drive that must hold data a year unpowered and an enterprise drive that need only manage three months.
Planar scaling destroyed endurance, and the numbers are measured
The reason this got so much worse through the 2000s is that the industry made the cells smaller, and a smaller cell holds fewer electrons.
A late planar 1x-nm MLC cell - the 15 to 19 nm node - holds roughly 100 electrons. That figure is from the Cai survey, and it is the whole story in one number: gaining or losing a handful of electrons out of a hundred moves the threshold voltage far enough to change what the cell reads back as. There is no clever encoding around a signal that small.
The measured endurance by node, from the same source:
| Technology | Measured endurance, P/E cycles per block |
|---|---|
| Older SLC | about 150,000 |
| 5x-nm planar MLC (50-59 nm) | about 10,000 |
| 1x-nm planar MLC (15-19 nm) | about 3,000 |
| 1x-nm planar TLC | about 1,000 |
Two process nodes cost planar MLC seventy per cent of its endurance, and the move from two bits per cell to three at the same node cost another two thirds. Those are the numbers behind the folklore that flash got worse every year, and the folklore was correct for about a decade.
3D NAND fixed it for two reasons, and the charge trap is the larger one
Stacking cells vertically reversed the trend, and the reason usually given for it is only the second-largest reason. An earlier version of this article said 3D NAND beat planar “because stacking let the cells grow physically larger again”. That is true and it is not the main effect.
The primary cause is the change of storage mechanism. Planar NAND stores charge on a conductive floating gate. 3D NAND stores it in an insulating charge trap layer, and the Cai survey states the consequence directly: the tunnelling oxide in charge-trap transistors “is less susceptible to breakdown than the oxide in floating-gate transistors during high-voltage operation”, and as a result “the endurance (i.e., the maximum P/E cycle count) for a 3D flash memory cell has increased by more than an order of magnitude”. An order of magnitude, from the material change alone.
The secondary cause is the feature size, and it fixes a different failure mode. 3D NAND uses feature sizes in the 30 to 50 nm range against 15 to 19 nm for late planar. More electrons per cell and more distance between neighbours, which is why the survey reports that read disturb and cell-to-cell program interference “are currently not major issues in 3D NAND flash memory reliability”. Those are disturb mechanisms, not wear mechanisms. They were what made late planar TLC so hard to program correctly, not what made it wear out.
The two effects compound in a way that shows up in the programming algorithm. Planar MLC needed two-step programming and planar TLC needed foggy-fine programming, both of which exist to place charge without disturbing the neighbours that have not been programmed yet. With interference that much lower, 3D NAND reverted to the older one-shot algorithm. A consequence worth knowing: under one-shot programming no cells are left partially programmed, which the survey notes significantly reduces the number of program errors.
Where the layer counts actually are in 2026
Layer count is the density lever now, and the capacity you can buy follows it directly. These are reported roadmaps rather than datasheet specifications, and roadmaps move - TrendForce’s June 2026 survey is the source, and one of its own data points is a target being revised downward:
| Vendor | In production | Announced next step |
|---|---|---|
| SK hynix | 321 layers, 9th generation, triple string stack | 375 layers by end-2026, revised down from 400 |
| Samsung | 286 layers, in production since April 2024 | 400-layer V10, cell-on-periphery, wafer-to-wafer bonding, 1 Tbit die, targeted H2 2026 |
| Kioxia and SanDisk | BiCS10 at 332 layers, sampling | - |
| Micron | 276 layers | - |
Kioxia’s BiCS10 is the one with published construction detail: a triple stack of strings of over 100 layers each, CMOS bonded to array, a 1 Tbit TLC die, a 4.8 Gbit/s interface, and claimed improvements against the 218-layer generation of 59 per cent bit density and 33 per cent NAND interface speed, with data-input power down 10 per cent and data-output power down 34 per cent. Those are the four figures the joint Kioxia and Sandisk ISSCC 2025 announcement publishes. Write and read power for the part are not among them, and figures circulating for those belong to a different, 9th-generation die.
String stacking is why the count keeps rising. Etching a single hole through 300 layers of alternating material and keeping it cylindrical is the hard limit; etching three holes through 100 layers each and aligning them is merely expensive. The alignment is what costs yield.
There is a ceiling in sight, and it was predicted before it arrived. Cai et al. projected in 2017 that 3D stacking would top out “in the range of 300-512 layers”, after which manufacturers would have to shrink the cells again and reintroduce the disturb mechanisms that vertical geometry had suppressed. Production is now past 300 layers. The drives on sale today are the experiment that tests that prediction, and nobody outside the fabs can tell you the result yet.
Bits per cell is a read-window budget
A flash cell stores charge, and reading it means deciding which of several voltage bands that charge falls into. One bit needs two distinguishable states. Each extra bit doubles the number of bands that must fit inside the same physical voltage window, and the window does not grow.
| Bits/cell | Voltage levels | Thresholds separating them | Minimum senses to resolve a cell | |
|---|---|---|---|---|
| SLC | 1 | 2 | 1 | 1 |
| MLC | 2 | 4 | 3 | 2 |
| TLC | 3 | 8 | 7 | 3 |
| QLC | 4 | 16 | 15 | 4 |
| PLC | 5 | 32 | 31 | 5 |
The last column is a binary search over the bands and is a floor, not a measurement. The real read cost is higher and depends on the encoding: with Gray coding a single page of a TLC block needs one, two or four sense operations depending on which bit position it holds, summing to the seven thresholds across the whole cell. That particular one-two-four split is a planar result and should not be carried forward as a constant - 3D NAND reverted to one-shot programming, the page-to-sense mapping is an encoding choice, and no vendor publishes it per product.
What does carry forward is the direction. More bands in the same window means narrower margins, more raw bit errors, stronger error correction needed, and more read-retry operations as the data ages. Read-retry - rereading the page with shifted reference voltages until the correction engine succeeds - has been standard in NAND since roughly the 30 nm planar node, and it is the mechanism behind the specific symptom of old data on a worn drive reading back slowly. Every retry is another full array sense.
Pseudo-SLC is the same die with a bigger budget
Running a TLC or QLC die in one-bit mode - pseudo-SLC - gets far more cycles out of it than its native rating, because it only has to keep two bands apart instead of eight or sixteen. The margins are enormous by comparison, so the same oxide damage that would push a TLC cell out of its band leaves a pSLC cell comfortably inside one.
Which is what every modern consumer drive does, constantly, and the SLC caching section later in this article is about what that costs.
What a current QLC die is really rated at
An older version of this article gave QLC as 100 to 1,000 P/E cycles. The bottom of that range belongs to early planar QLC and to first-generation marketing scepticism, and it is wrong for the 3D QLC in a drive you can buy today.
Here is the derivation, and the arithmetic is mine rather than a vendor’s, because no vendor publishes the cycle rating of the die in a finished drive. Solidigm’s D5-P5336 at 122.88 TB is rated 134.3 PBW:
134,300 TB written ÷ 122.88 TB capacity = 1,093 full drive writes
The datasheet’s endurance footnote says those writes are measured IU-aligned, at a write size matched to the drive’s indirection unit, which is the condition under which write amplification is closest to one. So the die is supporting at least about 1,093 cycles. If the vendor set the rating using the JEDEC guidance formula, which carries a factor-of-two guard band for wear-levelling imperfection, the underlying die rating is about twice that:
JEDEC guidance: TBW < (capacity × P/E cycles) ÷ (2 × WAF)
rearranged: P/E = (TBW × 2 × WAF) ÷ capacity
= (134,300 × 2 × 1.0) ÷ 122.88
= 2,186 cycles
So current datacentre 3D QLC sits somewhere between about 1,100 and about 2,200 P/E cycles, depending on how much guard band the vendor applied, which is one to two orders of magnitude above the figure still being quoted for it.
Run the identical sum on consumer QLC - the 2 TB, 400 TBW part from the conversion table above - and the answer is much lower:
lower bound 400 TB ÷ 2 TB = 200 full drive writes
upper bound (400 × 2 × 1.0) ÷ 2 = 400 cycles
Between about 200 and about 400 P/E cycles on the same assumptions, against 1,100 to 2,200 for the datacentre part - a factor of five and a half at both ends of the range, on what is nominally the same four-bits-per-cell technology. That says less about QLC as a technology than about which bins go into which product.
PLC has been demonstrated in research and announced in roadmaps. No PLC drive ships in volume on the general market, no vendor publishes an endurance figure for one, and this site will not estimate it.
JESD218 is four requirements at once, and the current revision is not free
If you want to check any of this yourself, two documents matter, and a guide that sends you to read them should tell you which revision and what it costs.
JESD218C.01, published August 2026, is the current revision of “Solid-State Drive (SSD) Requirements and Endurance Test Method”. It is the document that defines what a TBW rating means. It is no longer a free download: JEDEC charges non-members $90 for it. JESD219A.01, published June 2022, is “Solid-State Drive (SSD) Endurance Workloads”, the companion that defines what you write during the test, and it remains free with registration. It ships two supporting trace files, JESD219A_MT and JESD219A_TT.
Almost everything written about SSD endurance cites the JESD218B family, which is one revision generation behind. So does the interface specification: NVM Express Base Specification Revision 2.4, ratified 31 July 2026, points the reader at “the JEDEC JESD218B-02 standard for SSD device life and endurance measurement techniques” in its definition of Percentage Used. The interface spec is one JEDEC revision behind the JEDEC site, which is worth knowing before you treat a cross-reference as authoritative.
The table, and the duty cycles that make it comparable
JESD218 Table 1, reproduced in Seagate’s TP618 with JEDEC’s permission, sets four requirements per class:
| Client | Enterprise | |
|---|---|---|
| Workload | JESD219 client workload | JESD219 enterprise workload |
| Active use | 40 °C at 8 hours per day | 55 °C at 24 hours per day |
| Powered-off retention | 1 year at 30 °C | 3 months at 40 °C |
| Functional failure requirement | no more than 3 per cent | no more than 3 per cent |
| UBER | no more than 1 per 10^15 bits | no more than 1 per 10^16 bits |
The duty cycle rows are the ones usually dropped, and dropping them makes the temperatures look like an arbitrary 15 °C difference. They are not. Client is eight hours a day at 40 °C; enterprise is twenty-four hours a day at 55 °C. Integrate that over five years and the enterprise part is being asked to survive a thermal exposure in a different league, which is a large part of what the higher price buys.
Note also that the functional failure requirement is identical across the two classes. The class difference is retention, temperature, duty cycle and error rate - not the permitted failure rate.
Most published TBW figures were never measured
JESD218 permits two verification routes. Direct verification writes the entire TBW with the class workload and confirms the four requirements at the end. If that cannot be completed within 1000 hours, extrapolation methods are permitted.
Do the sum on what 1000 hours buys you. A drive taking a sustained 1 GB/s for the whole period writes 3.6 PB, which covers a consumer drive comfortably and does not come close to a 134 PBW QLC part. So the large endurance ratings on the market are extrapolated from accelerated stress, not measured to destruction, and that is a statement about the method rather than a criticism of it.
The stress conditions are specified. High-temperature retention stress is 96 hours at 66 °C or above, or 500 hours at 52 °C or above. Endurance stress uses a ramped or split-flow approach: a low leg at 25 °C or below and a high leg from 40 °C to the drive’s maximum for client parts, from 60 °C to maximum for enterprise.
And the guidance formula for setting the rating in the first place carries an explicit safety factor:
TBW < (capacity × NAND cycling capability) ÷ (2 × WAF)
The factor of two is a guard band for wear-levelling imperfection - the acknowledgement, written into the guidance, that not every block will have been cycled the same number of times when the first one runs out.
That formula is also a useful consistency check on a rating you are looking at. Take the 1 TB consumer TLC drive at 600 TBW and assume a 1,500-cycle die:
TBW = (1 TB × 1500) ÷ (2 × WAF)
600 = 1500 ÷ (2 × WAF)
WAF = 1500 ÷ 1200
= 1.25
A 600 TBW rating on a 1 TB TLC drive is exactly what you get from a 1,500-cycle die, the JEDEC guard band, and a write amplification of 1.25 under the client workload. That arithmetic is mine and the die rating is an assumption, not a disclosure. It is a consistency check, not a reverse engineering, and it is offered because it shows the rating is not arbitrary.
The two workloads are opposite in kind
An earlier version of this article described the JESD219 client workload as “a heavily random small-block trace”. That is the enterprise workload, and the two are backwards in a lot of writing on this subject. Western Digital’s endurance white paper describes both:
| Client workload | Enterprise workload | |
|---|---|---|
| Kind | A replayed reference trace | A synthetic mix |
| Origin | Real I/O recorded over an extended period on a PC with an operating system installed | Random writes of varying block lengths |
| Distribution | Whatever the machine actually did | Heavily biased to 4 KB and 8 KB |
The trace files contain only ATA I/O commands and carry no data payload, for a reason that is itself informative: the payload has to change between passes, or a controller that deduplicates or compresses would get an unrepresentatively easy ride.
This matters for reading a rating. A client TBW figure was measured against a realistic desktop trace with a lot of sequential traffic in it. If your workload is small random writes - a database, a metadata vdev, a busy mail spool - you are running the enterprise workload on a drive rated against the client one, and the budget burns faster than the counter suggests.
UBER, and what one error per 10^15 bits costs you
The uncorrectable bit error rate is the other half of the guarantee and it is quoted in a unit nobody has intuition for. Convert it:
1 error per 10^15 bits = 125 TB read client requirement
1 error per 10^16 bits = 1.25 PB read enterprise requirement
1 error per 10^17 bits = 12.5 PB read Solidigm's stated internal target
Solidigm’s D7-P5810 brief says so explicitly and names the standard it is beating: its test target is “1E-17 under a full range of conditions and cycle counts throughout the life of the drive, which is 10x higher than 1E-16 specified in the JEDEC Solid State Drive Requirements and Endurance Test Method (JESD218)”. Enterprise vendors design past the spec and say so in the datasheet, which is a level of disclosure the consumer market does not offer at all.
What makes those numbers achievable is the error correction, and the scale of it is worth stating. Modern engines correct data at a raw bit error rate between 1e-3 and 1e-2 and hand the host a post-correction rate of 1e-15. That is twelve to thirteen orders of magnitude of correction, using BCH or LDPC codes, with a CRC layered on top because a correction engine can silently mis-declare success on a codeword it has actually mangled.
Correction strength is not free, and it is paid for in capacity. The Cai survey works a 2.0 TB advertised, 2.4 TB physical drive through four coding configurations:
| ECC configuration | Over-provisioning left |
|---|---|
| Coding rate 0.93, no superpage parity | 11.6 per cent |
| Coding rate 0.93 with superpage parity | 8.1 per cent |
| Coding rate 0.90 | 8.0 per cent |
| Coding rate 0.90 with superpage parity | 4.6 per cent |
Stronger correction raises the P/E endurance the die can deliver and simultaneously raises write amplification, because it eats the over-provisioning that keeps write amplification down. There is no configuration that wins both, which is why two drives on the same die can be tuned into genuinely different products.
Write amplification, and why it is never one
The controller cannot overwrite a page in place. Flash is written in pages - 16 KB is typical - and erased only in whole blocks of hundreds of pages, so reusing a block holding a mix of stale and valid data means copying the valid pages elsewhere first. Those copies are NAND writes nobody asked for.
Write amplification factor is NAND bytes written divided by host bytes written.
The identity, and why it is not the steady state
If garbage collection picks a victim block in which a fraction u of pages are
still valid, erasing one block of P pages yields P(1−u) free pages at a cost
of uP copies. The earlier version of this article gave that as 1 + u ÷ (1 − u),
which is algebraically just:
WAF = 1 ÷ (1 − u)
That is the simpler form and it is worth memorising in that shape. But it is
a single-block accounting identity, not a steady-state result, and the earlier
version stopped there and never connected u to anything you can act on. Here is
the connection.
Under perfect wear levelling and uniformly random writes, with P physical bytes
and V bytes of live host data, every block is on average as full as the drive
is. So:
over-provisioning OP = (P − V) ÷ V
average valid fraction u = V ÷ P = 1 ÷ (1 + OP)
non-selective WAF = 1 ÷ (1 − u) = (1 + OP) ÷ OP
That last expression is the one to carry, and the derivation is mine. It is an upper bound, because it assumes garbage collection picks a block at random. Real greedy collection picks the emptiest block it can find and does better. The published steady-state curves for greedy collection under uniformly random writes all have the same shape: a steep fall through the first 10 to 20 per cent of over-provisioning, strongly diminishing returns beyond roughly 30 per cent, and convergence to 1 only asymptotically rather than at any finite spare fraction.
The shape is what to take from them. This article will not put a specific published figure on the bottom of that range, because the numbers in circulation for it are quoted far more often than they are sourced to a particular model, workload and spare factor - and all three change the answer. My own bound gives 11.1 at 9.95 per cent over-provisioning, which is the right order of magnitude for that end of the curve and is not a validation of anything.
The bound, tabulated:
| Over-provisioning | Non-selective WAF bound | Saved by the previous ten points |
|---|---|---|
| 10 per cent | 11.0 | - |
| 20 per cent | 6.0 | 5.0 |
| 30 per cent | 4.33 | 1.67 |
| 40 per cent | 3.50 | 0.83 |
| 50 per cent | 3.00 | 0.50 |
| 60 per cent | 2.67 | 0.33 |
| 100 per cent | 2.00 | - |
The third column is the whole design argument. The first ten points of over-provisioning are worth more than the next fifty combined, which is why every drive has some and why the expensive ones stop at about 37 per cent rather than going further.
Put it all together with the lifetime model the manufacturers use, which the Cai survey states in full:
Lifetime (years) = PEC × (1 + OP) ÷ (365 × DWPD × WA × R_compress)
where OP = (physical block count − logical block count) ÷ logical block count
Every lever in any endurance argument is one of those five terms. Worked on the 1 TB consumer TLC drive, with a 1,500-cycle die assumed, 9.95 per cent over-provisioning, its rated 0.33 DWPD, a measured write amplification of 2.2 and no compression:
= 1500 × 1.0995 ÷ (365 × 0.33 × 2.2 × 1)
= 1,649.3 ÷ 265.0
= 6.2 years
Which lands just past the five-year warranty, as it should. Change one term and the answer moves proportionally: double the write amplification and it is 3.1 years; run the drive at 1 DWPD instead of 0.33 and it is 2.1.
Over-provisioning, with the arithmetic corrected
Every SSD has some over-provisioning for free, from the gap between binary NAND and decimal marketing. The earlier version of this article got the consumer row wrong, in a way worth spelling out because the error is common.
512 GiB = 549,755,813,888 bytes
1024 GiB = 1,099,511,627,776 bytes
2048 GiB = 2,199,023,255,552 bytes
4096 GiB = 4,398,046,511,104 bytes
512 GiB sold as 512 GB 549.755813888 ÷ 512 = 1.0737 7.37 per cent
1024 GiB sold as 1 TB 1099.511627776 ÷ 1000 = 1.0995 9.95 per cent
2048 GiB sold as 2 TB 2199.023255552 ÷ 2000 = 1.0995 9.95 per cent
1024 GiB sold as 960 GB 1099.511627776 ÷ 960 = 1.1453 14.53 per cent
1024 GiB sold as 800 GB 1099.511627776 ÷ 800 = 1.3744 37.44 per cent
The familiar 7.37 per cent figure is correct only at the gigabyte boundary. Over the terabyte range the divergence is 9.95 per cent, because 1024 GiB is 1,099.51 GB sold as 1,000 GB. Consumer 1 TB and 2 TB drives get about a third more free over-provisioning than the old table credited them with, and the corrected figure changes the WAF bound from 14.6 to 11.1.
| Marketed capacity | Raw NAND | Over-provisioning | Positioning |
|---|---|---|---|
| 512 GB | 512 GiB | 7.37 per cent | Consumer |
| 1 TB / 2 TB / 4 TB | 1024 / 2048 / 4096 GiB | 9.95 per cent | Consumer |
| 480 / 960 / 1,920 / 3,840 GB | 512 / 1024 / 2048 / 4096 GiB | 14.53 per cent | Enterprise read-intensive |
| 400 / 800 / 1,600 / 3,200 GB | 512 / 1024 / 2048 / 4096 GiB | 37.44 per cent | Enterprise mixed-use or write-intensive |
The enterprise rows are exactly self-consistent at every capacity - 14.53 per cent in every row of the third line, 37.44 per cent in every row of the fourth - which is what told this site the consumer row was the one with the error in it.
That table is the most useful buying heuristic here: on the used enterprise market the capacity number tells you the over-provisioning before you find a datasheet. A 1.6 TB SAS SSD and a 1.92 TB one are the same silicon with different amounts hidden. The same holds at 400 against 480, 800 against 960, 3.2 against 3.84.
One honest qualification, which the earlier version did not make. The capacity ladder pins the over-provisioning exactly. It does not pin the DWPD, because the tier names have drifted: the 37.44 per cent ladder was a 10 DWPD write-intensive tier a generation ago and is commonly a 3 DWPD mixed-use tier now. Use the capacity to identify the tier, then look up the family for the rating. Enterprise drives works the same ladder from the other direction.
Your own free space is over-provisioning, if TRIM reaches the drive
This is the most actionable arithmetic in the article and it falls straight out of the formula above. Over-provisioning is not a factory setting. It is the ratio of physical bytes to live bytes, and you control the denominator.
A 1 TB consumer drive has 1,099.51 GB of physical NAND. Keep some of it empty, with a filesystem that issues discards and a path that carries them, and the controller sees the same picture it would see on an enterprise part:
| Live host data | Effective over-provisioning | Non-selective WAF bound |
|---|---|---|
| 1,000 GB, completely full | 9.95 per cent | 11.1 |
| 900 GB | 22.2 per cent | 5.5 |
| 800 GB | 37.4 per cent | 3.7 |
| 700 GB | 57.1 per cent | 2.8 |
| 500 GB | 119.9 per cent | 1.8 |
Keeping a consumer 1 TB drive twenty per cent empty puts it at exactly the 37.4 per cent over-provisioning of an enterprise write-intensive part, and cuts the write amplification bound by two thirds against the same drive run full. The arithmetic is mine; the effect is not controversial; the size of it on your particular drive depends on a garbage collection policy no datasheet discloses.
You can make the arrangement permanent rather than relying on discipline: partition to less than full capacity, or on NVMe create a namespace smaller than the device. The controller has to know the space is stale, so a factory-fresh or securely-erased drive is fine and a shrunk partition needs the freed range discarded.
TRIM, and the paths that swallow it
TRIM tells the drive which logical blocks the filesystem has stopped caring
about. Without it the controller cannot know a deleted file is deleted, believes
every LBA ever written is still live, and watches u climb toward one -
converging on the “completely full” row above no matter how empty the filesystem
looks. On Linux sudo fstrim -v / runs it once and fstrim.timer weekly, which
is the sane default.
Confirm the path actually carries it, because plenty do not:
$ sudo hdparm -I /dev/sda | grep -iE 'Model|Firmware|TRIM|ZEROs'
Model Number: ...
Firmware Revision: ...
* Data Set Management TRIM supported (limit 8 blocks)
* Deterministic read ZEROs after TRIM
$ lsblk -D
NAME DISC-ALN DISC-GRAN DISC-MAX DISC-ZERO
sda 0 512B 2G 0
DISC-GRAN and DISC-MAX of zero mean discards are not reaching the drive. USB
enclosures are the usual culprit - passthrough needs UASP and a bridge chip that
implements it - and several RAID layers swallow it too. “Deterministic read ZEROs
after TRIM” is the stronger of the two deterministic modes and is what lets a
filesystem treat a trimmed range as reliably zero rather than merely unspecified.
The indirection unit sets a floor the datasheet assumes you will respect
This is the modern failure mode and it is almost absent from consumer writing.
A controller does not map single sectors. It maps indirection units, and on large QLC drives the IU is much bigger than a 4 KiB sector - 16 KiB on most of the Solidigm D5-P5336 capacities and 32 KiB on the 122.88 TB one, according to its own endurance footnote. A host write smaller than the IU, or misaligned to it, forces a read-modify-write of the whole unit.
The arithmetic is unforgiving and it is mine:
16 KiB indirection unit, 4 KiB aligned host write
drive must read 16 KiB, merge 4 KiB, write 16 KiB
amplification floor 16 ÷ 4 = 4 before garbage collection
32 KiB indirection unit, 4 KiB aligned host write
amplification floor 32 ÷ 4 = 8 before garbage collection
Which is why the Solidigm endurance figures are footnoted as “IU-Aligned Endurance. Based on 100% Random Write 16KB for 16KB IU SKUs, and 100% Random Write 32KB for 32KB IU SKU”. That rating was measured with writes exactly matched to the drive’s indirection unit. Smaller or misaligned writes are not covered by the number on the datasheet, and nothing on the datasheet says by how much they are not covered.
You can find out what a drive wants. The OCP Datacenter NVMe SSD Specification
requires the device to report its indirection unit in one named field -
requirement NVMe-OPT-7, “the device shall report its Indirection Unit (IU) size
in the NPWG field in the Identify Namespace data structure” - so nvme id-ns
tells you the write size below which the drive must read-modify-write. Read NPWG,
the Namespace Preferred Write Granularity, and nothing else. The atomicity fields
next to it in the same structure are a different quantity and often a different
number, and the OCP specification does not name them here.
And the OCP extended SMART log counts how often you got it wrong: Unaligned I/O counts write commands whose start address is not aligned to the indirection unit, which is a direct measurement of the host doing the single most damaging thing it can do to a large-IU QLC drive.
Write amplification below one
The “NAND bytes divided by host bytes is at least one” rule that the rest of this arithmetic assumes has one clean exception: a controller that compresses. The SandForce controllers claimed a typical write amplification of 0.5 and a best case as low as 0.14 on the SF-2281. The endurance experiment described later in this article measured the effect directly on two otherwise identical drives, and it was worth roughly an extra petabyte of life on the one fed compressible data.
Nothing on the modern consumer market makes that claim, and R_compress in the
lifetime model is 1 unless you have a specific reason to think otherwise. It is
in here because it is the one case where the floor is not a floor.
Garbage collection, wear levelling and refresh are three algorithms in tension
The controller is running three background jobs against the same blocks, and they want different things.
Garbage collection wants to reclaim blocks cheaply, which means picking victims with as little valid data as possible and, ideally, keeping hot data and cold data in separate blocks so that a block of hot data goes almost entirely stale at once. The IBM Zurich write-amplification paper, which defines the over-provisioning factor as raw blocks over user blocks and the spare factor as the spare fraction of raw blocks, finds exactly that: separating static from dynamic data reduces write amplification.
Wear levelling wants the opposite. If cold data sits undisturbed in the same blocks forever, those blocks accumulate no cycles while the hot blocks take all of them, and the drive dies when the hot pool is exhausted with most of the flash barely used. Static wear levelling therefore deliberately relocates cold data into worn blocks and moves hot data into fresh ones - it mixes precisely what garbage collection wants separated, and every relocation is a write that no host asked for.
That tension has no clean resolution and it is where controller firmware earns its money. It is also why the factor-of-two guard band exists in the JEDEC TBW formula.
Refresh is the third job and the one most people do not know is running. Modern controllers do not wait for retention errors to appear. Remapping-based refresh periodically reads valid blocks, corrects them and rewrites them elsewhere before the error rate reaches the correction limit. In simulation the Cai survey reports that this raises SSD lifetime by an average of 9x, and a hybrid in-place scheme by an average of 31x. Refresh is not a maintenance overhead the drive pays for reliability. It is a large part of why the drive lasts as long as it does.
Read reclaim is the same mechanism aimed at a different stressor: repeated reads of a block slightly disturb the cells around the one being read, and “some flash vendors specify a maximum number of tolerable reads for a flash block”, after which the block is rewritten whether or not anything has been written to it.
Two consequences follow, and they are the practical ones.
The rates are adaptive, not fixed. Controllers poll temperature every few milliseconds and use the Arrhenius relation to estimate how fast retention errors are accumulating at the temperature the drive is actually running at. The same controller can refresh a block “once every year” at low P/E counts and “once every week” at high ones. A drive near end of life is doing far more background work than the same drive when new, which is one reason it feels slower.
None of it runs on a shelf. Every mechanism in this section requires power. That is the whole of the archive argument later in this article, stated as mechanism rather than as a warning.
SLC caching buys speed with a second write of every byte
Your TLC or QLC drive does not usually write in TLC or QLC. It writes a region of its NAND in pseudo-SLC mode - fast, one bit per cell - acknowledges the write, and folds that data down into full-density cells later, in the background. This is why a QLC drive can quote a sequential write figure it has no business quoting.
The cache is not free capacity: an SLC-mode block holds a third of a TLC block and a quarter of a QLC one, so it consumes raw cells at three or four times the rate of the data in it. Most consumer drives therefore size it dynamically, borrowing from whatever is free:
| 1 TB QLC drive | Free space | Raw cells available | Maximum pSLC cache |
|---|---|---|---|
| Empty | ~1,000 GB | ~1,000 GB | ~250 GB |
| Half full | ~500 GB | ~500 GB | ~125 GB |
| 90 per cent full | ~100 GB | ~100 GB | ~25 GB |
A nearly-full drive has almost no cache, which is why an SSD that felt fast for a year feels slow once you fill it. Divide by three rather than four for TLC. The figures are approximate - vendors cap the dynamic portion below the theoretical maximum and hold back headroom for the folding itself. The three arrangements in circulation are a static cache of fixed size, a dynamic one sized from free space, and a hybrid that is a small static region plus a dynamic extension; which one a given drive uses is not usually stated.
When the cache runs out mid-transfer, writes go direct to the full-density cells and throughput falls off a cliff. Take a drive that writes at 5,000 MB/s into cache and 150 MB/s direct, with 100 GB of cache available, and copy 300 GB onto it:
first 100 GB 100,000 MB ÷ 5,000 MB/s = 20 s
next 200 GB 200,000 MB ÷ 150 MB/s = 1,333 s
total 300,000 MB ÷ 1,353 s = 222 MB/s average
The box said 5,000. The transfer averaged 222, and the last two thirds ran at 150 - slower than a hard drive. Direct-to-NAND figures are almost never published, so the only place to find them is independent sustained-write testing that runs past the cache. Folding happens during idle, so a drive given a few minutes between transfers recovers its cache; back-to-back copies start on the cliff.
What the fold costs in endurance
Here is the part the performance discussion leaves out, and the arithmetic is mine.
Measured in media bytes written - the quantity the drive reports as physical media units written, and the numerator of the write amplification you can read off a log page - a byte that goes through the cache is written twice: once into the pSLC region and once into the TLC or QLC block it ends up in. SLC caching contributes a factor of 2 to write amplification for every byte that is cached and folded, on top of whatever garbage collection costs.
Measured in block erases, which is what the oxide actually experiences, it is
worse. Take a block that holds B bytes in TLC mode and B/3 in pSLC mode:
H bytes cached in pSLC, then folded to TLC
pSLC pass H bytes ÷ (B ÷ 3) = 3H ÷ B block erases
fold pass H bytes ÷ B = H ÷ B block erases
total 4H ÷ B block erases
direct TLC write, no cache H ÷ B block erases
Four times the block erases for the same host data. But those erases are not equivalent: three quarters of them happen on cells operating in one-bit mode, where the read window is enormous and the same oxide damage matters far less. No vendor publishes a pSLC-mode cycle rating for its own die, so the net effect on drive life is not calculable from the outside, and anyone who gives you a number for it has made one up. What is calculable is the media-bytes figure, and it is 2.
The practical reading: a workload of large sequential writes that mostly bypasses the cache is gentler on the drive than the same bytes dribbled in small enough to be cached, which is the opposite of what the performance numbers suggest and the same conclusion the write amplification section reaches by a different route.
DRAM, HMB, and what a DRAM-less drive is actually missing
The mapping table is the thing. A controller translating logical addresses to physical ones at 4 KiB granularity needs one entry per 4 KiB of capacity, and the rule of thumb that follows is a ratio worth deriving rather than memorising:
1 TB ÷ 4 KiB = 244,140,625 mapping entries
× 4 bytes per entry = 976,562,500 bytes
≈ 0.98 GB of DRAM per TB of capacity
Which is why the Cai survey states the industry ratio as roughly 1 GB of DRAM for every 1 TB of capacity. A 2 TB drive with a full map wants about 2 GB.
A DRAM-less drive does not have that. It has a small SRAM cache inside the controller and, on NVMe, the option of borrowing host memory. The Host Memory Buffer is reported in the Identify Controller structure as HMPRE, the preferred size, and HMMIN, the minimum useful size, both in 4 KiB units. A non-zero HMPRE is the definition of “this drive supports HMB”. Two properties of it matter:
- The drive must work without it. The NVMe specification requires that “the controller shall function properly without host memory resources”. HMB is an optimisation the drive may not depend on for correctness.
- It does not survive a controller-level reset. Whatever the drive was caching in host memory is gone, and it has to rebuild from the flash.
And then there is what the operating system will actually grant. Linux caps the
Host Memory Buffer at 128 MiB per controller by default, through the nvme
module parameter max_host_mem_size_mb. That is a hard cap regardless of how
much system RAM is free:
$ cat /sys/module/nvme/parameters/max_host_mem_size_mb
128
Against the 1 GB per TB ratio, that is stark:
| Capacity | Full map wants | Linux grants | Fraction of the map resident |
|---|---|---|---|
| 1 TB | ~0.98 GB | 128 MiB | about one seventh |
| 2 TB | ~1.95 GB | 128 MiB | about one fifteenth |
| 4 TB | ~3.91 GB | 128 MiB | about one twenty-ninth |
The map column is in decimal gigabytes, because that is what the derivation above produced. Rounding it to GiB instead is tempting, because 128 MiB is exactly one eighth of 1 GiB and the fractions come out clean - but 1 GiB is 1,073,741,824 bytes against the 976,562,500 the map actually needs, so the clean fractions understate the resident share by about ten per cent in every row. The same GB against GiB gap this article insists on elsewhere applies to the mapping table.
A DRAM-less 4 TB drive on Linux is running with about three and a half per cent of its mapping table in fast memory. The rest lives in flash and has to be read to resolve an address. On sequential work this barely matters, because consecutive addresses hit the same map pages. On scattered small random writes it matters a great deal: every miss is a flash read to find the map entry, and map updates are themselves writes that land on the NAND.
That is the endurance connection, and it is why “DRAM-less” belongs in an endurance article rather than only a performance one. A drive that has to page its own mapping table writes more metadata to flash than one that does not.
Power-loss protection changes what a synchronous write costs
Power-loss protection is usually described as a data-safety feature. It is also a performance and endurance feature, and the specification says why.
On a drive with real PLP, there is no volatile write cache to flush. The NVMe
Volatile Write Cache feature, Feature Identifier 06h, puts it plainly: “if the
controller is able to guarantee that data present in a write cache is written to
non-volatile storage media on loss of power, then that write cache is considered
non-volatile and this Feature does not apply”. So nvme id-ctrl reporting vwc
of 0 is a host-side test for power-loss protection that costs nothing to run.
The OCP Datacenter NVMe SSD Specification spells out the behavioural consequences in requirement NVMe-IO-3: because the device is power-fail safe, forced unit access “shall not incur a performance penalty”, flush cache “shall have no effect as the PLP makes any cache non-volatile”, and a Set Features command trying to disable the write cache “shall fail… as there is no volatile write cache”.
That is why a synchronous-write workload behaves like a completely different
workload on an enterprise drive. On a consumer drive, every fsync is a real
flush to NAND: small, immediate, and amplified. On a PLP drive the same fsync
is acknowledged out of a capacitor-backed buffer, which lets the controller
coalesce and align what it eventually writes. The endurance difference is not a
marketing claim; it is the difference between writing 4 KiB now and writing a
full indirection unit’s worth of coalesced data later.
The specification also requires the capacitors to be checked rather than assumed. OCP PLP-1 demands “full power-loss protection for all acknowledged data and metadata”. PLP-3 requires a health check at power-on before any writes are accepted and at least once per interval thereafter. PLP-6 is the one that matters for a used drive: the check “shall not just check for open/short capacitor conditions but shall measure the true available margin energy”. PLP-7 sets the factory default interval at 15 minutes.
And the result is readable. The OCP extended SMART log carries Capacitor Health, a percentage where “100% represents the passing hold up energy threshold when a device leaves manufacturing”, where “1% is the minimum hold up energy required to conduct a proper shutdown reliably”, and where the reserved value FFFFh means the device has no power-loss protection at all. A healthy new drive typically reads above 100. A used enterprise drive reading well below 100 has aged capacitors, which is invisible in every other field and is exactly the thing that makes it unfit for the job you bought it for.
On SATA the same information exists where the vendor chose to expose it: Intel and Solidigm datacentre SATA SSDs carry attribute 175 Power_Loss_Cap_Test, some at 201 instead, and Samsung datacentre SATA drives carry 201 Supercap_Status. Consumer drives carry nothing, because they have nothing to report.
The NVMe base specification also has a bit for the failure: Critical Warning bit 4, Volatile Memory Backup Failed, is the power-loss-protection failure flag. If it is set, the capacitors have already failed their own test.
Who genuinely does need to care
Everybody in this list, and the numbers get ugly fast. Same 600 TBW drive throughout.
ZFS SLOG. A separate log device absorbs every synchronous write to the pool - NFS with sync, iSCSI without writeback, databases opening files O_SYNC. A pool taking 200 MB/s of sync writes for eight hours a day puts 5.76 TB a day on the SLOG, and 600 ÷ 5.76 is 104 days.
There is a detail in the OpenZFS tunables that decides whether a SLOG sees
everything or only the large writes. zfs_immediate_write_sz defaults to 32 KiB,
and is usually described as the threshold above which data is written directly to
the pool rather than logged. The manual page adds the qualifier that settles it:
“in presence of SLOG this parameter is ignored, as if it was set to infinity,
storing all written data into ZIL to not depend on regular vdev latency”.
With a SLOG present, every synchronous write really does land on the log
device, regardless of size. It also needs power-loss protection, for the reason
in the section above and in enterprise drives.
L2ARC. This is where the earlier version of this article was wrong by a factor of four, and the correction matters because the conclusion flips.
Both OpenZFS 2.3 and 2.4 set l2arc_write_max to 33,554,432 bytes - 32 MiB -
per feed interval, with l2arc_feed_secs of 1. So the steady-state fill rate is
32 MiB/s, and l2arc_write_boost adds a further 32 MiB/s while the cache device
is still cold:
32 MiB/s × 86,400 s = 2,899,102,924,800 bytes = 2.9 TB per day
600 TBW ÷ 2.9 TB/day = 207 days
A read cache doing nothing but caching reads writes its way through a 600 TBW consumer drive in about seven months. The old figure of 2.3 years came from an 8 MiB/s default that is not what the current releases ship.
OpenZFS master, not yet in the 2.3 or 2.4 releases, adds l2arc_dwpd_limit
explicitly “to protect SSD endurance”. The default is 100, meaning 1.0 DWPD, and
the manual page suggests “30 = 0.3 DWPD for QLC SSDs”. It applies only after the
initial fill pass, and only when total L2ARC capacity is at least twice
arc_c_max. When that ships the answer changes again, which is a good reason
to read your own version’s manual page rather than any article, including this
one.
CCTV and NVR. Sixteen cameras at 4 Mbit/s each is 8 MB/s written continuously - 691 GB a day, 2.4 years to the rating. Surveillance kills consumer flash more reliably than anything else, because it never stops and nobody is watching the drive. See drives for a NAS for why this workload usually wants spinning rust for the bulk and flash only for metadata.
Write-heavy databases. Write-ahead logs plus checkpoints, small and random, so the NAND wears faster than the host counter implies - this is the JESD219 enterprise workload being run on a drive rated against the client one. Measure the actual instance; no rule of thumb survives contact with one.
Video capture and build servers. Both sit around the 500 GB-a-day row above, and CI adds the worse half: every job writes a working tree and an object cache and then discards them, exactly the turnover that drives write amplification up.
For all of these, buy an enterprise drive on DWPD - and the used market is unusually kind about it, for the over-provisioning reason above.
The wear is readable off a used drive before you pay for it
This is the part that decides whether a second-hand SSD is a bargain, and the first pass takes thirty seconds.
NVMe: the counters are standard, the two headline fields are not
The SMART / Health Information log page, log identifier 02h, is defined by the NVMe specification. The counters mean the same thing on every drive. The two fields everybody actually reads do not, and the earlier version of this article overstated the uniformity.
$ sudo nvme smart-log /dev/nvme0
critical_warning : 0
temperature : 41 C
available_spare : 100%
available_spare_threshold : 10%
percentage_used : 3%
data_units_read : 118,447,210
data_units_written : 47,281,940
host_read_commands : 1,204,883,712
host_write_commands : 688,192,404
controller_busy_time : 3,114
power_cycles : 412
power_on_hours : 21,904
unsafe_shutdowns : 38
media_errors : 0
num_err_log_entries : 0
percentage_used, byte 05, is defined by the specification as “a vendor
specific estimate of the percentage of NVM subsystem life used”. A value of 100
means the estimated endurance has been consumed “but may not indicate an NVM
subsystem failure”. The value is allowed to exceed 100, percentages above 254 are
reported as 255, and it is updated once per power-on hour when the controller is
not asleep. “Vendor specific estimate” is the operative phrase: the field is
standardised in its meaning and units, not in how it is computed.
Western Digital says what its own firmware bases it on, and it is not what most people assume: the percentage-life figure is “based on the actual PE cycles of the NAND utilized in the drive, not based on host writes”. That is why it can diverge from your own host-write arithmetic, and it is the more trustworthy of the two - it is measuring the thing that wears, rather than a proxy for it.
data_units_written, bytes 63:48, counts 512-byte data units written “as
part of processing a User Data Out Command”, reported in thousands and rounded
up. Two caveats the earlier version missed. The spec says “this value does not
include metadata”, so a namespace formatted with protection information
under-reports true host bytes. And a value of 0h means the counter is not
reported at all, not that nothing has been written.
TB written = data_units_written ÷ 1,953,125
= 47,281,940 ÷ 1,953,125
= 24.2 TB
Against a 600 TBW rating that is 4.0 per cent of the budget, which agrees near enough with the 3 per cent the controller reports. Cross-check it against the hours before you trust either:
24.2 TB ÷ 21,904 hours = 1.1 GB per hour = 26.5 GB per day
26.5 GB ÷ 1,000 GB capacity = 0.027 DWPD
A drive that has averaged under three hundredths of a drive write per day for two and a half years has been idling, whatever the seller said it was doing. The capacity in that sum is the example drive’s 1 TB, which is the same 600 TBW reference drive used everywhere else in this article.
media_errors - Media and Data Integrity Errors, bytes 175:160 - counts
“occurrences where the controller detected an unrecovered data integrity error”,
and the spec is explicit that “errors such as uncorrectable ECC, CRC checksum
failure, or LBA tag mismatch are included”. So it is not purely a NAND-wear
signal. A CRC failure on the link counts. Errors produced by a Write
Uncorrectable command may or may not be included. It should still be zero, and a
non-zero value is still a reason to ask a question, but it is not proof of worn
flash.
unsafe_shutdowns is what nvme-cli still prints. The specification now names
the field Unexpected Power Losses, bytes 159:144, and notes that “this field
was previously named Unsafe Shutdowns”. It increments only when main power is
lost before the controller reports that it is ready to be powered off.
critical_warning is a bit field, and two of its bits belong in an endurance
article:
| Bit | Meaning |
|---|---|
| 0 | Available spare has fallen below the threshold |
| 1 | Temperature outside an operating range |
| 2 | NVM subsystem reliability degraded |
| 3 | All Media Read-Only - all of the media has been placed in read only mode |
| 4 | Volatile Memory Backup Failed - the power-loss protection failure flag |
| 5 | Persistent Memory Region has become read-only |
Bit 3 is the end-of-life transition, visible in one byte, on a drive you have not bought yet.
The endurance group log is the page nobody runs
The earlier version of this article said that getting real write amplification off a drive needed an OCP datacentre extension and that “consumer drives do not offer this”. That is wrong twice, and the first half is the more useful correction.
The Endurance Group Information log page, Log Identifier 09h, is in the
standard NVMe specification. No vendor extension. nvme endurance-log reads it,
and it carries three fields that between them answer the questions the SMART page
cannot:
| Field | Bytes | What it is |
|---|---|---|
| Endurance Estimate | 47:32 | “an estimate of the total number of data bytes that may be written to the Endurance Group over the lifetime… assuming a write amplification of 1”, in units of 1,000,000,000 bytes |
| Data Units Written | 79:64 | host writes, excluding controller writes, in units of 1,000,000,000 bytes |
| Media Units Written | 95:80 | “includes data bytes written by both the host and the controller (e.g., due to garbage collection)” |
Dividing Media Units Written by Data Units Written gives lifetime write amplification straight out of the standard specification, with no OCP plugin and no vendor tooling:
media units written 53,258 × 10^9 bytes = 53.3 TB
data units written 24,208 × 10^9 bytes = 24.2 TB
lifetime WAF = 53.3 ÷ 24.2 = 2.2
And Endurance Estimate is the drive stating its own TBW in a machine-readable field. A value of 600,000 is 600 TB, which you can check against the listing title, the datasheet, and the host-write counter. A drive whose Endurance Estimate does not match the datasheet for the model it claims to be is a drive worth a second look.
Not every controller populates all three. A zero means not reported.
OCP datacentre drives publish the rest
Datacentre drives following the OCP Datacenter NVMe SSD Specification carry an extended SMART/Health log at C0h, and it is the richest wear disclosure available on any storage device:
$ sudo nvme ocp smart-add-log /dev/nvme0
The field the specification is most explicit about is Physical Media Units Written, bytes 15:0, “the number of bytes written to the media; this value includes both user and metadata written to the user and system areas”. The spec then states the intended use outright: “it shall be possible to use this attribute to calculate the Write Amplification Factor (WAF).”
The rest of what C0h carries, and what each field answers:
| Field | What it tells you |
|---|---|
| Bad User NAND Blocks | Normalised value is the percentage of user spare blocks still available; raw is blocks retired |
| XOR Recovery Count | How often the controller needed its parity to reconstruct data |
| Uncorrectable Read Error Count | Reads that correction could not fix |
| Soft ECC Error Count | Reads that correction did fix, which is the leading indicator |
| Refresh Counts | Blocks re-allocated to maintain data integrity, explicitly excluding garbage collection and wear levelling |
| User Data Erase Counts, min and max | The wear spread across the flash, which is wear levelling quality made visible |
| Per cent Free Blocks | Headroom the controller has left to work with |
| Unaligned I/O | Host writes that missed the indirection unit |
| PLP Start Count | How many times the capacitors have been called on |
| Capacitor Health | Hold-up energy margin, as above |
| Endurance Estimate | The drive’s own TBW, in bytes |
| Media Dies Offline | Whole dies the controller has given up on |
| NAND Average Erase Count | The wear figure itself |
The minimum and maximum user data erase counts are the single most informative pair on a used drive, because the gap between them is how well the wear levelling worked over the drive’s whole life. A narrow gap at a high average is a drive that was used hard and managed well. A wide gap is a drive whose hot blocks took the punishment.
nvme-cli 2.16 ships an OCP plugin whose subcommands map one to one onto the OCP requirements, and these are the commands worth running on a used datacentre NVMe drive:
nvme ocp smart-add-log the C0h log above
nvme ocp eol-plp-failure-mode the end-of-life policy, readable and settable
nvme ocp get-plp-health-check-interval how often the capacitors are tested
nvme ocp latency-monitor-log tail latency history
nvme ocp fw-activate-history every firmware the drive has run
nvme ocp error-recovery-log
nvme ocp device-capability-log
nvme ocp unsupported-reqs-log which OCP requirements this drive does not meet
That last one deserves attention on a listing. OCP compliance is a claim, not
a guarantee of every behaviour in the specification. Micron discloses of its
6550 ION - a 61.44 TB Gen5 drive on Micron’s G8 TLC NAND, doing 14.0 GB/s read and 7.0 GB/s write
inside 20 W - that it “complies with most, but not all, requirements of the Open
Compute Project (OCP) Datacenter NVMe SSD Specification 2.5”. unsupported-reqs-log
is where the drive tells you which ones.
SATA attributes are a minefield, and the standard names the reason
SMART attribute numbering above the low IDs is not standardised, and the smartctl manual page states the underlying problem in one sentence: “the conversion from Raw value to a quantity with physical units is not specified by the SMART standard”. Each vendor also uses its own algorithm to normalise the raw value into the 1-254 VALUE column. smartctl does not do that conversion. The firmware does, and smartctl prints names it gets from a per-model database rather than from the drive.
$ sudo smartctl -A /dev/sda
ID# ATTRIBUTE_NAME FLAGS VALUE WORST THRESH FAIL RAW_VALUE
5 Reallocated_Sector_Ct -O--CK 100 100 010 - 0
9 Power_On_Hours -O--CK 095 095 000 - 21904
12 Power_Cycle_Count -O--CK 099 099 000 - 412
173 Ave_Block-Erase_Count -O--CK 097 097 000 - 48
194 Temperature_Celsius -O---K 068 049 000 - 32 (Min/Max 19/51)
197 Current_Pending_Sector -O--CK 100 100 000 - 0
198 Offline_Uncorrectable ----CK 100 100 000 - 0
202 Percent_Lifetime_Remain ----CK 097 097 001 - 3
246 Total_LBAs_Written -O--CK 100 100 000 - 47281940512
Without a drive-database match, smartmontools falls back to a very short list of default meanings: 202 Data_Address_Mark_Errs, 231 Temperature_Celsius, 233 Media_Wearout_Indicator, 241 Total_LBAs_Written, 242 Total_LBAs_Read. Everything else you see printed came from a per-model preset, not from the drive.
Here is what those identifiers actually carry across the database, and it is worse than the earlier version of this article suggested:
| ID | Default meaning with no database match | What it is on other drives |
|---|---|---|
| 202 | Data_Address_Mark_Errs | Percent_Lifetime_Remain on Crucial and Micron; Exception_Mode_Status on Samsung datacentre parts |
| 231 | Temperature_Celsius | SSD_Life_Left |
| 233 | Media_Wearout_Indicator | Flash_Writes_GiB, Flash_Writes_32MiB, NAND_Writes_1GiB, NAND_GiB_Written, Lifetime_Nand_Writes, Lifetime_Wts_To_Flsh_GB, Total_LBAs_Written, Total_NAND_Writes_GiB, Percent_Lifetime_Remain, Remaining_Lifetime_Perc |
| 241 | Total_LBAs_Written, in 512-byte LBAs | Host_Writes_32MiB, Host_Writes_GiB, Lifetime_Writes_GiB |
| 246 | none | Total_LBAs_Written on Micron and Crucial; Total_Erase_Count; SLC_Writes_32MiB on Silicon Motion controllers; Timed_Workld_RdWr_Ratio on Samsung datacentre drives; Write_Protect_Detail; Heavy_Read_Retry_count |
Attribute 233 is the most overloaded identifier in the database, appearing as Media_Wearout_Indicator in twenty families and as at least ten other things elsewhere, several of which are byte counters rather than percentages. Reading it as a wear percentage on a drive where it is a NAND-write counter produces nonsense of a spectacular kind, and an earlier version of this article singled out 231 as the conflicted one. 233 is worse.
Attribute 246 is the trap the dump above is standing in. On the Micron and Crucial parts where smartmontools names it Total_LBAs_Written, it is host LBAs and the conversion below is right. On a Silicon Motion controller the same identifier is SLC_Writes_32MiB, which is a NAND-side counter in 32 MiB units. Reading that as host sectors is exactly the failure this section exists to prevent, and an earlier version of this article printed the attribute under a name - Total_Host_Sector_Write - that smartmontools does not use anywhere. The row above lists six meanings for that one identifier and is not exhaustive: one host counter, two NAND-side counters, a workload ratio, a status field and an error tally. A bare integer in the RAW_VALUE column looks identical in all six.
Read the direction before you read the number. On Crucial and Micron client SSDs the drive database records attribute 202 as “norm = max(100-raw,0); raw = percent_lifetime_used”, so the same identifier carries per cent used in the raw column and per cent remaining in the normalised column on the same drive. On the Micron 5100, 52x0, 5300 and 5400 families the database note reads “remaining endurance, trips at 10%”. In the dump above the same fact appears twice in opposite directions - normalised VALUE 097, raw 3 - and reading the wrong column turns a nearly-new drive into a dead one or the reverse.
Some SATA drives expose enough to compute write amplification directly, which the earlier version of this article said they did not:
| Family | Host side | NAND side | Write amplification |
|---|---|---|---|
| Micron and Crucial | 247 Host_Program_Page_Count | 248 Bckgnd_Program_Page_Cnt | (247 + 248) ÷ 247 |
| Samsung client | 241 Total_LBAs_Written | 249 NAND_Writes_1GiB | convert both to bytes, then divide |
| Silicon Motion | family-specific | 245 TLC_Writes_32MiB and 246 SLC_Writes_32MiB | the two NAND counters also split cached from direct traffic |
The Silicon Motion pair is the more interesting of the three for a different reason: 245 and 246 together tell you how much of the drive’s life went through the pSLC cache, which is the fold arithmetic from earlier in this article made observable.
Micron’s enterprise SATA drives go further and expose the background data retention scan itself:
| ID | Name | What it counts |
|---|---|---|
| 210 | RAIN_Success_Recovered | Parity recoveries that worked |
| 211 | Integ_Scan_Complete_Cnt | Periodic data integrity scans completed |
| 212 | Integ_Scan_Folding_Cnt | Blocks reallocated by integrity scans |
| 213 | Integ_Scan_Progress | Normalised is a percentage; raw is superblocks scanned by the current scan |
That is the refresh mechanism from the garbage collection section, counted. A rising 212 on a used drive is the controller telling you it is finding blocks that need rewriting.
Intel and Solidigm datacentre SATA parts add a workload timer that no other vendor exposes: 226 Workld_Media_Wear_Indic in units of per cent times 1024, 227 Workld_Host_Reads_Perc, and 228 Workload_Minutes. Together they say what mix of reads and writes the drive saw over a measured window, rather than over its whole life.
The unit conversions, and the sanity check that catches a wrong one
Once you know which unit you have:
512-byte LBAs → TB raw ÷ 1,953,125,000
32 MiB units → TB raw ÷ 29,802
1 GiB units → TB raw ÷ 931.3
1 GB units → TB raw ÷ 1,000
So 47,281,940,512 sectors × 512 = 24,208,353,542,144 bytes = 24.2 TB - the same drive as the NVMe example, which is the point: the NVMe data unit is exactly 1000 sectors, so the two divisors differ by a factor of a thousand and nothing else.
Sanity-check every conversion against power-on hours, because the check catches a wrong unit assumption immediately:
24.2 TB over 21,904 hours = 1.1 GB/hour plausible
same raw read as 32 MiB units:
47,281,940,512 ÷ 29,802 = 1,586,000 TB
1,586,000 TB ÷ 21,904 hours = 72 TB/hour
72 TB/hour ÷ 3,600 s = 20 GB/second impossible on SATA
SATA 6 Gbit/s tops out near 600 MB/s. The unit assumption was wrong, not the drive. Any conversion that implies more bandwidth than the interface has is arithmetic to redo, not a discovery.
Do the last division on paper rather than in your head. Stopping at 72 and writing GB/second beside it is the same class of error this whole check exists to catch: a figure carried across a unit boundary it was never converted through. It comes out 3.6 times too high and still points at the right conclusion, which is exactly what makes it easy to leave in.
The ATA Device Statistics log is standardised and underused
There is a standardised alternative on SATA, defined by the ATA command set rather than by the vendor: the Device Statistics log, General Purpose Log address 0x04, introduced in ACS-2.
$ sudo smartctl -l devstat /dev/sda
Page Offset Size Value Description
1 0x010 4 21904 Power-on Hours
1 0x018 6 47281940512 Logical Sectors Written
1 0x020 6 688192404 Number of Write Commands
1 0x028 6 118447210000 Logical Sectors Read
7 0x008 1 3 Percentage Used Endurance Indicator
Prefer this to the attribute table wherever the drive supports it. Page 1 carries Lifetime Power-On Resets, Power-on Hours, Logical Sectors Written, Number of Write Commands, Logical Sectors Read, Number of Read Commands and a date and time timestamp, with ACS-4 adding Pending Error Count, Workload Utilization, Utilization Usage Rate, Resource Availability and Random Write Resources Used. Page 7, Solid State Device Statistics, contains exactly one entry: Percentage Used Endurance Indicator. It counts up, like the NVMe field, and it means the same thing on every drive that reports it.
Sectors written divided by write commands gives the mean write size - here 68.7 sectors, about 34 KiB - which says roughly what the previous owner did with it. A mean write size near 4 KiB on a drive with a large indirection unit is a warning; a mean near 128 KiB is a drive that was fed sequential work.
SAS keeps its wear indicator in a log page, not an attribute
SAS and SCSI SSDs are the case people assume is hopeless, and it is the opposite: the wear figure is standardised and has been since SBC-3. It is simply not a SMART attribute.
SCSI log page 0x11, Solid State Media, parameter code 0x0001, is the “percentage used endurance indicator”, defined in SBC-3 section 6.3.6. smartctl reads it:
$ sudo smartctl -l ssd /dev/sdb
The manual page gives the scale: “a value of 0 indicates as new condition while 100 indicates the device is at the end of its lifetime as projected by the manufacturer. The value may reach 255.”
Host writes come from a different page. SCSI log page 0x19, General Statistics
and Performance, defined in SPC-6, reports the number of read commands, the
number of write commands, the number of logical blocks received and the number of
logical blocks transmitted. smartctl -l genstats reads it.
You do not have to remember either, because smartctl -x on a SAS device already
includes -l ssd -l genstats -l background -l envrep -l zdevstat. On a SAS SSD,
smartctl -x is the one command to ask a seller for, and it returns a
standardised wear percentage where SATA returns a vendor guess.
Enterprise drives covers the rest of what a SAS pull
needs from you.
Backing write amplification out of two attributes
Where a drive gives you an erase count and a host-write count and nothing else, the two together still estimate what it has been doing. The example drive is 1 TB, so its raw NAND is 1024 GiB, or 1.10 TB:
NAND written ≈ 48 erase cycles × 1.10 TB = 52.8 TB
WAF ≈ 52.8 TB ÷ 24.2 TB host = 2.2
The same two numbers give the P/E budget the controller thinks it has: 48 cycles at 3 per cent used implies about 1,600 cycles, which is a 1,500-cycle die - the same die the consistency check earlier in this article inferred from the 600 TBW rating alone, by a completely different route.
Both are estimates, and here are the specific reasons they are. The average erase count may or may not include pSLC blocks, which are erased three or four times as often per byte stored as the rest - the fold arithmetic above says so. The raw capacity is inferred from the marketed one. No vendor confirms either. Treat 2.2 as “roughly two”. Still more than the listing told you, and on an NVMe drive the endurance group log measures the same quantity properly rather than estimating it - 53.3 TB of media units against 24.2 TB of data units, which is 2.2 again on the figures used throughout this article. If your own drive gives you both routes and they disagree by a lot, the erase-count route is the one to distrust, for the two reasons in the paragraph above.
Percentage used, available spare and media errors answer three different questions
These three get conflated constantly, including by tools that summarise them into a single health percentage. They are not three views of one thing.
Percentage used is a wear estimate. It counts up, it is the vendor’s estimate of consumed endurance, and on Western Digital drives at least it is computed from actual NAND P/E cycles rather than from host writes. It is allowed to exceed 100 and the drive is allowed to keep working past it.
Available spare is an inventory. It is a normalised percentage of the reserve block pool, with no basis defined by the NVMe specification, falling as blocks are retired. Its threshold is where the drive says it is running out of replacements. A drive can be at 10 per cent used and 50 per cent spare, which means it has retired a lot of blocks early for reasons that are not wear.
Media errors are an event count. They are things the correction engine could not fix, including link CRC failures and LBA tag mismatches that have nothing to do with flash at all.
The relationship between the first two is only guaranteed on a drive that promises it. The OCP specification makes percentage used a defined function of bytes written, which is exactly what it is not on a consumer drive. Requirement ENDUD-1 requires the documentation to state the number of physical bytes writable “assuming a write amplification of 1”, in gigabytes of 10^9 bytes. ENDUD-3 requires Percentage Used to “track linearly with bytes written and at 100% it shall match the EOL value specified in ENDUD-1”.
ENDUD-2 even pins the preconditioning workload that gets a drive to end of life for the measurement, and it is worth reading as a description of what the number means: 50/50 read/write by I/O count, 4 KiB random reads aligned to 4 KiB, 128 KiB sequential writes aligned to 128 KiB, 100 per cent active range, 80 per cent full device, 0 per cent compressible data, 35 °C ambient, short-stroked if the device is 2 TB or larger.
And the two fields are required to stay in a sane relationship to each other. OCP EOL-6 requires the drive to carry enough spare blocks that Percentage Used reaches 100 per cent before Available Spare falls below its threshold. EOL-7 requires both fields to update at 1 per cent granularity. That ordering is the warning: on an OCP drive the wear estimate is guaranteed to exhaust itself before the spare pool does, so Percentage Used is the field that moves first and Available Spare crossing its threshold is already late.
Consumer drives carry no equivalent requirement. Nothing guarantees which of the two fields moves first on one, or that either moves in steps small enough to notice before it matters.
One more distinction, because it is the most common category error in this subject. MTBF and AFR are not endurance figures. Western Digital states it flatly: MTBF is “mostly unrelated to endurance”. An annualised failure rate predicts random electrical failure across a population. An endurance rating predicts an individual drive’s ability to read back what was stored on it. A drive past its endurance rating may keep running for years without any electrical failure at all, while its read error rates climb. Kioxia’s 2,500,000-hour MTTF figure and its 1 DWPD rating are answers to two unrelated questions.
What end of life actually does, and who implements it
An earlier version of this article said that “some enterprise drives fall back to read-only, most consumer drives carry on until they do not, and vendors decline to say which”. For datacentre NVMe drives that is now out of date, and the behaviour is specified, mandatory and readable before you buy.
The OCP Datacenter NVMe SSD Specification requires the transition and names the trigger. EOL-5: “the device shall switch to Read Only Mode (ROM) when the Available Spare field… reaches 0%”, setting Critical Warning bits 2 and 3 as it does so. Available Spare must report 0 per cent while spare blocks actually remain, precisely so that reads keep working after writes stop.
The policy is a settable feature rather than a fixed behaviour. Feature Identifier C2h, EOL/PLP Failure Mode, has three legal values:
| Value | Behaviour on PLP failure | Behaviour at end of life |
|---|---|---|
| 01b | Read-only | Read-only |
| 10b | Write-through | Read-only |
| 11b | Normal operation | Read-only |
Requirement ROWTM-1 sets the factory default: “the device shall default from
the factory to Read Only Mode (ROM) (01b)”. Every value ends in read-only at
end of life; what the three differ on is what happens when the capacitors fail.
nvme-cli exposes the whole thing as nvme ocp eol-plp-failure-mode, so on a used
datacentre drive you can read the policy the previous owner left it in.
Two related requirements are worth knowing because they contradict common assumptions. RETC-3 states that the device shall not throttle performance based on the endurance metric - a worn OCP drive is not allowed to get slower as a warning. And the base NVMe specification defines Critical Warning bit 3, All Media Read-Only, independently of OCP, so the flag exists on any compliant controller that chooses to set it.
On the consumer side most drives still do not document this, but “no consumer drive ever did” is wrong, and the counter-example is in the next section. The Intel 335 Series was designed to shift into read-only mode when its media wear indicator ran out and then to brick itself on the next power cycle, and it did exactly that, with a single reallocated sector to its name. That is a deliberate and arguably correct end-of-life design: stop writing, let the owner copy the data off, refuse to be trusted again after a reboot.
What the endurance-to-destruction runs proved, and what they did not
The best-known experiment ran six consumer drives until they died and published the numbers. It concluded on 12 March 2015, and it is worth saying up front that the original articles are no longer on the publisher’s site - the URLs return 404 and the domain now hosts unrelated content. Every figure below survives in the Internet Archive and in third-party summaries. A guide that cites it should say so.
The drives and where they stopped:
| Drive | Final host writes |
|---|---|
| Kingston HyperX 3K 240 GB, incompressible data | 728 TB |
| Intel 335 Series 240 GB | about 750 TB |
| Samsung 840 Series 250 GB, TLC | just under 1 PB |
| Corsair Neutron GTX 240 GB | about 1.2 PB |
| Kingston HyperX 3K 240 GB, compressible data | past 2.1 PB |
| Samsung 840 Pro 256 GB | over 2.4 PB |
Four things it proved, and they are still the four things worth taking from it.
Consumer drives massively exceed their ratings, on a sample of six. The 840 Pro’s 2.4 PB on a 256 GB drive is 9,375 full drive writes - roughly three times a 3,000-cycle planar MLC rating before write amplification is even considered.
Compression is worth about a petabyte. The two HyperX drives were the same model with different data. The incompressible one “wrote slightly more data to the flash than it received from the host, an expected result given the low write amplification of our sequential workload” - a write amplification of about one. Its compressible twin “wrote 28% less to the NAND” thanks to the SandForce controller’s compression, and went on for roughly another petabyte after the point that killed its sibling.
End-of-life behaviour was completely vendor-dependent, and only one drive did it properly. The Intel 335’s media wear indicator ran out shortly after 700 TB; it went read-only and bricked itself on the next power cycle, as designed. The Kingston HyperX refused to write after 728 TB with the data still readable, then did not respond after a reboot. The others simply died.
The SMART warnings were not reliable, and this is the finding that matters most for buying used. The Corsair, Intel and Kingston drives all issued SMART warnings before their deaths. Samsung’s own software “pronounced the 840 Series and 840 Pro to be in good health before their respective deaths”, and - the experimenters’ own emphasis - “worryingly, the 840 Series’ uncorrectable errors didn’t change that cheery assessment”. The 840 Pro accumulated over 7,000 reallocated sectors totalling 10.7 GB of flash on its way to 2.4 PB while the vendor tool reported it healthy, then died unattended during a week nobody was watching.
Now what it did not prove, because this experiment gets over-cited.
It is six drives, which is not a population. They were planar MLC and TLC from 2012 and 2013, a technology generation that no longer exists, with endurance characteristics that neither 3D TLC nor 3D QLC shares. The workload was largely sequential, which is the kindest case for write amplification and is not what a database does. And the useful result - that ratings are conservative - is a statement about warranty policy, not a licence to plan around three times the rating on a drive you depend on.
A follow-up series from 3DNews ran from January 2018, writing alternating small and large files to NTFS volumes and checking retention at rated milestones, and is widely cited as having found budget drives exceeding their ratings by an order of magnitude or more. The original tables could not be reached for this article, only summaries of them, so those figures are second-hand here and are not reproduced.
Unpowered retention, and why an SSD is a bad archive
A flash cell holds charge. Nothing about that charge is refreshed while the drive is unpowered, and it leaks - by trap-assisted tunnelling and by charge detrapping, which is where this article started.
The requirements are in the JESD218 table above: 1 year at 30 °C for client, 3 months at 40 °C for enterprise. Read the qualification before drawing conclusions: those are requirements at the end of rated life, on a drive that has consumed its full TBW. A drive at 3 per cent used is nowhere near that boundary, because retention and endurance trade directly against each other. The spec is written at end of endurance rather than at the start for exactly that reason.
That qualification is the source of a long-running piece of misinformation. The widely-circulated table showing retention collapsing to a week or so is the enterprise row, extrapolated to high storage temperatures, on a drive at 100 per cent of its endurance. Retention does roughly halve for every 5 to 10 °C - the exact figure depends on the activation energy assumed and is not a constant of nature - so the table is directionally right and quoted about the wrong drive.
The datacentre specification sets different numbers again, and the difference is worth seeing:
| Requirement | JEDEC JESD218 enterprise | OCP Datacenter NVMe |
|---|---|---|
| Powered-off retention at end of life | 3 months at 40 °C | at least 1 month at 40 °C (RETC-1) |
| Powered-on retention | not specified here | at least 7 years (RETC-2) |
| New in box, factory state | not specified here | fully functional after at least 5 years at 25 °C (SLIFE-1) |
The OCP end-of-life retention requirement is a third of JEDEC’s, which is not a contradiction - they are different specifications written for different customers - but it does mean “meets the standard” is an incomplete sentence. A datacentre drive designed to OCP RETC-1 has one month of unpowered retention guaranteed at end of life. That is a fleet requirement written by people who never intend to unplug anything.
The powered figure is the striking one. Seven years powered against one month unpowered is a factor of eighty-four, and it is the clearest statement anywhere of what the background refresh machinery is actually doing for you.
What follows for archives is not in dispute:
- Nothing refreshes an unpowered drive. Refresh, read reclaim and the temperature-adaptive scheduling that drives them all need the controller running. On a shelf, none of it is.
- More bits per cell means less retention headroom. QLC’s sixteen voltage bands sit in the same window as SLC’s two. It is the worst archive flash, and it is what cheap large drives are made of.
- 3D NAND added a retention failure mode planar did not have. Charge can escape a charge-trap cell in three dimensions rather than one, across the tunnel oxide and across the charge-trap insulator, causing rapid leakage “for only a few seconds after cell programming”. After those seconds the long-term behaviour resembles planar, and the survey reporting it notes that no mitigation specific to it had been designed.
- You cannot tell from the outside. A drive stored three years may read back perfectly or come back with uncorrectable sectors, and nothing on its label, its SMART data at the time of storage, or its price predicts which. This site will not give you a number of months, because nobody can.
Powering a stored drive up occasionally is better than not. Be honest about why: whether a consumer controller runs a full read-scrub during idle, and how long it needs, is documented by nobody, so it is a reasonable precaution rather than a guarantee. On a Micron enterprise SATA drive you can at least watch it happen, through attributes 211, 212 and 213.
A hard drive’s magnetic domains hold for decades on a shelf; its risks there are mechanical - stiction, lubricant migration, a bearing that will not spin up - and it needs verifying too. Tape is the medium designed for the job. Whatever you use, the rule that survives is multiple copies on more than one medium, checksummed and verified on a schedule. An SSD is an excellent working drive and a poor box in the loft - see HDD or SSD.
Datacentre QLC and the cheap QLC that made its reputation are different products
QLC deserves its own treatment because the folklore about it is a generation out of date in one direction and too generous in another.
The reference part is Solidigm’s D5-P5336, built on 192-layer QLC, and its product brief publishes the full set:
| Specification | Value |
|---|---|
| Capacities | 7.68 / 15.36 / 30.72 / 61.44 / 122.88 TB |
| Endurance | 0.42 to 0.60 DWPD over five years, 5.9 to 134.3 PBW |
| UBER | fewer than 1 sector per 10^17 bits read |
| MTBF | 2 million hours |
| Powered-off retention | 3 months at 40 °C |
| Maximum power | 25 W |
That is a QLC drive with a better error rate specification than the JESD218 enterprise requirement by an order of magnitude, holding 122 TB, rated for 134 petabytes of writes. The “QLC is disposable” position does not survive contact with that datasheet.
Solidigm publishes its own worst-case example, and it is the most useful single sentence in the brief: the 122.88 TB drive running 32 KB 100 per cent random writes at 100 per cent duty cycle, twenty-four hours a day for five continuous years, retains about 5 per cent of its life - roughly three months of remaining margin on a five-year rating, and that conversion is mine, not Solidigm’s. At 4 KB random writes it retains about 10 per cent.
Note the direction of those two figures, because it is not the direction you expect. The smaller write size leaves more life, not less. At a 100 per cent duty cycle the drive is bandwidth-limited, and 4 KB random writes move far fewer bytes per second than 32 KB ones, so fewer bytes reach the media over five years even though each one costs more in amplification. That is my reading of why the two figures fall that way; Solidigm publishes the outcome, not the reasoning.
Three things to carry to a listing page.
Datacentre QLC and consumer QLC are not the same product. The derivation earlier in this article puts the datacentre die somewhere between about 1,100 and 2,200 P/E cycles and the consumer die between about 200 and 400 on the same assumptions. A used 15.36 TB datacentre QLC drive is not a scaled-up version of the cheap 2 TB QLC drive that gave QLC its reputation.
The indirection unit is the whole risk. A large QLC drive rated on IU-aligned writes, fed 4 KiB random writes by a filesystem that does not know or care, is running at four to eight times the amplification its rating assumed. If the workload is small random writes, the capacity is not the reason to buy it and the endurance figure on the datasheet does not apply to you.
A model number does not pin the flash. Micron states in the small print of its consumer SSD pages that it “reserves the right to transition between NAND series during future production cycles”. Two used drives with the same label can have different NAND inside and therefore different endurance characteristics, which is one more reason to read the drive’s own counters rather than to look up the model.
The trap: “8TB written” is an odometer, not a capacity
This one cost this site real effort. A listing titled “Enterprise NVMe SSD - only 8TB written!” is not an 8 TB drive. The number is host writes - wear - and the seller is quoting it because it is low.
Read it that way and it is useful. Eight terabytes against a 600 TBW rating is 1.3 per cent of the budget. If the same drive reports 12,000 power-on hours - 500 days - it averaged 8 ÷ 500 = 16 GB a day, or 0.016 DWPD on a 1 TB drive. It has been idling.
The distinction that matters in a title:
| What the title says | What it means |
|---|---|
| “600 TBW” | The rating. What the drive is licensed to do |
| “8TB written”, “8TB host writes” | The odometer. What it has already done |
| “8TB”, alone, near a model number | Almost certainly capacity |
A seller who quotes host writes has read the SMART data, which is more than most
have done, and that is a mild positive signal in itself. A seller quoting “100 per
cent health” has read a vendor tool’s rounded summary and told you nothing - and
the destruction runs above showed exactly what a vendor tool’s cheery summary is
worth on a drive with accumulating uncorrectable errors. Ask for smartctl -x or
nvme smart-log output instead, and treat a refusal as an answer.
Where a title is genuinely ambiguous between capacity and wear, this site leaves the field empty rather than guessing. A wrong capacity would put a drive at the top of a price-per-terabyte table on the strength of a typo, which is exactly the failure mode that ranking invites - see buying used drives on eBay.
What cannot be known, and why this site does not guess it
Every section above has a boundary, and collecting them in one place is more useful than scattering caveats.
The P/E rating of the NAND in a finished drive is not published. Not by consumer vendors, not by enterprise vendors, not per product. Every cycle figure in this article is either a measurement from published research on a named node or a derivation from a TBW rating, and both are labelled.
The garbage collection policy is not published. Which means steady-state write amplification for your workload on your drive is not calculable in advance. The direction of every lever is certain; a number is not.
The SLC cache policy is rarely published. Static, dynamic or hybrid, its size, its cap, and the idle time it needs to fold are all typically undisclosed, which is why sustained-write behaviour has to be measured rather than looked up.
The pSLC-mode cycle rating is not published, so the endurance cost of the cache cannot be computed from the outside, only its media-bytes cost.
Percentage Used is a vendor-specific estimate unless the drive is OCP compliant, in which case it is required to track linearly with bytes written. On a consumer drive, two different models can report the same number for genuinely different amounts of wear.
Whether an average erase count includes pSLC blocks is not stated, which puts a real error bar on the erase-count route to write amplification.
Retention for a specific used drive is unknowable. The specification gives a requirement at end of life for a class of drive. It does not give a forecast for the drive in front of you at the wear level it is at, at the temperature your cupboard happens to be.
Which NAND is inside a given consumer model is not fixed, by the vendors’ own small print.
Most published TBW figures were extrapolated, not measured, under a method the standard explicitly permits when direct verification would exceed 1000 hours.
Where this site’s catalogue cannot source a figure, it records the field as not stated rather than inferring it, for the reasons on the methodology page.
What to do with this on a listing page
- Find your row in the perspective table by measuring, not by guessing. Read the host-write counter, wait a week, read it again, divide. In the top four rows, stop reading endurance figures entirely and sort all SSDs on price per terabyte.
- Below that line, shop enterprise and shop on the capacity number. 400 / 800 / 1,600 / 3,200 GB is 37.44 per cent over-provisioning; 480 / 960 / 1,920 / 3,840 GB is 14.53 per cent. The ladder pins the over-provisioning exactly and the tier approximately - check the family for the DWPD, because the same ladder was 10 DWPD a generation ago and is often 3 DWPD now. Start from the SSD listings or 1.6 TB parts.
- Demand the raw output before buying used, not after.
nvme smart-logplusnvme endurance-logon NVMe,smartctl -xon SAS,smartctl -x -l devstaton SATA. A seller who will not run one command has told you how much they know about the drive, and a “100 per cent health” screenshot is a tool’s rounded summary of a field it guessed the meaning of. - Convert the host writes yourself and check them three ways - against the rating, against the power-on hours, and against the interface bandwidth. Divide by 1,953,125 for NVMe data units and 1,953,125,000 for 512-byte LBAs. Any answer implying more than 600 MB/s sustained on SATA is a wrong unit assumption, not a remarkable drive.
- On a datacentre NVMe drive, run the OCP plugin before you bid.
nvme ocp smart-add-logfor Capacitor Health and the erase-count spread,nvme ocp eol-plp-failure-modefor the end-of-life policy, andnvme ocp unsupported-reqs-logfor the requirements the drive does not actually meet. Capacitor Health well below 100 on an enterprise SSD is a disqualification, not a discount. - Read the direction of every wear attribute before you read its value. A raw 3 on attribute 202 is three per cent used; a 3 on a drive reporting life remaining is three per cent left. Attribute 233 means at least eleven different things and 246 at least six. When the drive offers Percentage Used Endurance Indicator - devstat page 7 on SATA, log page 0x11 on SAS - use that instead.
- Match the drive to the write size, not just to the capacity. A large QLC
part on NVMe rated on 16 or 32 KiB IU-aligned
writes is superb for media and archive-style bulk and poor for a metadata vdev
or a busy database. Check
nvme id-nsfor the preferred write granularity before you assume the datasheet number applies to you. - Buy over-provisioning for free by not filling the drive. Twenty per cent
left empty on a consumer 1 TB drive, with working TRIM, is the same 37.4 per
cent over-provisioning the enterprise write-intensive ladder buys, and it cuts
the write amplification bound by two thirds. Verify with
lsblk -Dthat discards reach the device at all. - If the drive is going to be a SLOG or a sync-write target, buy power-loss
protection or do not buy.
nvme id-ctrlreportingvwcof 0 is the one-command test. Without it, everyfsyncis a real flush and the endurance arithmetic in this article is optimistic. - Exclude auctions while you are learning what a model is worth. Buy-it-now only compares prices somebody can pay today rather than bids that have not finished rising.
A used enterprise SSD at 10 per cent of its endurance has more writing left in it than a new consumer drive has in total, and usually sells for less per terabyte. That trade is only visible if you read the odometer - and after this article, you can read all four of the counters that make up the rating it is measured against.