Arrays with NVMe drives: how-to, do's and don'ts
By Harry Saarinen ·
Count the lanes before you count the drives, and for most home builds stop at a two-drive mirror. An NVMe array fails at the platform long before it fails at the drive, and everything past a mirror - parity, wider stripes, a cache tier, a controller card - buys complexity first and a benefit second.
An NVMe array is not a SATA array with faster drives
Most NVMe arrays that disappoint were built by someone who had built a good SATA array first. From a distance the parts look identical - several drives, a redundancy scheme, a filesystem over the top - so the habits come across intact: a controller card in a spare slot, a stripe across every drive present, a chunk size copied from a forum thread that predates the first consumer NVMe SSD. Those habits were correct for a device that sat behind a host bus adapter and delivered 550 MB/s. They are being applied to a device that is a PCIe endpoint in its own right, that the CPU addresses directly, and that delivers thirteen to twenty-five times that on its own.
The drive stopped hanging off an HBA
Under AHCI the operating system talks to a controller, and the controller talks to the drives. The controller is what the driver binds to; the drives are ports on it, reached through it, arbitrated by it and capped by its single link. Under NVMe the drive itself is the PCIe function. It has its own configuration space, its own BAR, its own MSI-X vectors and its own doorbell registers, and in a directly attached slot nothing sits between the CPU and the flash controller except the root port. A PCIe switch or a retimer in the path is the exception, and it is one of the five questions below. Which physical socket carries which protocol is a separate question, covered in SATA, SAS, NVMe and M.2: what actually plugs into what.
That changes the submission path more than it changes the wire. The NVM Express Base Specification Revision 2.4 lists these key attributes: the path “does not require uncacheable / MMIO register reads in the command submission or completion path”, “a maximum of one MMIO register write or one 64B message is necessary in the command submission path”, and “all information to complete a 4 KiB read request is included in the 64B command itself”. The queues live in host memory. Issuing work is a write to a doorbell. Nothing in the fast path stalls a core on a read back across the bus.
One queue of 32 against 65,535 queues of 65,536
| AHCI / SATA | NVMe | |
|---|---|---|
| Queues | one command list per port | up to 65,535 I/O queues |
| Entries per queue | up to 32 slots | 65,536 slots, 65,535 usable |
| Fast-path register read | required | none |
| Submission cost | driver-mediated | at most one MMIO write |
| Arbitration | host driver | round robin, required; weighted round robin with an urgent priority class, optional |
The AHCI figure is the Serial ATA AHCI Specification Revision 1.3.1: “Up to 32 slots (i.e. entries) are supported in a command list.” The NVMe figures are from Base Specification 2.4, which describes “supporting up to 65,535 I/O Queues with up to 65,535 outstanding commands per I/O Queue”, puts the maximum size of an I/O queue at 65,536 slots “limited by the maximum queue size supported by the controller that is reported in the CAP.MQES field”, and explains the off-by-one: “One slot in each queue is not available for use due to Head and Tail entry pointer definition.”
Multiplied out, that ceiling is about 4.29 billion outstanding commands. No shipping device implements it and no workload would ask for it, so the number is the wrong thing to be impressed by. What the model actually buys is that queue pairs are per core. Each core owns a submission and completion queue, so two cores issuing I/O need share no lock, bounce no cache line and contend for no common tail pointer; Linux blk-mq maps hardware queues to CPUs on exactly that assumption. The 32-slot AHCI command list is one shared structure per port, and the cores issuing to that port take their turn at it.
Be honest about the distance between the specification and the silicon. Real
controllers advertise far fewer and far smaller queues than the architectural
maximum, in CAP.MQES and in the Number of Queues feature. nvme show-regs and
nvme get-feature -f 7 report what a particular drive supports, and that, not
the headline, is what an array is built on.
The bottleneck moved inside the machine
PCIe per lane, after the 128b/130b line code (GB = 10^9 bytes)
Gen 3 8.0 GT/s x 128/130 = 7.877 Gb/s = 0.985 GB/s
Gen 4 16.0 GT/s x 128/130 = 15.754 Gb/s = 1.969 GB/s
Gen 5 32.0 GT/s x 128/130 = 31.508 Gb/s = 3.938 GB/s
x4 link ceiling: Gen 3 3.938 GB/s
Gen 4 7.877 GB/s
Gen 5 15.754 GB/s
TLP headers, DLLPs and flow control take a few per cent
more, which this arithmetic does not count.
SATA 6Gb/s, after 8b/10b: 6.0 x 0.8 = 4.8 Gb/s = 600 MB/s
about 550 MB/s in practice
Four SATA SSDs striped, best case 4 x 550 MB/s = 2.2 GB/s
One Gen4 x4 NVMe SSD, datasheet about 7.0 to 7.4 GB/s
One Gen5 x4 NVMe SSD, datasheet about 14 to 14.9 GB/s
7.2 / 2.2 = 3.3x 14 / 2.2 = 6.4x
Both NVMe rows are vendor datasheet claims for large sequential reads rather than measurements, and random work looks nothing like them.
On sequential reads, one drive now outruns the array it replaced by a factor of three to six. The four SATA SSDs never troubled the storage stack, because 2.2 GB/s asks little of a current CPU. A single Gen5 drive does ask something, and the limits it meets are all on the host side of the connector.
One core at 3.5 GHz = 3.5 x 10^9 cycles/s
1,000,000 IOPS on that core = 3,500 cycles per I/O
Three thousand five hundred cycles has to cover the syscall, the filesystem, the block layer, the driver, the doorbell write, the interrupt and the completion. Any global lock, any per-I/O allocation, any single kernel thread in the path consumes the budget outright, which is why a filesystem’s own locking becomes a storage specification. Memory bandwidth is a second wall, and it arrives sooner than people expect:
Four Gen5 drives reading flat out 4 x 14 = 56 GB/s of DMA into RAM
A buffered read copies that again,
one read plus one write per byte + 112 GB/s
= 168 GB/s of memory traffic
Dual-channel DDR5-6000, theoretical peak
6000 MT/s x 8 B x 2 channels = 96 GB/s
Direct I/O drops the copy and leaves the 56 GB/s of DMA, which is still more than half that peak, and the peak is theoretical: no memory subsystem delivers its rated figure under a mixed load. The array is now a memory-subsystem problem, which is a sentence no SATA build ever had to say. Channel count and rate are worth checking before the drive count is (DDR5 modules and rates).
Which leaves redundancy as the reason to build one
On SATA, striping was how bandwidth was obtained, because one drive could not supply it. That argument has expired. Seven gigabytes per second is 56 Gb/s; the network in front of a home server is 1, 2.5 or 10 Gb/s, and 10 Gb/s Ethernet asks for 1.25 GB/s, under a fifth of one drive. A second drive adds no bandwidth a single drive was failing to deliver.
What it adds is that the array survives one of them dying.
For most home builds, a two-drive mirror is the whole of the right answer. It covers the failure an array can do something about, costs no parity arithmetic, and needs no journal device, no stripe cache tuning, no parity write hole to mitigate and no chunk-size decision. The shopping list is two drives and a second slot to put the other one in, and the NVMe listings sort by price per terabyte by default (M.2 only). Everything past a mirror - parity, wider stripes, a cache tier, a controller card - buys complexity first and a benefit second, and that benefit should be nameable before the money is spent. A mirror is also not a backup: it defends against a device failing, not against deleting the wrong directory.
What the rest of this guide settles
Five questions, in the order a build meets them. How many drives the platform can attach and at what width, which is lanes, bifurcation, switches, cabling and form factor rather than free slots. Which redundancy implementation - Linux md, tri-mode hardware RAID, VROC - and what parity actually costs against flash, measured in write amplification and endurance rather than in throughput. Which filesystem, and where each one stops. The point at which the honest answer stops being one box. And how endurance and thermal limits set the real drive choice, on top of what SSD endurance already covers.
PCIe lanes are the budget, and it is spent before the drives arrive
The drives are almost never the constraint. The board is. A current PCIe 4.0 x4 NVMe SSD reads at something near 7 GB/s. Four of them in four M.2 sockets on a consumer board do not read at 28 GB/s, because three of those sockets usually sit behind a single link a third as wide as the demand they present to it. Nothing on the drive’s specification sheet says so. The motherboard manual does.
A lane is one differential pair in each direction. PCIe is full duplex, so every figure below is per lane, per direction, and any vendor number labelled “bidirectional” is twice it. How many lanes reach a slot is set by firmware, and the speed is agreed during link training, so by the time the operating system sees a drive the allocation has already happened. A link can retrain later, and usually downwards.
What one drive asks for
x4 is the norm. The M.2 M-key socket (Socket 3) defines PCIe x4, and the U.2 SFF-8639 pinout allocates up to four PCIe lanes to the device, drawn from a pool of high-speed paths the same connector also uses for SAS and SATA. Budget x4 per drive and the arithmetic stays simple.
x2 and x1 NVMe devices exist and are sold without saying so. Two separate mechanisms produce them. The first is keying: a module notched B+M fits either socket, but connector documentation puts only two PCIe lanes on the B-key pin group, so a B+M device is capped at x2 whatever it is plugged into. The second is the controller, which may implement two lanes in an otherwise ordinary M-key package. Neither fact reaches the listing, and both halve the ceiling.
The trained link is the honest source. lspci -vv prints LnkCap, what the
device is capable of, and LnkSta, what it actually negotiated. A drive
reporting “Speed 8GT/s, Width x2” is on a Gen3 x2 link at roughly 1.97 GB/s
regardless of what the seller wrote, and that is the check to run on arrival
rather than the benchmark. The catalogue’s
M.2 NVMe rows are where most of these turn up.
The arithmetic, derived rather than asserted
Per lane, per direction. GB = 10^9 bytes.
Gen1 2.5 GT/s, 8b/10b
2.5 x (8/10) = 2.000 Gb/s / 8 = 0.250 GB/s 20% to the code
Gen2 5.0 GT/s, 8b/10b
5.0 x (8/10) = 4.000 Gb/s / 8 = 0.500 GB/s
Gen3 8.0 GT/s, 128b/130b
every 130 bits on the wire carry 128 of payload plus a 2-bit sync
header, so the code costs 2/130 = 1.54 per cent
8.0 x (128/130) = 7.877 Gb/s / 8 = 0.985 GB/s 985 MB/s, familiar
x4 lanes = 3.938 GB/s
Gen4 16.0 GT/s, 128b/130b
16.0 x (128/130) = 15.754 Gb/s / 8 = 1.969 GB/s x4 = 7.877 GB/s
Gen5 32.0 GT/s, 128b/130b
32.0 x (128/130) = 31.508 Gb/s / 8 = 3.938 GB/s x4 = 15.754 GB/s
Gen6 64.0 GT/s, PAM4 at 32 GBd (two bits per symbol), no line code
64.0 x (1/1) = 64.000 Gb/s / 8 = 8.000 GB/s raw
then the FLIT tax. A PCIe 6.0 FLIT is a fixed 256 bytes:
236 B TLP payload
6 B data link payload
8 B CRC
6 B FEC
256 B total
CRC and FEC cost 14 of 256 bytes = 5.47 per cent
8.000 x (242/256) = 7.563 GB/s x4 = 30.25 GB/s
counting the data link payload as overhead too leaves a TLP field of
8.000 x (236/256) = 7.375 GB/s x4 = 29.50 GB/s
| Generation | Rate per lane | Line code | Payload per lane | At x4 |
|---|---|---|---|---|
| PCIe 1.0 | 2.5 GT/s | 8b/10b | 0.250 GB/s | 1.000 GB/s |
| PCIe 2.0 | 5.0 GT/s | 8b/10b | 0.500 GB/s | 2.000 GB/s |
| PCIe 3.0 | 8.0 GT/s | 128b/130b | 0.985 GB/s | 3.938 GB/s |
| PCIe 4.0 | 16.0 GT/s | 128b/130b | 1.969 GB/s | 7.877 GB/s |
| PCIe 5.0 | 32.0 GT/s | 128b/130b | 3.938 GB/s | 15.754 GB/s |
| PCIe 6.0 | 64.0 GT/s | PAM4, FLIT with FEC | 7.563 GB/s | 30.25 GB/s |
Several Gen6 numbers circulate and each is correct on its own terms. PCI-SIG’s own press material quotes the raw rate, up to 256 GB/s bidirectionally at x16. The 7.563 GB/s per lane in the table is that same link after FLIT framing, counting only the CRC and FEC bytes as overhead; sources that treat the six-byte data link payload as overhead as well publish 7.375 GB/s instead. Label which one you mean. FLIT mode also removes the per-packet data link overhead of earlier generations, so small-transfer efficiency improves by more than the headline implies.
Every figure above is a link-layer ceiling, reached before the 12 to 16 byte headers on each transaction layer packet (TLP), acknowledgements, flow-control credits and the platform’s maximum payload size take their share. A Gen4 x4 drive rated at 7.4 GB/s sequential read is running at about 94 per cent of its 7.877 GB/s ceiling, which is close to the practical limit of the link rather than a modest claim.
A released specification is not a product. PCIe 6.0 went to PCI-SIG members in January 2022 and PCIe 7.0 on 11 June 2025; PCIe 8.0 was announced on 5 August 2025 at 256 GT/s, targeted for release to members by 2028. Trade press through 2026 reports the first PCIe 6.0 enterprise SSDs in mass production, Micron’s 9650 from February and Samsung’s PM1763 from July, while the first Gen6 host silicon, AMD’s EPYC Venice on SP7, is in production with partner platforms due in the fourth quarter of 2026; Intel’s Diamond Rapids is stated for 2027. This guide will not claim retail Gen6 host availability it cannot source. For a machine you can order and boot today, Gen5 is the ceiling.
Lane budgets by class of platform
Orders of magnitude, because the model numbers move faster than a guide can:
- Consumer desktop, roughly 20 to 28 CPU lanes, of which about 24 reach devices. Vendor platform documents give AM5 28 PCIe 5.0 lanes, four of them spent on a chipset link that runs at Gen4, and Intel’s LGA1851 20 PCIe 5.0 plus 4 PCIe 4.0. The conventional split is x16 to graphics, x4 to one M.2 socket, x4 left over.
- Workstation, roughly 48 to 128. Threadripper PRO 9000WX on WRX90 provides 128 PCIe 5.0 lanes; Xeon W-3500 provides up to 112.
- Server, roughly 88 to 128 per socket. EPYC 9005 on SP5 provides 128 PCIe 5.0 lanes per socket; Xeon 6 provides 88 on the 6700P and 96 on the 6900P, with CXL devices competing for the same lanes.
One drive is x4, so a 24-bay Gen5 U.2 chassis wants 96 lanes of drive attachment before a network card or an accelerator gets one. That single sum explains the shape of every dense NVMe server: a high-lane socket, two sockets, or a PCIe switch. The catalogue’s U.2 NVMe listings are thin, but the lane arithmetic is identical whether the drives are new or pulled from one of those chassis.
The chipset uplink, and what oversubscription costs an array
Extra M.2 sockets almost always hang off the chipset, and the chipset is a switch with a fixed uplink. AM5’s Promontory reaches the CPU over one PCIe 4.0 x4 link, and a board with a second chipset die daisy-chains it behind the first rather than adding a second uplink; Intel’s DMI 4.0 link is x8 on the top LGA1851 chipset, Z890, and x4 on B860 and H810. Four chipset M.2 sockets present 16 lanes of demand to an uplink four to eight lanes wide, shared with the USB controllers, the SATA ports and the onboard network.
Oversubscription is not automatically a mistake. It is a mistake for the workloads an array actually runs. A rebuild, a scrub, a replication send and a nightly backup each ask every member for sequential throughput at the same moment, which is exactly the case an uplink cannot serve. Small random work behaves differently: it exhausts the drives’ own latency budget long before it saturates an x4 uplink, which is why a chipset-attached array can look fine in a desktop workload and then take a day to rebuild.
The trap that lives in the manual, not in the specification
Boards borrow. Populating the second M.2 socket may disable two SATA ports, or drop the primary x16 slot to x8, or force a third socket to Gen3 x2. This is per-board routing and not a property of the platform: the same CPU on a different board shares differently, and a mid-range board and its flagship sibling often differ in precisely this. The block diagram and the slot-configuration table in the motherboard manual are the authority, and they are documents to read before ordering drives rather than after. What plugs into what covers the connector half of the same question.
Worked example: how many drives can this board feed at full rate?
Board: consumer AM5, 24 usable CPU lanes, chipset on PCIe 4.0 x4
Plan: eight PCIe 4.0 x4 NVMe drives, headless, no graphics card
Given: CPU, board and firmware all split the x16 slot x4/x4/x4/x4.
Plenty of consumer boards do not, which ends the exercise at
two drives fed at full rate. Check the manual, not the chipset.
CPU-attached
x16 slot bifurcated x4/x4/x4/x4 ..... 4 drives at 7.877 GB/s each
CPU M.2 socket, x4 .................. 1 drive at 7.877 GB/s
lanes spent ......................... 20 of 24
Chipset-attached
3 further M.2 sockets, x4 each ...... 23.63 GB/s of demand
uplink, PCIe 4.0 x4 ................. 7.877 GB/s of supply
ratio ............................... 3:1, before USB, SATA and the NIC
Note the four spare lanes: on AM5 they are the second CPU-attached
x4 group, and a board that exposes them as an M.2 socket feeds a
sixth drive at full rate. This example assumes it does not.
Answer
drives attached and working ......... 8
drives fed at full rate, together ... 5
aggregate with all eight busy ....... 5 x 7.877 + 7.877 = 47.3 GB/s
the sum of eight x4 ceilings ........ 8 x 7.877 = 63.0 GB/s
The gap between those last two lines is the whole argument of this section. It is not a defect in the drives and no filesystem or RAID level recovers it. Count the lanes first, then count the drives. The same eight drives on a workstation socket are eight independent x4 links with 96 lanes still unspent, and the arithmetic stops mattering. That, rather than clock speed, is what the money buys at the top of this market.
What physically fits
The cheapest NVMe on the second-hand market is in the form factor your machine has no socket for. Decommissioned U.2 enterprise SSDs sell cheaply per terabyte because the buyer pool is small: nearly every consumer board ships M.2 sockets and no 2.5-inch NVMe bay at all. An array on the desktop form factor pays a premium for convenience; an array on the enterprise one pays in cables, a 12 V feed and an hour of homework. Where the connector question outgrows this article, SATA, SAS, NVMe and M.2: what plugs into what takes it further, and Enterprise drive pulls covers where the cheap U.2 stock comes from in the first place.
M.2, and what an M-key socket does not promise
M.2 storage is 22 mm wide. 2280 is the default length; 22110 is the server and workstation length, and the extra 30 mm is usually a bank of power-loss-protection capacitors. A socket carries standoffs only for the lengths it supports.
The keying trap costs money. The M key occupies pins 59 to 66 and may route SATA, PCIe x4 and SMBus. Keying describes the socket’s wiring permissions, not the silicon sitting behind it, so an M-key socket is sometimes a SATA-only socket, and an NVMe drive dropped into one does not appear at all. A B+M drive fits either socket and is capped at SATA or PCIe x2.
Thickness is specified rather than folklore: single-sided classes S1 to S3, double-sided D1 to D5. A thin laptop bay, a NAS cache slot or the socket under a graphics card often has clearance on one side only, and a double-sided module either will not seat or seats with its underside against the board. Length, key and sidedness are the three fields to hold in mind while filtering NVMe in the M.2 form factor.
M.2 is also 3.3 V only, screw-retained and not designed for hot-plug, which is why hyperscale left it. SNIA’s own EDSFF explainer says M.2 110 mm “was popular in hyperscale data centers” but had “challenges in terms of hotplug/serviceability, thermals and overheating”.
U.2: the cheap drive and the bay it wants
U.2 is a 2.5-inch drive on the SFF-8639 connector, usually 15 mm thick and rated to 25 W at that thickness. The PCI-SIG module pin usage gives it a dedicated x4 PCIe path alongside separate SAS and SATA wiring, 12 V and 5 V rails plus 3.3 V auxiliary on E3, and staged mating for hot-plug: the device plug’s P13 is “+12 V Precharge”, a pin that mates early so the drive charges before the signal pins land.
Three ways exist to attach one to a machine with no backplane. A host adapter card with OCuLink, MiniSAS HD, SlimSAS or MCIO cables to individual drives. An enclosure or backplane fed by the same cables. Or an M.2-to-U.2 adapter, which needs an external power feed, because M.2 has no 12 V rail and nothing like a 25 W budget. The catalogue carries only a few at a time, which is the market signal itself: NVMe in U.2.
U.3 and SFF-TA-1001: compatibility runs one way
SFF-TA-1001 Rev 1.1 (28 May 2018), “Universal x4 Link Definition for SFF-8639”, is the same connector with a different lane map. U.2 keeps SAS0 on its own pins, separate from PCIe0 to PCIe3; U.3 overlays PCIe0/SAS0 through PCIe3/SAS3 onto a shared set. PRSNT#, IfDet# and IfDet2# on E6 encode what arrived, and the drive reads HPT0 to set its own lane map. The Scope states the direction plainly:
“This specification does not mandate nor imply any SFF-TA-1001 slot compatibility with Quad PCIe devices. This specification does mandate that SFF-TA-1001 devices are compatible with both Quad PCIe slots and SFF-TA-1001 slots.”
A U.3 drive drops into a U.2 bay. A U.2 drive in a U.3 bay is listed as “No link” in the specification’s own interoperation table. Chassis vendors do wire backplanes as dual-mode, which is why field reports disagree with the standard. The standard mandates one direction; a given backplane may do both. Check the vendor’s documentation for the exact part before buying U.2 drives for a U.3 machine. “Tri-mode” is the host-side name for the same idea, one port speaking SAS, SATA or NVMe with detection pins deciding: bay flexibility, not simultaneous bandwidth.
EDSFF is where the industry actually went
| Form factor | Size | Thickness and power | Lanes | Spec |
|---|---|---|---|---|
| E1.S | 31.5 x 111.49 mm | 5.9 mm/12 W to 25 mm/25 W | x4, x8 | SFF-TA-1006 |
| E1.L | 38.4 x 318.75 mm | 9.5 mm/25 W, 18 mm/40 W | x4, x8 | SFF-TA-1007 |
| E3.S | 76 x 112.75 mm | 7.5 mm/25 W, 16.8 mm/40 W | x4, x8, x16 | SFF-TA-1008 |
| E3.L | 76 x 142.2 mm | 7.5 mm/40 W, 16.8 mm/70 W | x4, x8, x16 | SFF-TA-1008 |
E1.S is the 1U replacement for M.2 22110. E1.L, the “ruler”, maximises capacity per rack unit. E3.S replaces U.2 in 2U at roughly the footprint of a 2.5-inch drive, and E3.L exists for parts that draw real power. All four share SFF-TA-1002, a protocol-agnostic connector in 1C, 2C and 4C sizes of 56, 84 and 140 contacts for x4, x8 and x16. The case for EDSFF is that power envelope, 12 W to 70 W with defined airflow profiles against M.2’s unenclosed 3.3 V module, plus hot-plug as a design goal. That envelope is also the warning for a second-hand buyer: the bay is assumed, and so is the airflow behind it. Buy EDSFF only with the chassis, or after confirming both.
Add-in cards, and cables that are components
An add-in card is either one controller on a x4 or x8 board, or a carrier holding two or four M.2 sockets. A passive carrier is wiring: without x4x4x4x4 bifurcation only the drive on lanes 0 to 3 enumerates. Bifurcation is set by firmware at power-on, before link training, and needs the root complex, the board’s routing and a BIOS option to agree. A carrier with an onboard switch needs no bifurcation support and costs what the switch costs.
Internal cabling is SlimSAS SFF-8654 (Rev 1.2, 27 April 2018), 0.6 mm pitch, 38 positions for 4X and 74 for 8X, an internal SAS part the industry repurposed for PCIe; and MCIO SFF-TA-1016 (Rev 1.3, 15 November 2024), 0.6 mm pitch, 38, 74, 124 or 148 contacts, with hybrid plugs that terminate a cable straight into an EDSFF E1 or E3 connector. PCI-SIG’s CopprLink specifications (1 May 2024) put numbers on reach: 32.0 and 64.0 GT/s internally over the SFF-TA-1016 form factor to a maximum of 1 m, the same rates externally over SFF-TA-1032 to 2 m.
SFF-8654 Rev 1.2 removed its speed characteristics and states no line rate. SFF-TA-1016 does set mated-connector loss and crosstalk limits for rates from 25 GT/s to 112 GT/s, but it puts cable assembly signal integrity out of scope and names no PCIe reach, so the rates above belong to CopprLink rather than to the connector form factors CopprLink borrows. A Gen4-era cable does not become a Gen5 cable by fitting the socket, because the channel is budgeted bump to bump. Astera Labs puts the total insertion loss budget at 32 GT/s at 36 dB, taken at the 16 GHz Nyquist frequency, and describes roughly 10 to 20 per cent held back as design margin, which is a rule of thumb and not a specified value. Subtract the second from the first and the remainder is this article’s arithmetic, not a published pair of figures:
bump-to-bump loss budget, 32 GT/s 36.0 dB at 16 GHz Nyquist
less design margin, 10 to 20 per cent 3.6 to 7.2 dB (derived)
leaves both packages, the board,
connectors and cable to share 28.8 to 32.4 dB (derived)
A cable is a signal-integrity component from Gen4 upward, and a Gen5 cabled hop spends part of a budget it shares with both packages, the boards and every connector on the path. Cable is normally the cheaper medium per unit length, which is why it is used at all; what runs out is total reach, and the longer cabled Gen5 runs become a retimer application. The ceiling of two retimers per link is what caps how long a chain of riser, cable and backplane can get.
How the drives actually attach
The cheapest four-drive NVMe card is the commonest wasted purchase in this whole subject. A passive quad-M.2 carrier is wiring. It contains no PCIe device at all. Plugged into a board that cannot split its x16 slot, exactly one drive appears - the one on lanes 0 to 3 - and the other three are invisible, with no error message anywhere to say why. The card is not faulty and neither are the drives. The board never offered four links for the card to carry.
Bifurcation is a firmware decision made before link training
Bifurcation splits one root-port lane group into several independent links, each with its own root port, its own LTSSM and its own link training. It is not switching: nothing is routed, nothing is fanned out, and no lanes are created. An x16 slot set to x4x4x4x4 is four x4 links and nothing more. Common granularities are x16 to x8x8, x8x4x4 and x4x4x4x4, with x2 granularity on some server root ports.
The timing is what makes this firmware rather than something the operating system can arrange. The split is applied at power-on, before link training, so it lives in BIOS/UEFI and needs a reboot. Three things must agree: the CPU or root complex must support it, the board must route the lanes to that physical slot, and the firmware must expose the option. Miss any one of the three and there is no usable bifurcation.
Finding out is a manual-reading exercise, made harder by inconsistent naming. Intel server firmware exposes it per IIO stack, labelled IOU0/IOU1/IOU2 with values x4x4x4x4, x4x4x8, x8x8 and x16. AMD platforms expose a per-slot “PCIe lane bifurcation” item. Board vendors add their own names on top of both. If neither the board manual nor the BIOS menu names the slot and a split, assume the board cannot do it. The absence is the evidence. No card corrects it, and the drives you already bought sit there unenumerated, which is why the board manual comes before the third and fourth M.2 drives.
The second trap is the chipset. Extra M.2 sockets on a desktop board usually hang off the chipset, and that uplink is narrow: four PCIe 4.0 lanes on AM5, a Gen 4 link by design rather than Gen 5 lanes negotiated down, and DMI 4.0 x8 on Intel’s LGA1851. Either way it is shared with USB, SATA and the network interface. Four chipset sockets present sixteen lanes of demand across that one uplink, and the aggregate throughput of the array is then capped by the uplink rather than by the drives.
One NVMe drive = x4 lanes.
AM5 24 usable PCIe 5.0 lanes + a PCIe 4.0 x4 chipset link
16 to graphics leaves 8 = two drives at full width
LGA1851 20 PCIe 5.0 + 4 PCIe 4.0 = 24 CPU lanes
TR PRO 9000WX 128 PCIe 5.0 lanes
EPYC 9005 128 PCIe 5.0 lanes per socket
8 drives x4 = 32 lanes -> past every consumer socket
24 drives x4 = 96 lanes -> a high-lane server socket, or a switch
Switch cards: the option that works on any host
A PCIe switch is a real device - one upstream port, N downstream ports, routing TLPs. Buy one when you need more endpoints than the host has root ports, when you need per-port hot-plug, or when the host cannot bifurcate at all. Current parts: Microchip’s Switchtec PFX Gen 5 family comes in 28, 36, 52, 68, 84 and 100-lane parts, with virtual switch partitions and a hot- and surprise-plug controller on every port at 32 GT/s; the 52 ports its datasheet quotes are the family’s maximum, not a figure every variant reaches. Broadcom’s PEX89000 covers 24 to 144 lanes.
What a switch does not do is create bandwidth. Twenty-four drives at x4 behind an x16 uplink is 6:1 oversubscription, and that is a design choice rather than a defect: a capacity array where no two clients read at once is well served by it, a database that wants every drive at once is not. The costs are real - silicon price, tens of watts of heat that needs directed airflow, and a latency hop. Vendors describe the hop as “cut-through” without publishing a figure, so this guide will not print one.
| Element | Protocol aware | Resets the loss budget | Adds ports |
|---|---|---|---|
| Switch | Yes, routes TLPs | Yes, it re-originates the link | Yes |
| Retimer | Yes, joins the LTSSM | Yes | No |
| Redriver | No, analogue only | No | No |
Retimers, redrivers and long paths
A retimer is spec-defined, formalised with PCIe 4.0. It terminates one electrical segment and originates another, recovers the clock, participates in Detect, Recovery and equalisation, corrects lane-to-lane skew, and resets the jitter and insertion-loss budget. A maximum of two retimers per link is permitted, which is precisely why long cable plus backplane plus drive topologies fail at Gen5. A redriver is an analogue conditioner: no clock recovery, no protocol participation, no skew correction, and it amplifies the noise along with the signal.
The budget decides which you need, and the previous section already spent it: 36 dB bump to bump at 32 GT/s, less the 10 to 20 per cent of design margin, leaves a long backplane trace, a riser and a cabled Gen5 hop competing for one remainder. That is the entire case for a retimer, and the two-retimer ceiling is why riser plus cable plus backplane is a topology Gen5 refuses. A redriver in the same position amplifies a signal that has already lost its eye.
Tri-mode HBAs, and what they are actually for
Broadcom’s tri-mode MegaRAID line speaks SAS, SATA and NVMe on one card, and the 9600 product brief prints the number that matters: 240 SAS or SATA physical drives against 32 NVMe. What the parity path on that silicon costs is the next section’s argument, host link included. What decides the card’s reach is width: sixteen lanes are four x4 links’ worth, and anything past that reaches the drives through switching in the backplane, shared. Buy a tri-mode card for a mixed estate, bootability and managed enclosures, where SAS SSDs and enterprise pulls sit behind the same backplane as flash; do not buy one expecting it to give the drives’ own performance back. What physically mates with what is covered in SATA, SAS, NVMe and M.2.
Hot-plug and surprise removal
An M.2 socket is not hot-pluggable. It is screw-retained, 3.3 V only, has no
staged mating, and the slot usually does not even advertise itself as hot-plug
capable. U.2 on SFF-8639, U.3 on SFF-TA-1001, and EDSFF E1.S and E3.S are
designed for it, with precharge and presence pins that mate first - if the build
needs a drive pulled while the array is running, you are shopping in
U.2, not M.2. Beyond the connector, the slot must advertise
hot-plug in firmware, the OS needs native control granted through ACPI _OSC,
and bus numbers and MMIO windows must be reserved ahead of the hot-add or the
new device enumerates with no resources. Linux handles the removal in pciehp;
reads to a yanked device time out and return a fabricated all-ones response,
which drivers have to recognise as absence rather than data. On Intel platforms the
Volume Management Device (VMD), taken up under hardware RAID below, supplies
hot-plug, LED management and error isolation inside its own PCI domain, and
toggling it in BIOS changes how drives enumerate - never flip it
under an existing array. Surprise-removal races were still drawing unmerged RFC
patches on the linux-pci list in September 2026, so treat “surprise removal
works” as a claim about one kernel on one platform, not a property of NVMe.
The decision, in order
Drives Board bifurcates? Do this
------- -------------------- ------------------------------------------
1 to 2 irrelevant CPU-attached M.2 sockets, nothing else
3 to 4 yes, x4x4x4x4 passive carrier in the CPU x16 slot
3 to 4 no switched carrier, or single-drive risers
5 to 8 yes, lanes to spare two bifurcated slots, watch the chipset
5 to 8 no spare lanes switched card, accept the uplink as the cap
9 to 24 either server board, U.2 or E3.S backplane, switch
25+ either a chassis decision, not a card decision
Bifurcation first, because it is free. A switch when the board will not split, and priced honestly against what it adds in heat and oversubscription. A real server chassis when the drive count passes what either can carry, at which point the backplane, not the card, is the product you are buying.
Hardware RAID is mostly over for NVMe
The RAID-on-chip in a used server was sized for a shelf of SAS drives, and eight NVMe drives ask it for roughly eight times what it was built to move. That is not a firmware problem, and no setting fixes it.
The arithmetic the card cannot win
The design target of a 12 Gb/s SAS controller is visible in the drives it was drawn around. Taking the narrow-port figure from the 9300-8i user guide, which is an HBA of that generation rather than a RAID card but carries the same SAS-3 ports, and twelve drives as the illustration:
SAS-3 narrow port, per drive 1,200 MB/s
Twelve SAS SSDs 12 x 1,200 MB/s = 14.4 GB/s
PCIe 5.0 x4 NVMe SSD, sequential read ~14 GB/s (one drive)
Eight of them 8 x 14 GB/s = ~112 GB/s
The host link is the harder wall, and it is arithmetic anyone can check. PCIe throughput per lane after 128b/130b framing is 1.969 GB/s at Gen4 and 3.938 GB/s at Gen5:
Gen4 x16 host link 16 x 1.969 GB/s = 31.5 GB/s
Gen5 x16 host link 16 x 3.938 GB/s = 63.0 GB/s
Drives to fill a Gen5 x16 card 63.0 / 14 = 4.5
Direct x4 drives, 32 device lanes 32 / 4 = 8
Four to five current drives saturate the host link of a single x16 card, and a 9760W-32i carries two x16 device connectors, so it can attach eight at x4 and feed about half of them. Broadcom’s 9600 series comparison table states the rest plainly: 240 SAS or SATA physical drives, 32 NVMe, with a footnote pointing anyone who wants more at a sales representative. The 32 is reached through PCIe switching in the backplane, which the drives then share.
Then the controller’s own ceiling, from the same product brief:
| Broadcom 9600 series, MegaRAID | Published figure |
|---|---|
| JBOD path, 4K random read | 6M to 6.4M IOPS |
| RAID 0/1, 4K random read | 4.5M to 6.4M IOPS |
| RAID 5, 4K random write | 900K to 1.1M IOPS |
The parity figure is about one sixth of the JBOD figure through the same card. Broadcom publishes no RAID 5 read number, so that ratio is a random write set against a random read, not one workload measured two ways.
Every byte through the ASIC
The queue model set out at the top of this guide - tens of thousands of queues against AHCI’s single 32-slot command list, and at most one MMIO write to submit - exists so that each core can talk to the drive without a shared intermediary.
A RAID-on-chip puts that intermediary back. Every byte of a RAID volume is staged through the controller’s memory and every command is serialised by its firmware. Hardware RAID for SAS moved work off a weak host CPU; hardware RAID for NVMe reintroduces the middleman the command set was written to remove.
Tri-mode, and the ceiling it admits
The Gen5 answer is real hardware: the 9700 series, 9760W-32i and 9760W-16i, x16 Gen5 host, SFF-TA-1016 connectors. Two changes are genuine architecture. The cache is integrated into the ROC, so the adapter “no longer requires onboard DDR for data caching” where previous generations needed up to 18 separate DDR devices, and the energy backup is an embedded supercap on the card (CVPM33) rather than a tethered module. Broadcom claims a 2x increase in random reads and a 5x increase in random write RAID performance over previous generations.
The 9700 brief prints no IOPS figures and no maximum NVMe drive count, deferring to techdocs. This guide will not supply one. Where the previous brief led with those rows and the new one omits them, the omission is the thing to notice, not a gap to fill with a guess.
VMD, VROC, and the AMD equivalent
Intel’s answer is not a card. VMD is a CPU feature, a secondary PCI host bridge
that pulls a set of root ports into its own domain and supplies surprise
hot-plug, LED management and error isolation; the Linux driver has been in the
tree since 4.5 and at drivers/pci/controller/vmd.c since 4.18. VROC sits on
top for bootable arrays, and on Linux VROC is md: IMSM external metadata driven
by mdadm, with PPL as the RAID5 write-hole mechanism. The data path is the
kernel’s, not a co-processor’s. What the platform adds is pre-OS configuration
and a boot volume the firmware understands.
The limits are platform limits. Pass-Thru is the unlicensed default and does no NVMe RAID at all; Standard (part VROCSTANMOD) adds bootable NVMe RAID 0/1/10 and Premium (VROCPREMMOD) adds RAID 5, up to 48 NVMe SSDs per platform and 24 per RAID 0/5 array. Toggling VMD in firmware changes how drives enumerate, so it is not a switch to flip under a live array. And as of 8 June 2026 VROC is no longer Intel’s to maintain: StorageReview reports stewardship moved to Graid Technology with a 24-month roadmap, rollout from Q3 2026 and UEFI-based licensing replacing hardware keys for new deployments. Date any sentence written about it.
AMD’s RAIDXpert2 is the same shape with a worse tail: configured pre-boot, driven
in Linux by the out-of-tree rcraid DKMS module, array at /dev/rcraid0, RAID 5
only on Threadripper-class parts. A closed binary shim tied to particular kernel
and distribution combinations is a root volume a kernel upgrade can leave
unbootable.
A native md superblock assembles on any Linux machine with mdadm. An array that needs a vendor driver assembles where that driver loads.
What it still buys, and the verdict
One thing, and it is a real one. A battery- or flash-backed write cache absorbs
partial-stripe writes and coalesces them into full-stripe writes, converting the
four-operation RAID5 read-modify-write - six for RAID6 - into a single N-wide
write, which on flash also cuts the write amplification charged against
the drives’ DWPD: a small RAID 5 write costs two drive writes and a RAID 6 one
costs three, where a full stripe costs N plus one writes for N blocks of data.
Broadcom’s policy set is honest about the dependency:
WT write through, WB write back, and AWB always write back. Write back
drops to write through when the energy backup is not ready; always write back
does not, which is the setting that loses the cache on a power cut. Without
backup the correct setting is
write-through, and write-through on parity pays the full RMW every time. The
software equivalent is a PLP SSD configured as an md write journal in
write-back mode.
For NVMe the default should be software, and hardware RAID should have to give a reason. The two reasons that survive scrutiny are a boot volume the firmware can find and an enclosure management stack the OEM already wired.
A buying note follows from all of it: when a used server is cheap, its controller is sometimes why. A 9600-class card in front of a shelf of SAS drives is still the right part for the job it was designed around; in front of eight NVMe SSDs it is the component the array waits on. Enterprise drive pulls covers IT mode and which cards pass drives through untouched.
Linux md is the right answer, and half its defaults are wrong for flash
The code is mature, the metadata is documented, no licence key stands between
you and your array, and any Linux kernel with md assembles it. The defaults
were chosen for disks with a seek penalty. One of them, the speed_limit_max
rebuild ceiling, sits at 200000 KiB/s and can hold a degraded NVMe array there
for most of a day. The other cost is not a default at all: parity writes carry a
read-modify-write multiplier that no setting removes.
The read path already knows what a non-rotational device is
drivers/md/raid1.c states the rule in a comment: “If all disks are rotational,
choose the closest disk. If any disk is non-rotational, choose the disk with
less pending request even the disk is rotational.” It is gated on
conf->nonrot_disks, set per member as the device is added, and RAID10 carries
the equivalent in has_nonrot_disk and min_pending. On flash md dispatches by
queue depth, not head position.
RAID10 here is a native personality, not mirrors glued under a stripe, so it
takes an arbitrary device count with near, far and offset layouts.
Mirroring costs half the capacity and buys a write path with no parity
arithmetic at all: no read-modify-write, no stripe cache, no write hole. On NVMe
that is usually decisive.
mdadm --create /dev/md0 --level=10 --raid-devices=8 --chunk=256K \
--write-zeroes /dev/nvme{0,1,2,3,4,5,6,7}n1
--write-zeroes (mdadm 4.3) zeroes every member so the initial resync becomes
unnecessary; mdadm’s page says the result behaves “as if –assume-clean was
specified”, while calling bare --assume-clean “not recommended”.
A partial-stripe write is four operations, and the unit is 4 KiB
For a write that does not fill a stripe, handle_stripe_dirtying() chooses
between a read-modify-write and a reconstruct write. For one dirty unit:
RAID5, N devices, N-1 of them data:
full-stripe write 0 reads + N writes
RMW 2 reads (old data, old P)
+ 2 writes (new data, new P) = 4 operations
RCW N-2 reads + 2 writes
RAID6, N devices, N-2 of them data:
RMW 3 reads (old data, old P, old Q)
+ 3 writes (new data, new P, new Q) = 6 operations
RCW N-3 reads + 3 writes
The kernel does not hardcode “four”: it counts the devices each strategy must
read and takes the cheaper, ties broken by md/rmw_level - 0 disables RMW, 1
enables it, 2 prefers it. RAID5 defaults to enabled; RAID6 only where the
selected raid6 implementation offers an xor_syndrome() routine, so RAID6
small-write behaviour is architecture-dependent.
The re-read is 4 KiB, not the chunk. md’s stripe head unit is
DEFAULT_STRIPE_SIZE, defined as 4096 in raid5.h, so a 4 KiB host write reads
4 KiB of old data and 4 KiB of old parity, then writes 4 KiB of each. The
commonly repeated claim that a small write re-reads the whole 512 KiB chunk is
wrong. The multiplier is real all the same: a partial-stripe write puts roughly
twice the host bytes on the drives for RAID5 and three times for RAID6, before
the filesystem’s amplification and the FTL’s. That is a charge against DWPD, and
SSD endurance shows what it does to a rated figure.
Three block sizes want to agree, and one is undocumented
The default chunk on --create is 512 KiB; RAID4, 5, 6 and 10 require a power
of two, minimum 4 KiB. md publishes the geometry as queue limits - io_min is
the chunk, io_opt the chunk times the data-disk count - which mkfs.xfs and
mke2fs read to set stripe unit and width. Do not hand-set sunit/swidth to
contradict what md exported.
The third size is the SSD’s indirection unit, the granularity at which the drive’s flash translation layer (FTL) maps logical blocks to physical ones. An IU larger than the write forces a read-modify-write inside the NAND, underneath md’s own. Consumer datasheets rarely print it, and this guide will not put a number on a drive whose datasheet does not state one; enterprise-class SSDs are the ones that publish it.
A write-intent bitmap does not close the write hole
md(4) is blunt: “interruption of write operations (system crash, etc.) to RAID456 array can lead to inconsistent parity and data loss”. Some members took the new data, the parity member did not.
The bitmap solves a different problem well: it records which regions may be out
of sync before the write is honoured, so an unclean shutdown resyncs only the
set bits and a briefly-absent member can be --re-added with only the dirty
regions recovered. Recent mdadm no longer adds one on its own: ask
with --bitmap=internal, whose chunk defaults to 64 MiB. It says nothing about
whether parity was consistent.
Drive-level power-loss protection does not close it either. PLP capacitors guarantee the drive’s own buffer reaches NAND; the hole is a cross-device atomicity failure, and no per-device guarantee addresses it.
Two mechanisms do. A journal (raid5-cache, Linux 4.4) writes data and parity
to a separate device first and replays on recovery; md/journal_mode defaults
to write-through, and write-back makes it a real write cache needing a PLP
drive under it. If the journal device dies the array goes read-only. PPL
(Linux 4.12, RAID5 only) stores partial parity - the XOR of the stripe chunks
this write did not modify - in the member metadata area, so no journal device is
needed. The kernel documentation states the cost: write performance “reduced by
up to 30%-40%”, a 64-disk maximum, bitmap and PPL cannot be used together, and
it “does not protect from losing in-flight data, only from silent data
corruption”.
mdadm --create /dev/md0 --level=5 --raid-devices=6 --consistency-policy=ppl /dev/nvme[0-5]n1
mdadm --create /dev/md0 --level=5 --raid-devices=6 --write-journal /dev/nvme9n1 /dev/nvme[0-5]n1
echo write-back > /sys/block/md0/md/journal_mode
The stripe cache and the RAID5 thread are the all-flash ceiling
md/stripe_cache_size is measured in pages per device and defaults to 256, with
memory_consumed = system_page_size * nr_disks * stripe_cache_size and an md(4)
warning that raising it too far ends in an out-of-memory condition; watch
md/stripe_cache_active first. The single md<N>_raid5 thread is the classic
bottleneck, and md/group_thread_cnt splits stripe handling across worker
groups - core-count and workload dependent, so measure. md/skip_copy drops the
payload copy into the cache page and sets stable-writes on the queue: a real
win, and a real requirement on the filesystem.
echo 8192 > /sys/block/md0/md/stripe_cache_size
echo 4 > /sys/block/md0/md/group_thread_cnt
echo 2000000 > /proc/sys/dev/raid/speed_limit_max
That last line matters most. speed_limit_max defaults to 200000 KiB/s, roughly
200 MB/s, and it is per device rather than per array. On NVMe it is the
difference between a rebuild measured in minutes and one that runs most of a day
degraded.
Discard reaches the drives on some levels and not others
| Level | TRIM handling |
|---|---|
| RAID0 | Native, raid0_handle_discard() |
| RAID1 | Passed through; write-behind not used for discard |
| RAID10 | raid10_handle_discard(), rewritten in Linux 5.13 |
| RAID4/5/6 | Off by default |
The RAID10 rewrite matters on older kernels: the patch series reports mkfs.xfs
on an md RAID10 falling from 4 min 40 s to under 1 s.
Parity levels ship with devices_handle_discard_safely = false, and the reason
is correctness. If a discarded region does not read back as zeroes, a later RMW
computes parity over garbage and a later rebuild reconstructs corruption. The
device-side fact that settles it is DLFEAT in Identify Namespace, where 001b
means deallocated blocks read as zeroes. Only 001b justifies the knob.
nvme id-ns /dev/nvme0n1 | grep -i dlfeat
echo Y > /sys/module/raid456/parameters/devices_handle_discard_safely
md can also only discard a whole stripe - raid5.c puts it as “It doesn’t make
sense to discard data disk but write parity disk” - so a five-drive RAID5 at a
512 KiB chunk works out at a 2 MiB granularity and drops anything smaller. Treat
the rest as version-dependent: RAID5 discard validation was reworked in 7.2-rc1,
and that rework shipped with a regression in which a discard issued during a
reshape never completed. The fix landed in 7.3, and the sequence is
recent enough that the release notes for the kernel you are running are worth
more than this paragraph.
The monitor is the part nobody configures until it is too late
md degrades quietly. It keeps serving reads, and nothing says so out loud. A month later the second drive goes.
mdadm --monitor --scan --daemonise --test
echo check > /sys/block/md0/md/sync_action
cat /sys/block/md0/md/mismatch_cnt
mdadm /dev/md0 --replace /dev/nvme3n1 --with /dev/nvme9n1
--test sends a message at start-up, the only way to prove the alert path
rather than assume it. Set MAILADDR or PROGRAM in mdadm.conf, whose path
is distribution-dependent, and confirm the packaged unit is enabled. check
scrubs without repairing and counts into mismatch_cnt. --replace --with
swaps a suspect member while it can still be read, so the array never goes
degraded - the right move once a drive starts warning you in SMART, and
what SMART tells you and what it cannot covers
which attributes to believe.
ZFS on NVMe: what genuinely works
Two of the decisions below are permanent for the life of a vdev, and the single most consequential tunable is none of the knobs that forum posts paste at you. ashift is fixed when the vdev is created, RAIDZ geometry can only be added to, and recordsize decides how much work a small write costs. Almost everything else here is reversible on a live pool.
ashift: the one number you cannot change
ashift is the base-2 logarithm of the smallest unit ZFS will write to a
vdev. ashift=12 means 4 KiB. It is fixed at vdev creation and no command
changes it afterwards; the remedy is to destroy the vdev and rebuild from a
backup or a send stream.
Set it too low - ashift=9 on a drive whose internal write unit is 4 KiB - and every small write becomes a read-modify-write inside the SSD, invisible to the host and paid for in latency and flash wear. Set it too high and every block, metadata included, is padded up to that size.
NVMe namespaces frequently report 512-byte logical blocks by default while
also supporting a 4096-byte LBA format. nvme id-ns /dev/nvme0n1 lists the
supported formats with a relative performance ranking for each. ashift=12 is
the defensible default: it matches every drive with a 4 KiB format, and a
device with a larger internal page is no worse served by it than by any other
host writing 4 KiB. ashift=13 is justified only where the
vendor documents an 8 KiB write unit, and it wastes space on small blocks and
on metadata.
Mirrors or RAIDZ, and why flash tilts the argument
Random IOPS scale with the number of top-level vdevs, not with the number of drives. A RAIDZ vdev of any width serves a small random read from its data columns as one logical operation, so eight drives in a single RAIDZ2 vdev deliver roughly one vdev’s worth of random-read IOPS. The same eight drives as four mirrors deliver four vdevs’ worth, and each read can come from either side of its mirror.
Resilver is the second argument, and it is the one people meet late. OpenZFS
can rebuild a mirror or a dRAID vdev sequentially - zpool replace -s -
streaming allocated regions in device order; checksums are not verified on the
way, so a scrub starts when the resilver finishes. dRAID is RAIDZ with the spare
capacity spread across every member instead of parked in an idle drive, and it
returns under rebuilds below. zpool-replace(8) is blunt
about the limit: sequential reconstruction is not supported for raidz. A RAIDZ
vdev performs a healing resilver instead, walking the block pointer tree,
which is a metadata-driven random-read workload. Flash removes the seek
penalty that makes that catastrophic on spinning disks; it does not change the
fact that the order is dictated by the tree rather than by the device.
Mirrors also grow and shrink a pair at a time, and a top-level mirror vdev can be removed outright, provided the pool holds no RAIDZ or dRAID vdev, every top-level vdev has the same ashift, and the keys for any encrypted datasets are loaded. A RAIDZ vdev cannot be removed at all. Against that, mirrors cost 50 per cent of raw capacity. A home pool holding mostly cold data reaches a different answer from a VM host, and the drive-buying half of that decision is in what differs from a desktop.
RAIDZ expansion, described accurately
OpenZFS 2.3 added zpool attach for RAIDZ vdevs: one disk at a time, online,
with a background reflow that relocates every allocated sector into the new,
wider layout. What the reflow does not do is re-stripe the data.
Existing blocks keep the data-to-parity ratio they were written with. The
zpool-attach(8) man page states it plainly: old blocks retain their old
data-to-parity ratio, redistributed across the larger set of disks, and new
blocks are written with the new one. A 5-wide RAIDZ1 expanded to 6-wide
therefore stores old blocks at 4 data to 1 parity and new blocks at 5 to 1.
The new drive’s capacity does appear, but less of it is usable than naive
arithmetic predicts, because the old parity overhead is still on disk:
5-wide raidz1, 1 MiB of user data: 1 MiB x (5/4) = 1.25 MiB occupied
6-wide raidz1, same data rewritten: 1 MiB x (6/5) = 1.20 MiB occupied
the reflow alone recovers none of that 0.05 MiB
The accounting lags as well: the vdev’s assumed parity ratio does not change, so the man page warns that slightly less space than expected may be reported for newly written blocks. The only way to collect the better ratio is to rewrite the blocks: copy the files in place, or send and receive the dataset to a fresh one. Expansion also does not change the RAIDZ level - a raidz1 stays a raidz1, however wide it gets, and a 10-wide raidz1 is a bad place to be during a resilver.
recordsize and volblocksize: the consequential tunable
recordsize is a maximum, not a fixed size. A file smaller than recordsize gets a single block rounded up to the sector size; the default maximum is 128 KiB. volblocksize is not a maximum. A zvol’s block size is fixed at creation, and OpenZFS 2.2 raised its default from 8 KiB to 16 KiB.
Because ZFS is copy-on-write and checksums whole records, a write smaller than the record forces the entire record to be read, merged and written somewhere new:
recordsize=1M, application writes 8 KiB at a random offset
read 1 MiB (unless the record is already in ARC)
write 1 MiB to newly allocated space
amplification against NAND: 1048576 / 8192 = 128x
recordsize=16K, same 8 KiB write
read 16 KiB, write 16 KiB, amplification 2x
That arithmetic is the whole argument. Large records for data written once and read in bulk; a record close to the page the application rewrites for anything doing random overwrites, which points at 8 KiB for PostgreSQL and 16 KiB for InnoDB, matching their default page sizes. Treat those two as a starting point rather than a measured result: with compression on, a larger record sometimes wins on PostgreSQL, and the workload decides. The same trap bites zvols on RAIDZ: a small volblocksize on a wide RAIDZ vdev at ashift=12 loses far more to parity and sector padding than the nominal ratio suggests.
Compression is a throughput win, not a tax
The mechanism is arithmetic, not opinion: compressed records move fewer bytes
across PCIe and write fewer bytes to NAND, so a pool that compresses 1.5:1
reads and writes about a third less. compression=on selects lz4 wherever the
lz4_compress feature is enabled, which is every pool a current release
creates, and lz4 gives up early on data that will not compress, so the cost on
already-compressed media files is small. The ARC, ZFS’s adaptive replacement
cache in RAM, holds records in their compressed form as well, so the same RAM
holds more. zstd (OpenZFS 2.0)
buys ratio with CPU; at NVMe speeds the CPU is what runs out first, which
makes lz4 or a low zstd level the honest choice for a pool expected to stream
gigabytes per second.
The special vdev: the best-value NVMe trick here
A special allocation class vdev pulls all metadata - the block pointer tree,
dnodes, spacemaps - off the main vdevs. Set special_small_blocks on a
dataset (default 0, meaning off) and data blocks at or below that size land
there too. On a hybrid pool this converts the metadata random-read workload,
the thing that makes a big hard drive pool feel slow on directory
traversal and scrub, into NVMe reads, for the price of two small
M.2 drives.
A special vdev is not a cache. Lose it and the pool is gone. Mirror it to at least the redundancy of the data vdevs. When it fills, new allocations spill back to the normal vdevs rather than failing.
SLOG and L2ARC are usually the wrong purchase on all-flash
The ZIL records synchronous writes only, and it is read only during import after a crash. A separate log device earns its slot when it is meaningfully faster or safer than the pool vdevs, which on an all-NVMe pool it generally is not. A consumer NVMe SLOG in front of enterprise SSDs with power-loss protection is a downgrade. The exception is real and narrow: a small, high-endurance, PLP-equipped, low-latency device in front of consumer QLC drives carrying an NFS, iSCSI or database workload that is genuinely sync-heavy.
L2ARC carries a RAM cost, because every buffer it holds needs a header in ARC, so a large L2ARC shrinks the ARC that was serving hits for free. On an all-flash pool the second tier is the same class of device as the first tier’s backing store. Persistent L2ARC (OpenZFS 2.0) survives reboot, which narrows the case rather than widening it.
autotrim, and what sync actually promises
autotrim is off by default. Without trim the SSD’s garbage collector cannot
know which LBAs ZFS has freed, and write amplification climbs as the pool
fills. zpool set autotrim=on batches frees in the background; zpool trim
runs a full pass. Ranges below zfs_trim_extent_bytes_min, 32 KiB by default,
are skipped unless they fall inside a larger range that was chunked.
sync=standard honours what the application asked for. sync=always makes
every write synchronous. sync=disabled ignores fsync: the pool stays
consistent, because each transaction group commits atomically, but
acknowledged writes can vanish back to the last commit, and zfs_txg_timeout
defaults to 5 seconds. Most benchmarks “fixed” by sync=disabled were
measuring precisely what the application had asked for.
ARC sizing is a RAM question
ARC is where ZFS read performance comes from. The default cap has moved: with
zfs_arc_max=0, current OpenZFS documents the limit as the larger of all
system memory minus 1 GiB and 5/8 of system memory, where older releases on
Linux took half of RAM, and distributions ship their own override in
/etc/modprobe.d. Read zfs(4) for the version you are running rather than a
forum figure. ARC is not the page cache and it does not appear as free memory.
The “1 GB of RAM per TB of pool” line is dedup-era folklore and is not a
requirement. Two things are true instead: a metadata-heavy workload wants ARC
for the same reason it wants a special vdev, and compressed ARC means each
byte of RAM holds more than a byte of data. For a pool of several NVMe drives
serving many clients, the cheapest performance left on the table is usually
more DDR4 or DDR5, not another drive.
Where ZFS costs more than it gives
The previous section made the case. This one is the bill, and it arrives in four currencies: capacity, endurance, CPU and flexibility.
On a ten-wide RAIDZ2 vdev, an 8 KiB block occupies 24 KiB. Not the 10 KiB the 8+2 geometry implies. Three times the logical size, permanently, and the pool will not mention it until it is nearly full.
The padding is structural, not a tuning mistake
The allocator is vdev_raidz_psize_to_asize() in OpenZFS’s vdev_raidz.c,
named vdev_raidz_asize() in older trees, and it is three lines:
D = ceil(psize / 2^ashift) data sectors
P = nparity * ceil(D / (cols - nparity)) parity sectors
asize = roundup(D + P, nparity + 1) * 2^ashift
The third line is the one nobody plans for. Every allocation rounds up to a
multiple of nparity + 1 sectors so that freeing it cannot leave a hole too
small to hold the smallest legal block. Those pad sectors carry no data; OpenZFS
may issue an optional write across them to keep a stripe contiguous, but the
capacity is gone either way.
Why 4TB shows as 3.64TB sets out the separate
decimal-against-binary gap and the ZFS slop reserve, which are different
losses from this one. Here is what it costs a flash array, at
ashift=12, ten drives, RAIDZ2:
4 KiB block D=1 P=2*ceil(1/8)=2 3 -> roundup(3,3)= 3 sec = 12 KiB 3.00x
8 KiB block D=2 P=2*ceil(2/8)=2 4 -> roundup(4,3)= 6 sec = 24 KiB 3.00x
16 KiB block D=4 P=2*ceil(4/8)=2 6 -> roundup(6,3)= 6 sec = 24 KiB 1.50x
128 KiB block D=32 P=2*ceil(32/8)=8 40 -> roundup(40,3)=42 sec = 168 KiB 1.31x
The 1.25x that the width promises is an asymptote approached only by large
records. A zvol at volblocksize=8K, which was the OpenZFS default until 2.2
raised it to 16K, returns 33 per cent of raw capacity before compression, not
80. That is also why zpool list and zfs list disagree: the first counts raw
sectors including parity and padding, the second counts what a dataset can store.
Write amplification stacks, and the layers multiply
Three multipliers sit on top of each other, none aware of the others.
Copy-on-write is first. ZFS cannot modify a record in place, so a 4 KiB overwrite inside a 128 KiB record forces the whole record to be read, modified and rewritten elsewhere. RAIDZ padding is second, from the arithmetic above. The drive’s own flash translation layer is third: NAND programs by page and erases by block, and its garbage collector adds a multiplier that the host never sees.
4 KiB application overwrite, recordsize=128K, 10-wide RAIDZ2, ashift=12
copy-on-write 4 KiB -> 128 KiB record rewritten 32x
RAIDZ allocation 168 KiB written to the vdev 1.31x
------------------------------------------------------------
allocated across the vdev 42 KiB per KiB of application 42x
... and the FTL's own amplification then applies on top
Endurance budget, ten 1.92 TB drives each rated 1 DWPD over 5 years:
1.92 TB * 365 * 5 = 3504 TB per drive, 35040 TB over the ten
35040 / 42 = 834 TB of application data
same sum at recordsize=16K (6x): 5840 TB
Lowering recordsize to 16K buys back seven times the endurance on that
workload and costs eight times as many block pointers. Neither number can be
found by reading a spec sheet; both follow from the allocator. How those DWPD and
TBW ratings are defined, and what a used drive’s SMART counters reveal about how
much of one has already been spent, is the subject of
SSD endurance - size the drive against the amplification you
have derived, not against the host writes you expect. Parity arrays on flash are
where the extra headroom of enterprise stock stops being optional; the
catalogue’s enterprise-class SSDs are the place to
compare price per terabyte.
Fragmentation, and what actually happens past eighty per cent
ZFS cannot defragment free space, and never could. The FRAG column in
zpool list reports free-space fragmentation, not file fragmentation. OpenZFS
2.3.4 added zfs rewrite, which reallocates a file’s blocks and so can undo the
second; the first still ends at a zfs send into a fresh pool.
Behind the eighty per cent rule of thumb sits the metaslab allocator. Within a
metaslab it uses first fit while there is room and falls back to best fit when
there is not, the switch governed by metaslab_df_free_pct, which defaults to 4
per cent free, and metaslab_df_alloc_threshold, which defaults to a largest
free segment of 128 KiB. Both thresholds are per metaslab, not per pool, so
eighty per cent is a margin operators choose rather than the point at which the
allocator changes gear. Best fit walks a size-sorted tree and costs more
CPU per allocation. The compounding part is worse than the CPU: as long runs
disappear, large allocations fail more often and ZFS falls back to gang blocks,
which are several smaller allocations and so pay the padding above several times
over. Capacity loss and allocation cost feed each other. ZFS holds back slop
space as well, governed by spa_slop_shift, precisely because a copy-on-write
filesystem needs free space in order to delete a file.
CPU and memory
On an NVMe array the CPU is the ceiling, and ZFS finds it first. fletcher4
is what checksum=on currently selects and it is vectorised, with the
implementation selectable through zfs_fletcher_4_impl; compression=on means
lz4 wherever the lz4_compress feature is enabled, and lz4 early-aborts on
incompressible data. Both are cheap per gigabyte. What binds is not the algorithm
but where it runs: compression and checksumming happen per record inside the ZIO
issue taskqs, on the host’s cores, so the ceiling is set by how much of that
pipeline a single writer can keep parallel rather than by the drives. Selecting
checksum=sha256, the algorithm dedup=on implies, or compression=zstd-19
moves that ceiling down a long way. The zstd project’s published Silesia table
quotes lz4 1.10.0 at 675 MB/s and zstd 1.5.7 level 1 at 510 MB/s, and gives the
speed of its top levels only as a chart rather than a number, so treat any
single figure quoted for zstd-19 with suspicion. This guide will not print a
GB/s per core either, because the figure is meaningless without naming the exact
part; the method is to measure single-record throughput on your own core and
multiply by the taskq width.
Memory is more often misdescribed than short. The “1 GB of RAM per TB of pool”
rule is deduplication folklore that escaped its context. ARC is adaptive and
returns memory under pressure, but on Linux it is not page cache, so free
reports it as used rather than cached, and zfs_arc_max is a cap you set rather
than a hint. Deduplication is the genuine demand: at the commonly cited 320 bytes
per unique block held in core, a terabyte of unique 128 KiB records needs roughly
2.4 GB of DDT, and the same terabyte at 16 KiB records needs roughly 19.5 GB.
OpenZFS 2.3’s Fast Dedup reworked the on-disk layout and the eviction path, so
that figure should be rechecked against the running version rather than carried
forward.
What it still cannot do
| Operation | Status |
|---|---|
| Remove a mirror or single-disk top-level vdev | Possible since OpenZFS 0.8, and it leaves a permanent indirect mapping, held in memory, for the space that moved |
| Remove a RAIDZ top-level vdev | Not possible |
| Remove a top-level data vdev from a pool that contains a RAIDZ vdev | Not possible; cache, log and spare devices can still be removed |
| Shrink a vdev or a pool | Not possible |
| Widen a RAIDZ vdev by one disk | zpool attach, since OpenZFS 2.3 |
| Narrow a RAIDZ vdev | Not possible |
| Change parity level, RAIDZ1 to RAIDZ2 | Not possible |
Change ashift after vdev creation |
Not possible |
The one row that reads better than it behaves is the expansion, for the reason
already given: the reflow widens the vdev without re-striping, so reported free
space rises while the old data keeps the parity overhead it was written with.
Every row marked “not possible” has the same workaround, zfs send to a new
pool, which needs the capacity twice over and a maintenance window.
When to use mdadm and XFS instead
If the workload is 4 KiB to 16 KiB random overwrites - a database, a busy VM
store, a metadata-heavy queue - RAIDZ’s padding and copy-on-write amplification
shrink with a smaller recordsize but never reach zero, and ZFS’s own answer
is mirrors, which surrenders half the capacity outright. md RAID10 under XFS
costs two writes per
write and pads nothing. md RAID5’s small-write cost is two reads plus two writes
at the 4 KiB page granularity of md’s stripe cache, RAID6 three and three: a 2x
to 3x multiplier against the drives’ rated endurance, against the 42x derived
above.
Choose mdadm plus XFS when the data is reproducible, when the workload is small
random overwrites, or when the operator knows md and does not know ZFS. What
you give up is precise and not small: end-to-end checksums, snapshots, send and
receive, transparent compression, and the detection of silent corruption - md’s
check pass counts mismatches in mismatch_cnt but cannot say which drive is
wrong, and XFS checksums its metadata and not your data. Make that trade
deliberately, rather than drifting into whichever stack the tutorial you read
happened to use.
The other filesystems, and where each one stops
A filesystem gives you redundancy only if it manages more than one device
itself. btrfs and ZFS do. XFS, ext4 and f2fs do not: an external log or
journal device, and the several devices mkfs.f2fs -c concatenates into one
volume, add capacity or separation, never a second copy. Formatting XFS across
a single NVMe namespace and calling the result an array is a common and costly
error: what you have is one failure domain and nothing but a checksum
on the metadata alone. Redundancy for those three comes from md or LVM RAID
underneath, and everything the block layer cannot do - tell you which copy is
correct, snapshot cheaply, rebuild only the blocks in use - stays undone.
btrfs: the mirrored profiles are good, the parity ones are not
btrfs allocates in chunks and gives each chunk a profile, so data and metadata
can differ. raid1 is two copies on two distinct devices, raid1c3 three,
raid1c4 four, and raid10 stripes those mirrored pairs. None of them demands
equal-sized members: the allocator places copies on whichever devices have the
most unallocated space, which is why a home array of four mismatched drives is
comfortable in btrfs and a nuisance in md.
The payoff is the checksum. btrfs checksums data as well as metadata (crc32c by
default, with xxhash, sha256 and blake2 selectable at mkfs time), so a read
that fails verification is retried against the other copy and the bad copy
rewritten. md cannot do this. echo check > /sys/block/md0/md/sync_action
counts disagreements into mismatch_cnt and has no means of deciding which side
is right.
btrfs RAID5 and RAID6 carry a write hole, and the project documents it itself. The upstream status page marks RAID56 unstable, and the mount-options manual carries a known-problems list with the write hole on it, next to the advice to keep the feature to evaluation and testing rather than production. Read that as the authors’ own recommendation and keep anything you care about off it. The absence of a stability claim from the people who wrote the code is the evidence.
What btrfs is genuinely good at on flash is the rest: reflink copies, near-free
snapshots, send/receive backups that ship only changed extents, transparent
zstd compression, online device add and remove, and profile conversion by
balance - a single-device filesystem becomes a mirror without a reformat, and
a mirror shrinks. Use discard=async, which batches freed extents and
rate-limits the trim instead of issuing it synchronously in the commit path; the
btrfs manual gives it as the default since 6.2 on devices that support it, and
/proc/mounts will say what your kernel actually mounted with.
The trade is copy-on-write fragmentation on randomly overwritten files, meaning
VM images and database files. chattr +C on the containing directory turns COW
off for files created in it afterwards, and takes the data checksums and
compression with it.
LVM RAID: md underneath, with a better front end
lvcreate --type raid1|raid10|raid5|raid6 builds dm-raid, which drives the same
md personalities. The geometry, the read-modify-write cost and the parity write
hole are md’s, unchanged. What you buy is management: several logical volumes at
different RAID levels over one pool of physical volumes, online conversion,
lvconvert --repair, thin pools on top. The capability plain md has no
equivalent for is --raidintegrity y, which inserts a dm-integrity layer
beneath each RAID image and gives the stack per-block checksums, turning a
mismatch into a correctable read rather than a coin toss. dm-integrity can be
stacked under md by hand; what LVM sells is having it done for you. It is not
free - integrity metadata is written alongside every block - so budget the
endurance for it before you buy, and read
SSD endurance: TBW, DWPD, how worried to be first.
XFS on md or LVM: the boring high-performance answer
XFS is the default on RHEL and its derivatives, for good reasons. Allocation
groups let many cores allocate in parallel, which is the same shape as NVMe’s
many-queue submission model; delayed allocation and extents keep large
sequential writes contiguous; and mkfs.xfs reads io_min and io_opt from the
device and sets sunit/swidth to the array’s chunk and stripe width without
being told, whenever the layer beneath reports them, which md does and a good
many hardware controllers do not.
Its limits are specific rather than vague. Metadata is CRC-protected, data is
not, so a silently corrupted block is handed to the application looking correct.
And XFS shrinks only at the margin: xfs_growfs -D can hand back unused
space in the last allocation group, the man page records that this is the only
part implemented and that a filesystem cannot be taken down to a single group,
and nothing moves data away from the end of the device for you. Reflink gives
cheap file copies, not filesystem snapshots. It has no redundancy of its own, so
the entire write-hole question belongs to the layer beneath it.
ext4, and where it still wins
For a boot volume or a small mirror, ext4 remains the right answer: metadata
checksums via metadata_csum, on by default in current e2fsprogs; a fast,
extremely well-exercised fsck; a smaller in-memory footprint than the
copy-on-write options; an external journal device if you want one; and the thing
XFS all but refuses, an offline shrink with resize2fs. On a two-drive md RAID1
holding /, that combination is hard to improve on, and nothing more ambitious
adds a property a boot volume needs. It has no data checksums either, and it
does not scale across cores the way XFS does. It is the wrong choice for the
bulk pool and the right one for the pair of small drives the bulk pool boots
from.
f2fs: a real niche, and this is not it
f2fs is log-structured, written for devices whose FTL strongly prefers sequential writes: eMMC and UFS in phones, SD and USB flash. In front of an enterprise NVMe controller with DRAM and a mature FTL, much of that advantage is already being performed by the drive, and this guide will not put a number on what is left. It has no multi-device redundancy at all. Its serious modern application is zoned storage, where sequential writing within a zone is a requirement of the device rather than an optimisation. For an array, look elsewhere; for a zoned namespace device, look here.
bcachefs: the right mechanism, no longer in the kernel tree
The design answers the problem head-on: copy-on-write, checksummed data,
replicas= set per filesystem, directory or file, erasure coding, and native
tiering, with fast devices as foreground targets in front of slow background
ones and no bcache or dm-cache bolted underneath. It is no longer part of the
Linux kernel. It was merged in 6.7 flagged experimental, marked externally
maintained in the 6.17 MAINTAINERS file, and removed from the tree outright in
6.18, since when it has been distributed out of tree as a DKMS module with its
own tools package. The question is no longer which kernel carries it: check that
the module builds and loads against the exact kernel you intend to run, and that
whoever packages it for you will keep rebuilding it, on the day you build the
array. Storage you must keep for five years is a poor place to bet on a
distribution mechanism that has already changed twice.
The candidates, against the properties that matter
| Stack | Data checksums | Own redundancy | Snapshots | Shrink | Write hole |
|---|---|---|---|---|---|
| ZFS mirror or raidz | yes | yes | yes | limited, no raidz | no |
| btrfs raid1 / 1c3 / 10 | yes | yes | yes | yes | no |
| btrfs raid5 / raid6 | yes | yes | yes | yes | yes, documented |
| bcachefs with replicas | yes | yes | yes | yes | no |
| XFS on md RAID10 | metadata only | no | no | last AG only | none, no parity |
| XFS on md RAID5/6 | metadata only | no | no | last AG only | yes, unless a write journal, or PPL on RAID5 |
| ext4 on md RAID1 | metadata only | no | no | yes, offline | none, no parity |
| LVM RAID + integrity | yes | yes | yes | not while integrity is on | parity levels inherit md’s |
| f2fs, single device | metadata only | no | no | no | not applicable |
Read the table as one sentence. Decide the data checksum column first: five
of these nine stacks have it, and the three that get it from the filesystem -
ZFS, btrfs and bcachefs - are a mkfs decision you cannot revisit later, while
dm-integrity can be added to an LVM RAID volume that already exists with
lvconvert --raidintegrity y. Then decide the hardware: the
power-loss-protected enterprise parts an md write journal or a dm-integrity
layer really wants are in the SSD listings.
When local RAID stops being the answer
Everything to this point has treated the drive as the thing that fails. Past a certain size that stops being true. A twenty-four drive NVMe array in one chassis shares one power domain, one memory subsystem, one kernel and one motherboard, and md protects you against none of them. RAID6 across those drives survives two drive failures and zero chassis failures, and a chassis is a single object that can be gone for three days waiting for a board. The question that decides whether local RAID has run out is not how many drives the array can lose. It is what happens on the morning the box does not power on.
The threshold is a failure domain, not a capacity
Three things push a design out of one chassis, and only one is about size:
- The outage budget. If a day of unavailability is an inconvenience, one box with tested restores is the right answer. If the budget is minutes, no amount of parity inside that box reaches it.
- Fan-out. More hosts need the data than one server can serve, or the single export in front of the array becomes the thing that fails.
- Physical limits. The working set no longer fits the lanes, bays and thermal envelope of one machine at any sane drive price.
Most readers meet none of the three, and reason 3 alone is usually answered by a larger chassis and bigger enterprise drives, which cost a fraction of a cluster and need no new skills.
Crossing the line carries a cost no tuning removes. Once an acknowledged write must reach a second machine, its latency belongs to the network and to the software at both ends, not to the drive. Distributed storage buys availability and aggregate throughput. It does not buy latency, and a correct deployment gets slower per operation on the day it becomes a cluster.
Ceph, and what one OSD per drive assumes
Ceph’s BlueStore writes to the raw block device with no filesystem underneath, keeping metadata and its write-ahead log in RocksDB. The Luminous release notes made it the default for newly created OSDs in 2017, and on flash it matters because it removes the double-write path FileStore had.
The habit to question is one OSD per drive. An OSD is a single process with a bounded set of shards and threads: against a SATA SSD it saturates the device without trying, against a Gen4 or Gen5 NVMe drive it may not, which is the argument for putting two OSDs on one device. Ceph’s current hardware recommendations draw that line far more narrowly than the habit does, suggesting the split only for PCIe Gen4-and-later SSDs larger than 30 TB. The payoff is release-dependent and this guide will not put a number on it. The cost of the split is not release-dependent:
Ceph's documented default osd_memory_target = 4 GiB per OSD
24 NVMe drives, one OSD each = 96 GiB for OSDs alone
24 NVMe drives, two OSDs each = 192 GiB for OSDs alone
Ceph's own sizing rule is stricter than that:
total server RAM > number of OSDs * osd_memory_target * 2
so 192 GiB and 384 GiB for the two lines above, before MON, MDS,
page cache, or anything the node actually serves
CPU is the other half. Ceph’s hardware recommendations ask for one thread per HDD-backed OSD as a minimum and three as the recommendation, against four and six for an NVMe-backed one, and on fast flash the OSD goes CPU-bound before the drive does. Sizing a cluster on drive throughput and then finding the core count is the ceiling is the standard first mistake.
Latency is where expectations break. A replicated pool defaults to three copies, and that write path is not one hop:
Local two-way mirror, 4 KiB write:
one PCIe round trip, the slower of two drives tens of microseconds
Ceph replicated pool, size 3, same 4 KiB write:
client -> primary OSD hop 1
primary -> both replicas hop 2
replicas -> primary (commit acks) hop 3
primary -> client (ack) hop 4
plus a RocksDB commit on three nodes, and queueing in each OSD
Two network round trips before the software is counted, and the software is the larger term: a switch port adds a fraction of a microsecond, while the stacks at either end cost far more than the wire between them. A three-node cluster on NVMe will have worse single-threaded write latency than the mirrored pair it replaced. It wins in aggregate, across many clients at once, which is the workload it was built for.
Three copies also triples host writes before RocksDB and BlueStore metadata are counted, so the drives under a cluster are bought on DWPD and power-loss protection rather than price per terabyte. That arithmetic is in SSD endurance, and the stock here carrying a real duty-cycle rating is under enterprise SSDs.
NVMe over Fabrics: the one you could actually deploy
The other shape keeps the block device and moves it across the network. NVMe over Fabrics carries the same command set and queue model to a remote target, so the initiator sees an ordinary namespace.
| Transport | What the network must provide | Realistic here |
|---|---|---|
| NVMe/TCP | any Ethernet NIC and switch, no fabric config | yes |
| RoCEv2 | RDMA NICs both ends, PFC and ECN on every switch | rarely |
| iWARP | RDMA NICs; TCP in hardware, so loss is tolerated | narrow choice |
| FC-NVMe | a Fibre Channel fabric that already exists | only if it does |
RoCEv2’s requirement is the underestimated one. It wants near-lossless Ethernet, and one mis-set switch port rarely produces an obvious error - it shows up as pause-frame and discard counters, and as collapse under load. NVMe/TCP is the version most readers could deploy. Its host and target drivers have been in mainline Linux since 5.0, the target is configured through configfs, and no vendor software is involved. Its cost is a software network stack in front of a device whose enterprise datasheet 4 KiB read latency is commonly quoted well under 100 microseconds.
One warning, because it is the failure people hit first. Once the namespace appears on the initiator everything earlier in this guide applies to it - md, LVM, a filesystem - but a fabric timeout looks exactly like a dead drive, and md will eject the member and rebuild onto a disk that was never broken.
Disaggregation, and the part that does not scale down
Hyperscalers separate compute from flash for reasons that are all arithmetic. Drives stranded in a busy compute node are capacity nobody can use; flash and CPU depreciate on different schedules; a host reboot should not take its storage offline; and a pool hands out 800 GB where a physical drive hands out all of itself - buy a 15.36 TB part and something has to find a use for 15.36 TB. A Gen5 drive rated at around 14 GB/s of sequential read sits idle most of the time behind one host. It works because fabric latency is now small next to the drive’s own, and there is no home-scale version: with three machines there is nothing to pool.
SPDK is the ceiling, and it is not a home project
At the very top the kernel’s per-I/O cost becomes the limit. SPDK binds the drive
to a userspace driver, vfio-pci where an IOMMU is available and
uio_pci_generic where it is not, and runs a polled-mode driver: no interrupts,
no syscalls on the data path, and a core pinned at 100 per cent whether or not
there is work. The drive then vanishes from the kernel - no block device, no md, no
filesystem, none of the tooling in the rest of this guide - so the trade pays
only when the storage application is yours to write.
Build the cluster when a chassis outage costs more than a second chassis. Below that line, one good box, tested restores and a cold spare mainboard recover a total loss in hours, for a fraction of the money and none of the new latency.
Endurance and power-loss protection in an array
The drive that fails you in an array is rarely the one that dies. It is the one that was never rated for the workload an array produces. Endurance under stacked amplification and power-loss protection are the two properties a consumer NVMe drive quietly does not have, and neither is visible in a thirty-second benchmark.
What a DWPD rating already contains, and what it does not
Three layers sit between an application’s write and a NAND program. The filesystem adds metadata and journal. The RAID level turns one logical write into several device writes. The FTL then erases and rewrites blocks to reclaim space.
The RAID multiplier is the one derived under md above, charged at md’s 4 KiB stripe head rather than at the chunk: for writes smaller than a full stripe, parity RAID costs about 2x host bytes on RAID5 and 3x on RAID6. A full-stripe write is much cheaper, because the parity is computed from data the array already holds. What the sum below adds is what that multiplier does to a rating expressed in host bytes.
Application writes 1.00
Filesystem metadata and journal x1.1
RAID6 partial-stripe parity x3
--------------------------------------------------------
Host bytes written per application byte ~3.3
Drive: 1.92 TB, rated 1 DWPD over 5 years
rated host writes = 1 x 1.92 TB x 365 x 5 = 3,504 TB = 3.5 PB
application budget = 3,504 / 3.3 = 1,062 TB
effective rating = 1,062 / 3,504 = 0.30 DWPD
The FTL layer is deliberately absent from that sum, and the reason matters. A DWPD or TBW figure is expressed in host bytes, and the qualification already contains the drive’s own amplification under the JEDEC endurance workload the part was rated against. What it does not contain is your workload. An array kept 90 per cent full and written randomly can amplify inside the FTL harder than the qualification workload did, in which case 0.30 DWPD is the optimistic end of the range; a large-sequential workload can run the other way, because the JEDEC enterprise workload is itself severe. The figure is a starting point, not a floor. SSD endurance: TBW, DWPD, how worried to be works through the ratings themselves.
The SLC cache is a burst budget
Consumer TLC and QLC drives program a pool of blocks at one bit per cell, which is fast, then fold them back into their dense mode in the background. Most of that pool is dynamic: it is carved out of free blocks, so it shrinks as the array fills. When host writes outrun it, the controller is doing two jobs at once, accepting new data and folding old, and the drive falls to the native program rate of its dense mode minus the folding traffic.
An advertised sequential figure is a burst figure. An array meets the steady-state figure. A rebuild, a resilver, a scrub that repairs, or a bulk ingest is a sustained write to every member at once, so the whole array leaves its cache inside the same minute, and a parity array then runs at the steady-state rate of its slowest member. That is also the moment the array has the least redundancy it will ever have. Few consumer datasheets publish a sustained figure at all; the drives whose datasheets do publish one are telling you something by doing it.
What power-loss protection actually guarantees
PLP is a bank of capacitors holding enough energy to finish the NAND programs in flight and push the controller’s DRAM write buffer to media after the rail collapses. The guarantee is narrow and exact: data the drive has already acknowledged stays acknowledged. It does nothing for data still in the host’s page cache, and it does not close the RAID5/6 write hole, which is a cross-device atomicity failure and needs a journal, a copy-on-write filesystem, or, on RAID5 alone, md’s partial parity log.
The cost of not having it is paid on every flush. Identify Controller carries a
VWC field saying whether the drive has a volatile write cache at all, and
Feature Identifier 06h, read with nvme get-feature /dev/nvme0 -f 6, says
whether that cache is enabled. A drive with PLP either reports no volatile
cache or completes the Flush command straight away, because its buffer is
already power-safe. A drive without PLP has to program NAND first.
Single-threaded synchronous writer, latency bound.
The flush latencies below are assumed, not measured:
sync writes per second = 1 / flush latency
PLP drive, flush satisfied from a power-safe buffer
~20 us -> ~50,000 sync writes/s
no-PLP drive, flush waits on a TLC program
~800 us -> ~1,250 sync writes/s
Those are order-of-magnitude figures to show the shape, not datasheet values;
program times are part- and mode-dependent. The shape is the point. Every fsync,
every ZFS ZIL commit, every PostgreSQL commit with synchronous_commit on, and
every md RAID5 journal write in write-back mode lands on the second line. It
is why an array that benchmarks well asynchronously collapses as a database
host, and why the kernel’s md documentation warns that write-back journal mode
acknowledges a write as soon as it reaches the cache device, so that device
failing takes the data with it. PLP is close to a product-line property rather
than a tunable, so the honest way to shop for it is
to filter for it: enterprise-class SSDs. Then read the
datasheet for “enhanced power loss protection” or the vendor’s equivalent
phrase. Its absence from a datasheet is the answer.
Identical drives wear out identically
Mirror members take identical writes and parity members take near-identical ones, so eight drives from one order arrive at the same wear point in the same week. They also share a firmware image. HPE’s customer bulletin of November 2019 is the cleanest published example: a supplier firmware defect made certain SAS SSDs fail unrecoverably at 32,768 power-on hours, which HPE put at 3 years, 270 days and 8 hours, unless firmware HPD8 had been applied first. HPE published no mechanism beyond the defect itself; 32,768 is 2^15, so a counter overflow is an inference, not HPE’s account. HPE’s warning was that drives put into service together would fail close to together. Same model, same batch, same firmware is one failure domain wearing several badges.
Split the order across two vendors or two production batches, check that the serial ranges are not contiguous, or run two models of the same capacity - a list of enterprise-class drives of any kind ranked by price per terabyte makes the second model easy to pick. The price is real: the array runs at the slower model’s steady state and sizes to its smallest member. Keeping the spare cold is the cheapest way to stagger power-on hours.
Over-provisioning is free if you take it before you partition
Spare area is the room the FTL works in, and more of it lowers garbage-collection amplification. The gap between binary NAND and decimal labels is where the standard figures come from, and the same silicon is sold at three prices because of it.
NAND on the dies: 1024 GiB = 1024 x 2^30 = 1,099,511,627,776 B
spare ratio = (NAND - exposed) / exposed
exposed 1,000,204,886,016 B -> 9.9 per cent consumer 1 TB
exposed 960,000,000,000 B -> 14.5 per cent 1 DWPD part
exposed 800,000,000,000 B -> 37.4 per cent 3 DWPD part
Those ratios are arithmetic from a drive built out of exactly 1024 GiB of NAND,
the usual construction for that family rather than a spare-area figure any
vendor publishes. A reader can take the middle or top ratio without paying for
it: run blkdiscard over the whole namespace so the FTL knows every block is
free, create a partition short of the end, and never grow the array into the
gap. Where the drive supports NVMe Namespace Management, deleting and recreating
the namespace at a smaller size is cleaner still, because the reserved LBAs are
never exposed to the host at all.
Read the wear counters before the money moves
The NVMe SMART / Health Information log, page 02h, carries Percentage Used, Data Units Written, Power On Hours, Unsafe Shutdowns, and Available Spare against its threshold. Data Units Written counts thousands of 512-byte units, so the conversion is easy to get wrong by three orders of magnitude.
Data Units Written 2,734,375,000 units
Bytes per unit 1000 x 512 = 512,000
Host bytes written 2,734,375,000 x 512,000 = 1.4 x 10^15 B = 1.4 PB
Against a 3.5 PBW rating: 1.4 / 3.5 = 40 per cent of the rating consumed
If Percentage Used reads 8 next to that, the two are not automatically in conflict: Percentage Used is the vendor’s estimate of life consumed from actual wear, and a workload gentler than the qualification workload leaves it well below the share of the rated host writes already spent. A gap that wide is still worth a question, because a reset counter looks the same from outside. Both are host-side figures in any case; the NAND-side number is not in the standard log. The OCP Datacenter NVMe SSD specification adds a physical media units written field in an extended SMART log, so a drive built to it reports both numbers, and the ratio of the two is the drive’s lifetime write amplification, one of the most useful numbers a used enterprise drive can hand you. What SMART tells you and what it cannot covers the rest of the log.
Heat and power
Two things quietly hold a home NVMe array below its datasheets: heat, and an idle draw several times what the builder budgeted. Neither shows up in a thirty-second benchmark, which is roughly why neither shows up in most build guides.
What a hot drive actually does
An NVMe controller reports Composite Temperature in the SMART / Health Information log, log identifier 02h, in Kelvin, alongside up to eight individual sensors. Identify Controller carries the warning and critical composite temperature thresholds, WCTEMP and CCTEMP; a third pair, TMT1 and TMT2, is set by the host through the Host Controlled Thermal Management feature, identifier 10h.
There are two gears, and the wording of the SMART counters gives them away. Past TMT1 the controller has “transitioned to lower power active power states or performed vendor specific thermal management actions while minimizing the impact on performance”. Past TMT2 it does the same thing regardless of the impact on performance. The first gear costs throughput. The second parks the drive in a low active power state and leaves it there for as long as the heat has nowhere to go, which in a sealed case with no directed air is indefinitely.
The drive keeps a receipt. Warning and Critical Composite
Temperature Time count minutes, the two Thermal Management Temperature
Transition Counts count events, and Total Time For Thermal Management
Temperature 1 and 2 count seconds. Of those last four the specification says
that a value of zero means either that the transition never happened or that
the field is not implemented, so a row of zeroes is not by itself proof of a
cool life. All of it comes out of nvme smart-log, and on a second-hand
drive it reads alongside the wear fields in
What SMART tells you and what it cannot.
Four drives on a card in still air
A quad M.2 carrier puts four modules side by side behind one plate, in an x16 slot that in most ATX cases sits below the graphics card, where front-to-back airflow has already been spent. Each module is a heat source and a wall for its neighbour. The inner drives set the throttle point for the whole array, and nothing in the array’s own reporting says so. When a home array disappoints, suspect the carrier and the still air around it before the filesystem or the RAID level. For a drive that has to live in a bay like that, pick on sustained power draw and a metal label rather than peak sequential figures: NVMe in the M.2 form factor.
The controller is hot, the NAND may be too cold
Two different problems share one board, and conflating them is why heatsink advice is so often wrong.
The controller trips Composite Temperature. NAND has the opposite pair of failure modes. Hot NAND loses retention, charge loss being an Arrhenius process, and JEDEC’s JESD218 specifies power-off retention for client SSDs as one year at 30°C and for enterprise SSDs as three months at 40°C at the rated end of endurance (see SSD endurance). Run an array hot and those figures are optimistic.
Cold NAND is the counter-intuitive half. Programming at low temperature distorts the threshold voltage distribution, and the IEEE device literature on 2D and 3D arrays describes a cross-temperature effect in which the raw bit error rate rises with the gap between the temperature a block was written at and the temperature it is read at. An array that writes at 65°C through a long ingest and reads back at 25°C a week later is that case. The target is not a cold drive, it is a narrow spread.
Heatsinks, airflow, and the form factors that solved this
A heatsink is a thermal resistance to ambient plus a thermal capacitance. The capacitance decides how long the drive runs before throttling; the resistance decides where it settles. Benchmarks are short and reward the capacitance; arrays are sustained and are paid only by the resistance term, which in still air falls with fin area rather than with mass, and slowly. Mass buys the benchmark, air buys the array.
A 2.5-inch U.2 drive in a hot-swap cage is the easier version of the same problem: a metal enclosure, a fan pulling through the cage, and a 12 V rail. EDSFF then made airflow a specification parameter. SFF-TA-1023, the EDSFF thermal characterisation specification, rates a device by the air it needs: a MaxTherm level, the minimum airflow at a stated approach air temperature for which the device runs at its full rated power, and DTherm levels below it, each naming the airflow for which it runs at a stated reduced performance. SNIA’s E1.S figures put the power envelope at about 12 W for the 5.9 mm case and 25 W at 25 mm, because thickness is what buys the spreader and the channel.
The power arithmetic
M.2 has no 12 V rail; it is a 3.3 V part. The CEM specification gives a PCIe add-in card slot 3.0 A on 3.3 V, which is 9.9 W. The rest of a x16 slot’s 75 W comes from 12 V at up to 5.5 A, and only once power management configuration lifts it past the 25 W the slot starts at; a x4 or x8 slot stays at 25 W.
Four M.2 drives, Samsung 990 PRO 2 TB datasheet figures (Rev 2.0)
average active read 4 x 6.1 W = 24.4 W
that current at 3.3 V = 7.4 A
PCIe slot 3.3 V budget = 3.0 A
Those are Samsung’s average active figures under its own stated test conditions, not peaks, and they come from the current 990 PRO datasheet, Revision 2.0 of July 2023, which revised the power values of the October 2022 original. On that arithmetic a carrier that takes 3.3 V straight from the slot, with no 12 V to 3.3 V regulator of its own, has nothing like the headroom for four drives active at once. It can feed one, which is what a benchmark asks for.
Eight drives, all active, continuous
consumer M.2 at 5.5 W write each = 44.0 W
enterprise U.2 at 15 W read each = 120.0 W
Solidigm’s brief quotes 15 W average active read for the 3.84 TB D7-P5520. A hundred and twenty watts of continuous drive load is a different power supply from the one sized for eight spinning disks, and unlike disks it never spins down.
Idle power is where the NAS budget actually goes
The 990 PRO datasheet quotes idle with APST on at 55 mW for the 2 TB model and 50 mW for the 1 TB, and 5 mW in L1.2, and stops its operating range at 70°C with “Proper airflow recommended” printed beside it. Those idle figures are conditional. APST moves the controller into a low power state; ASPM moves the PCIe link into L1 and its L1.1 and L1.2 substates. Separate mechanisms, both required, and ASPM is often not running at all: one device that declines it keeps the link awake, and the ACPI FADT carries a bit declaring that the platform does not support PCIe ASPM, which Linux honours. The cost is not the drive’s milliwatts, it is the package C-state the CPU can no longer reach. One documented Z590 build measured the whole system, with no cards fitted, at 28 W on BIOS defaults, 18 W with ASPM and the deep C-states enabled but the package stuck at C2, and 16 W once runtime tuning got the package to C8; refitting a ConnectX-3 NIC or an LSI HBA, neither of which will take ASPM, put it back to 25 W to 26 W.
Check the platform before blaming the drives.
lspci -vvv | grep -i aspm # "ASPM Disabled" in LnkCtl is the bad answer
cat /sys/module/pcie_aspm/parameters/policy
nvme smart-log /dev/nvme0 # composite temperature, sensors, throttle counters
sensors # the nvme hwmon driver exports the same sensors
A drive’s idle figure is a promise the platform has to keep.
Failure, rebuild and the arithmetic that changes on flash
The argument that killed RAID 5 was an argument about hard drives. It rested on one datasheet line - the consumer SATA bound of no more than one unrecoverable read error per 10^14 bits read - and on rebuilds that ran for a day and a half. Neither input survives the move to NVMe intact, so the conclusion should be recomputed rather than carried across.
Consumer SATA UBER bound 1 error per 1e14 bits read
1e14 bits / 8 = 1.25e13 bytes = 12.5 TB
Rebuild of a 6 x 2 TB RAID 5 reads the 5 survivors in full
5 x 2 TB = 1.0e13 bytes = 8.0e13 bits
P(at least one error) = 1 - (1 - 1e-14) ^ 8.0e13 = 0.55
A coin flip, with the array degraded for the whole of it. That was a fair description of the arrays the 2007 argument anticipated; 2 TB drives did not ship until 2009.
Three of the four inputs changed
The UBER figure is a specification bound, not a measured rate. A vendor promises the error rate will not exceed it; nothing says the drive sits at it. Bairavasundaram and colleagues at NetApp, who examined 1.53 million drives over 32 months for SIGMETRICS 2007, reported that latent sector errors are not independent events and show both temporal and spatial locality, which is precisely what the independent-trial exponent above assumes away. The arithmetic is wrong in both directions: rarer than the bound on average, far more clustered when it does happen.
Enterprise flash is specified two to three orders of magnitude tighter. Data-centre NVMe datasheets quote one unrecoverable sector per 10^17 bits read, Samsung’s PM1725b among them. Recompute the block above at 1e-17 and the same rebuild sits at about one chance in 1,250. Read the datasheet for the part in front of you rather than trusting the class figure, and treat the number as the vendor’s promised ceiling, because the published flash figures are all vendor figures.
A single bad block need not kill the rebuild. Linux md keeps a 4 KiB bad block list per member, so an unreadable block during recovery is recorded and the rebuild continues; you lose a block, not an array. Hardware controllers vary and some still abort. That behaviour is a property of the implementation, not a law of parity.
Where the risk actually went
Into correlation, which parity arithmetic models badly.
Wear is shared. Every member of a mirror takes byte-identical host writes, and
every member of a parity set takes near-identical writes plus its share of
parity, so percentage_used climbs in lockstep across the batch. The rebuild is
the heaviest read the survivors have ever served and the heaviest write the
replacement will ever take, arriving at the moment the whole set is most worn.
SSD endurance has the DWPD arithmetic that parity
inflates: a small write costs two device writes on RAID 5 and three on RAID 6.
Firmware is shared. A set bought on one purchase order carries one firmware image and one bug. HPE issued two customer advisories a few months apart: one in November 2019 for SAS SSDs running firmware older than HPD8, which failed at 32,768 power-on hours, and one in March 2020 for a second group running firmware older than HPD7, which failed at 40,000. HPE traced both to a supplier defect rather than its own code, and warned that drives put into service together would fail at close to the same time. The 32,768 figure is the one met earlier: 3.74 years of continuous power, and a suspiciously exact 2^15, though the advisories still describe the fix rather than the cause. Either way, no parity level survives its entire membership failing on a Tuesday. Buy in two batches from two sellers and tolerate mixed firmware revisions instead of engineering them away; the enterprise SSD listings are ranked by price per terabyte, which makes a split order cheap to plan.
The failure mode is different. A worn hard drive dies. A worn SSD with decent firmware stops accepting writes and keeps serving reads. Every byte is still readable, and md ejects the member anyway, because md sees write errors. The array degrades over a drive that is not, in the ordinary sense, broken.
Rebuild time, and the default that ruins it
Member 4.0 TB = 4.0e12 bytes
md stock speed_limit_max 200,000 KiB/s = 204.8 MB/s
4.0e12 / 204.8e6 = 19,531 s = 5 h 25 min
Cap lifted, bounded instead by the replacement's sustained
write rate past its SLC cache, say 3.0 GB/s
4.0e12 / 3.0e9 = 1,333 s = 22 min
The 3.0 GB/s is a stand-in. Substitute the sustained write figure from the datasheet of the drive you are fitting, because past the SLC cache it varies by an order of magnitude across the NVMe drives listed here.
/proc/sys/dev/raid/speed_limit_max defaults to 200000 KiB/s and applies per
device, not per array. On NVMe it is the most damaging stock value in md: it
turns a twenty-minute rebuild into most of a working day, for the benefit of a
disk that is not in the machine. Raise it, then find the real ceiling, which on
RAID 5 and RAID 6 is usually the single md0_raid5 or md0_raid6 thread -
md/group_thread_cnt spreads that work, and the right value depends on core
count and workload, so measure rather than copy a number.
A write-intent bitmap changes the question from how fast to how much: a member
that dropped out briefly is re-added with only the dirty regions recovered. The
lockless bitmap merged in Linux 6.18, and exposed by mdadm 4.6 as
--bitmap=lockless, syncs only regions that were ever written, so a fresh array
needs no initial sync at all. Both are newer than what most distributions
ship today, so check your versions before planning around it.
Reading a drive that is on its way out
nvme smart-log /dev/nvme0 # critical_warning, percentage_used,
# available_spare, media_errors
nvme error-log /dev/nvme0 # the controller's own recent errors
nvme id-ns /dev/nvme0n1 # nsattr bit 0: namespace write protected
smartctl -a /dev/nvme0 # the same log, different presentation
lsblk -o NAME,RO,SIZE # RO=1 is the read-only transition, visible
cat /proc/mdstat # which member md ejected
critical_warning is a bit field: bit 0 available spare below threshold, bit 1
temperature past a critical threshold, bit 2 reliability degraded by media or
internal errors, bit 3 media placed in read-only mode, bit 4 volatile memory
backup failed. Bit 3 is the one that ejects an otherwise healthy member.
percentage_used is the vendor’s estimate against rated endurance, is permitted
to exceed 100, and predicts nothing on its own. available_spare is a
normalised percentage of remaining spare blocks and moves first: a drive at 100
per cent spare and 140 per cent used is fine, a drive at 8 per cent spare is
not. media_errors counts unrecovered data-integrity errors, and any nonzero
value on a drive bought second-hand deserves an explanation. These are the same
four fields to read before money changes hands, which
what SMART tells you and what it cannot covers
for drives with a history.
Scrubbing, spares, and the thing an array is not
A scrub reads every block and checks it against parity or its mirror.
echo check > /sys/block/md0/md/sync_action # find and count, change nothing
cat /sys/block/md0/md/mismatch_cnt
echo repair > /sys/block/md0/md/sync_action # fix
Debian and Ubuntu package a checkarray script, and mdadm now ships timers of
its own: mdcheck_start.timer opens a check on the first Sunday of the month
and mdcheck_continue.timer carries it across successive nights. Find out which
of the two your distribution enabled before assuming either runs. Monthly suits
flash, where reads cost no endurance. A nonzero mismatch_cnt on RAID 1 or
RAID 10 is not by itself corruption: md(4) states that the count “can not be
interpreted very reliably on RAID1 or RAID10, especially when the device is used
for swap”, since a page can be rewritten in memory between the two copies going
out. On RAID 5 or RAID 6 it wants investigating.
A hot spare is an idle drive that contributes nothing until the day it does, and it rebuilds onto one device, so recovery is bounded by that device’s write rate. Distributed spare capacity spreads the spare space across every member, so rebuild writes land on all of them in parallel and the window shrinks with the member count divided by the stripe width, which is how OpenZFS describes the scaling of its own sequential resilver. Mainline md has no such thing; OpenZFS has had it since 2.1 as dRAID, with spare space integrated into the vdev. What md offers instead is better than a spare for planned work:
mdadm /dev/md0 --replace /dev/nvme3n1 --with /dev/nvme9n1
That swaps a drive showing wear while full redundancy is still intact. The array never goes degraded, which is the entire point.
A redundant array is not a backup. Every write reaches every copy at the same
instant, including the wrong ones: a mistaken rm -rf, an encrypting process, a
bad migration, a filesystem bug. Redundancy covers a device stopping. It covers
nothing else.
Home builds, from two drives upward
Start with the uncomfortable part. An array buys availability and nothing else. It does not buy capacity, it does not buy durability, and a second copy on the same power supply behind the same filesystem is not a backup. Each tier below is worth building only once you can say which of those you are buying, and at the low end the honest answer is usually none of them.
The lane arithmetic is identical at every tier, so state it once. A PCIe link from Gen3 to Gen5 carries 128 bits of payload in every 130-bit block, a 1.54 per cent line-code overhead, so per lane per direction:
Gen3 8.0 GT/s x 128/130 / 8 = 0.985 GB/s x4 = 3.938 GB/s
Gen4 16.0 GT/s x 128/130 / 8 = 1.969 GB/s x4 = 7.877 GB/s
Gen5 32.0 GT/s x 128/130 / 8 = 3.938 GB/s x4 = 15.754 GB/s
One NVMe drive is x4. Every question that follows is how many x4 groups the board can really deliver, and from which silicon.
| Tier | Drives | Lanes reaching the drives | Thermal arrangement | Filesystem that fits |
|---|---|---|---|---|
| Mini-PC, 2-slot NAS | 2 x M.2 | 8, CPU-attached | still air, 3.3 V only | ZFS mirror or btrfs raid1 |
| Desktop tower | 4 to 6 x M.2 | 24 CPU lanes, or 4:1 behind the chipset | carrier plus a slot fan | mirrors, or raidz1 with the endurance cost counted |
| Ex-enterprise workstation | 8 to 16 x U.2 | 112 to 128, no longer the constraint | hot-swap cage, fan wall | ZFS mirrors or raidz2 |
| Past the ceiling | 24 and up | 96 needed, or a switch at 6:1 | designed airflow, rack noise | any of them; the CPU is the limit now |
Two M.2 slots, and why that is where most people should stop
The mini-PC or the two-slot M.2 NAS. Both sockets are M-key Socket 3, both usually hang off the CPU root complex, and the recipe is a mirror. The ceiling:
2 x M.2 Socket 3, CPU-attached, Gen4 x4 each:
mirror read ~ 2 x 7.877 = 15.75 GB/s (md picks the member with fewer
pending requests once any device
reports non-rotational)
mirror write = 7.877 GB/s (both copies are written)
10GbE client = 1.25 GB/s
2.5GbE client = 0.3125 GB/s
The read figure is a ceiling rather than a promise. Each request goes to one member, so the pair is twice as quick only when enough requests are in flight to keep both busy; a single sequential reader sees roughly one drive.
One drive’s x4 link already outruns 10 gigabit Ethernet by a factor of six. Anything you add above two drives at this tier is spent on a link that cannot carry it, unless the consumer is the machine itself.
Thermally M.2 is the weak form factor, and the reason is in the pinout: no 12 V rail, power off 3.3 V, an unenclosed module, and no staged mating, so the socket is not hot-pluggable either. The reasons hyperscale abandoned it, quoted earlier, are the reasons it bites here. In a fanless mini-PC the two modules often sit on opposite faces of the board with no defined airflow, and throttling arrives as a collapse in sustained write rate rather than as an error anyone logs.
Filesystem: take the checksums. A ZFS mirror or btrfs raid1 verifies every
block read and repairs from the good copy. md RAID1 under ext4 or XFS returns
whatever the member handed it, and a check can only tell you the two members
disagree, not which one is right.
Cost, honestly: consumer M.2 prices run close enough to linear that two 2 TB drives often come to about what one 4 TB drive costs, and leave you with 2 TB. That is the full price of double the capacity, spent on uptime. Rank the candidates by price per terabyte across every drive before deciding the mirror is the cheaper answer, because that ratio moves week to week. Usually it is the more expensive one.
The tower, four to six drives, and the chipset trap
Consumer sockets are lane-poor in a very specific way. Socket AM5 provides 28 CPU PCIe lanes of which 24 are usable at PCIe 5.0, the chipset link running at PCIe 4.0:
AM5: x16 graphics slot 16
CPU-attached M.2 4
second CPU group 4
--
usable CPU lanes 24
chipset link 4 (28 total, not yours)
Intel LGA1851 is 20 PCIe 5.0 plus 4 PCIe 4.0, also 24 CPU lanes, with the chipset behind a DMI 4.0 link, x8 on Z890 and x4 further down the range. Either way, on a board that routes and exposes x4x4x4x4 on that slot, drop the graphics card and six drives sit on the root complex with no sharing.
That slot is the bifurcation question from the attachment section, unchanged: root complex, board routing and a firmware option, all three or none. Without them the passive quad carrier shows one drive, and the switched carrier costs what the switch costs.
The chipset M.2 slots are where builds quietly go wrong:
Four chipset M.2 slots, AM5 Promontory (Gen4 x4 uplink):
drive demand 4 x 7.877 GB/s = 31.5 GB/s
uplink supply 1 x 7.877 GB/s = 7.9 GB/s
ratio = 4:1, shared with USB, SATA and the NIC
That ratio is the optimistic case rather than the worst one: on boards with two daisy-chained chipset dies the second hangs off the first, so the ports on both still share the one PCIe 4.0 x4 link back to the CPU.
Aggregate array throughput here is set by the uplink, not by the drives. Four chipset drives in a stripe are one drive wide.
Thermals: a quad carrier parks four modules in the dead-air pocket behind the graphics card. Put it in the top slot, accept a slot fan or a carrier with its own blower, and treat any card whose heatsink is a flat plate as a two-drive card.
Filesystem: four to six drives is where parity tempts, and parity on flash has two costs. The write hole is real on md RAID5, and the fixes are PPL, which needs no journal device but whose kernel documentation puts the write-performance cost at up to 30 to 40 per cent, or a journal on a drive with power-loss protection. ZFS raidz1 avoids the hole by construction, since it never does a read-modify-write. The second cost is endurance: md’s stripe head unit is 4 KiB, so a small RAID5 write is roughly 2x host bytes to the devices and RAID6 is 3x, straight off the rated DWPD.
The ex-enterprise workstation, where U.2 becomes the cheap option
Lanes stop being scarce abruptly. Threadripper PRO 9000WX on WRX90 carries 128 PCIe 5.0 lanes; Xeon W-3500 up to 112; an EPYC 9005 socket, 128. At x4 per drive, 112 lanes is 28 drives of attachment before the NIC gets a look in. Slots, cables, power and airflow become the constraint instead.
The form factor changes here. U.2 on the SFF-8639 connector brings 12 V and 5 V rails, staged mating for hot-plug - the device plug’s P13 is literally “+12 V Precharge” - and power-loss protection as a fitted norm rather than a premium. Second-hand datacentre 2.5-inch NVMe is usually the cheapest honest terabyte in this section: the NVMe listings filtered to the 2.5-inch parts are where they surface. One direction of that compatibility catches buyers. SFF-TA-1001 mandates that U.3 devices work in U.2 slots, and its interoperation table gives “No link” for a U.2 device in a U.3 slot. Some backplanes are wired dual-mode anyway, so check the vendor’s documentation rather than the spec.
Cabling follows generation by practice rather than by rating. SFF-8654 (SlimSAS, Rev 1.2) specifies mechanics and gives no GT/s figure of its own, but it is the internal cable the Gen4 era settled on; MCIO, SFF-TA-1016 Rev 1.3, is what Gen5 designs use. A Gen4 cable does not become Gen5 by fitting: the Gen5 bump-to-bump insertion-loss budget of 36 dB at 16 GHz is shared with both packages, the boards and every connector on the path, so a long Gen5 cabled run becomes a retimer application, and the spec permits two retimers per link. Buying used Gen3 or Gen4 U.2 sidesteps that entire problem, which is a reason to do it rather than a consolation for doing it.
This is the first tier where the thermal arrangement is designed rather than improvised: a hot-swap cage with a fan wall behind it, which is the forced air enterprise U.2 assumes. The same drive in a still desktop case will throttle.
Filesystem: ZFS, mirrors for rebuild speed and IOPS or raidz2 for capacity. If
you use md, change one default before anything else. The stock
/proc/sys/dev/raid/speed_limit_max is 200000 KiB/s per device - about 200
MB/s, against drives that sustain several times that - and it stretches a
rebuild the hardware could finish inside an hour into most of a day. Put the
working set on this array and leave the bulk capacity on spinning disk, for the
reasons set out in where each still wins.
The ceiling, and what pushes a build past it
Count the lanes for a 24-bay Gen5 chassis and the answer is immediate:
24 drives x 4 lanes = 96 lanes of drive attachment alone,
before NIC, HBA or accelerator
That is a 128-lane socket, a dual-socket board, or a PCIe switch - and a switch
routes packets, it does not create bandwidth. Twenty-four drives at x4 behind an
x16 uplink is 6:1 oversubscription. Past roughly eight to twelve drives the
bottleneck also stops being the fabric and becomes the host CPU: md RAID5 stripe
handling, which stays on one thread until group_thread_cnt is raised, ZFS
checksum and compression work, the interrupt load.
A home build ends where the spare parts begin. The markers are not capacity.
They are a dedicated circuit, noise that needs its own room, a spare chassis on
the shelf because a failure you cannot fix this week is an outage, a written
rebuild procedure, and a management plane - UBM enclosure signalling, ledmon
and ledctl for drive-locate LEDs, Intel VMD for hot-plug. Once those are
required for correct operation, the thing is an installation with an owner, and
it should be budgeted as one.
Enterprise and the top of the scale
A 24-bay NVMe chassis needs 96 lanes of PCIe before anyone plugs in a network card. That one sum explains why server sockets carry lane counts a desktop builder finds absurd, and why most of them still are not enough.
The lane budget is the design
Supermicro’s June 2026 storage brochure lists the two shapes that dominate: a 2U with 24 hot-swap 2.5-inch U.2 bays, and a 2U with 32 hot-swap E3.S bays, the E3.S entries specified as PCIe 5.0 x4 per bay. Four lanes a bay is the U.2 norm too. Count what one populated chassis asks for, with a pair of x16 network cards as the illustration:
24 drive bays x 4 lanes = 96 lanes
2 network cards at x16, illustrative = 32 lanes
boot device, BMC, miscellaneous = 4 lanes
---------
132 lanes
AMD EPYC 9005 (SP5), per socket = 128 PCIe 5.0 lanes
Intel Xeon 6900P, per socket = 96 PCIe 5.0 lanes
Intel Xeon 6700P, per socket = 88 PCIe 5.0 lanes
No single socket covers that budget. The 128-lane EPYC comes closest and is four lanes short: the drives and both cards fit, the boot path does not. Two sockets solve it and run a NUMA boundary through the array. Most vendors take the other option.
Switches in the backplane, and the oversubscription decision
A PCIe switch routes transaction-layer packets from one upstream port to many downstream. Microchip lists its Switchtec PFX Gen 5 family from 28 to 100 lanes with up to 52 ports; Broadcom’s PEX89000 covers 24 to 144. What neither creates is bandwidth:
PCIe 5.0, 32 GT/s with 128b/130b = 3.938 GB/s per lane
x4, one drive = 15.75 GB/s
x16, the switch uplink = 63.0 GB/s
24 drives into one x16 uplink = 96 lanes into 16 = 6 : 1
Six to one is not automatically wrong. It is a bet that the drives will not all run flat out at once, and the switch buys per-port hot-plug and error isolation root ports do not give. That bet is harder now: KIOXIA rates the CM9-V at 3.4 million random 4 KiB reads a second, near 13.9 GB/s, so even small-block work now fills most of a Gen5 x4 link. A sequential scan of the whole array reaches 63 GB/s and stops. Decide which of those you are before buying the chassis; a backplane is not field-replaceable.
Dual-port is half a link, and it fixes the other failure
Real high availability needs three things and most builds have one: two hosts, an enclosure wiring each bay to both, and a drive that accepts two hosts at once. The drive is the part people assume.
KIOXIA’s enterprise data sheet gives the interface of its CM7, CM9 and LC9 NVMe families as “PCIe Gen5 single x4, dual x2”. Dual-port is not a second link. It is the same four lanes split in two, one x2 per host, 7.88 GB/s each instead of 15.75, and the sheet’s footnote concedes that its NVMe figures are “based on single-port mode (single x4)”. The drive must also present one NVM subsystem with two PCIe ports and a namespace shared between two controllers, the hosts arbitrating through NVMe reservations rather than hope. A single-port drive in a dual-port bay is a second cable to nothing.
Note the failure this covers: a dead host, HBA or cable. The drive dying is still RAID’s problem. SAS got here first without the halving, and KIOXIA’s PM7 is a “dual-port 24G SAS” drive whose published figures are measured in dual-port mode at 18 W. Where the availability requirement is real and the bandwidth requirement is not, dual-ported SAS SSDs remain the cheap way to put two hosts on one drive.
EDSFF, and why the 2.5-inch bay is ending
A PCIe 5.0 enterprise NVMe drive is a 25 W part, KIOXIA’s typical figure for CM7, CM9 and the 2.5-inch LC9 alike. Twenty-four of them is 600 W of drives in a 2U before the processors draw anything.
| Form factor | Thickness | Width | Length | Spec |
|---|---|---|---|---|
| 2.5-inch U.2 | 15.0 mm | 69.85 mm | 100.45 mm | SFF-8639 |
| E3.S | 7.5 mm | 76.0 mm | 112.75 mm | SFF-TA-1008 |
| E3.L | 7.5 mm | 76.0 mm | 142.2 mm | SFF-TA-1008 |
| E1.S | 5.9 to 25 mm | 31.5 mm | 111.49 mm | SFF-TA-1006 |
E3.S is half the thickness of the bay it replaces, which is how the same 2U frontage goes from 24 drives to 32. SNIA pairs each thickness with a power envelope and an airflow profile, from E1.S at 5.9 mm and 12 W up to the double-thickness 16.8 mm E3.L at 70 W, so cooling is specified by the form factor instead of improvised by the chassis. M.2 lost this argument on power. SNIA’s account of what went wrong with M.2 22110 in hyperscale racks, quoted earlier, is that argument in one line: a 3.3 V module with no 12 V rail, no enclosure and a retaining screw was never going to carry 25 W.
The endurance ladder that is actually bought
Read-intensive, mixed-use and write-intensive are marketing names for one engineering decision: how much flash to withhold as spare. KIOXIA’s CM7-R and CM7-V are the same BiCS FLASH generation 5 silicon.
CM7-R 3,840 GB user, 1 DWPD, 5 yr: 3.84 x 1 x 365 x 5 = 7,008 TB
CM7-V 3,200 GB user, 3 DWPD, 5 yr: 3.20 x 3 x 365 x 5 = 17,520 TB
640 GB withheld = 16.7 per cent less capacity, 2.5x the lifetime writes
Three times the DWPD is two and a half times the bytes, because DWPD is scaled by the capacity it just shrank. Read the TB figure, not the rung; the guide to TBW, DWPD and how worried to be takes that arithmetic further. Below it sits a QLC capacity rung, the LC9, 122.88 TB in 2.5-inch and 245.76 TB in E3.L, rated 0.3 DWPD at a 16 KiB indirection unit but 0.075 DWPD at 4 KiB. Above mixed-use KIOXIA’s sheet lists nothing, and arithmetic explains the missing rung: 3 DWPD on a 12.8 TB drive allows 38.4 TB a day, where 10 DWPD on a 1.6 TB drive, the shape high-endurance parts used to take, would allow 16.
Firmware and fleet management, which has no home equivalent
A fleet does not accept the firmware that arrived. It qualifies a revision, stages it across failure domains, and keeps the previous image in the drive’s second slot. The OCP Datacenter NVMe SSD Specification, v2.0 in 2021 through the v2.7 of November 2025 and its January 2026 revision, has mandated since at least v2.5 the log pages that needs: Firmware Slot Information (03h), Telemetry Host-Initiated (07h) and Controller-Initiated (08h), Device Self-test (06h), Persistent Event Log (0Dh). NVMe-MI carries them out of band to the BMC, so a drive answers while its host is down. A fleet buys the ability to know, not the ability to repair, and none of that apparatus follows a drive out of the rack.
Why any of this reaches you
Enterprise flash is bought on a three to five year refresh and then leaves, whatever is left in it. Take that 3.2 TB mixed-use drive rated for 17,520 TB and run it four years at 0.2 DWPD, a fifteenth of what it is rated to take:
0.2 x 3.2 TB = 0.64 TB/day
0.64 x 365 x 4 = 934 TB written
934 / 17,520 = 5.3 per cent of the rating consumed
Ninety-five per cent of the endurance, sold used. That is the trade this catalogue exists for, and a good one if you read the number instead of the story, which is what the enterprise-class SSDs here are priced against and what the guide to enterprise drive pulls works through field by field. What stays behind is everything above: no support contract means no vendor firmware, a U.2 drive wants 12 V and an SFF-8639 bay no desktop has, E1.S and E3.S want a backplane that exists nowhere outside a server, and a dual-port drive in a single-port adapter is a single-port drive. Ranking enterprise drives of every kind by price per terabyte is how you test whether the pull still wins once the adapter is counted.
The do’s and don’ts
Nothing here is new. Every line compresses a mechanism argued above to the length you can photograph, and if one survives the screenshot, make it this one: an NVMe array fails at the platform long before it fails at the drive.
Do
Count the lanes before you count the drives. One NVMe device is an x4 link, and a current mainstream socket leaves about 24 usable CPU lanes: 24 of the 28 on AM5, the other four being a Gen4 link to the chipset, or 20 Gen5 plus 4 Gen4 on Intel’s LGA1851. Extra M.2 slots hang off a chipset whose uplink is four lanes on AM5 or a DMI x8 on Intel, shared with USB, SATA and the network, so the uplink caps the array before the drives do.
1 NVMe device = 4 lanes
AM5 usable CPU lanes = 24
graphics at x16 = 16 -> 8 left -> 2 drives
graphics at x8 = 8 -> 16 left -> 4 drives
4 chipset M.2 slots = 16 lanes of demand
chipset uplink = 4 lanes -> 4:1 oversubscribed
Match model and firmware where you can, and split the purchase where you can afford to. Identical drives keep stripe alignment, discard granularity and flush behaviour predictable. Identical manufacturing dates put every member on the same wear-out curve, which is why two sellers on the NVMe listings beat one sealed box of five.
Leave over-provisioning space unpartitioned. Capacity the host never claims stays available to the garbage collector, which is the cheapest way to hold steady-state write latency down and the endurance figure up, as SSD endurance works through.
Put moving air over every M.2 module. M.2 is a bare 3.3 V board with no defined airflow profile, while EDSFF publishes a power envelope per thickness: E1.S runs from 5.9 mm at 12 W to 25 mm at 25 W, and the thicker E3 variants are rated higher still, to 70 W for E3.L 2T. A module that throttles holds up every write in a mirror, because the write is not finished until the slow member says so.
Prefer mirrors for anything latency-sensitive. A partial-stripe write to RAID 5 costs two reads and two writes; RAID 6 costs three of each. A mirror costs one write per copy and no read at all.
Set ashift and recordsize before the first byte lands. ashift is fixed for the life of the vdev, and a recordsize change applies only to data written after it. Both are creation-time decisions dressed as tunables.
Enable autotrim or schedule fstrim, then confirm it reaches the drives.
Linux md refuses discard on RAID 4, 5 and 6 by default, because a member that
does not return zeroes from a discarded region will poison the next parity
calculation. The drive-side field that settles it is DLFEAT in nvme id-ns.
Watch percentage used and available spare. Those two fields from
nvme smart-log move before anything else does, and on a second-hand drive they
are the only honest account of what the seller did with it, read as
what SMART tells you sets out.
Scrub on a schedule and read the number it returns. A scrub nobody checks
is a cron job, not a policy: md/mismatch_cnt after a check, or the repaired
byte count from a ZFS scrub, is what separates an array quietly correcting from
one quietly rotting.
Keep a copy that is not in the array, and restore from it once. An untested backup is a plan. The restore is the test.
Don’t
Do not buy a passive bifurcation carrier without reading the board manual. A passive quad-M.2 card is wiring, and it needs CPU support, board routing and a firmware option all present at once. If any of the three is missing, only the drive on lanes 0 to 3 appears, as the interfaces guide sets out.
Do not put a drive without power-loss protection under a synchronous workload. Without the capacitors the drive must push each flush through to NAND before acknowledging it, and sustained sync write rates land far below the headline. Datasheet PLP is the line to look for on the enterprise SSDs.
Do not use the btrfs parity profiles for data you care about. The project’s own documentation marks RAID 5 and RAID 6 unstable, says plainly not to use them for metadata, and records that the write journal which would close the write hole is not implemented - the same hole md closes with PPL or a journal. Use the mirrored profiles or a different filesystem.
Do not fill a pool past about eighty per cent. Eighty is a rule of thumb, not a number OpenZFS publishes: its own tuning guide asks instead for more than ten per cent free, so that metaslabs stay clear of the four per cent mark where the allocator gives up first-fit for best-fit. The mechanism is the same one: a copy-on-write filesystem needs free space to place a full stripe in one piece, and as that space fragments the allocator works harder for every write, so latency climbs on a pool that has changed in no other way.
Do not read the box figure as a sustained figure. A Gen4 x4 link ceiling is 7.877 GB/s once the 128b/130b sync header is paid for, real sequential transfers land near 7 GB/s, and a client drive’s headline write number is normally an SLC cache number that ends when the cache does.
Do not build RAID 5 across drives that wear identically. Same model, same batch, same write stream means the members reach their limits together, and the second failure arrives during the rebuild from the first.
Do not trust hot-plug on consumer hardware. A bare M.2 socket is not designed for it; U.2 on SFF-8639, U.3 and the EDSFF bays are, with staged power and presence pins. The firmware still has to advertise the slot as hot-plug capable and hand native control to the operating system.
Do not let a rebuild be the first time you look at the spare. Read the whole
device through once a quarter: a cold spare that fails its first full read is
worse than no spare. Raise /proc/sys/dev/raid/speed_limit_max from its stock
200000 KiB/s too, or the rebuild runs at hard-drive pace on flash.
Do not put a special vdev, journal or log device on a single drive. Losing an unmirrored ZFS special vdev loses the pool, and md states plainly that a failed write journal forces the array into read-only mode. The metadata device inherits the redundancy requirement of everything that points at it.
Do not call redundancy a backup. A mirror replicates a mistaken deletion in microseconds and ransomware at line rate. Redundancy buys availability through a hardware failure and nothing at all against the other ways data leaves.
Buying the drives
Buy the counters, not the drive. An NVMe SSD is the only part of this build that arrives carrying a written record of how hard its last owner used it, and on the second-hand market that record is worth more than the model number. An array makes the point sharper than a single drive does: eight members bought blind are eight independent chances that one of them was pulled from a write-saturated database node.
Ask for the health log before you pay
Every NVMe device exposes the SMART / Health Information log, log page 02h in the
NVM Express Base Specification. Six fields carry the purchase decision: Critical
Warning, Percentage Used, Available Spare against Available Spare Threshold, Data
Units Written, Power On Hours and Unsafe Shutdowns. On Linux the command is
nvme smart-log /dev/nvme0; smartctl and CrystalDiskInfo print the same fields
under their own labels, and a phone photograph of any of them is enough.
Ask for it in those words - “please paste or photograph the output of
nvme smart-log, or a CrystalDiskInfo screenshot”. A seller with the drive in a
machine produces it in a minute. A seller who will not has either never powered
it or would rather you did not look. The refusal is the reading.
Data Units Written is the field to do arithmetic on, because Percentage Used is a vendor estimate rather than a measurement. The specification defines a data unit as 1,000 units of 512 bytes. The sum, with figures for illustration:
Data Units Written 412,551,873
one data unit 1,000 x 512 B = 512,000 B
host bytes written 412,551,873 x 512,000 = 211,226,558,976,000 B
= 211.2 TB
rated endurance, 1.92 TB drive at 1 DWPD over 5 years:
1.92 x 1 x 365 x 5 = 3,504 TB
211.2 / 3,504 = 6.0 per cent of rated life consumed
If the drive’s own Percentage Used reports 40 while that sum says 6, check the drive’s actual rating first: Percentage Used is computed against whatever the vendor rated the part at, so a 0.3 DWPD read-intensive drive reaches 40 on far fewer host bytes than the 1 DWPD assumed above. If the rating does match, the usual remaining explanation is internal write amplification: the host wrote little, the controller wrote a great deal moving it around, and the flash counts programme/erase cycles rather than host bytes. That would be a drive that lived on small random writes. SSD endurance: TBW, DWPD, how worried to be works through what the two numbers mean against each other, and what SMART tells you and what it cannot covers the fields a seller can quietly omit.
The traps, one line each
| Trap | What actually happens |
|---|---|
| B+M keyed M.2 drive | Links at PCIe x2 or SATA, never x4 |
| M.2 SATA sold as “M.2 SSD” | Same board outline as NVMe, and a B+M notch fits either, so only the model number settles the protocol |
| M.2 22110 in a 2280 board | 30 mm too long; enterprise M.2 with power-loss protection is often 22110 |
| U.2 drive, U.3 backplane | SFF-TA-1001 says no link; the reverse is mandated to work, and some backplanes are wired both ways, so check the vendor document |
| E1.S or E3.S | SFF-TA-1002 connector; no EDSFF bay in a desktop, so it needs a passive PCIe carrier card |
| Passive quad-M.2 card | Without x4x4x4x4 bifurcation in firmware, one drive enumerates |
| SAS SSD in a 2.5-inch photograph | Not NVMe, and no NVMe host will see it |
SATA, SAS, NVMe and M.2: what plugs into what has the keying detail if a listing photograph is ambiguous.
Enterprise pulls, and what the host owes them
Pulls are the value play, and the reason is structural: enterprise write ratings run high enough that a 3 DWPD drive retired at 20 per cent of its life can have more writing left in it than a new consumer drive has in total, as the sums at the end of this section show for one such pair. The host has to meet three conditions. Power: a U.2 drive expects 12 V and, depending on the part, draws up to about 25 W, well beyond anything the 3.3 V-only M.2 socket can supply; the figure for a given drive is on its datasheet. Airflow: SFF-TA-1023 characterises EDSFF cooling as the airflow a device needs at a given inlet temperature, and U.2 enterprise parts assume forced air; in a silent desktop case they throttle and stay throttled. Attachment: the cable is a component and not a wire, and which one a Gen4 or a Gen5 run needs was settled above. Buying a Gen3 or Gen4 pull is how most readers walk past that problem rather than solve it. Enterprise drive pulls goes through the provenance side, and the current enterprise-class SSDs in the catalogue are where the pulls surface first.
The sector format problem
A pull arrives in whatever LBA format the previous array wanted.
nvme id-ns /dev/nvme0n1 lists every supported format with its metadata size and
Relative Performance; a drive carrying 8 bytes of metadata for T10 protection
information presents 520- or 4104-byte sectors, and plenty of hosts will not
touch it. 4Kn, meaning 4096 bytes with metadata size 0, is usually the
best-performing format the drive offers, though Relative Performance is the
drive’s own declaration rather than a rule, so read the field rather than assume
it.
nvme format /dev/nvme0n1 --lbaf=N changes it, destroying the namespace
contents. It works on many drives, but this guide will not promise it for any
specific part: OEM-badged firmware is widely reported to refuse the operation or
to offer only metadata-bearing formats, and a drive that refuses is a drive you
now own. If the listing shows a Dell, HPE or NetApp part number, ask the seller
to confirm that nvme id-ns lists a format with ms:0 before money moves.
Why the cheap large drive is cheap
Three explanations cover nearly all of them. It is counterfeit, with a controller reporting capacity the flash does not have, which is why a full write and read-back verify comes before the drive joins anything. It is QLC sold on its burst figure, where the headline write speed is the SLC cache and the post-cache rate is a fraction of it - and a rebuild is precisely a sustained full-capacity write, so that drive is slowest at the one moment the array depends on it. Or it is at eighty per cent of rated life, which Percentage Used would say, which is why the listing does not show it.
Ranking by the number that matters
Every NVMe SSD in the catalogue is ranked so you can compare parts rather than adverts, and every drive by price per terabyte puts NVMe next to SAS and spinning disk on one scale. Price per terabyte is still the wrong sole metric for an array member. Divide by the writing the drive has left, as in these two illustrative parts:
A: new 2 TB consumer TLC, 1,200 TBW rated, 90 EUR
90 / 1200 = 0.075 EUR per TB written
B: 1.92 TB enterprise pull, 3 DWPD over 5 y = 10,512 TBW rated,
Percentage Used 18 -> 8,620 TBW left, 140 EUR
140 / 8620 = 0.016 EUR per TB written
The pull costs about half as much again on the shelf, and its cost per terabyte written is a little over a fifth of the new drive’s. Sort by price per terabyte to build the shortlist, then re-sort that shortlist by the second figure before buying.
Which returns to where this started. An array of NVMe drives is not a way to go faster; a single modern drive already outruns most of what a desktop asks of it. It is a way to keep serving after a device stops answering, and every decision above - the counters, the format, the airflow, the endurance left - is a decision about how long the array survives its first failure.