Drive capacity explained: why 4TB shows up as 3.64TB
By Harry Saarinen ·
Nothing is missing. The drive holds four trillion bytes and a little over, and Windows is dividing those bytes by 1,024 four times while writing “TB” on the result.
The manufacturer counts in powers of ten, the operating system counts in powers of two, and only one of them labels its units honestly. That single conversion is the whole of the 9%, and the things people usually blame - formatting, the partition table, a hidden recovery area - are real but tiny beside it.
The rest of this page is the part that is usually left out, and one of those omissions is expensive. Two layers below the unit conversion are worth more than a percentage point each: formatting a SAS drive with T10 Protection Information costs 2.03% of it permanently until you reformat, and the choice between ext4 and XFS at format time is worth 6.55% of a 4 TB label. Neither is a rounding error, and neither is visible on a listing page.
The arithmetic, and the standard that fixes the byte count
A 4 TB drive is not approximately four trillion bytes. It is an exact number, fixed by an industry standard so that any vendor’s 4 TB drive can replace any other vendor’s in an array without the array shrinking.
The standard is IDEMA’s LBA1-03, and the first correction on this page is that most explanations cite the wrong one. LBA1-02 is the document usually quoted; LBA1-03 supersedes it, and says so in its own opening, because LBA1-02 “was intended for only IDE disk drives”. LBA1-03 widened the scope to SATA and SAS, to 4,096-byte Large Data Sector drives, and to SAS drives formatted with T10 Protection Information. The algorithm itself did not change:
512-byte sectors
LBA count = 97,696,368 + 1,953,504 × (reported capacity in GBytes − 50.0)
4 TB: 97,696,368 + 1,953,504 × 3,950 = 7,814,037,168 sectors
× 512 bytes = 4,000,787,030,016 bytes
The second correction is smaller and more persistent. The 50.0 is an anchor constant inside the expression, not a scope threshold. The formula does not “apply to drives above 50 GB”; LBA1-03 states its scope by form factor and interface instead. Section 6.0 lists four cases, and the fourth is usually dropped when the scope is paraphrased: 2.5-inch SATA at 80 GB or greater, 2.5-inch at 320 GB or greater with no interface named, 2.5-inch SAS at 73 GB or greater, and 3.5-inch SAS at 450 GB or greater. That second clause is almost certainly a typo for 3.5-inch in the standard itself, because as printed the scope covers no 3.5-inch SATA drive at all - which is the drive every worked example on this page uses. Subtracting 50 from the capacity is arithmetic, not eligibility.
The 4,096-byte form is the same expression with the first two constants divided by eight, which is exactly what you would expect of a formula whose output is a count of sectors eight times larger:
4,096-byte sectors
LBA count = 12,212,046 + 244,188 × (reported capacity in GBytes − 50.0)
4 TB: 12,212,046 + 244,188 × 3,950 = 976,754,646 sectors
× 4,096 bytes = 4,000,787,030,016 bytes
The byte total is identical. A 4 TB drive formatted 4Kn and a 4 TB drive formatted 512e hold the same 4,000,787,030,016 bytes; only the unit of address differs. That fact does more work later on this page than it looks like it should.
The margin, and why it shrinks as drives grow
IDEMA says the formula is designed to “provide 0.02% margin” over the nominal decimal capacity, so a 4 TB drive is very slightly more than 4 × 10¹². Run the margin out across the range and it drifts:
| Label | LBA count | Exact bytes | Over nominal |
|---|---|---|---|
| 250 GB | 488,397,168 | 250,059,350,016 | 0.0237% |
| 500 GB | 976,773,168 | 500,107,862,016 | 0.0216% |
| 1 TB | 1,953,525,168 | 1,000,204,886,016 | 0.0205% |
| 2 TB | 3,907,029,168 | 2,000,398,934,016 | 0.0199% |
| 3 TB | 5,860,533,168 | 3,000,592,982,016 | 0.0198% |
| 4 TB | 7,814,037,168 | 4,000,787,030,016 | 0.0197% |
| 5 TB | 9,767,541,168 | 5,000,981,078,016 | 0.0196% |
| 6 TB | 11,721,045,168 | 6,001,175,126,016 | 0.0196% |
| 8 TB | 15,628,053,168 | 8,001,563,222,016 | 0.0195% |
The drift is structural. 97,696,368 is a fixed addend, so its contribution to the proportional margin falls as the multiplied term grows. Nothing is going wrong; the standard simply converges on 1,953,504 sectors per GB, which is 0.0194% above 10⁹ ÷ 512, and the table’s last row has not quite got there.
The 6 TB and 8 TB rows are not theory. Seagate’s Archive HDD product manual publishes guaranteed sector counts of 11,721,045,168 for the ST6000AS0012 and 15,628,053,168 for the ST8000AS0002, which is the formula’s output to the sector.
Above eight terabytes the formula stops working
This is the part no explanation of drive capacity seems to carry, and it invalidates every capacity table on the internet that extends past 8 TB by running IDEMA’s expression further.
Seagate’s Exos X24 product manuals, both SATA and SAS, state it in one line: “LBA Counts for drive capacities greater than 8TB are calculated based upon the SFF-8447 standard publication.” Anything above 8 TB computed from IDEMA is simply the wrong number. A table saying a 16 TB drive holds 16,003,115,606,016 bytes is 2.2 GB out, and a 20 TB row computed the same way is 3.3 GB out.
Here are the real figures, taken from Seagate’s Exos X24 product manual and Western Digital’s Ultrastar DC HC580 specification:
| Label | 512-byte LBAs | 4,096-byte LBAs | Exact bytes | GiB | Over nominal |
|---|---|---|---|---|---|
| 10 TB | 19,532,873,728 | 2,441,609,216 | 10,000,831,348,736 | 9,314.00 | 0.0083% |
| 12 TB | 23,437,770,752 | 2,929,721,344 | 12,000,138,625,024 | 11,176.00 | 0.0012% |
| 16 TB | 31,251,759,104 | 3,906,469,888 | 16,000,900,661,248 | 14,902.00 | 0.0056% |
| 20 TB | 39,063,650,304 | 4,882,956,288 | 20,000,588,955,648 | 18,627.00 | 0.0029% |
| 22 TB | 42,970,644,480 | 5,371,330,560 | 22,000,969,973,760 | 20,490.00 | 0.0044% |
| 24 TB | 46,875,541,504 | 5,859,442,688 | 24,000,277,250,048 | 22,352.00 | 0.0012% |
Two things jump out of that table.
The margin collapses. IDEMA’s 0.02% becomes 0.001% to 0.008%. A modern large drive holds barely more than the number on the box, where a 1 TB drive from the IDEMA era held about 205 MB more than its label. The slack that used to absorb a vendor’s manufacturing variation is gone.
Every GiB figure is a whole number. Not approximately, exactly: 22,352 GiB with no remainder at 24 TB, 14,902 at 16 TB, 9,314 at 10 TB. SFF-8447 is not a public document, so what follows is a derivation from the datasheets rather than a rule read out of the specification: the expression that reproduces all six published points, and only that expression, is
capacity in GiB = ceil(nominal decimal bytes ÷ 2³⁰)
24 TB: 24 × 10¹² ÷ 1,073,741,824 = 22,351.74... -> 22,352 GiB
22,352 × 1,073,741,824 = 24,000,277,250,048 bytes
÷ 512 = 46,875,541,504 sectors
Applied to the capacities that table does not cover, the rule predicts 14 TB at 13,039 GiB (27,344,764,928 sectors), 18 TB at 16,764 GiB (35,156,656,128 sectors, 18,000,207,937,536 bytes), 26 TB at 24,215 GiB, 28 TB at 26,078 GiB, 30 TB at 27,940 GiB and 32 TB at 29,803 GiB.
Two of those six have since stopped being predictions. Western Digital’s Ultrastar DC HC550 specification publishes 27,344,764,928 sectors at 14 TB and 35,156,656,128 at 18 TB, which is the rule reproducing, to the sector, two points it was not fitted against. The remaining four - 26, 28, 30 and 32 TB - are mine, not a vendor’s, and are unconfirmed against any specification. Treat them as a sanity check on a listing, not as a figure to argue with a seller about.
The free integrity check this hands you
The derivation gives a used-drive test that costs one command. A hard drive above 8 TB should report a whole number of gibibytes. If a drive claiming to be 18 TB reports 16,763.4 GiB rather than 16,764, something has clipped it, and the sections on Protection Information and Host Protected Areas below are the two likely culprits. The check is not proof of health, but a fractional GiB count on a large modern drive is an anomaly worth explaining before the return window closes - see buying used drives on eBay for how long that window is.
The interchangeability promise survived the change of standard
The point of fixing capacity by standard is that a replacement drive fits. That still holds across the SFF-8447 boundary, and it is checkable: Seagate’s Exos X24 at 24 TB and Western Digital’s Ultrastar DC HC580 at 24 TB both publish 46,875,541,504 sectors at 512 bytes and 5,859,442,688 at 4,096, for an identical 24,000,277,250,048 bytes. Two vendors, two documents written independently, one number to the byte. Buying a mixed-vendor array of the same nominal capacity is safe on hard drives for exactly this reason, and is not safe on SSDs, which is a distinction the array section returns to.
Two percentages for one fact, and quoting the wrong one is the usual mistake
Now divide by 1,024 four times, which is what Windows does before printing “TB”:
4,000,787,030,016 ÷ 1,024 = 3,907,018,584 KiB
÷ 1,024 ≈ 3,815,448 MiB
÷ 1,024 ≈ 3,726 GiB
÷ 1,024 ≈ 3.64 TiB
Or in one step: 4,000,787,030,016 ÷ 1,099,511,627,776 = 3.6387.
The ratio is 10¹² ÷ 2⁴⁰ = 0.9094947017729282, and it never changes. It is a
property of the two number systems, identical for every drive ever made and
every drive that will ever be made.
| Prefix | Decimal | Binary | Binary unit is larger by | Reported number is smaller by |
|---|---|---|---|---|
| kilo | 10³ | 2¹⁰ = 1,024 | 2.400% | 2.344% |
| mega | 10⁶ | 2²⁰ = 1,048,576 | 4.858% | 4.633% |
| giga | 10⁹ | 2³⁰ = 1,073,741,824 | 7.374% | 6.868% |
| tera | 10¹² | 2⁴⁰ = 1,099,511,627,776 | 9.951% | 9.051% |
| peta | 10¹⁵ | 2⁵⁰ = 1,125,899,906,842,624 | 12.590% | 11.182% |
Those are two different numbers for one fact, and swapping them is the single most common error in this subject. A TiB is 9.951% larger than a TB; a capacity expressed in TiB is 9.051% smaller than the same capacity expressed in TB. The gap compounds a further 2.4% at each prefix, which is why the effect is invisible on a USB stick and irritating on a petabyte array.
| On the label | Standard | Exact bytes | GiB | TiB | What Windows prints |
|---|---|---|---|---|---|
| 500 GB | IDEMA LBA1-03 | 500,107,862,016 | 465.76 | 0.455 | 465 GB |
| 1 TB | IDEMA LBA1-03 | 1,000,204,886,016 | 931.51 | 0.910 | 931 GB |
| 2 TB | IDEMA LBA1-03 | 2,000,398,934,016 | 1,863.02 | 1.819 | 1.81 TB |
| 4 TB | IDEMA LBA1-03 | 4,000,787,030,016 | 3,726.02 | 3.639 | 3.63 TB |
| 8 TB | IDEMA LBA1-03 | 8,001,563,222,016 | 7,452.04 | 7.277 | 7.27 TB |
| 10 TB | SFF-8447 | 10,000,831,348,736 | 9,314.00 | 9.096 | 9.09 TB |
| 12 TB | SFF-8447 | 12,000,138,625,024 | 11,176.00 | 10.914 | 10.91 TB |
| 16 TB | SFF-8447 | 16,000,900,661,248 | 14,902.00 | 14.553 | 14.55 TB |
| 20 TB | SFF-8447 | 20,000,588,955,648 | 18,627.00 | 18.190 | 18.19 TB |
| 22 TB | SFF-8447 | 22,000,969,973,760 | 20,490.00 | 20.010 | 20.00 TB |
| 24 TB | SFF-8447 | 24,000,277,250,048 | 22,352.00 | 21.828 | 21.82 TB |
Windows truncates the last column rather than rounding it, and that is where
the two competing numbers for a 4 TB drive come from. The drive is 3.6387 TiB.
fdisk rounds and prints 3.64 TiB; Explorer cuts and prints 3.63 TB. Both
describe the same 4,000,787,030,016 bytes, and the truncation is consistent all
the way down the column - 465 GB for 500 GB, 931 GB for 1 TB, 7.27 TB for 8 TB.
If you arrived here searching for either number, neither is missing capacity and
neither is a different drive.
The units have had correct names since 1998 and nobody uses them
The conversion is honest arithmetic. The labelling is not, and it has not been for a quarter of a century, because a fix has existed the whole time.
IEC 60027-2 Amendment 2, adopted in December 1998 and published in January 1999, was the first international standard to define kibi (Ki), mebi (Mi), gibi (Gi), tebi (Ti), pebi (Pi) and exbi (Ei). The 2005 third edition added zebi and yobi. The prefixes now live in ISO/IEC 80000-13, which cancels and replaces the relevant subclauses of IEC 60027-2:2005, and the 2025 edition of IEC 80000-13 extended the series with robi (Ri, 2⁹⁰) and quebi (Qi, 2¹⁰⁰) to match the new SI prefixes ronna and quetta. The series advances ten powers of two at a rung, which is the detail most reproductions of it get wrong: kibi 2¹⁰ through yobi 2⁸⁰, then 2⁹⁰ and 2¹⁰⁰, with nothing in between.
IEEE 1541 did the same job for engineering practice. Issued as a trial-use
standard in 2002, elevated to full use on 19 March 2005, reaffirmed on 27 March
2008 and revised as IEEE 1541-2021, it mandates b for bit, B for byte, o
for octet, and that SI prefixes are never used for binary multiples.
NIST publishes the table and states the arithmetic without qualification: 1 GiB is 2³⁰ bytes is 1,073,741,824 bytes, 1 GB is 10⁹ bytes is 1,000,000,000 bytes, and SI prefixes refer strictly to powers of ten. The BIPM expressly prohibits using SI prefixes for binary multiples and points at the IEC prefixes instead. The EU has required the IEC prefixes since 2007, and CENELEC harmonised them as HD 60027-2:2003-03.
So the drive vendor is following the standard and the operating system is not. That is the opposite of the way the argument is usually put.
JEDEC is the reason the memory side is different
JEDEC’s JESD100B.01, “Terms, Definitions, and Letter Symbols for Microcomputers, Microprocessors, and Memory Integrated Circuits”, defines kilo as 2¹⁰, mega as 2²⁰ and giga as 2³⁰ for semiconductor storage capacity. It is the one standards body that blesses the binary reading of an SI prefix, which is why a “16 GB” DIMM really is 17,179,869,184 bytes and is correctly labelled.
Two details of that document get dropped when it is cited. It says the power-of-two definitions “are included only to reflect common usage” - it is recording practice, not endorsing it. And it defines only kilo, mega and giga that way. JEDEC has never defined a binary tera. A “1 TB” memory module has no standards basis for being 2⁴⁰ bytes at all.
Why the IEC prefixes failed, in Microsoft’s own words
Microsoft’s stated reason for never adopting them is comprehension rather than correctness, and Raymond Chen put it plainly on the Old New Thing: “Every document on the Internet (to within experimental error) which talks about memory and storage uses the terms kilobyte/KB, megabyte/MB, gigabyte/GB… In other words, the entire computing industry has ignored the guidance of the IEC.” And: “If Explorer were to switch to the term kibibyte, it would merely be showing users information in a form they cannot understand, and for what purpose?”
That is a defensible position for a file manager and an indefensible one for a specification, and it is why Windows is the last major platform still printing binary values under decimal labels. macOS switched to decimal in Mac OS X 10.6 in 2009, which is why a 1 TB drive appeared to gain about 70 GB when people upgraded. Android and iOS are decimal.
Three settlements about disclosure, one ruling about the unit
Decimal capacity labelling has been litigated repeatedly and now stands. Safier v. Western Digital settled in 2006, with software worth roughly $30 to buyers between 22 March 2001 and 15 February 2006. Cho v. Seagate settled in 2007 with a cash refund or software. Vroegh v. Kodak, filed in 2004 over flash, settled on the same shape of terms: the manufacturers kept the decimal definition and undertook to print it on the packaging and on their web sites. Then in Dinan v. SanDisk the court ruled that GB meaning 1,000,000,000 bytes is not deceptive, and the Ninth Circuit affirmed in February 2021.
The settlements were about disclosure. The ruling was about the unit. The unit won.
The vendor accused of lying is the one using the prefixes properly
Western Digital’s Ultrastar DC HC580 specification carries a glossary that gets all four right in one place: “GB 1,000,000,000 bytes”, “TB 1,000,000,000,000 bytes”, “KiB 1,024 bytes”, “MiB 1,048,576 bytes”. Seagate prints a boilerplate on every product manual: “When referring to drive capacity, one gigabyte, or GB, equals one billion bytes and one terabyte, or TB, equals one trillion bytes. Your computer’s operating system may use a different standard of measurement and report a lower capacity.”
The industry’s inconsistency is not marketing, it is addressing. A DRAM chip or a NAND die is selected by binary address lines, so its capacity is a power of two and calling 17,179,869,184 bytes “16 GB” describes something real. A hard drive is a count of sectors, and there is no electrical reason for that count to be a power of anything - a platter holds whatever number of sectors the areal density gives it. Storage went decimal because storage is not binary underneath.
This site inherits that split rather than resolving it. Capacities here are decimal, because decimal is the unit on the label and the unit the seller set the price against. Pick one and stay in it.
What every tool on your own machine means by “G”
The Linux answer is usually given as “df is binary, fdisk prints both”, which is wrong in a way that matters, because the most decimal tool in the set is missing from it. Run against one 4 TB drive and one root filesystem on Ubuntu 26.04:
| Tool and flag | Unit it means | What it printed |
|---|---|---|
df -h |
powers of 1,024 | 468G |
df -H |
powers of 1,000 | 502G |
df (no flag) |
1K-blocks, binary | 1K-blocks column |
POSIXLY_CORRECT=1 df |
512B-blocks | 512B-blocks column |
lsblk |
powers of 1,024, printed as K/M/G | 476.9G |
lsblk -b |
bytes | 512110190592 |
fdisk -l header |
TiB, correctly labelled | 3.64 TiB, 4000787030016 bytes |
fdisk -l size column |
powers of 1,024, printed as T | 3.6T |
parted print |
powers of 1,000 | 4001GB, start 1049kB |
parted unit B print |
bytes | 4000787030016B, start 1048576B |
ls -lh, du -h |
powers of 1,024 | 977K for a 1,000,000-byte file |
ls --si, du --si |
powers of 1,000 | 1.0M for the same file |
btrfs filesystem usage |
TiB, correctly labelled | 3.64TiB |
One machine, one drive, four different numbers out of one binary. GNU
coreutils df alone prints 468G, 502G, a 1K-block count and a 512B-block count
for the same filesystem depending on flags and environment.
Three entries deserve singling out. parted is decimal by default and
contradicts every other partitioning tool on the system, so a capacity read out
of parted print and compared against one read out of fdisk will disagree by
9% with no error anywhere. lsblk’s own manual admits the offence outright:
sizes are “in units that are powers of 1024 bytes” and “the formal abbreviations
for these units (KiB, MiB, GiB, …) are further shortened to just their first
letter: K, M, G”. That is a gibibyte wearing a decimal letter, which is precisely
what Windows is criticised for. And btrfs-progs is the one tool in the list that
prints the IEC unit correctly, which is a small mercy given what df does to
btrfs.
When a number has to be exact, ask for bytes. The three that cannot be
misinterpreted are lsblk -b, parted unit B print, and
blockdev --getsize64.
Reading the capacity off the drive yourself
Everything above is checkable in about thirty seconds, and on a used drive it should be. The number to compare against is the IDEMA or SFF-8447 figure for the nominal capacity, from the tables above.
sudo blockdev --getsize64 /dev/sdb # bytes, no units to misread
lsblk -b /dev/sdb # bytes, device and partitions
sudo fdisk -l /dev/sdb # bytes and sector count together
sudo hdparm -I /dev/sdb | grep -iE 'device size|sector size|Model|Serial'
sudo smartctl -i /dev/sdb # "User Capacity" and "Sector Sizes"
hdparm -I reports the logical and physical sector sizes separately, which is
how you tell 512n from 512e from 4Kn without guessing, and smartctl -i prints
both alongside the model and serial. On NVMe the fields are named by the
specification rather than by the tool:
sudo nvme id-ns /dev/nvme0n1 -H | grep -iE 'nsze|ncap|nuse|lbaf'
sudo nvme id-ctrl /dev/nvme0n1 -H | grep -iE 'tnvmcap|unvmcap'
| Field | What it is |
|---|---|
| NSZE | Namespace size, in logical blocks |
| NCAP | Namespace capacity, the blocks actually allocated to it |
| NUSE | Namespace utilisation, the blocks currently written |
| TNVMCAP | Total NVM capacity of the whole controller, in bytes |
| UNVMCAP | Unallocated NVM capacity, in bytes |
NSZE and TNVMCAP disagreeing is not a fault, it is namespace management. A host can create a namespace smaller than the drive, and the remainder then shows up in UNVMCAP. That is the one way capacity genuinely disappears rather than merely going unpartitioned, and it is the cleanest way to add over-provisioning deliberately. It is also a way a used enterprise NVMe arrives smaller than advertised, and it is recoverable by deleting and recreating the namespace.
4Kn, 512e and 512n: the sector size changes the platter, not the byte count
Three format types exist, and only two numbers distinguish them:
| Format | Logical sector | Physical sector |
|---|---|---|
| 512n | 512 bytes | 512 bytes |
| 512e | 512 bytes | 4,096 bytes |
| 4Kn | 4,096 bytes | 4,096 bytes |
None of them changes the capacity. The 4 TB arithmetic above gave 4,000,787,030,016 bytes under both IDEMA expressions. What the sector size changes is how much of the platter is available to hold those bytes in the first place, and the gain there is the reason Advanced Format exists.
The capacity motive is format efficiency
Every sector on a platter carries overhead that is not data: a gap, a sync pattern, an address mark and an error correction code. That overhead is roughly fixed per sector, so making sectors eight times larger spreads it eight times further. Dell’s 512e and 4Kn white paper gives the layout:
512-byte sector 512 data + 50 ECC, in a 577-byte footprint
512 ÷ 577 = 88.7% format efficiency
4,096-byte sector 4,096 data + 100 ECC, in a 4,211-byte footprint
4,096 ÷ 4,211 = 97.3% format efficiency
Dell’s own summary: “This yields a format efficiency of 97 percent, almost a 10 percent improvement.” Reported gains in usable platter area run 7 to 11 per cent. Note that the ECC field doubles rather than growing eightfold, because a longer codeword corrects more efficiently - the density gain and the error correction improvement come from the same change.
The 512e penalty is read-modify-write, and it is a write penalty only
A 512e drive tells the host it has 512-byte sectors and does not. The media cannot be written in anything smaller than 4,096 bytes, so a host write that is misaligned, smaller than 4K, or not a multiple of 4K forces the drive to read the whole physical sector, merge the change and write it back - which, as Western Digital’s Advanced Format technical brief puts it, “can require additional revolutions of the hard disk”. IDEMA and vendor analysis put roughly 5 to 10 per cent of writes in a typical business PC environment into that category on an unaligned volume.
Reads cost nothing. This is a write problem, exactly as shingled recording is a write problem for a different reason - see CMR vs SMR.
The alignment rule is arithmetic, not judgement
A partition on a 512e drive must start on a multiple of eight logical sectors. There are eight possible positions for LBA 0 within a 4K physical sector; Dell calls the correct one “Alignment 0”, and the other seven all produce read-modify-write on every straddling write.
start LBA mod 8 = 0 aligned
start LBA = 2,048 the modern default, = 1 MiB
2,048 mod 8 = 0 satisfied, and satisfied for any physical
sector size up to 1 MiB
The 1 MiB default that every current partitioner uses is deliberately far more generous than the eight-sector rule needs. It exists so that the same alignment also suits RAID stripe units, SSD erase blocks and hypervisor block sizes, none of which are 4 KiB. The historical failure case is a Windows XP-era partition starting at sector 63, which is odd and therefore misaligned by construction.
Converting between them takes seconds and moves no data
Modern drives change sector size in the field with SET SECTOR CONFIGURATION EXT
(opcode B2h, defined in ACS-4). Seagate’s Exos X24 manual: “The selected sector
size change occurs immediately upon command completion. Default shipping format
is 512E.” Western Digital adds the detail that matters for a listing: “Changing
the block size does not change the HDD Model Number reported by the drive.”
So the sector format is not a property of the model number, it is a property of how the last owner left the drive, and it is not a reason to reject a used enterprise drive. Enterprise drives covers the reformat in practice, including the far nastier 520-byte case.
T10 Protection Information: the 2% that only bites SAS buyers
This is the most likely reason a used drive is genuinely short of capacity, and it does not appear in general explanations of the subject at all.
SAS drives can be formatted with T10 Protection Information: eight extra bytes appended to every logical block, carrying a guard CRC, an application tag and a reference tag, so that corruption is detectable end to end rather than only on the media. The drive stores 520-byte blocks instead of 512, or 4,160 instead of 4,096. The extra eight bytes come out of the capacity.
Seagate’s Exos X24 SAS product manual publishes both guaranteed operating points:
24 TB without PI 46,875,541,504 blocks × 512 = 24,000,277,250,048 bytes
24 TB with PI 45,923,434,496 blocks × 512 = 23,512,798,461,952 bytes
loss = 487,478,788,096 bytes
= 2.03%
16 TB without PI 31,251,759,104 blocks × 512 = 16,000,900,661,248 bytes
16 TB with PI 30,616,322,048 blocks × 512 = 15,675,556,888,576 bytes
= 2.03% again
8 ÷ 520 = 1.538%, and 8 ÷ 512 = 1.563%, and the observed loss is 2.03%. The extra is the drive reserving whole blocks rather than fractions and rounding the guaranteed count down to a published figure. The ratio is identical at both capacities, which is what tells you it is a formatting rule rather than a per-model decision.
There is a documented contradiction here and it is worth stating rather than smoothing over. IDEMA LBA1-03 says the eight PI bytes are protocol overhead rather than user-addressable space, and therefore “the number of LBA count on a T10 PI formatted drives must be the same as their non-T10 PI counterparts for the same reported capacity”. Real drives above 8 TB do not behave that way. The Exos X24 figures above are from the vendor’s own manual and they differ by 2.03%. Trust the datasheet over the standard here; the standard is describing the world below the SFF-8447 boundary.
The practical consequences:
- A used SAS drive sold as 24 TB that arrives reporting 23.5 TB is almost certainly PI-formatted, not clipped and not counterfeit.
- Supported block sizes on the Exos X24 are 512 and 520 for the 512e family, 4,096 and 4,160 for 4Kn. A block size of 520 or 4,160 is the tell.
- Reformatting to 512 or 4,096 returns the capacity. It destroys the data and it takes hours on a large drive, which is why the decision belongs before you build the array rather than after.
- SATA drives cannot do any of this. If you are buying SAS, check; otherwise the section does not apply to you.
The two ways a drive can hide capacity, and only one of them is HPA
A drive tells the host how large it is, and that answer can be reduced in two independent places. Most write-ups cover the first and not the second.
Host Protected Area, defined by T13 in ATA-4 and published as ANSI NCITS 317-1998, reserves sectors at the end of the drive that the host cannot address. Laptop vendors used it for recovery images; sellers set it by accident. It survives a reformat and a repartition, because it is below both.
sudo hdparm -N /dev/sdb
The man page describes the output precisely: “the current setting, which is
reported as two values: the first gives the current max sectors setting, and the
second shows the native (real) hardware limit for the disk”. Two different
numbers mean an HPA. Prepending p to a new value makes the change permanent,
and hdparm’s own warning on that path is “VERY DANGEROUS, DATA LOSS IS EXTREMELY
LIKELY”.
Device Configuration Overlay is a second, separate clip, and hdparm -N does
not reveal it. DCO arrived four years after the HPA, in ATA-6, and it works one
level lower: it can reduce the reported native capacity itself, so the two
numbers -N prints can agree with each other while the drive is still short. It
is queried and reset with different flags:
sudo hdparm --dco-identify /dev/sdb # what the vendor or OEM disabled
sudo hdparm --dco-restore /dev/sdb # DO NOT RUN THIS CASUALLY
--dco-identify will “query and dump information regarding drive configuration
settings which can be disabled by the vendor or OEM installer”. --dco-restore
will “reset all drive settings, features, and accessible capacities back to
factory defaults and full capabilities”, and hdparm labels it “EXTREMELY
DANGEROUS” and says “DO NOT USE THIS COMMAND”. Both warnings are the tool
author’s, not this page’s, and both are justified.
A drive can be short by HPA, by DCO, or by both at once. On a used drive that
reports less than the standard figure, check -N first, then --dco-identify,
and only then conclude the drive is misdescribed.
MBR, which is a table limit rather than a drive limit
If a drive over 2.2 TB reports as exactly 2.2 TB, this is why. An MBR partition table stores both the starting LBA and the sector count in 32-bit fields, and the largest value a 32-bit count can hold is 2³² − 1, not 2³²:
(2³² − 1) × 512 = 4,294,967,295 × 512 = 2,199,023,255,040 bytes
2 TiB = 2,199,023,255,552 bytes
short by 512 bytes
The round ceiling of 2,199,023,255,552 is quoted almost everywhere. It is one sector too high.
The same fields with 4,096-byte logical sectors reach (2³² − 1) × 4,096 = 17,592,186,040,320 bytes, one 4 KiB sector short of 16 TiB by the same rule -
the identical off-by-one, eight times larger. That is the reason a 512e drive
reformatted to 4Kn can sometimes be made to work past 2.2 TB under MBR on a
system that cannot boot GPT. Everything past the ceiling is unaddressable by the table, not
by the drive. Convert to GPT and it returns. The same 32-bit ceiling turns up in
cheap USB-SATA bridges and old RAID controllers, and there the enclosure has to
go - drive interfaces covers which bridges do it.
Where the rest of it goes: a measured ext4 ledger
Take that 4 TB drive through an ordinary Linux install, GPT plus ext4 with
mke2fs defaults. Every figure below comes from creating the filesystem on a
real 4,000,785,964,544-byte partition with mke2fs 1.47.2 and reading it back with
dumpe2fs, rather than from the arithmetic people usually reproduce.
Drive 4,000,787,030,016 B 3.6387 TiB
GPT
protective MBR, header, entry array 17,408 B LBA 0-33
alignment padding to LBA 2048 1,031,168 B LBA 34-2047
backup header and array at the end 16,896 B last 33 LBAs
partition 4,000,785,964,544 B 3.6387 TiB
ext4, mke2fs 1.47.2 defaults
partition ÷ 4,096 = 976,754,385.875 last 3,584 B unused
inode table 244,195,328 × 256 B 62,514,003,968 B 58.22 GiB
journal 262,144 blocks × 4,096 B 1,073,741,824 B 1.00 GiB
bitmaps, group descriptors, reserved GDT,
superblock backups, 29,809 groups 378,552,320 B 361 MiB
total overhead, 15,616,772 clusters 63,966,298,112 B 59.57 GiB
df "Size" 3,936,819,662,848 B 3.5805 TiB
root reserve, 48,837,719 blocks 200,039,297,024 B 186.30 GiB
df "Avail" on an empty filesystem 3,736,780,365,824 B 3.3986 TiB
That ledger subtracts, and a published one that does not was assembled rather
than measured. df “Size” is the partition less the 3,584 unused bytes and
less the 63,966,298,112 of overhead; df “Avail” is “Size” less the root
reserve, exactly. Check any capacity breakdown you find against that test before
believing it.
Four things in it the arithmetic will not tell you and only a dumpe2fs will.
The journal is 1 GiB, not 128 MiB. ext2fs_default_journal_size has a ladder
- 1,024 blocks under 128 MB, 4,096 under 1 GB, 8,192 under 2 GB, 16,384 under
16 GB, 32,768 under 32 GB, 65,536 under 64 GB, 131,072 under 128 GB, and 262,144
above that. At 4 KiB blocks the top rung is exactly 1 GiB, and
dumpe2fson the freshly created filesystem reports “Total journal size: 1024M”. The widely repeated 128 MiB figure has been wrong for years.
The bitmaps and descriptors are 361 MiB, not 1.5 GiB. The figure usually
given is about four times too large, and it happens to cancel against the
understated journal, which is why tables built from both still produce a
plausible df total. Two errors that agree are harder to find than one.
The inode table is 97.7% of all filesystem metadata. Everything else - journal, bitmaps, group descriptors, reserved GDT blocks, superblock backups across 29,809 block groups - is the remaining 2.3%. Anyone optimising ext4 overhead who is not changing the inode ratio is optimising the wrong thing.
A partition is rarely a whole number of filesystem blocks. 4,000,785,964,544 ÷ 4,096 = 976,754,385.875, so the last 3,584 bytes are simply unused. It is a trivial amount and it is a real layer.
The partition table costs about a megabyte, and anyone blaming it is wrong
GPT is 34 sectors at the front and 33 at the back, and those 34 are routinely
described as alignment padding in full. They are not. Running
sgdisk on the image reports “Main partition table begins at sector 2 and ends
at sector 33. First usable sector is 34”, and “Total free space is 2014 sectors
(1007.0 KiB)”. So:
LBA 0 protective MBR 512 B
LBA 1 primary GPT header 512 B
LBA 2-33 partition entry array, 32 LBAs 16,384 B
LBA 34-2047 alignment padding, 2,014 LBAs 1,031,168 B
last 33 LBAs backup array and header 16,896 B
total 1,065,472 B
0.0000266% of the drive
The 34/33 split is a consequence of the UEFI specification reserving at least 16,384 bytes for the GUID Partition Entry Array - 128 entries of 128 bytes. At 512-byte blocks that is 32 LBAs, so 1 + 1 + 32 = 34 at the front and 32 + 1 = 33 at the back.
On a 4Kn drive the array is only 4 blocks, but GPT costs more in bytes, not less. Six blocks at the front and five at the back is 45,056 bytes, against 34,304 at 512. The count went down and the footprint went up, which is the sort of thing that makes a per-sector intuition fail.
The 5% root reserve is the biggest single recoverable chunk, and it may not be 5%
186 GiB of that 4 TB is not overhead at all. It is ordinary free space that only
root may allocate, and mke2fs sets it from the profile with a hardcoded
fallback of reserved_ratio = 5.0. The man page gives the reason: it “avoids
fragmentation, and allows root-owned daemons, such as syslogd(8), to continue to
function correctly after non-privileged processes are prevented from writing to
the file system.” Root-owned is the load-bearing word. The reserve is useful
precisely because the processes that must keep writing when the disk fills run as
root and can allocate into it.
On a dedicated data drive neither reason applies:
sudo tune2fs -m 1 /dev/sdb1 # 5% -> 1%, instant, no unmount needed
sudo tune2fs -l /dev/sdb1 | grep -E 'Block count|Reserved block count'
Check what you actually have before assuming there is 5% to recover.
Distributions edit the defaults, and the reserve is edited more often than the
inode ratio is. On an Ubuntu 26.04 root filesystem stat -f / reports
122,512,118 blocks with 116,393,871 free and 115,142,264 available, which is a
reserve of 1,251,607 blocks - about 1.02% of the block count, not 5%.
What format-time options actually recover
Run mke2fs -n against the same 4 TB partition with different flags and the
inode table moves by tens of gigabytes:
| Options | Inodes | Inode table | Recovered against default |
|---|---|---|---|
| defaults | 244,195,328 | 58.22 GiB | - |
-i 65536 |
61,048,832 | 14.56 GiB | 43.67 GiB |
-i 262144 |
15,262,208 | 3.64 GiB | 54.58 GiB |
-T largefile |
3,815,552 | 0.91 GiB | 57.31 GiB |
-T largefile4 |
953,888 | 0.23 GiB | 57.99 GiB |
-I 128 halves the table without changing the inode count, at the cost of
nanosecond timestamps and post-2038 dates, which is a bad trade on anything built
in 2026.
Combined, format-time choices are worth more than every other loss on this page put together:
4 TB, mke2fs defaults Avail 3,736,780,365,824 B
4 TB, mke2fs -i 262144 -m 0 Avail 3,995,426,541,568 B
gain 258,646,175,744 B
= 258.6 GB = 6.47% of the label
Two flags. None of it is recoverable afterwards - the inode ratio is fixed at
format time - which is the argument for deciding it before the drive has data on
it. Set it low only if you know the filesystem will hold large files, and read
the inode column of that table before choosing. -T largefile4 leaves 953,888
inodes on this partition, so a million small files exhaust the table on a drive
showing terabytes free, and the error message will not say so clearly. One inode
per 256 KiB leaves 15,262,208 of them, which a mail spool or a package cache can
still reach and an ordinary filesystem will not.
The ext4 inode-ratio cliff punishes the bigger drive
mke2fs picks its usage type from the filesystem size, and the branch points are
in parse_fs_type in mke2fs.c. They are binary thresholds, and nobody expects
that. The source computes meg as the number of blocks in one MiB, then
branches: under 3 MiB “floppy”, under 512 MiB “small”, under 4*1024*1024*meg
(4 TiB) the default, under 16*1024*1024*meg (16 TiB) “big”, otherwise “huge”.
The ratios that follow from /etc/mke2fs.conf on e2fsprogs 1.47.2 are
inode_ratio 16,384 by default, 32,768 for big, 65,536 for huge.
A 4 TB drive is 3.639 TiB. It is below the first threshold, so it gets the densest inode ratio in the table. A 5 TB drive is 4.548 TiB and does not:
| Drive | Size in TiB | Usage type | Inodes | Inode table |
|---|---|---|---|---|
| 4 TB | 3.639 | default | 244,195,328 | 58.22 GiB |
| 5 TB | 4.548 | big | 152,621,056 | 36.39 GiB |
| 16 TB | 14.553 | big | 488,308,736 | 116.42 GiB |
| 17 TB | 15.462 | big | 518,815,744 | 123.70 GiB |
| 18 TB | 16.371 | huge | 274,661,376 | 65.48 GiB |
The rows above 8 TB are built on the SFF-8447 byte counts from earlier on this page, not on IDEMA run past its boundary, and the distinction is not cosmetic. The IDEMA figure for 16 TB is 2.2 GB larger, which is enough to add seventeen block groups and 69,632 inodes that the drive does not have. Any inode table you find quoted for a drive above 8 TB is worth checking for that, because the wrong capacity produces a plausible-looking wrong answer rather than an obvious one.
A 5 TB drive carries 21.8 GiB less inode table than a 4 TB one, in absolute terms, not as a proportion. The same inversion repeats at the top: an 18 TB drive gets 43.8% fewer inodes than a 16 TB one, and 50.9 GiB less table. The crossover is at 16 TiB, which is 17.592 TB, so every drive on sale between 16 TB and 17.5 TB is on the dense side of it and every drive from 18 TB up is not.
None of this is a bug. It is a sensible heuristic - large filesystems usually hold large files - expressed in units the buyer of a “5 TB” drive has no reason to be thinking in. It is the clearest case on this page of a binary threshold producing a result that looks wrong in decimal.
Choosing the filesystem is worth more than everything else on this page
Put ext4, XFS and btrfs on the identical 4,000,785,964,544-byte partition and read back what each one took before a single user byte was written:
| Filesystem | Metadata at format | Free at format | As % of the 4 TB label |
|---|---|---|---|
| ext4, defaults, root reserve applied | 59.57 GiB | 3,736,780,365,824 B | 93.42% |
| ext4, defaults, as root | 59.57 GiB | 3,936,819,662,848 B | 98.42% |
| XFS, defaults | 1.82 GiB | 3,998,832,308,224 B | 99.97% |
| btrfs, defaults | 2.02 GiB | roughly 3.64 TiB reported | - |
XFS free at mkfs 3,998,832,308,224 B
ext4 Avail at mkfs 3,736,780,365,824 B
difference 262,051,942,400 B = 262.1 GB = 6.55% of the label
One 4 TB drive, one decision, 262 GB. That is larger than every other loss on this page combined, and it is not lost capacity at all - it is front-loaded metadata policy.
The comparison flatters XFS and the honest version says so. XFS allocates inodes
on demand rather than up front, so its advantage closes as the filesystem fills
with small files and the metadata it did not build at format time gets built
later. xfs_db on that 4 TB filesystem reports dblocks 976,754,385 - identical
to ext4’s block count, as it should be - an internal log of 476,930 blocks
(1,953,505,280 bytes, 1.819 GiB), imaxpct 5, isize 512 with CRCs on by
default, agcount 4 and agblocks 244,188,597. The imaxpct defaults are
documented as “25% for filesystems under 1TB, 5% for filesystems under 50TB and
1% for filesystems over 50TB”, so XFS has a ceiling on inode space rather than a
reservation of it.
btrfs duplicates its metadata, and df is the wrong tool on it
mkfs.btrfs on the same image allocates “Data: single 8.00MiB, Metadata: DUP
1.00GiB, System: DUP 8.00MiB” with a node size of 16,384 - 2.02 GiB before a byte
of user data, because DUP is the default metadata profile on a single device
and metadata is therefore written twice.
The reporting is the bigger problem. btrfs has a two-stage allocator: chunks are
allocated to data or metadata first, then filled. The documentation is blunt that
df “cannot tell how much unallocated disk space is available”, and a filesystem
with all space allocated but metadata near full will show free space in df and
still return ENOSPC. Use btrfs filesystem usage, and read its data ratio -
1.0 for single, 2.0 for DUP and RAID1.
ZFS keeps a slop reserve, and it is capped
The commonly quoted figure is that ZFS holds back 1/32 of a pool. The OpenZFS
zfs(4) documentation says it precisely: “Normally, we don’t allow the last 3.2%
(1/2^spa_slop_shift) of space in the pool to be consumed”, with spa_slop_shift
defaulting to 5.
The fraction is right and the implication is misleading on the pools this site’s
readers build, because the reserve is floored at 128 MiB and capped at 128 GiB,
and never exceeds half the pool. A 200 TB pool holds back 128 GiB, not 6.25 TB.
Nor is it wholly unavailable: file removal and most administrative actions may
use half of it, and operations almost certain to free space, such as zfs destroy, up to three quarters.
The other ZFS surprise is documented just as explicitly. From zpoolprops(7):
“the zpool free property is not generally useful for this purpose, and can be
substantially more than the zfs available space”, and “The space usage properties
report actual physical space available to the storage pool. The physical space
can be different from the total amount of space that any contained datasets can
actually use.” zpool list counts parity; zfs list and df do not. People
compare the two and conclude a pool lost space it never had.
NTFS reserves 12.5% for the MFT, and does not take it away
Windows reserves an MFT zone of 12.5% of the volume by default, set by
NtfsMftZoneReservation under
HKEY_LOCAL_MACHINE\System\CurrentControlSet\Control\FileSystem, a REG_DWORD
valid from 1 to 4 with a default of 1. That reservation is not missing
capacity, and Microsoft says so directly: “The MFT Zone is not subtracted from
available (free) drive space used for user data files, it is only space that is
used last.” Microsoft also declines to publish what values 2 to 4 mean, on the
grounds that the ratios “are undocumented because they are not standardized and
may change in future releases.”
NTFS’s cluster size sets the volume ceiling, because the cluster count field is 32 bits and the limit is 2³² − 1 clusters:
| Cluster size | Maximum volume and file size |
|---|---|
| 4 KB (default) | 16 TB |
| 8 KB | 32 TB |
| 16 KB | 64 TB |
| 32 KB | 128 TB |
| 64 KB | 256 TB |
| 128 KB | 512 TB |
| 1,024 KB | 4 PB |
| 2,048 KB | 8 PB |
Those are the figures for Windows Server 2019 and later and Windows 10 1709 and later. A 20 TB drive formatted NTFS at the default 4 KB cluster will not take a single volume, which is a live problem now that single drives exceed 16 TB, and the fix is a larger cluster at format time. VSS-based backup caps the practical volume at 64 TB regardless.
APFS free space is a container property, not a volume property
On macOS the question “how much space is free” has no single answer by design.
Free space belongs to the container, not to any volume inside it: it is the
container’s capacity less what each of its volumes uses, shared among all of
them. macOS then folds “purgeable” space - mostly APFS and Time Machine snapshots
plus caches - into the figure that Finder, About This Mac and Disk Utility each
call “available”, which is why those three disagree with each other and with
df. diskutil apfs list gives the container figure, and it is the one to
believe.
The container is also not the whole disk. On a 2 TB Mac SSD the Macintosh HD container holds about 1.995 TB, with iSCPreboot at roughly 524 MB and Recovery at roughly 5.4 GB taking the rest - Apple’s own hidden volumes, before any user data exists.
And below all of them sits slack
A filesystem with 4 KiB blocks wastes on average half a block per file:
1,000,000 files × 2,048 B average waste = 2,048,000,000 B = 1.91 GiB
That is why a backup of a mail spool never matches the size the source reported.
exFAT makes it much worse if misconfigured: the specification permits clusters up
to 32 MB, via SectorsPerClusterShift of at most 25 minus BytesPerSectorShift,
so a badly chosen cluster size on a card full of small files can waste orders of
magnitude more than any figure on this page.
SSD over-provisioning: where the binary NAND went
An SSD inverts the problem. NAND is built in binary quantities and the drive presents a decimal capacity, so the gap between the two is not lost. It is kept by the controller, and it is doing work.
Before the arithmetic, the connection that is usually missed: SSDs follow the same IDEMA table as hard drives. Seagate’s BarraCuda SSD product manual, publication 100835666 Rev A, cites it by name - “Sector Size: 512 Bytes. User-addressable LBA count = ((97696368) + (1953504 x (Desired Capacity in Gb-50.0)) From International Disk Drive Equipment and Materials Association (IDEMA) (LBA1-03_standard.doc)”. An SSD is not a separate capacity regime. It is the same one, with a different physical substrate underneath.
The 7.4% figure is wrong for most drives on sale
The standard explanation says a consumer SSD carries about 7.37% spare, from
2³⁰ ÷ 10⁹. That is only true when the label number is itself a power of
two. Run it against real IDEMA user capacities and the common case is
noticeably different:
| Sold as | NAND on board | User bytes (IDEMA) | Spare | Over-provisioning |
|---|---|---|---|---|
| 512 GB | 512 GiB | 512,110,190,592 | 37,645,623,296 | 7.35% |
| 500 GB | 512 GiB | 500,107,862,016 | 49,647,951,872 | 9.93% |
| 480 GB | 512 GiB | 480,103,981,056 | 69,651,832,832 | 14.51% |
| 400 GB | 512 GiB | 400,088,457,216 | 149,667,356,672 | 37.41% |
| 1,000 GB | 1 TiB | 1,000,204,886,016 | 99,306,741,760 | 9.93% |
| 960 GB | 1 TiB | 960,197,124,096 | 139,314,503,680 | 14.51% |
| 800 GB | 1 TiB | 800,166,076,416 | 299,345,551,360 | 37.41% |
| 2,000 GB | 2 TiB | 2,000,398,934,016 | 198,624,321,536 | 9.93% |
| 1,920 GB | 2 TiB | 1,920,383,410,176 | 278,639,845,376 | 14.51% |
| 3,840 GB | 4 TiB | 3,840,755,982,336 | 557,290,528,768 | 14.51% |
| 7,680 GB | 8 TiB | 7,681,501,126,656 | 1,114,591,895,552 | 14.51% |
| 120 GB | 128 GiB | 120,034,123,776 | 17,404,829,696 | 14.50% |
Most consumer SSDs today carry decimal labels - 500 GB, 1 TB, 2 TB - so the common baseline is 9.93%, not 7.4%. Only a drive labelled 512 GB or 256 GB, where the number is a power of two, sits at 7.35%.
That correction changes the enterprise comparison too. A 960 GB drive and a 1 TB drive of the same family are usually the same silicon with a different firmware capacity, and the premium buys 14.51% against 9.93% - about 4.6 percentage points of extra spare, not the 7 the old arithmetic implied. The 800 GB against 1 TB comparison is the one that is genuinely dramatic: 37.41% against 9.93%.
The formula itself is uncontroversial, and stated identically by SNIA-adjacent vendor material and by the general literature:
OP (%) = (physical capacity − user capacity) ÷ user capacity × 100
Samsung's worked example: ((128 − 120) ÷ 120) × 100 = 6.7%
Samsung’s own note on that example - “Typically Samsung DC SSD provides 6.7 percent of factory OP as its default” - is a good illustration of why the denominator matters. The same drive described against its 128 GiB of NAND and its 120 GB of user space with exact IDEMA bytes gives 14.50%, because 128 GiB is 137,438,953,472 bytes rather than 128 × 10⁹. Two vendors quoting over-provisioning are not necessarily using the same denominator, and the formula above is the one to re-run yourself.
What the spare buys, and the identity that cannot be simplified
NAND erases in blocks far larger than the pages it writes, so freeing a block means relocating the live pages inside it first, and those relocations are writes nobody asked for. That ratio is write amplification, and SSD endurance derives its floor from block geometry.
Hu et al. at IBM Zurich Research Laboratory set the terms precisely in “Write
Amplification Analysis in Flash-Based Solid State Drives” (SYSTOR 2009). For a
raw capacity of t blocks of which u are user-visible, the over-provisioning
factor is Of = t/u and the spare factor is Sf = (t − u)/t. Their conclusions:
write amplification falls as over-provisioning rises, and separating static from
dynamic data reduces it further. Their model holds for a uniformly distributed
random workload of 4 KiB writes and is not claimed beyond it, which is exactly
the honesty a reader needs here.
Samsung publishes the endurance identity rather than a write-amplification formula:
WAF = physical write amount ÷ host write amount
DWPD = (NAND P/E cycles × raw density)
÷ (logical density × 365 × warranty years × WAF)
Samsung’s own worked estimates for the 860 DCT show what that buys: a 960 GB drive at 0% additional user over-provisioning rates 0.34 DWPD over three years; at 10% user OP (872 GB exposed) 0.71; at 20% (800 GB exposed) 0.96. Nearly tripling endurance for 17% of the capacity. Those three are calculated from the identity above and the paper says so in terms - “a calculated value of each SSD, not a guaranteed value” - so treat them as the shape of the curve rather than as measurements of a part you might buy. Industry practice has settled on four levels - 0% additional, 7%, 14% and 28% - which is what produces the capacity ladders in the table above.
Reported real write amplification values, from vendor and press measurement rather than from this page: the Intel X25-M in 2008 as low as 1.1; SandForce SF-1000 at 0.5 with compression; SF-2281 best case 0.14. Without compression, write amplification cannot drop below 1, and a figure under 1 tells you the controller is not writing everything the host sent.
Adding your own, and the condition that is usually skipped
You can add over-provisioning by leaving part of the drive unpartitioned, on one condition.
The controller has to know those blocks are free. It sees logical block addresses and whether each has been written, not partitions. Leaving the last 20% unpartitioned on a drive that has been filled at some point achieves nothing, because every one of those addresses still holds data the controller must preserve. Erase first, then partition short:
sudo blkdiscard /dev/nvme0n1 # TRIM the whole device
# or, for a full reset of the mapping tables:
sudo nvme format /dev/nvme0n1 --ses=1
sudo sgdisk -n 1:0:+800G /dev/nvme0n1 # leave the rest untouched
nvme format also changes the LBA data size and metadata size, so it is the
command that converts an NVMe drive between 512-byte and 4,096-byte logical
blocks. Namespace management is cleaner still where the drive supports it,
because the capacity genuinely disappears from NSZE rather than merely going
unused.
Before any of that, check the cheaper option:
systemctl status fstrim.timer
sudo fstrim -av
Free space inside a trimmed filesystem is already over-provisioning. A drive
kept 40% empty with fstrim running weekly has 40% spare from the controller’s
point of view, at no cost. Manual over-provisioning is insurance against the
drive being filled later, not a substitute for TRIM working now.
So it is worth doing on a consumer drive at the 9.93% baseline that will run near full under sustained random writes - a VM host, a busy database, a write cache. It is not worth doing on a drive that will sit half empty, or on an enterprise drive already shipping at 14.51% or 37.41%.
Price per terabyte, and why delivered cost is the only honest input
Per-terabyte cost is the comparison that matters because capacity is the thing being bought, and the arithmetic is one division:
price per TB = delivered cost ÷ capacity in decimal TB
Both halves go wrong in practice. Take three listings, in arbitrary units:
| Capacity | Item | Shipping | Delivered | Per TB | |
|---|---|---|---|---|---|
| A | 8 TB | 100 | 0 | 100 | 12.50 |
| B | 8 TB | 92 | 18 | 110 | 13.75 |
| C | 6 TB | 70 | 12 | 82 | 13.67 |
Sorted by item price: C, B, A. Sorted by what you pay per terabyte: A, C, B. The cheapest-looking listing is not the cheapest and the most expensive-looking one is the best. Shipping a 3.5-inch drive is a real cost that some sellers fold into the price and some do not, which is why this site computes per-TB from delivered cost and leaves the figure blank rather than guessing when shipping cannot be determined - the methodology page sets out when that happens.
The unit error is the other half, and it is larger than the spread you are shopping for. Price A per tebibyte instead of per terabyte:
A: 100 ÷ 7.2774 TiB = 13.74 per TiB
B: 110 ÷ 8 TB = 13.75 per TB
Identical, for a drive that is 10% cheaper. Mixing TB and TiB in a comparison erases exactly the size of advantage most people are hunting for. Use the label capacity on both sides, every time. It is also the only capacity knowable before the drive arrives, since what a filesystem will report depends on decisions not yet made.
How this site derives capacity, and where that can be wrong
This site stores capacity as decimal GB by construction, and says so in its
own code rather than only in prose. Drive.capacityGb is documented as “Capacity
in DECIMAL gigabytes, as the manufacturer markets it”, and the conversion
constant is named rather than written as a literal:
GB_PER_TB = 1_000
with the comment that this exists “because the whole point of the unit discipline in types.ts is that the conversion is visible and stated rather than a literal somebody later corrects to 1024”. The methodology page commits to the same thing publicly: “Price per terabyte is the delivered cost divided by the total capacity on offer, in decimal terabytes - the unit drives are sold in, where 1 TB is 1,000 GB rather than 1,024.”
Now the honest part. The capacity on every row of this site comes from the
seller’s title text, not from the drive. statedCapacitiesGb() matches
(\d+(?:\.\d+)?)\s*(tb|gb|mb), multiplies terabytes by 1,000 and takes the
largest surviving match. It already excludes the things that look like capacities
and are not: link speeds (1.5, 3, 6, 12 and 22.5 Gb followed by SAS or SATA),
endurance odometers (“8TB written” on a 512 GB drive, which once made that drive
the cheapest SSD on the site), serial numbers with embedded digit-letter runs,
and compatibility lists naming three or more distinct capacities.
There is no step anywhere that converts binary to decimal. The number in the title is taken at face value, and that produces two errors in opposite directions.
A title written in TiB overstates price per terabyte by about 10%
A seller who types what Windows reported instead of what the label said writes “3.63TB” for a 4 TB drive:
delivered 60, capacity read as 3.63 TB = 16.53 per TB
delivered 60, capacity read as 4 TB = 15.00 per TB
error = +10.2% on the ranking
The underlying unit gap is 9.95%; the rounding in “3.63” supplies the rest. Either way the listing sorts worse than it deserves, by more than the spread most of the ranking turns on. The same seller behaviour produces 931GB for 1 TB, 465GB for 500 GB, 7.27TB for 8 TB and 14.55TB for 16 TB.
A title written “1024GB” understates it by 2.34%
The error runs the other way when a seller writes the binary number as though it
were the label. The site’s eBay capacity bands already normalise this one
spelling - the 1,000 GB band lists ['1 TB', '1TB', '1000GB', '1024GB'], mapping
1024GB onto 1,000 GB - but the free-text title parser does not do the same. A
title reading “1024GB” parses as 1,024 GB, which is 2.4% more than the drive
holds and understates price per TB by 2.34%. It also means that row will not
appear under 1 TB drives, because the filter matches
exactly.
That is not a straightforward bug to fix, because COMMON_CAPACITIES_GB.SSD
lists 1,024 and 2,048 as legitimate SSD capacities, which is true for some
vendors and false for others.
The ambiguity that cannot be resolved at all
The catalogue documents the sharpest case in its own source, and the comment is worth quoting because it is the correct answer rather than a limitation:
“1.8 TB IS THE REASON THIS LIST MAY NOT BE USED TO SNAP A VALUE. It is a real 10K SAS SKU, and it is also how a careless seller writes the size Windows reports for a 2 TB drive. Both readings are in this catalogue’s input, no rule tells them apart.”
A nearest-neighbour correction would silently rewrite one into the other, and there is no evidence in a title that distinguishes them. So the stated number stands, and the reader is told. The same trap sits at 931GB against 1TB, 465GB against 500GB, 3.63TB against 4TB and 7.27TB against 8TB - in each pair, one reading is a real product and the other is an operating system’s arithmetic.
Lots: total capacity is not the capacity you get
“10 TB” in a title can be one 10 TB drive or five 2 TB drives, and per-terabyte cost cannot tell them apart. Everything else can.
| 1 × 10 TB | 5 × 2 TB | |
|---|---|---|
| Drive bays / SATA ports | 1 | 5 |
| Idle power, at about 6 W each | about 6 W | about 30 W |
| Over a year | about 53 kWh | about 263 kWh |
| Sequential throughput | one drive’s worth | up to 5× striped |
| A single failure costs | 10 TB | 2 TB |
| Chance of at least one failure at 1% AFR | 1.00% | 4.90% |
| Expected TB lost per year | 0.10 | 0.10 |
The last two rows are the interesting pair. Five drives are almost five times as
likely to produce a failure event - 1 − 0.99⁵ = 4.90% - but each event costs a
fifth as much, so the expected bytes lost per year is identical. What differs
is the shape of the loss, and which shape you want depends entirely on whether
the drives are in an array. Five independent drives with no redundancy is the
worst arrangement available: five chances to lose something, and no parity.
The power difference is 210 kWh a year; multiply by your own tariff. Spin-up is separate again - a 3.5-inch drive pulls a couple of amps on the 12 V rail coming up to speed, and five doing it at once is what staggered spin-up exists for.
Older drives are also slower in proportion to their age, since sequential rate tracks areal density. Seagate’s Exos X24 manual publishes a maximum sustained transfer rate of 272 MiB/s at the outer diameter and 285 MB/s maximum; a 2 TB drive from the generation when 2 TB was the flagship publishes roughly half that. Five striped still beat one modern drive on throughput and lose badly on everything per watt.
Three other datasheet numbers from that manual belong here, because they are specification bounds rather than measured field rates and are routinely quoted as though they were the latter: non-recoverable read errors at 1 sector per 10¹⁵ bits read, an annualised failure rate of 0.35% based on 8,760 power-on hours, and a maximum workload rate of 550 TB per year, where “Workload Rate = TB transferred × (8760 / recorded power on hours)”. For contrast, the BarraCuda SSD product manual gives UBER as 1 error in 10¹⁶ bits read and TBW of 120, 249, 485 and 1,067 for its 250, 500, 1,000 and 2,000 GB models, with the warranty running “five years, or when the device reaches Host TBW, whichever happens first”. The un-round numbers are the tell that they were read off the manual; round ones quoted for the same drives have been through a marketing page first. SSD UBER is an order of magnitude better than HDD URE on these sheets, which is a real difference and still an upper bound rather than a measurement.
And read the quantity field before any of this applies: a title saying “10TB” beside a quantity of five may mean the price buys one drive, not the lot - buying used drives on eBay covers how that goes wrong in both directions.
Advertised against usable in an array
Parity costs capacity in a way that is exactly predictable, unlike everything else
on this page. For n drives of capacity C with p parity drives, usable is
(n − p) × C, and overhead is p ÷ n. Worked through for 4 TB members:
| Layout | Drives | Raw | Usable | Usable as TiB | Overhead | Survives |
|---|---|---|---|---|---|---|
| RAID 0 | 4 | 16 TB | 16 TB | 14.55 | 0% | nothing |
| RAID 1, 2-way | 2 | 8 TB | 4 TB | 3.64 | 50% | 1 |
| RAID 1, 3-way | 3 | 12 TB | 4 TB | 3.64 | 66.7% | 2 |
| RAID 10 | 4 | 16 TB | 8 TB | 7.28 | 50% | 1, sometimes 2 |
| RAID 5 / RAIDZ1 | 4 | 16 TB | 12 TB | 10.91 | 25% | 1 |
| RAID 5 / RAIDZ1 | 8 | 32 TB | 28 TB | 25.47 | 12.5% | 1 |
| RAID 6 / RAIDZ2 | 6 | 24 TB | 16 TB | 14.55 | 33.3% | 2 |
| RAID 6 / RAIDZ2 | 8 | 32 TB | 24 TB | 21.83 | 25% | 2 |
| RAIDZ3 | 12 | 48 TB | 36 TB | 32.74 | 25% | 3 |
Overhead falls as the array widens while the protection stays fixed, which is the whole argument for RAID 6 over RAID 5 at width and against very wide RAID 5. “Sometimes 2” for RAID 10 means a second failure is survivable only in the other mirror: with four drives, two of the three possible second failures are.
Four things that table cannot show.
RAID-Z does not hit those numbers, and the rule that says why is misattributed
ZFS writes variable-width stripes, and the allocation constraint that explains
the shortfall is routinely cited to zpoolprops(7), which does not contain it -
the man page says only that “the amount of space used in a raidz configuration
depends on the characteristics of the data being written”. The rule is Matt
Ahrens’s, from his analysis of RAID-Z space accounting: RAID-Z “requires that
each allocation be a multiple of (p+1), so that when it is freed it does not
leave a free segment which is too small to be used”, where the unit is a sector
of 2^ashift bytes. Allocations round up, producing up to three pad sectors per
block whose contents are never used. ashift is a pool property, valid from 9 to
16, defaulting to 0 meaning autodetect. Cite the analysis, not the man page -
a reader who goes to zpoolprops(7) looking for the p+1 sentence will not find
it.
Working the rule through at ashift=12 - 4 KiB sectors, which is what a modern
drive wants - gives efficiencies that differ from the naive (n − p)/n by several
points:
| Layout | Efficiency at 128 KiB recordsize | Naive (n−p)/n | Efficiency at 4 KiB recordsize |
|---|---|---|---|
| raidz1, 4 wide | 72.7% | 75.0% | 50.0% |
| raidz1, 5 wide | 80.0% | 80.0% | 50.0% |
| raidz1, 8 wide | 84.2% | 87.5% | 50.0% |
| raidz2, 6 wide | 66.7% | 66.7% | 33.3% |
| raidz2, 8 wide | 71.1% | 75.0% | 33.3% |
| raidz2, 10 wide | 76.2% | 80.0% | 33.3% |
| raidz3, 12 wide | 72.7% | 75.0% | 25.0% |
Those figures are computed from that p+1 rule, not measured on a pool. Two results in the table are worth carrying away.
At a 4 KiB record, every raidz1 is 50% efficient, every raidz2 is 33% and every
raidz3 is 25%, regardless of width. A block that small cannot be spread, so each
one costs a full parity sector per parity level. A pool of databases or VM images
with a small recordsize does not get the parity efficiency the width suggests;
a pool of large media files lands close to the naive table.
The widths that come out even are not accidental. Ahrens’s recommendation for 512-byte-sector devices is to “use at least 5 disks with RAIDZ1; use at least 6 disks with RAIDZ2; and use at least 11 disks with RAIDZ3”, and the rows above show why - 5-wide raidz1 and 6-wide raidz2 are the widths where the p+1 rounding costs nothing at 128 KiB. He also notes that “each RAID-Z group has approximately the performance of a single disk in the group” for random IOPS.
The overhead figure attached to all this is quoted in at least three conventions, and mixing them is how an 8-wide raidz2 ends up being compared with RAID 5. Do it once, in sectors, with the working shown:
8-wide raidz2, ashift=12, 128 KiB record
128 KiB ÷ 4 KiB = 32 data sectors
32 across 6 data columns = 6 rows, five of 6 sectors and one of 2
parity, 2 sectors a row = 12 sectors
32 + 12 = 44, rounded up to a multiple of p+1 = 3
= 45 sectors allocated
45 − 32 = 13 sectors of parity and padding
32 ÷ 45 = 71.1% efficient, which is the table row above
13 ÷ 32 = 40.6% on top of the data
2 ÷ 8 = 25% by the p ÷ n convention the array table uses
So that array spends about 41 sectors of parity and padding for every 100 sectors of data, where an 8-wide RAID 6 - the same protection at the same width - spends 33, and where the array table above states the same layout as 25%. The comparison is only meaningful once both sides use the same denominator, and RAID 5 is not the other side of it at all: single parity across eight drives costs 12.5% of raw.
ZFS will not show you the difference either. The deflation ratio it applies
to a raidz vdev is computed once from a 128 KiB block - vdev_set_deflate_ratio
in the OpenZFS source does it with 1 << 17 - so a pool whose datasets use a
small recordsize reports free space it cannot actually hand out.
Array metadata takes a slice before parity does
Linux md with a 1.2 superblock offsets the data start, commonly by 264,192 sectors:
264,192 × 512 = 135,266,304 B = 129 MiB per member
4-drive RAID5 of 4 TB members 3 × 135,266,304 = 405,798,912 B short of 12 TB
8-drive RAID6 of 20 TB members 6 × 135,266,304 = 811,597,824 B short of 120 TB
The offset covers the superblock itself plus room for a write-intent bitmap and
reshape headroom, and it is not a fixed constant - mdadm scales it with device
size, so check mdadm --examine rather than assuming 129 MiB.
LVM adds its own. pvcreate(8) documents that “with default settings, the first
physical extent (PE), which contains LV data, is 1 MiB from the start of the
device”, controlled by default_data_alignment in lvm.conf, with the metadata
area starting one page (usually 4 KiB) in. The default physical extent size is
4 MiB, and logical volumes round up to whole extents, so a stack of many small
LVs loses up to 4 MiB each to rounding.
A NAS charges twice, before and after parity
Synology’s own RAID calculator states the two losses that sit either side of the parity arithmetic: “Each drive in the RAID must reserve approximately 10 GB of system space”, and “Volumes with Btrfs file system reserved 4% of the capacity for metadata, and 2% for volumes with ext4 file system”.
8-bay SHR-2 of 8 TB drives
raw 8 × 8 TB = 64 TB
system space 8 × 10 GB = −80 GB
two drives of parity 6/8 of the rest ≈ 47.94 TB
btrfs metadata reserve −4% ≈ 46.02 TB
That is a worked model using Synology’s published figures, not a Synology quotation. The order matters: the 10 GB is per drive and comes off before parity, the 4% comes off the volume after it. Drives for a NAS covers the rest of the sizing decision.
The array is sized by its smallest member
IDEMA’s and SFF-8447’s fixed sector counts are why this is a non-issue between
hard drives of the same nominal capacity from different vendors -
the Exos X24 and Ultrastar DC HC580 agreeing to the byte at 24 TB is the
demonstration. SSDs are less consistent, and a replacement a few thousand
sectors short is simply refused. Compare blockdev --getsize64 across used
drives before committing rather than after.
There is a further reason a model number does not settle what you are buying: vendors second-source components under an unchanged SKU. Solidigm’s Product Change Notification 0000023684-00, dated 2 July 2025, adds new DRAM part numbers from SK hynix and Nanya and a new enclosure supplier to the D5-P5430 line, with the existing parts marked end of life, all under the same product SKUs. The capacity does not change; the hardware behind the number does.
One thing the capacity table cannot show is that rebuild time scales with member
size and array throughput and nothing else: 20 TB ÷ 200 MB/s = 100,000 s = 27.78 hours at full speed with no other load, and days on a loaded array, all of
it degraded. Whether a second drive is likely to fail in that window is the
argument for double parity, and it turns on unrecoverable read error rates that
are datasheet upper bounds rather than measured field rates. That one is
genuinely unsettled, and drives for a NAS takes it apart
honestly.
Parity is not backup, and a per-terabyte comparison that ignores the redundancy tier compares two different products. Decide the layout first, then sort on price per terabyte for the capacity that layout actually needs.
What to do with this on a listing page
- Sort by price per TB, and read the capacity column as decimal. Everything on this site is in the unit the label uses. A listing whose title reads 3.63TB, 931GB, 7.27TB or 14.55TB has been typed off a Windows screen, and its price per TB is about 10% worse than the drive deserves - that is a discount hiding in a rounding error, and it is the single most exploitable thing on this page.
- Filter on the round capacity, then search the binary spelling separately. 4 TB hard drives matches titles that say 4TB or 4000GB; it does not match one that says 3.63TB or 3726GB, because the parser takes the title at face value. The mismatched rows are where the bargains sit.
- On SAS, assume T10 Protection Information until the seller says otherwise. A 24 TB SAS drive reporting 23.5 TB is PI-formatted, not short, and a reformat to 512 or 4,096 byte blocks returns 2.03% of it. Price the listing on the label capacity and budget the hours.
- Check a large drive reports a whole number of GiB. Above 8 TB, SFF-8447
capacities land on exact gibibyte boundaries - 22,352 GiB at 24 TB, 18,627 at
20 TB, 14,902 at 16 TB. A fractional figure on arrival means HPA, DCO or a
namespace clip, and
hdparm -N,hdparm --dco-identifyandnvme id-nsare the three commands that say which. - Buy the capacity the array needs, not the raw total. Parity overhead is
p ÷ nexactly for md and hardware RAID, and worse than that for RAID-Z at small record sizes. Work out usable first, then filter. - On SSDs, read over-provisioning off the label. 480, 960, 1920 and 3840 GB are 14.51% spare; 400, 800 and 1600 are 37.41%; the round decimal numbers are 9.93%. The smaller number in each pair is the same silicon with more held back, and it is the endurance tier stated in the one field every listing carries.
- Decide the filesystem before the drive arrives. XFS offers 262 GB more
than default ext4 on a 4 TB drive, and
mke2fs -i 262144 -m 0recovers 258 GB of that within ext4 itself. Neither choice can be made after the data is on it. - Compare delivered cost, never item price, and never mix TB with TiB on either side of the division. Both errors are about 10%, which is larger than the spread you are shopping in.