HardDiskIndex

Drive capacity explained: why 4TB shows up as 3.64TB

By Harry Saarinen ·

Nothing is missing. The drive holds four trillion bytes and a little over, and Windows is dividing those bytes by 1,024 four times while writing “TB” on the result.

The manufacturer counts in powers of ten, the operating system counts in powers of two, and only one of them labels its units honestly. That single conversion is the whole of the 9%, and the things people usually blame - formatting, the partition table, a hidden recovery area - are real but tiny beside it.

The rest of this page is the part that is usually left out, and one of those omissions is expensive. Two layers below the unit conversion are worth more than a percentage point each: formatting a SAS drive with T10 Protection Information costs 2.03% of it permanently until you reformat, and the choice between ext4 and XFS at format time is worth 6.55% of a 4 TB label. Neither is a rounding error, and neither is visible on a listing page.

The arithmetic, and the standard that fixes the byte count

A 4 TB drive is not approximately four trillion bytes. It is an exact number, fixed by an industry standard so that any vendor’s 4 TB drive can replace any other vendor’s in an array without the array shrinking.

The standard is IDEMA’s LBA1-03, and the first correction on this page is that most explanations cite the wrong one. LBA1-02 is the document usually quoted; LBA1-03 supersedes it, and says so in its own opening, because LBA1-02 “was intended for only IDE disk drives”. LBA1-03 widened the scope to SATA and SAS, to 4,096-byte Large Data Sector drives, and to SAS drives formatted with T10 Protection Information. The algorithm itself did not change:

512-byte sectors
LBA count = 97,696,368 + 1,953,504 × (reported capacity in GBytes − 50.0)

4 TB:   97,696,368 + 1,953,504 × 3,950  =  7,814,037,168 sectors
                            × 512 bytes  =  4,000,787,030,016 bytes

The second correction is smaller and more persistent. The 50.0 is an anchor constant inside the expression, not a scope threshold. The formula does not “apply to drives above 50 GB”; LBA1-03 states its scope by form factor and interface instead. Section 6.0 lists four cases, and the fourth is usually dropped when the scope is paraphrased: 2.5-inch SATA at 80 GB or greater, 2.5-inch at 320 GB or greater with no interface named, 2.5-inch SAS at 73 GB or greater, and 3.5-inch SAS at 450 GB or greater. That second clause is almost certainly a typo for 3.5-inch in the standard itself, because as printed the scope covers no 3.5-inch SATA drive at all - which is the drive every worked example on this page uses. Subtracting 50 from the capacity is arithmetic, not eligibility.

The 4,096-byte form is the same expression with the first two constants divided by eight, which is exactly what you would expect of a formula whose output is a count of sectors eight times larger:

4,096-byte sectors
LBA count = 12,212,046 + 244,188 × (reported capacity in GBytes − 50.0)

4 TB:   12,212,046 + 244,188 × 3,950  =    976,754,646 sectors
                          × 4,096 bytes  =  4,000,787,030,016 bytes

The byte total is identical. A 4 TB drive formatted 4Kn and a 4 TB drive formatted 512e hold the same 4,000,787,030,016 bytes; only the unit of address differs. That fact does more work later on this page than it looks like it should.

The margin, and why it shrinks as drives grow

IDEMA says the formula is designed to “provide 0.02% margin” over the nominal decimal capacity, so a 4 TB drive is very slightly more than 4 × 10¹². Run the margin out across the range and it drifts:

Label LBA count Exact bytes Over nominal
250 GB 488,397,168 250,059,350,016 0.0237%
500 GB 976,773,168 500,107,862,016 0.0216%
1 TB 1,953,525,168 1,000,204,886,016 0.0205%
2 TB 3,907,029,168 2,000,398,934,016 0.0199%
3 TB 5,860,533,168 3,000,592,982,016 0.0198%
4 TB 7,814,037,168 4,000,787,030,016 0.0197%
5 TB 9,767,541,168 5,000,981,078,016 0.0196%
6 TB 11,721,045,168 6,001,175,126,016 0.0196%
8 TB 15,628,053,168 8,001,563,222,016 0.0195%

The drift is structural. 97,696,368 is a fixed addend, so its contribution to the proportional margin falls as the multiplied term grows. Nothing is going wrong; the standard simply converges on 1,953,504 sectors per GB, which is 0.0194% above 10⁹ ÷ 512, and the table’s last row has not quite got there.

The 6 TB and 8 TB rows are not theory. Seagate’s Archive HDD product manual publishes guaranteed sector counts of 11,721,045,168 for the ST6000AS0012 and 15,628,053,168 for the ST8000AS0002, which is the formula’s output to the sector.

Above eight terabytes the formula stops working

This is the part no explanation of drive capacity seems to carry, and it invalidates every capacity table on the internet that extends past 8 TB by running IDEMA’s expression further.

Seagate’s Exos X24 product manuals, both SATA and SAS, state it in one line: “LBA Counts for drive capacities greater than 8TB are calculated based upon the SFF-8447 standard publication.” Anything above 8 TB computed from IDEMA is simply the wrong number. A table saying a 16 TB drive holds 16,003,115,606,016 bytes is 2.2 GB out, and a 20 TB row computed the same way is 3.3 GB out.

Here are the real figures, taken from Seagate’s Exos X24 product manual and Western Digital’s Ultrastar DC HC580 specification:

Label 512-byte LBAs 4,096-byte LBAs Exact bytes GiB Over nominal
10 TB 19,532,873,728 2,441,609,216 10,000,831,348,736 9,314.00 0.0083%
12 TB 23,437,770,752 2,929,721,344 12,000,138,625,024 11,176.00 0.0012%
16 TB 31,251,759,104 3,906,469,888 16,000,900,661,248 14,902.00 0.0056%
20 TB 39,063,650,304 4,882,956,288 20,000,588,955,648 18,627.00 0.0029%
22 TB 42,970,644,480 5,371,330,560 22,000,969,973,760 20,490.00 0.0044%
24 TB 46,875,541,504 5,859,442,688 24,000,277,250,048 22,352.00 0.0012%

Two things jump out of that table.

The margin collapses. IDEMA’s 0.02% becomes 0.001% to 0.008%. A modern large drive holds barely more than the number on the box, where a 1 TB drive from the IDEMA era held about 205 MB more than its label. The slack that used to absorb a vendor’s manufacturing variation is gone.

Every GiB figure is a whole number. Not approximately, exactly: 22,352 GiB with no remainder at 24 TB, 14,902 at 16 TB, 9,314 at 10 TB. SFF-8447 is not a public document, so what follows is a derivation from the datasheets rather than a rule read out of the specification: the expression that reproduces all six published points, and only that expression, is

capacity in GiB = ceil(nominal decimal bytes ÷ 2³⁰)

24 TB:  24 × 10¹² ÷ 1,073,741,824  =  22,351.74...  ->  22,352 GiB
        22,352 × 1,073,741,824      =  24,000,277,250,048 bytes
        ÷ 512                       =  46,875,541,504 sectors

Applied to the capacities that table does not cover, the rule predicts 14 TB at 13,039 GiB (27,344,764,928 sectors), 18 TB at 16,764 GiB (35,156,656,128 sectors, 18,000,207,937,536 bytes), 26 TB at 24,215 GiB, 28 TB at 26,078 GiB, 30 TB at 27,940 GiB and 32 TB at 29,803 GiB.

Two of those six have since stopped being predictions. Western Digital’s Ultrastar DC HC550 specification publishes 27,344,764,928 sectors at 14 TB and 35,156,656,128 at 18 TB, which is the rule reproducing, to the sector, two points it was not fitted against. The remaining four - 26, 28, 30 and 32 TB - are mine, not a vendor’s, and are unconfirmed against any specification. Treat them as a sanity check on a listing, not as a figure to argue with a seller about.

The free integrity check this hands you

The derivation gives a used-drive test that costs one command. A hard drive above 8 TB should report a whole number of gibibytes. If a drive claiming to be 18 TB reports 16,763.4 GiB rather than 16,764, something has clipped it, and the sections on Protection Information and Host Protected Areas below are the two likely culprits. The check is not proof of health, but a fractional GiB count on a large modern drive is an anomaly worth explaining before the return window closes - see buying used drives on eBay for how long that window is.

The interchangeability promise survived the change of standard

The point of fixing capacity by standard is that a replacement drive fits. That still holds across the SFF-8447 boundary, and it is checkable: Seagate’s Exos X24 at 24 TB and Western Digital’s Ultrastar DC HC580 at 24 TB both publish 46,875,541,504 sectors at 512 bytes and 5,859,442,688 at 4,096, for an identical 24,000,277,250,048 bytes. Two vendors, two documents written independently, one number to the byte. Buying a mixed-vendor array of the same nominal capacity is safe on hard drives for exactly this reason, and is not safe on SSDs, which is a distinction the array section returns to.

Two percentages for one fact, and quoting the wrong one is the usual mistake

Now divide by 1,024 four times, which is what Windows does before printing “TB”:

4,000,787,030,016 ÷ 1,024 = 3,907,018,584 KiB
                  ÷ 1,024 ≈     3,815,448 MiB
                  ÷ 1,024 ≈         3,726 GiB
                  ÷ 1,024 ≈          3.64 TiB

Or in one step: 4,000,787,030,016 ÷ 1,099,511,627,776 = 3.6387.

The ratio is 10¹² ÷ 2⁴⁰ = 0.9094947017729282, and it never changes. It is a property of the two number systems, identical for every drive ever made and every drive that will ever be made.

Prefix Decimal Binary Binary unit is larger by Reported number is smaller by
kilo 10³ 2¹⁰ = 1,024 2.400% 2.344%
mega 10⁶ 2²⁰ = 1,048,576 4.858% 4.633%
giga 10⁹ 2³⁰ = 1,073,741,824 7.374% 6.868%
tera 10¹² 2⁴⁰ = 1,099,511,627,776 9.951% 9.051%
peta 10¹⁵ 2⁵⁰ = 1,125,899,906,842,624 12.590% 11.182%

Those are two different numbers for one fact, and swapping them is the single most common error in this subject. A TiB is 9.951% larger than a TB; a capacity expressed in TiB is 9.051% smaller than the same capacity expressed in TB. The gap compounds a further 2.4% at each prefix, which is why the effect is invisible on a USB stick and irritating on a petabyte array.

On the label Standard Exact bytes GiB TiB What Windows prints
500 GB IDEMA LBA1-03 500,107,862,016 465.76 0.455 465 GB
1 TB IDEMA LBA1-03 1,000,204,886,016 931.51 0.910 931 GB
2 TB IDEMA LBA1-03 2,000,398,934,016 1,863.02 1.819 1.81 TB
4 TB IDEMA LBA1-03 4,000,787,030,016 3,726.02 3.639 3.63 TB
8 TB IDEMA LBA1-03 8,001,563,222,016 7,452.04 7.277 7.27 TB
10 TB SFF-8447 10,000,831,348,736 9,314.00 9.096 9.09 TB
12 TB SFF-8447 12,000,138,625,024 11,176.00 10.914 10.91 TB
16 TB SFF-8447 16,000,900,661,248 14,902.00 14.553 14.55 TB
20 TB SFF-8447 20,000,588,955,648 18,627.00 18.190 18.19 TB
22 TB SFF-8447 22,000,969,973,760 20,490.00 20.010 20.00 TB
24 TB SFF-8447 24,000,277,250,048 22,352.00 21.828 21.82 TB

Windows truncates the last column rather than rounding it, and that is where the two competing numbers for a 4 TB drive come from. The drive is 3.6387 TiB. fdisk rounds and prints 3.64 TiB; Explorer cuts and prints 3.63 TB. Both describe the same 4,000,787,030,016 bytes, and the truncation is consistent all the way down the column - 465 GB for 500 GB, 931 GB for 1 TB, 7.27 TB for 8 TB. If you arrived here searching for either number, neither is missing capacity and neither is a different drive.

The units have had correct names since 1998 and nobody uses them

The conversion is honest arithmetic. The labelling is not, and it has not been for a quarter of a century, because a fix has existed the whole time.

IEC 60027-2 Amendment 2, adopted in December 1998 and published in January 1999, was the first international standard to define kibi (Ki), mebi (Mi), gibi (Gi), tebi (Ti), pebi (Pi) and exbi (Ei). The 2005 third edition added zebi and yobi. The prefixes now live in ISO/IEC 80000-13, which cancels and replaces the relevant subclauses of IEC 60027-2:2005, and the 2025 edition of IEC 80000-13 extended the series with robi (Ri, 2⁹⁰) and quebi (Qi, 2¹⁰⁰) to match the new SI prefixes ronna and quetta. The series advances ten powers of two at a rung, which is the detail most reproductions of it get wrong: kibi 2¹⁰ through yobi 2⁸⁰, then 2⁹⁰ and 2¹⁰⁰, with nothing in between.

IEEE 1541 did the same job for engineering practice. Issued as a trial-use standard in 2002, elevated to full use on 19 March 2005, reaffirmed on 27 March 2008 and revised as IEEE 1541-2021, it mandates b for bit, B for byte, o for octet, and that SI prefixes are never used for binary multiples.

NIST publishes the table and states the arithmetic without qualification: 1 GiB is 2³⁰ bytes is 1,073,741,824 bytes, 1 GB is 10⁹ bytes is 1,000,000,000 bytes, and SI prefixes refer strictly to powers of ten. The BIPM expressly prohibits using SI prefixes for binary multiples and points at the IEC prefixes instead. The EU has required the IEC prefixes since 2007, and CENELEC harmonised them as HD 60027-2:2003-03.

So the drive vendor is following the standard and the operating system is not. That is the opposite of the way the argument is usually put.

JEDEC is the reason the memory side is different

JEDEC’s JESD100B.01, “Terms, Definitions, and Letter Symbols for Microcomputers, Microprocessors, and Memory Integrated Circuits”, defines kilo as 2¹⁰, mega as 2²⁰ and giga as 2³⁰ for semiconductor storage capacity. It is the one standards body that blesses the binary reading of an SI prefix, which is why a “16 GB” DIMM really is 17,179,869,184 bytes and is correctly labelled.

Two details of that document get dropped when it is cited. It says the power-of-two definitions “are included only to reflect common usage” - it is recording practice, not endorsing it. And it defines only kilo, mega and giga that way. JEDEC has never defined a binary tera. A “1 TB” memory module has no standards basis for being 2⁴⁰ bytes at all.

Why the IEC prefixes failed, in Microsoft’s own words

Microsoft’s stated reason for never adopting them is comprehension rather than correctness, and Raymond Chen put it plainly on the Old New Thing: “Every document on the Internet (to within experimental error) which talks about memory and storage uses the terms kilobyte/KB, megabyte/MB, gigabyte/GB… In other words, the entire computing industry has ignored the guidance of the IEC.” And: “If Explorer were to switch to the term kibibyte, it would merely be showing users information in a form they cannot understand, and for what purpose?”

That is a defensible position for a file manager and an indefensible one for a specification, and it is why Windows is the last major platform still printing binary values under decimal labels. macOS switched to decimal in Mac OS X 10.6 in 2009, which is why a 1 TB drive appeared to gain about 70 GB when people upgraded. Android and iOS are decimal.

Three settlements about disclosure, one ruling about the unit

Decimal capacity labelling has been litigated repeatedly and now stands. Safier v. Western Digital settled in 2006, with software worth roughly $30 to buyers between 22 March 2001 and 15 February 2006. Cho v. Seagate settled in 2007 with a cash refund or software. Vroegh v. Kodak, filed in 2004 over flash, settled on the same shape of terms: the manufacturers kept the decimal definition and undertook to print it on the packaging and on their web sites. Then in Dinan v. SanDisk the court ruled that GB meaning 1,000,000,000 bytes is not deceptive, and the Ninth Circuit affirmed in February 2021.

The settlements were about disclosure. The ruling was about the unit. The unit won.

The vendor accused of lying is the one using the prefixes properly

Western Digital’s Ultrastar DC HC580 specification carries a glossary that gets all four right in one place: “GB 1,000,000,000 bytes”, “TB 1,000,000,000,000 bytes”, “KiB 1,024 bytes”, “MiB 1,048,576 bytes”. Seagate prints a boilerplate on every product manual: “When referring to drive capacity, one gigabyte, or GB, equals one billion bytes and one terabyte, or TB, equals one trillion bytes. Your computer’s operating system may use a different standard of measurement and report a lower capacity.”

The industry’s inconsistency is not marketing, it is addressing. A DRAM chip or a NAND die is selected by binary address lines, so its capacity is a power of two and calling 17,179,869,184 bytes “16 GB” describes something real. A hard drive is a count of sectors, and there is no electrical reason for that count to be a power of anything - a platter holds whatever number of sectors the areal density gives it. Storage went decimal because storage is not binary underneath.

This site inherits that split rather than resolving it. Capacities here are decimal, because decimal is the unit on the label and the unit the seller set the price against. Pick one and stay in it.

What every tool on your own machine means by “G”

The Linux answer is usually given as “df is binary, fdisk prints both”, which is wrong in a way that matters, because the most decimal tool in the set is missing from it. Run against one 4 TB drive and one root filesystem on Ubuntu 26.04:

Tool and flag Unit it means What it printed
df -h powers of 1,024 468G
df -H powers of 1,000 502G
df (no flag) 1K-blocks, binary 1K-blocks column
POSIXLY_CORRECT=1 df 512B-blocks 512B-blocks column
lsblk powers of 1,024, printed as K/M/G 476.9G
lsblk -b bytes 512110190592
fdisk -l header TiB, correctly labelled 3.64 TiB, 4000787030016 bytes
fdisk -l size column powers of 1,024, printed as T 3.6T
parted print powers of 1,000 4001GB, start 1049kB
parted unit B print bytes 4000787030016B, start 1048576B
ls -lh, du -h powers of 1,024 977K for a 1,000,000-byte file
ls --si, du --si powers of 1,000 1.0M for the same file
btrfs filesystem usage TiB, correctly labelled 3.64TiB

One machine, one drive, four different numbers out of one binary. GNU coreutils df alone prints 468G, 502G, a 1K-block count and a 512B-block count for the same filesystem depending on flags and environment.

Three entries deserve singling out. parted is decimal by default and contradicts every other partitioning tool on the system, so a capacity read out of parted print and compared against one read out of fdisk will disagree by 9% with no error anywhere. lsblk’s own manual admits the offence outright: sizes are “in units that are powers of 1024 bytes” and “the formal abbreviations for these units (KiB, MiB, GiB, …) are further shortened to just their first letter: K, M, G”. That is a gibibyte wearing a decimal letter, which is precisely what Windows is criticised for. And btrfs-progs is the one tool in the list that prints the IEC unit correctly, which is a small mercy given what df does to btrfs.

When a number has to be exact, ask for bytes. The three that cannot be misinterpreted are lsblk -b, parted unit B print, and blockdev --getsize64.

Reading the capacity off the drive yourself

Everything above is checkable in about thirty seconds, and on a used drive it should be. The number to compare against is the IDEMA or SFF-8447 figure for the nominal capacity, from the tables above.

sudo blockdev --getsize64 /dev/sdb      # bytes, no units to misread
lsblk -b /dev/sdb                       # bytes, device and partitions
sudo fdisk -l /dev/sdb                  # bytes and sector count together
sudo hdparm -I /dev/sdb | grep -iE 'device size|sector size|Model|Serial'
sudo smartctl -i /dev/sdb               # "User Capacity" and "Sector Sizes"

hdparm -I reports the logical and physical sector sizes separately, which is how you tell 512n from 512e from 4Kn without guessing, and smartctl -i prints both alongside the model and serial. On NVMe the fields are named by the specification rather than by the tool:

sudo nvme id-ns /dev/nvme0n1 -H | grep -iE 'nsze|ncap|nuse|lbaf'
sudo nvme id-ctrl /dev/nvme0n1 -H | grep -iE 'tnvmcap|unvmcap'
Field What it is
NSZE Namespace size, in logical blocks
NCAP Namespace capacity, the blocks actually allocated to it
NUSE Namespace utilisation, the blocks currently written
TNVMCAP Total NVM capacity of the whole controller, in bytes
UNVMCAP Unallocated NVM capacity, in bytes

NSZE and TNVMCAP disagreeing is not a fault, it is namespace management. A host can create a namespace smaller than the drive, and the remainder then shows up in UNVMCAP. That is the one way capacity genuinely disappears rather than merely going unpartitioned, and it is the cleanest way to add over-provisioning deliberately. It is also a way a used enterprise NVMe arrives smaller than advertised, and it is recoverable by deleting and recreating the namespace.

4Kn, 512e and 512n: the sector size changes the platter, not the byte count

Three format types exist, and only two numbers distinguish them:

Format Logical sector Physical sector
512n 512 bytes 512 bytes
512e 512 bytes 4,096 bytes
4Kn 4,096 bytes 4,096 bytes

None of them changes the capacity. The 4 TB arithmetic above gave 4,000,787,030,016 bytes under both IDEMA expressions. What the sector size changes is how much of the platter is available to hold those bytes in the first place, and the gain there is the reason Advanced Format exists.

The capacity motive is format efficiency

Every sector on a platter carries overhead that is not data: a gap, a sync pattern, an address mark and an error correction code. That overhead is roughly fixed per sector, so making sectors eight times larger spreads it eight times further. Dell’s 512e and 4Kn white paper gives the layout:

512-byte sector    512 data + 50 ECC, in a 577-byte footprint
                   512 ÷ 577    =  88.7% format efficiency

4,096-byte sector  4,096 data + 100 ECC, in a 4,211-byte footprint
                   4,096 ÷ 4,211  =  97.3% format efficiency

Dell’s own summary: “This yields a format efficiency of 97 percent, almost a 10 percent improvement.” Reported gains in usable platter area run 7 to 11 per cent. Note that the ECC field doubles rather than growing eightfold, because a longer codeword corrects more efficiently - the density gain and the error correction improvement come from the same change.

The 512e penalty is read-modify-write, and it is a write penalty only

A 512e drive tells the host it has 512-byte sectors and does not. The media cannot be written in anything smaller than 4,096 bytes, so a host write that is misaligned, smaller than 4K, or not a multiple of 4K forces the drive to read the whole physical sector, merge the change and write it back - which, as Western Digital’s Advanced Format technical brief puts it, “can require additional revolutions of the hard disk”. IDEMA and vendor analysis put roughly 5 to 10 per cent of writes in a typical business PC environment into that category on an unaligned volume.

Reads cost nothing. This is a write problem, exactly as shingled recording is a write problem for a different reason - see CMR vs SMR.

The alignment rule is arithmetic, not judgement

A partition on a 512e drive must start on a multiple of eight logical sectors. There are eight possible positions for LBA 0 within a 4K physical sector; Dell calls the correct one “Alignment 0”, and the other seven all produce read-modify-write on every straddling write.

start LBA mod 8 = 0        aligned
start LBA = 2,048          the modern default, = 1 MiB
2,048 mod 8 = 0            satisfied, and satisfied for any physical
                           sector size up to 1 MiB

The 1 MiB default that every current partitioner uses is deliberately far more generous than the eight-sector rule needs. It exists so that the same alignment also suits RAID stripe units, SSD erase blocks and hypervisor block sizes, none of which are 4 KiB. The historical failure case is a Windows XP-era partition starting at sector 63, which is odd and therefore misaligned by construction.

Converting between them takes seconds and moves no data

Modern drives change sector size in the field with SET SECTOR CONFIGURATION EXT (opcode B2h, defined in ACS-4). Seagate’s Exos X24 manual: “The selected sector size change occurs immediately upon command completion. Default shipping format is 512E.” Western Digital adds the detail that matters for a listing: “Changing the block size does not change the HDD Model Number reported by the drive.”

So the sector format is not a property of the model number, it is a property of how the last owner left the drive, and it is not a reason to reject a used enterprise drive. Enterprise drives covers the reformat in practice, including the far nastier 520-byte case.

T10 Protection Information: the 2% that only bites SAS buyers

This is the most likely reason a used drive is genuinely short of capacity, and it does not appear in general explanations of the subject at all.

SAS drives can be formatted with T10 Protection Information: eight extra bytes appended to every logical block, carrying a guard CRC, an application tag and a reference tag, so that corruption is detectable end to end rather than only on the media. The drive stores 520-byte blocks instead of 512, or 4,160 instead of 4,096. The extra eight bytes come out of the capacity.

Seagate’s Exos X24 SAS product manual publishes both guaranteed operating points:

24 TB without PI   46,875,541,504 blocks × 512  =  24,000,277,250,048 bytes
24 TB with PI      45,923,434,496 blocks × 512  =  23,512,798,461,952 bytes
                                        loss    =     487,478,788,096 bytes
                                                =   2.03%

16 TB without PI   31,251,759,104 blocks × 512  =  16,000,900,661,248 bytes
16 TB with PI      30,616,322,048 blocks × 512  =  15,675,556,888,576 bytes
                                                =   2.03% again

8 ÷ 520 = 1.538%, and 8 ÷ 512 = 1.563%, and the observed loss is 2.03%. The extra is the drive reserving whole blocks rather than fractions and rounding the guaranteed count down to a published figure. The ratio is identical at both capacities, which is what tells you it is a formatting rule rather than a per-model decision.

There is a documented contradiction here and it is worth stating rather than smoothing over. IDEMA LBA1-03 says the eight PI bytes are protocol overhead rather than user-addressable space, and therefore “the number of LBA count on a T10 PI formatted drives must be the same as their non-T10 PI counterparts for the same reported capacity”. Real drives above 8 TB do not behave that way. The Exos X24 figures above are from the vendor’s own manual and they differ by 2.03%. Trust the datasheet over the standard here; the standard is describing the world below the SFF-8447 boundary.

The practical consequences:

  • A used SAS drive sold as 24 TB that arrives reporting 23.5 TB is almost certainly PI-formatted, not clipped and not counterfeit.
  • Supported block sizes on the Exos X24 are 512 and 520 for the 512e family, 4,096 and 4,160 for 4Kn. A block size of 520 or 4,160 is the tell.
  • Reformatting to 512 or 4,096 returns the capacity. It destroys the data and it takes hours on a large drive, which is why the decision belongs before you build the array rather than after.
  • SATA drives cannot do any of this. If you are buying SAS, check; otherwise the section does not apply to you.

The two ways a drive can hide capacity, and only one of them is HPA

A drive tells the host how large it is, and that answer can be reduced in two independent places. Most write-ups cover the first and not the second.

Host Protected Area, defined by T13 in ATA-4 and published as ANSI NCITS 317-1998, reserves sectors at the end of the drive that the host cannot address. Laptop vendors used it for recovery images; sellers set it by accident. It survives a reformat and a repartition, because it is below both.

sudo hdparm -N /dev/sdb

The man page describes the output precisely: “the current setting, which is reported as two values: the first gives the current max sectors setting, and the second shows the native (real) hardware limit for the disk”. Two different numbers mean an HPA. Prepending p to a new value makes the change permanent, and hdparm’s own warning on that path is “VERY DANGEROUS, DATA LOSS IS EXTREMELY LIKELY”.

Device Configuration Overlay is a second, separate clip, and hdparm -N does not reveal it. DCO arrived four years after the HPA, in ATA-6, and it works one level lower: it can reduce the reported native capacity itself, so the two numbers -N prints can agree with each other while the drive is still short. It is queried and reset with different flags:

sudo hdparm --dco-identify /dev/sdb     # what the vendor or OEM disabled
sudo hdparm --dco-restore  /dev/sdb     # DO NOT RUN THIS CASUALLY

--dco-identify will “query and dump information regarding drive configuration settings which can be disabled by the vendor or OEM installer”. --dco-restore will “reset all drive settings, features, and accessible capacities back to factory defaults and full capabilities”, and hdparm labels it “EXTREMELY DANGEROUS” and says “DO NOT USE THIS COMMAND”. Both warnings are the tool author’s, not this page’s, and both are justified.

A drive can be short by HPA, by DCO, or by both at once. On a used drive that reports less than the standard figure, check -N first, then --dco-identify, and only then conclude the drive is misdescribed.

MBR, which is a table limit rather than a drive limit

If a drive over 2.2 TB reports as exactly 2.2 TB, this is why. An MBR partition table stores both the starting LBA and the sector count in 32-bit fields, and the largest value a 32-bit count can hold is 2³² − 1, not 2³²:

(2³² − 1) × 512  =  4,294,967,295 × 512  =  2,199,023,255,040 bytes
2 TiB            =                        2,199,023,255,552 bytes
                                    short by 512 bytes

The round ceiling of 2,199,023,255,552 is quoted almost everywhere. It is one sector too high.

The same fields with 4,096-byte logical sectors reach (2³² − 1) × 4,096 = 17,592,186,040,320 bytes, one 4 KiB sector short of 16 TiB by the same rule - the identical off-by-one, eight times larger. That is the reason a 512e drive reformatted to 4Kn can sometimes be made to work past 2.2 TB under MBR on a system that cannot boot GPT. Everything past the ceiling is unaddressable by the table, not by the drive. Convert to GPT and it returns. The same 32-bit ceiling turns up in cheap USB-SATA bridges and old RAID controllers, and there the enclosure has to go - drive interfaces covers which bridges do it.

Where the rest of it goes: a measured ext4 ledger

Take that 4 TB drive through an ordinary Linux install, GPT plus ext4 with mke2fs defaults. Every figure below comes from creating the filesystem on a real 4,000,785,964,544-byte partition with mke2fs 1.47.2 and reading it back with dumpe2fs, rather than from the arithmetic people usually reproduce.

Drive                                          4,000,787,030,016 B   3.6387 TiB

GPT
  protective MBR, header, entry array               17,408 B    LBA 0-33
  alignment padding to LBA 2048                  1,031,168 B    LBA 34-2047
  backup header and array at the end                16,896 B    last 33 LBAs
  partition                                  4,000,785,964,544 B   3.6387 TiB

ext4, mke2fs 1.47.2 defaults
  partition ÷ 4,096 = 976,754,385.875           last 3,584 B unused
  inode table   244,195,328 × 256 B         62,514,003,968 B   58.22 GiB
  journal       262,144 blocks × 4,096 B      1,073,741,824 B    1.00 GiB
  bitmaps, group descriptors, reserved GDT,
    superblock backups, 29,809 groups           378,552,320 B     361 MiB
  total overhead, 15,616,772 clusters        63,966,298,112 B   59.57 GiB
  df "Size"                                   3,936,819,662,848 B   3.5805 TiB

  root reserve, 48,837,719 blocks               200,039,297,024 B  186.30 GiB
  df "Avail" on an empty filesystem           3,736,780,365,824 B   3.3986 TiB

That ledger subtracts, and a published one that does not was assembled rather than measured. df “Size” is the partition less the 3,584 unused bytes and less the 63,966,298,112 of overhead; df “Avail” is “Size” less the root reserve, exactly. Check any capacity breakdown you find against that test before believing it.

Four things in it the arithmetic will not tell you and only a dumpe2fs will.

The journal is 1 GiB, not 128 MiB. ext2fs_default_journal_size has a ladder

  • 1,024 blocks under 128 MB, 4,096 under 1 GB, 8,192 under 2 GB, 16,384 under 16 GB, 32,768 under 32 GB, 65,536 under 64 GB, 131,072 under 128 GB, and 262,144 above that. At 4 KiB blocks the top rung is exactly 1 GiB, and dumpe2fs on the freshly created filesystem reports “Total journal size: 1024M”. The widely repeated 128 MiB figure has been wrong for years.

The bitmaps and descriptors are 361 MiB, not 1.5 GiB. The figure usually given is about four times too large, and it happens to cancel against the understated journal, which is why tables built from both still produce a plausible df total. Two errors that agree are harder to find than one.

The inode table is 97.7% of all filesystem metadata. Everything else - journal, bitmaps, group descriptors, reserved GDT blocks, superblock backups across 29,809 block groups - is the remaining 2.3%. Anyone optimising ext4 overhead who is not changing the inode ratio is optimising the wrong thing.

A partition is rarely a whole number of filesystem blocks. 4,000,785,964,544 ÷ 4,096 = 976,754,385.875, so the last 3,584 bytes are simply unused. It is a trivial amount and it is a real layer.

The partition table costs about a megabyte, and anyone blaming it is wrong

GPT is 34 sectors at the front and 33 at the back, and those 34 are routinely described as alignment padding in full. They are not. Running sgdisk on the image reports “Main partition table begins at sector 2 and ends at sector 33. First usable sector is 34”, and “Total free space is 2014 sectors (1007.0 KiB)”. So:

LBA 0        protective MBR                       512 B
LBA 1        primary GPT header                   512 B
LBA 2-33     partition entry array, 32 LBAs    16,384 B
LBA 34-2047  alignment padding, 2,014 LBAs  1,031,168 B
last 33 LBAs backup array and header          16,896 B
                                      total  1,065,472 B
                                             0.0000266% of the drive

The 34/33 split is a consequence of the UEFI specification reserving at least 16,384 bytes for the GUID Partition Entry Array - 128 entries of 128 bytes. At 512-byte blocks that is 32 LBAs, so 1 + 1 + 32 = 34 at the front and 32 + 1 = 33 at the back.

On a 4Kn drive the array is only 4 blocks, but GPT costs more in bytes, not less. Six blocks at the front and five at the back is 45,056 bytes, against 34,304 at 512. The count went down and the footprint went up, which is the sort of thing that makes a per-sector intuition fail.

The 5% root reserve is the biggest single recoverable chunk, and it may not be 5%

186 GiB of that 4 TB is not overhead at all. It is ordinary free space that only root may allocate, and mke2fs sets it from the profile with a hardcoded fallback of reserved_ratio = 5.0. The man page gives the reason: it “avoids fragmentation, and allows root-owned daemons, such as syslogd(8), to continue to function correctly after non-privileged processes are prevented from writing to the file system.” Root-owned is the load-bearing word. The reserve is useful precisely because the processes that must keep writing when the disk fills run as root and can allocate into it.

On a dedicated data drive neither reason applies:

sudo tune2fs -m 1 /dev/sdb1     # 5% -> 1%, instant, no unmount needed
sudo tune2fs -l /dev/sdb1 | grep -E 'Block count|Reserved block count'

Check what you actually have before assuming there is 5% to recover. Distributions edit the defaults, and the reserve is edited more often than the inode ratio is. On an Ubuntu 26.04 root filesystem stat -f / reports 122,512,118 blocks with 116,393,871 free and 115,142,264 available, which is a reserve of 1,251,607 blocks - about 1.02% of the block count, not 5%.

What format-time options actually recover

Run mke2fs -n against the same 4 TB partition with different flags and the inode table moves by tens of gigabytes:

Options Inodes Inode table Recovered against default
defaults 244,195,328 58.22 GiB -
-i 65536 61,048,832 14.56 GiB 43.67 GiB
-i 262144 15,262,208 3.64 GiB 54.58 GiB
-T largefile 3,815,552 0.91 GiB 57.31 GiB
-T largefile4 953,888 0.23 GiB 57.99 GiB

-I 128 halves the table without changing the inode count, at the cost of nanosecond timestamps and post-2038 dates, which is a bad trade on anything built in 2026.

Combined, format-time choices are worth more than every other loss on this page put together:

4 TB, mke2fs defaults          Avail  3,736,780,365,824 B
4 TB, mke2fs -i 262144 -m 0    Avail  3,995,426,541,568 B
                              gain      258,646,175,744 B
                                    =  258.6 GB  =  6.47% of the label

Two flags. None of it is recoverable afterwards - the inode ratio is fixed at format time - which is the argument for deciding it before the drive has data on it. Set it low only if you know the filesystem will hold large files, and read the inode column of that table before choosing. -T largefile4 leaves 953,888 inodes on this partition, so a million small files exhaust the table on a drive showing terabytes free, and the error message will not say so clearly. One inode per 256 KiB leaves 15,262,208 of them, which a mail spool or a package cache can still reach and an ordinary filesystem will not.

The ext4 inode-ratio cliff punishes the bigger drive

mke2fs picks its usage type from the filesystem size, and the branch points are in parse_fs_type in mke2fs.c. They are binary thresholds, and nobody expects that. The source computes meg as the number of blocks in one MiB, then branches: under 3 MiB “floppy”, under 512 MiB “small”, under 4*1024*1024*meg (4 TiB) the default, under 16*1024*1024*meg (16 TiB) “big”, otherwise “huge”. The ratios that follow from /etc/mke2fs.conf on e2fsprogs 1.47.2 are inode_ratio 16,384 by default, 32,768 for big, 65,536 for huge.

A 4 TB drive is 3.639 TiB. It is below the first threshold, so it gets the densest inode ratio in the table. A 5 TB drive is 4.548 TiB and does not:

Drive Size in TiB Usage type Inodes Inode table
4 TB 3.639 default 244,195,328 58.22 GiB
5 TB 4.548 big 152,621,056 36.39 GiB
16 TB 14.553 big 488,308,736 116.42 GiB
17 TB 15.462 big 518,815,744 123.70 GiB
18 TB 16.371 huge 274,661,376 65.48 GiB

The rows above 8 TB are built on the SFF-8447 byte counts from earlier on this page, not on IDEMA run past its boundary, and the distinction is not cosmetic. The IDEMA figure for 16 TB is 2.2 GB larger, which is enough to add seventeen block groups and 69,632 inodes that the drive does not have. Any inode table you find quoted for a drive above 8 TB is worth checking for that, because the wrong capacity produces a plausible-looking wrong answer rather than an obvious one.

A 5 TB drive carries 21.8 GiB less inode table than a 4 TB one, in absolute terms, not as a proportion. The same inversion repeats at the top: an 18 TB drive gets 43.8% fewer inodes than a 16 TB one, and 50.9 GiB less table. The crossover is at 16 TiB, which is 17.592 TB, so every drive on sale between 16 TB and 17.5 TB is on the dense side of it and every drive from 18 TB up is not.

None of this is a bug. It is a sensible heuristic - large filesystems usually hold large files - expressed in units the buyer of a “5 TB” drive has no reason to be thinking in. It is the clearest case on this page of a binary threshold producing a result that looks wrong in decimal.

Choosing the filesystem is worth more than everything else on this page

Put ext4, XFS and btrfs on the identical 4,000,785,964,544-byte partition and read back what each one took before a single user byte was written:

Filesystem Metadata at format Free at format As % of the 4 TB label
ext4, defaults, root reserve applied 59.57 GiB 3,736,780,365,824 B 93.42%
ext4, defaults, as root 59.57 GiB 3,936,819,662,848 B 98.42%
XFS, defaults 1.82 GiB 3,998,832,308,224 B 99.97%
btrfs, defaults 2.02 GiB roughly 3.64 TiB reported -
XFS  free at mkfs   3,998,832,308,224 B
ext4 Avail at mkfs  3,736,780,365,824 B
     difference       262,051,942,400 B  =  262.1 GB  =  6.55% of the label

One 4 TB drive, one decision, 262 GB. That is larger than every other loss on this page combined, and it is not lost capacity at all - it is front-loaded metadata policy.

The comparison flatters XFS and the honest version says so. XFS allocates inodes on demand rather than up front, so its advantage closes as the filesystem fills with small files and the metadata it did not build at format time gets built later. xfs_db on that 4 TB filesystem reports dblocks 976,754,385 - identical to ext4’s block count, as it should be - an internal log of 476,930 blocks (1,953,505,280 bytes, 1.819 GiB), imaxpct 5, isize 512 with CRCs on by default, agcount 4 and agblocks 244,188,597. The imaxpct defaults are documented as “25% for filesystems under 1TB, 5% for filesystems under 50TB and 1% for filesystems over 50TB”, so XFS has a ceiling on inode space rather than a reservation of it.

btrfs duplicates its metadata, and df is the wrong tool on it

mkfs.btrfs on the same image allocates “Data: single 8.00MiB, Metadata: DUP 1.00GiB, System: DUP 8.00MiB” with a node size of 16,384 - 2.02 GiB before a byte of user data, because DUP is the default metadata profile on a single device and metadata is therefore written twice.

The reporting is the bigger problem. btrfs has a two-stage allocator: chunks are allocated to data or metadata first, then filled. The documentation is blunt that df “cannot tell how much unallocated disk space is available”, and a filesystem with all space allocated but metadata near full will show free space in df and still return ENOSPC. Use btrfs filesystem usage, and read its data ratio - 1.0 for single, 2.0 for DUP and RAID1.

ZFS keeps a slop reserve, and it is capped

The commonly quoted figure is that ZFS holds back 1/32 of a pool. The OpenZFS zfs(4) documentation says it precisely: “Normally, we don’t allow the last 3.2% (1/2^spa_slop_shift) of space in the pool to be consumed”, with spa_slop_shift defaulting to 5.

The fraction is right and the implication is misleading on the pools this site’s readers build, because the reserve is floored at 128 MiB and capped at 128 GiB, and never exceeds half the pool. A 200 TB pool holds back 128 GiB, not 6.25 TB. Nor is it wholly unavailable: file removal and most administrative actions may use half of it, and operations almost certain to free space, such as zfs destroy, up to three quarters.

The other ZFS surprise is documented just as explicitly. From zpoolprops(7): “the zpool free property is not generally useful for this purpose, and can be substantially more than the zfs available space”, and “The space usage properties report actual physical space available to the storage pool. The physical space can be different from the total amount of space that any contained datasets can actually use.” zpool list counts parity; zfs list and df do not. People compare the two and conclude a pool lost space it never had.

NTFS reserves 12.5% for the MFT, and does not take it away

Windows reserves an MFT zone of 12.5% of the volume by default, set by NtfsMftZoneReservation under HKEY_LOCAL_MACHINE\System\CurrentControlSet\Control\FileSystem, a REG_DWORD valid from 1 to 4 with a default of 1. That reservation is not missing capacity, and Microsoft says so directly: “The MFT Zone is not subtracted from available (free) drive space used for user data files, it is only space that is used last.” Microsoft also declines to publish what values 2 to 4 mean, on the grounds that the ratios “are undocumented because they are not standardized and may change in future releases.”

NTFS’s cluster size sets the volume ceiling, because the cluster count field is 32 bits and the limit is 2³² − 1 clusters:

Cluster size Maximum volume and file size
4 KB (default) 16 TB
8 KB 32 TB
16 KB 64 TB
32 KB 128 TB
64 KB 256 TB
128 KB 512 TB
1,024 KB 4 PB
2,048 KB 8 PB

Those are the figures for Windows Server 2019 and later and Windows 10 1709 and later. A 20 TB drive formatted NTFS at the default 4 KB cluster will not take a single volume, which is a live problem now that single drives exceed 16 TB, and the fix is a larger cluster at format time. VSS-based backup caps the practical volume at 64 TB regardless.

APFS free space is a container property, not a volume property

On macOS the question “how much space is free” has no single answer by design. Free space belongs to the container, not to any volume inside it: it is the container’s capacity less what each of its volumes uses, shared among all of them. macOS then folds “purgeable” space - mostly APFS and Time Machine snapshots plus caches - into the figure that Finder, About This Mac and Disk Utility each call “available”, which is why those three disagree with each other and with df. diskutil apfs list gives the container figure, and it is the one to believe.

The container is also not the whole disk. On a 2 TB Mac SSD the Macintosh HD container holds about 1.995 TB, with iSCPreboot at roughly 524 MB and Recovery at roughly 5.4 GB taking the rest - Apple’s own hidden volumes, before any user data exists.

And below all of them sits slack

A filesystem with 4 KiB blocks wastes on average half a block per file:

1,000,000 files × 2,048 B average waste  =  2,048,000,000 B  =  1.91 GiB

That is why a backup of a mail spool never matches the size the source reported. exFAT makes it much worse if misconfigured: the specification permits clusters up to 32 MB, via SectorsPerClusterShift of at most 25 minus BytesPerSectorShift, so a badly chosen cluster size on a card full of small files can waste orders of magnitude more than any figure on this page.

SSD over-provisioning: where the binary NAND went

An SSD inverts the problem. NAND is built in binary quantities and the drive presents a decimal capacity, so the gap between the two is not lost. It is kept by the controller, and it is doing work.

Before the arithmetic, the connection that is usually missed: SSDs follow the same IDEMA table as hard drives. Seagate’s BarraCuda SSD product manual, publication 100835666 Rev A, cites it by name - “Sector Size: 512 Bytes. User-addressable LBA count = ((97696368) + (1953504 x (Desired Capacity in Gb-50.0)) From International Disk Drive Equipment and Materials Association (IDEMA) (LBA1-03_standard.doc)”. An SSD is not a separate capacity regime. It is the same one, with a different physical substrate underneath.

The 7.4% figure is wrong for most drives on sale

The standard explanation says a consumer SSD carries about 7.37% spare, from 2³⁰ ÷ 10⁹. That is only true when the label number is itself a power of two. Run it against real IDEMA user capacities and the common case is noticeably different:

Sold as NAND on board User bytes (IDEMA) Spare Over-provisioning
512 GB 512 GiB 512,110,190,592 37,645,623,296 7.35%
500 GB 512 GiB 500,107,862,016 49,647,951,872 9.93%
480 GB 512 GiB 480,103,981,056 69,651,832,832 14.51%
400 GB 512 GiB 400,088,457,216 149,667,356,672 37.41%
1,000 GB 1 TiB 1,000,204,886,016 99,306,741,760 9.93%
960 GB 1 TiB 960,197,124,096 139,314,503,680 14.51%
800 GB 1 TiB 800,166,076,416 299,345,551,360 37.41%
2,000 GB 2 TiB 2,000,398,934,016 198,624,321,536 9.93%
1,920 GB 2 TiB 1,920,383,410,176 278,639,845,376 14.51%
3,840 GB 4 TiB 3,840,755,982,336 557,290,528,768 14.51%
7,680 GB 8 TiB 7,681,501,126,656 1,114,591,895,552 14.51%
120 GB 128 GiB 120,034,123,776 17,404,829,696 14.50%

Most consumer SSDs today carry decimal labels - 500 GB, 1 TB, 2 TB - so the common baseline is 9.93%, not 7.4%. Only a drive labelled 512 GB or 256 GB, where the number is a power of two, sits at 7.35%.

That correction changes the enterprise comparison too. A 960 GB drive and a 1 TB drive of the same family are usually the same silicon with a different firmware capacity, and the premium buys 14.51% against 9.93% - about 4.6 percentage points of extra spare, not the 7 the old arithmetic implied. The 800 GB against 1 TB comparison is the one that is genuinely dramatic: 37.41% against 9.93%.

The formula itself is uncontroversial, and stated identically by SNIA-adjacent vendor material and by the general literature:

OP (%) = (physical capacity − user capacity) ÷ user capacity × 100

Samsung's worked example:  ((128 − 120) ÷ 120) × 100  =  6.7%

Samsung’s own note on that example - “Typically Samsung DC SSD provides 6.7 percent of factory OP as its default” - is a good illustration of why the denominator matters. The same drive described against its 128 GiB of NAND and its 120 GB of user space with exact IDEMA bytes gives 14.50%, because 128 GiB is 137,438,953,472 bytes rather than 128 × 10⁹. Two vendors quoting over-provisioning are not necessarily using the same denominator, and the formula above is the one to re-run yourself.

What the spare buys, and the identity that cannot be simplified

NAND erases in blocks far larger than the pages it writes, so freeing a block means relocating the live pages inside it first, and those relocations are writes nobody asked for. That ratio is write amplification, and SSD endurance derives its floor from block geometry.

Hu et al. at IBM Zurich Research Laboratory set the terms precisely in “Write Amplification Analysis in Flash-Based Solid State Drives” (SYSTOR 2009). For a raw capacity of t blocks of which u are user-visible, the over-provisioning factor is Of = t/u and the spare factor is Sf = (t − u)/t. Their conclusions: write amplification falls as over-provisioning rises, and separating static from dynamic data reduces it further. Their model holds for a uniformly distributed random workload of 4 KiB writes and is not claimed beyond it, which is exactly the honesty a reader needs here.

Samsung publishes the endurance identity rather than a write-amplification formula:

WAF  = physical write amount ÷ host write amount

DWPD = (NAND P/E cycles × raw density)
       ÷ (logical density × 365 × warranty years × WAF)

Samsung’s own worked estimates for the 860 DCT show what that buys: a 960 GB drive at 0% additional user over-provisioning rates 0.34 DWPD over three years; at 10% user OP (872 GB exposed) 0.71; at 20% (800 GB exposed) 0.96. Nearly tripling endurance for 17% of the capacity. Those three are calculated from the identity above and the paper says so in terms - “a calculated value of each SSD, not a guaranteed value” - so treat them as the shape of the curve rather than as measurements of a part you might buy. Industry practice has settled on four levels - 0% additional, 7%, 14% and 28% - which is what produces the capacity ladders in the table above.

Reported real write amplification values, from vendor and press measurement rather than from this page: the Intel X25-M in 2008 as low as 1.1; SandForce SF-1000 at 0.5 with compression; SF-2281 best case 0.14. Without compression, write amplification cannot drop below 1, and a figure under 1 tells you the controller is not writing everything the host sent.

Adding your own, and the condition that is usually skipped

You can add over-provisioning by leaving part of the drive unpartitioned, on one condition.

The controller has to know those blocks are free. It sees logical block addresses and whether each has been written, not partitions. Leaving the last 20% unpartitioned on a drive that has been filled at some point achieves nothing, because every one of those addresses still holds data the controller must preserve. Erase first, then partition short:

sudo blkdiscard /dev/nvme0n1              # TRIM the whole device
# or, for a full reset of the mapping tables:
sudo nvme format /dev/nvme0n1 --ses=1
sudo sgdisk -n 1:0:+800G /dev/nvme0n1     # leave the rest untouched

nvme format also changes the LBA data size and metadata size, so it is the command that converts an NVMe drive between 512-byte and 4,096-byte logical blocks. Namespace management is cleaner still where the drive supports it, because the capacity genuinely disappears from NSZE rather than merely going unused.

Before any of that, check the cheaper option:

systemctl status fstrim.timer
sudo fstrim -av

Free space inside a trimmed filesystem is already over-provisioning. A drive kept 40% empty with fstrim running weekly has 40% spare from the controller’s point of view, at no cost. Manual over-provisioning is insurance against the drive being filled later, not a substitute for TRIM working now.

So it is worth doing on a consumer drive at the 9.93% baseline that will run near full under sustained random writes - a VM host, a busy database, a write cache. It is not worth doing on a drive that will sit half empty, or on an enterprise drive already shipping at 14.51% or 37.41%.

Price per terabyte, and why delivered cost is the only honest input

Per-terabyte cost is the comparison that matters because capacity is the thing being bought, and the arithmetic is one division:

price per TB = delivered cost ÷ capacity in decimal TB

Both halves go wrong in practice. Take three listings, in arbitrary units:

Capacity Item Shipping Delivered Per TB
A 8 TB 100 0 100 12.50
B 8 TB 92 18 110 13.75
C 6 TB 70 12 82 13.67

Sorted by item price: C, B, A. Sorted by what you pay per terabyte: A, C, B. The cheapest-looking listing is not the cheapest and the most expensive-looking one is the best. Shipping a 3.5-inch drive is a real cost that some sellers fold into the price and some do not, which is why this site computes per-TB from delivered cost and leaves the figure blank rather than guessing when shipping cannot be determined - the methodology page sets out when that happens.

The unit error is the other half, and it is larger than the spread you are shopping for. Price A per tebibyte instead of per terabyte:

A:  100 ÷ 7.2774 TiB  =  13.74 per TiB
B:  110 ÷ 8 TB        =  13.75 per TB

Identical, for a drive that is 10% cheaper. Mixing TB and TiB in a comparison erases exactly the size of advantage most people are hunting for. Use the label capacity on both sides, every time. It is also the only capacity knowable before the drive arrives, since what a filesystem will report depends on decisions not yet made.

How this site derives capacity, and where that can be wrong

This site stores capacity as decimal GB by construction, and says so in its own code rather than only in prose. Drive.capacityGb is documented as “Capacity in DECIMAL gigabytes, as the manufacturer markets it”, and the conversion constant is named rather than written as a literal:

GB_PER_TB = 1_000

with the comment that this exists “because the whole point of the unit discipline in types.ts is that the conversion is visible and stated rather than a literal somebody later corrects to 1024”. The methodology page commits to the same thing publicly: “Price per terabyte is the delivered cost divided by the total capacity on offer, in decimal terabytes - the unit drives are sold in, where 1 TB is 1,000 GB rather than 1,024.”

Now the honest part. The capacity on every row of this site comes from the seller’s title text, not from the drive. statedCapacitiesGb() matches (\d+(?:\.\d+)?)\s*(tb|gb|mb), multiplies terabytes by 1,000 and takes the largest surviving match. It already excludes the things that look like capacities and are not: link speeds (1.5, 3, 6, 12 and 22.5 Gb followed by SAS or SATA), endurance odometers (“8TB written” on a 512 GB drive, which once made that drive the cheapest SSD on the site), serial numbers with embedded digit-letter runs, and compatibility lists naming three or more distinct capacities.

There is no step anywhere that converts binary to decimal. The number in the title is taken at face value, and that produces two errors in opposite directions.

A title written in TiB overstates price per terabyte by about 10%

A seller who types what Windows reported instead of what the label said writes “3.63TB” for a 4 TB drive:

delivered 60, capacity read as 3.63 TB   =  16.53 per TB
delivered 60, capacity read as 4 TB      =  15.00 per TB
                                  error  =  +10.2% on the ranking

The underlying unit gap is 9.95%; the rounding in “3.63” supplies the rest. Either way the listing sorts worse than it deserves, by more than the spread most of the ranking turns on. The same seller behaviour produces 931GB for 1 TB, 465GB for 500 GB, 7.27TB for 8 TB and 14.55TB for 16 TB.

A title written “1024GB” understates it by 2.34%

The error runs the other way when a seller writes the binary number as though it were the label. The site’s eBay capacity bands already normalise this one spelling - the 1,000 GB band lists ['1 TB', '1TB', '1000GB', '1024GB'], mapping 1024GB onto 1,000 GB - but the free-text title parser does not do the same. A title reading “1024GB” parses as 1,024 GB, which is 2.4% more than the drive holds and understates price per TB by 2.34%. It also means that row will not appear under 1 TB drives, because the filter matches exactly.

That is not a straightforward bug to fix, because COMMON_CAPACITIES_GB.SSD lists 1,024 and 2,048 as legitimate SSD capacities, which is true for some vendors and false for others.

The ambiguity that cannot be resolved at all

The catalogue documents the sharpest case in its own source, and the comment is worth quoting because it is the correct answer rather than a limitation:

“1.8 TB IS THE REASON THIS LIST MAY NOT BE USED TO SNAP A VALUE. It is a real 10K SAS SKU, and it is also how a careless seller writes the size Windows reports for a 2 TB drive. Both readings are in this catalogue’s input, no rule tells them apart.”

A nearest-neighbour correction would silently rewrite one into the other, and there is no evidence in a title that distinguishes them. So the stated number stands, and the reader is told. The same trap sits at 931GB against 1TB, 465GB against 500GB, 3.63TB against 4TB and 7.27TB against 8TB - in each pair, one reading is a real product and the other is an operating system’s arithmetic.

Lots: total capacity is not the capacity you get

“10 TB” in a title can be one 10 TB drive or five 2 TB drives, and per-terabyte cost cannot tell them apart. Everything else can.

1 × 10 TB 5 × 2 TB
Drive bays / SATA ports 1 5
Idle power, at about 6 W each about 6 W about 30 W
Over a year about 53 kWh about 263 kWh
Sequential throughput one drive’s worth up to 5× striped
A single failure costs 10 TB 2 TB
Chance of at least one failure at 1% AFR 1.00% 4.90%
Expected TB lost per year 0.10 0.10

The last two rows are the interesting pair. Five drives are almost five times as likely to produce a failure event - 1 − 0.99⁵ = 4.90% - but each event costs a fifth as much, so the expected bytes lost per year is identical. What differs is the shape of the loss, and which shape you want depends entirely on whether the drives are in an array. Five independent drives with no redundancy is the worst arrangement available: five chances to lose something, and no parity.

The power difference is 210 kWh a year; multiply by your own tariff. Spin-up is separate again - a 3.5-inch drive pulls a couple of amps on the 12 V rail coming up to speed, and five doing it at once is what staggered spin-up exists for.

Older drives are also slower in proportion to their age, since sequential rate tracks areal density. Seagate’s Exos X24 manual publishes a maximum sustained transfer rate of 272 MiB/s at the outer diameter and 285 MB/s maximum; a 2 TB drive from the generation when 2 TB was the flagship publishes roughly half that. Five striped still beat one modern drive on throughput and lose badly on everything per watt.

Three other datasheet numbers from that manual belong here, because they are specification bounds rather than measured field rates and are routinely quoted as though they were the latter: non-recoverable read errors at 1 sector per 10¹⁵ bits read, an annualised failure rate of 0.35% based on 8,760 power-on hours, and a maximum workload rate of 550 TB per year, where “Workload Rate = TB transferred × (8760 / recorded power on hours)”. For contrast, the BarraCuda SSD product manual gives UBER as 1 error in 10¹⁶ bits read and TBW of 120, 249, 485 and 1,067 for its 250, 500, 1,000 and 2,000 GB models, with the warranty running “five years, or when the device reaches Host TBW, whichever happens first”. The un-round numbers are the tell that they were read off the manual; round ones quoted for the same drives have been through a marketing page first. SSD UBER is an order of magnitude better than HDD URE on these sheets, which is a real difference and still an upper bound rather than a measurement.

And read the quantity field before any of this applies: a title saying “10TB” beside a quantity of five may mean the price buys one drive, not the lot - buying used drives on eBay covers how that goes wrong in both directions.

Advertised against usable in an array

Parity costs capacity in a way that is exactly predictable, unlike everything else on this page. For n drives of capacity C with p parity drives, usable is (n − p) × C, and overhead is p ÷ n. Worked through for 4 TB members:

Layout Drives Raw Usable Usable as TiB Overhead Survives
RAID 0 4 16 TB 16 TB 14.55 0% nothing
RAID 1, 2-way 2 8 TB 4 TB 3.64 50% 1
RAID 1, 3-way 3 12 TB 4 TB 3.64 66.7% 2
RAID 10 4 16 TB 8 TB 7.28 50% 1, sometimes 2
RAID 5 / RAIDZ1 4 16 TB 12 TB 10.91 25% 1
RAID 5 / RAIDZ1 8 32 TB 28 TB 25.47 12.5% 1
RAID 6 / RAIDZ2 6 24 TB 16 TB 14.55 33.3% 2
RAID 6 / RAIDZ2 8 32 TB 24 TB 21.83 25% 2
RAIDZ3 12 48 TB 36 TB 32.74 25% 3

Overhead falls as the array widens while the protection stays fixed, which is the whole argument for RAID 6 over RAID 5 at width and against very wide RAID 5. “Sometimes 2” for RAID 10 means a second failure is survivable only in the other mirror: with four drives, two of the three possible second failures are.

Four things that table cannot show.

RAID-Z does not hit those numbers, and the rule that says why is misattributed

ZFS writes variable-width stripes, and the allocation constraint that explains the shortfall is routinely cited to zpoolprops(7), which does not contain it - the man page says only that “the amount of space used in a raidz configuration depends on the characteristics of the data being written”. The rule is Matt Ahrens’s, from his analysis of RAID-Z space accounting: RAID-Z “requires that each allocation be a multiple of (p+1), so that when it is freed it does not leave a free segment which is too small to be used”, where the unit is a sector of 2^ashift bytes. Allocations round up, producing up to three pad sectors per block whose contents are never used. ashift is a pool property, valid from 9 to 16, defaulting to 0 meaning autodetect. Cite the analysis, not the man page - a reader who goes to zpoolprops(7) looking for the p+1 sentence will not find it.

Working the rule through at ashift=12 - 4 KiB sectors, which is what a modern drive wants - gives efficiencies that differ from the naive (n − p)/n by several points:

Layout Efficiency at 128 KiB recordsize Naive (n−p)/n Efficiency at 4 KiB recordsize
raidz1, 4 wide 72.7% 75.0% 50.0%
raidz1, 5 wide 80.0% 80.0% 50.0%
raidz1, 8 wide 84.2% 87.5% 50.0%
raidz2, 6 wide 66.7% 66.7% 33.3%
raidz2, 8 wide 71.1% 75.0% 33.3%
raidz2, 10 wide 76.2% 80.0% 33.3%
raidz3, 12 wide 72.7% 75.0% 25.0%

Those figures are computed from that p+1 rule, not measured on a pool. Two results in the table are worth carrying away.

At a 4 KiB record, every raidz1 is 50% efficient, every raidz2 is 33% and every raidz3 is 25%, regardless of width. A block that small cannot be spread, so each one costs a full parity sector per parity level. A pool of databases or VM images with a small recordsize does not get the parity efficiency the width suggests; a pool of large media files lands close to the naive table.

The widths that come out even are not accidental. Ahrens’s recommendation for 512-byte-sector devices is to “use at least 5 disks with RAIDZ1; use at least 6 disks with RAIDZ2; and use at least 11 disks with RAIDZ3”, and the rows above show why - 5-wide raidz1 and 6-wide raidz2 are the widths where the p+1 rounding costs nothing at 128 KiB. He also notes that “each RAID-Z group has approximately the performance of a single disk in the group” for random IOPS.

The overhead figure attached to all this is quoted in at least three conventions, and mixing them is how an 8-wide raidz2 ends up being compared with RAID 5. Do it once, in sectors, with the working shown:

8-wide raidz2, ashift=12, 128 KiB record

  128 KiB ÷ 4 KiB          =  32 data sectors
  32 across 6 data columns =   6 rows, five of 6 sectors and one of 2
  parity, 2 sectors a row  =  12 sectors
  32 + 12                  =  44, rounded up to a multiple of p+1 = 3
                           =  45 sectors allocated
  45 − 32                  =  13 sectors of parity and padding

  32 ÷ 45  =  71.1% efficient, which is the table row above
  13 ÷ 32  =  40.6% on top of the data
   2 ÷ 8   =  25% by the p ÷ n convention the array table uses

So that array spends about 41 sectors of parity and padding for every 100 sectors of data, where an 8-wide RAID 6 - the same protection at the same width - spends 33, and where the array table above states the same layout as 25%. The comparison is only meaningful once both sides use the same denominator, and RAID 5 is not the other side of it at all: single parity across eight drives costs 12.5% of raw.

ZFS will not show you the difference either. The deflation ratio it applies to a raidz vdev is computed once from a 128 KiB block - vdev_set_deflate_ratio in the OpenZFS source does it with 1 << 17 - so a pool whose datasets use a small recordsize reports free space it cannot actually hand out.

Array metadata takes a slice before parity does

Linux md with a 1.2 superblock offsets the data start, commonly by 264,192 sectors:

264,192 × 512 = 135,266,304 B = 129 MiB per member

4-drive RAID5 of 4 TB members    3 × 135,266,304  =   405,798,912 B short of 12 TB
8-drive RAID6 of 20 TB members   6 × 135,266,304  =   811,597,824 B short of 120 TB

The offset covers the superblock itself plus room for a write-intent bitmap and reshape headroom, and it is not a fixed constant - mdadm scales it with device size, so check mdadm --examine rather than assuming 129 MiB.

LVM adds its own. pvcreate(8) documents that “with default settings, the first physical extent (PE), which contains LV data, is 1 MiB from the start of the device”, controlled by default_data_alignment in lvm.conf, with the metadata area starting one page (usually 4 KiB) in. The default physical extent size is 4 MiB, and logical volumes round up to whole extents, so a stack of many small LVs loses up to 4 MiB each to rounding.

A NAS charges twice, before and after parity

Synology’s own RAID calculator states the two losses that sit either side of the parity arithmetic: “Each drive in the RAID must reserve approximately 10 GB of system space”, and “Volumes with Btrfs file system reserved 4% of the capacity for metadata, and 2% for volumes with ext4 file system”.

8-bay SHR-2 of 8 TB drives
  raw                            8 × 8 TB          =  64 TB
  system space                   8 × 10 GB         =  −80 GB
  two drives of parity           6/8 of the rest   ≈  47.94 TB
  btrfs metadata reserve         −4%               ≈  46.02 TB

That is a worked model using Synology’s published figures, not a Synology quotation. The order matters: the 10 GB is per drive and comes off before parity, the 4% comes off the volume after it. Drives for a NAS covers the rest of the sizing decision.

The array is sized by its smallest member

IDEMA’s and SFF-8447’s fixed sector counts are why this is a non-issue between hard drives of the same nominal capacity from different vendors - the Exos X24 and Ultrastar DC HC580 agreeing to the byte at 24 TB is the demonstration. SSDs are less consistent, and a replacement a few thousand sectors short is simply refused. Compare blockdev --getsize64 across used drives before committing rather than after.

There is a further reason a model number does not settle what you are buying: vendors second-source components under an unchanged SKU. Solidigm’s Product Change Notification 0000023684-00, dated 2 July 2025, adds new DRAM part numbers from SK hynix and Nanya and a new enclosure supplier to the D5-P5430 line, with the existing parts marked end of life, all under the same product SKUs. The capacity does not change; the hardware behind the number does.

One thing the capacity table cannot show is that rebuild time scales with member size and array throughput and nothing else: 20 TB ÷ 200 MB/s = 100,000 s = 27.78 hours at full speed with no other load, and days on a loaded array, all of it degraded. Whether a second drive is likely to fail in that window is the argument for double parity, and it turns on unrecoverable read error rates that are datasheet upper bounds rather than measured field rates. That one is genuinely unsettled, and drives for a NAS takes it apart honestly.

Parity is not backup, and a per-terabyte comparison that ignores the redundancy tier compares two different products. Decide the layout first, then sort on price per terabyte for the capacity that layout actually needs.

What to do with this on a listing page

  1. Sort by price per TB, and read the capacity column as decimal. Everything on this site is in the unit the label uses. A listing whose title reads 3.63TB, 931GB, 7.27TB or 14.55TB has been typed off a Windows screen, and its price per TB is about 10% worse than the drive deserves - that is a discount hiding in a rounding error, and it is the single most exploitable thing on this page.
  2. Filter on the round capacity, then search the binary spelling separately. 4 TB hard drives matches titles that say 4TB or 4000GB; it does not match one that says 3.63TB or 3726GB, because the parser takes the title at face value. The mismatched rows are where the bargains sit.
  3. On SAS, assume T10 Protection Information until the seller says otherwise. A 24 TB SAS drive reporting 23.5 TB is PI-formatted, not short, and a reformat to 512 or 4,096 byte blocks returns 2.03% of it. Price the listing on the label capacity and budget the hours.
  4. Check a large drive reports a whole number of GiB. Above 8 TB, SFF-8447 capacities land on exact gibibyte boundaries - 22,352 GiB at 24 TB, 18,627 at 20 TB, 14,902 at 16 TB. A fractional figure on arrival means HPA, DCO or a namespace clip, and hdparm -N, hdparm --dco-identify and nvme id-ns are the three commands that say which.
  5. Buy the capacity the array needs, not the raw total. Parity overhead is p ÷ n exactly for md and hardware RAID, and worse than that for RAID-Z at small record sizes. Work out usable first, then filter.
  6. On SSDs, read over-provisioning off the label. 480, 960, 1920 and 3840 GB are 14.51% spare; 400, 800 and 1600 are 37.41%; the round decimal numbers are 9.93%. The smaller number in each pair is the same silicon with more held back, and it is the endurance tier stated in the one field every listing carries.
  7. Decide the filesystem before the drive arrives. XFS offers 262 GB more than default ext4 on a 4 TB drive, and mke2fs -i 262144 -m 0 recovers 258 GB of that within ext4 itself. Neither choice can be made after the data is on it.
  8. Compare delivered cost, never item price, and never mix TB with TiB on either side of the division. Both errors are about 10%, which is larger than the spread you are shopping in.
Affiliate disclosure:We are a member of the eBay Partner Network and earn a commission from qualifying purchases made through links to eBay on this site. Prices and availability are captured periodically and may have changed - the live price is always the one shown on eBay.