Buying used drives on eBay: what SMART tells you, and what it can't
By Harry Saarinen ·
You cannot make it safe. You can make it cheap, reversible and quickly falsifiable, and that is the entire game.
Buy the SMART report, not the drive. The photograph, the condition word and the seller’s account of where the thing came from are unverifiable or already priced in; the drive’s own record of its life is the only part of the transaction carrying information. The decision you are actually making is whether to spend the money before you have seen that record or after.
And then, since early 2025, a second decision: whether to believe the record at all. In January and February 2025, used Seagate nearline drives with roughly 25,000 hours of service on them were sold as new into Germany, Switzerland, Austria, Luxembourg, the UK, the Czech Republic, Japan and the USA, eBay included, with the SMART power-on hours counter reset to near zero. SMART is firmware-reported and unauthenticated; it has always been forgeable in principle. In 2025 it was forged at scale. The counter-measure is a second, separate lifetime log on Seagate drives that the forgers did not touch, and it is the single most useful thing added to this guide.
A drive has a consumable budget - bearings, an actuator, a surface flown over at a height measured in nanometres, or for solid state a fixed number of program and erase cycles - and unlike a stick of memory it keeps a log. Most of what follows is about reading that log properly, which is harder than it looks because almost none of it is standardised. The rest is about the fraud the log cannot catch, which you have about thirty days to find.
What the listing withholds
eBay gives you a capacity, an interface, a form factor, a condition code and a price. This site turns those into a price per terabyte and ranks them, which is not the same as saying the drive is good.
| Unknown from the listing | Where it lives | How you get it |
|---|---|---|
| Power-on hours | SMART attribute 9, ATA Device Statistics page 0x01, FARM page 1 | smartctl -a, -l devstat, -l farm |
| Whether the hours are real | FARM log, Seagate only | smartctl -l farm |
| Sectors remapped, pending, or unreadable | SMART 5, 197, 198; Pending Defects log | smartctl -a, -l defects |
| Which LBAs are bad | ATA Pending Defects log, self-test log | smartctl -l defects, -l selftest |
| Lifetime bytes read and written | Device Statistics page 0x01; NVMe health log | smartctl -l devstat, nvme smart-log |
| Head-level wear, depopulated heads | FARM reliability pages, Seagate only | smartctl -l farm |
| Recording technology | Sometimes the smartmontools database; FARM on Seagate | smartctl -i, -l farm |
| Whether a host protected area clips it | ATA READ NATIVE MAX vs IDENTIFY | hdparm -N |
| Whether the capacity is real | Nowhere. It is not reported | A full write-and-verify pass |
That last row is the one people get wrong, and it is worth stating as hard as it deserves. No attribute, log page, self-test or vendor diagnostic proves a drive holds the capacity it claims. The claimed size is the drive’s own answer to a question the operating system asks, the operating system has no way to check it, and establishing it is your job and costs you a day.
Asking for SMART, and what a refusal is worth
Ask the seller for the output of smartctl -x on the exact drive, serial number
visible. A photograph of a terminal is fine; a CrystalDiskInfo window is what most
Windows sellers will send, and it is worse for a reason covered below. Four
things make the answer useful rather than decorative.
- The serial must be in the output, and must match the drive that arrives. A
report with the serial cropped is a report about some drive and you cannot tell
which. Check it against the label photograph before buying, against
smartctl -ion arrival, and against the manufacturer’s warranty lookup, which is the only authoritative external check on a drive’s identity and age that exists. Seagate runs one at seagate.com/support/warranty-and-replacements/ and Western Digital at support-en.wd.com; both want the serial and the model or part number. A drive sold through an unauthorised channel, or as an OEM system-builder part, frequently returns no entitlement at all, which is information rather than a red flag. - Attributes 5, 9, 187, 188, 197, 198 and 199 at minimum, plus the self-test log, plus - on any Seagate - the FARM log. A screenshot cropped to a green “Good” banner says only that no pre-fail attribute has crossed its vendor-set threshold, which is far weaker than it looks and is explained in the next section.
- Ask for
-x, not-a. The extended report adds the Device Statistics log, the error log with timestamps, the SCT error recovery settings and the SATA capability words. Those are the vendor-neutral parts, and they are exactly what a GUI screenshot throws away. - On a lot, ask whether the price is per drive or for the lot. Lots are where a price per terabyte most often looks extraordinary and most often turns out to be the seller’s arithmetic rather than a bargain. See capacity in lots for how this site computes the total.
A refusal proves nothing. A recycler moving four hundred pulls a week cannot transcribe a report for each; a seller with three listings can. Treat a refusal as a price rather than a verdict, and buy into unreported lots cheap enough that a dud is an annoyance rather than a dispute. What you should not accept is a substitute: “tested, no bad sectors” means the seller ran something and liked the result, and Google’s own fleet engineers documented the failure mode behind that sentence - they observed “situations where a drive tester consistently ‘green lights’ a unit that invariably fails in the field”.
The commands to send, so the answer is complete
Send these rather than asking for “a SMART report”, because most sellers will run what you paste and nothing more.
sudo smartctl -x /dev/sdX # everything: attributes, logs, devstat
sudo smartctl -l devstat /dev/sdX # standardised counters, log 0x04
sudo smartctl -l selftest /dev/sdX # test history with lifetime timestamps
sudo smartctl -l defects /dev/sdX # the LBAs of pending sectors
sudo smartctl -l farm /dev/sdX # Seagate only; the tamper cross-check
Use smartmontools 7.5 or newer if you are the one running them. Version 7.4 introduced FARM decoding, and 7.5 - released 30 April 2025 - added hybrid USB-bridge autodetection and per-namespace NVMe health output. An older build will silently omit the two things most worth having.
The three columns, and which one SMART judges by
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED RAW_VALUE
1 Raw_Read_Error_Rate 0x000f 083 064 044 Pre-fail Always 201546824
5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always 0
7 Seek_Error_Rate 0x000f 089 060 045 Pre-fail Always 823410293
9 Power_On_Hours 0x0032 055 055 000 Old_age Always 39820
12 Power_Cycle_Count 0x0032 100 100 020 Old_age Always 58
187 Reported_Uncorrect 0x0032 100 100 000 Old_age Always 0
188 Command_Timeout 0x0032 100 100 000 Old_age Always 0
197 Current_Pending_Sector 0x0012 100 100 000 Old_age Always 0
198 Offline_Uncorrectable 0x0010 100 100 000 Old_age Offline 0
199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age Always 0
Attribute 1 reports 201,546,824 read errors. The drive is fine.
Three columns are worth separating. VALUE is a normalised health figure, conventionally starting at 100 or 200, higher is better, and the vendor owns the entire mapping from physical reality to that number. THRESH is the vendor’s own condemnation line. RAW_VALUE is 48 bits of vendor-encoded state that may be a count, a rate, a packed pair of counts, a temperature triple, or a duration in units the vendor chose.
What SMART itself judges by is VALUE against THRESH, and only for Pre-fail
attributes. When a pre-fail attribute’s VALUE reaches its THRESH, the drive’s
answer to the SMART RETURN STATUS command flips, and that alone is what makes
smartctl -H print FAILED. Old_age attributes can sit below threshold
indefinitely without changing the verdict.
Two consequences follow, and both cut against the buyer.
The first is that PASSED is a very low bar. It is one boolean from firmware,
and on ATA even that is fragile: the smartmontools manual notes the return value
“may be unknown due to limitations or bugs in some layer (e.g. RAID controller or
USB bridge firmware) between disk and operating system”, in which case smartctl
falls back to checking whether any pre-fail attribute has reached its threshold.
On NVMe it is weaker still - an NVMe PASSED is a single byte, the Critical
Warning byte of the SMART/Health log, and nothing else.
The second is that the normalised column can only ever tell you what the vendor decided to be alarmed about. It is not a measurement you can compare across brands, and it is not a percentage of anything. The raw column is where the physics is, and the raw column is a minefield.
Nothing in SMART is standardised, including the parts that look standardised
The common framing is that low attribute IDs are standard and high ones are vendor-specific. That is wrong, and correcting it is the precondition for reading any of this properly. Seagate’s own tool documentation for openSeaChest states it without hedging: “every drive vendor can implement their own attributes to track whatever the vendor thinks is necessary to determine drive health. There are no requirements that a specific attribute is implemented or tracked in any specific way. This is vendor specific and always has been.” The same page concludes that “straight comparisons between vendors and SMART attributes does not make sense on its own”.
Every attribute at every ID is vendor-defined. The low IDs are conventional, not standardised, and the convention is maintained by smartmontools rather than by any standards body. The attribute names printed above are smartmontools’ names, not the drive’s; the drive reports a number and a blob.
Raw values are bytes, not integers
The raw field is 48 bits, and smartctl’s -v flag takes a byte-order string that
selects which of those bytes to print and in what order. The manual gives the
alphabet as characters from the set 012345rvwz, where 0 to 5 select byte 0
to byte 5 of the raw value, and the default order is 543210 - the little-endian
reading of the six stored bytes, with 012345 available for the big-endian one -
which is why an attribute that is really two packed counters prints as one
enormous meaningless integer.
This is the mechanism behind a trick that circulates for Seagate drives without its explanation attached:
smartctl -A -v 1,raw48:54 /dev/sdX
5 and 4 select the top two bytes and discard the rest, printing the 16-bit
error count and throwing away the 32-bit operation count sitting underneath it.
The properly decoded form is better. smartmontools’ own drive database assigns
raw24/raw32 to attributes 1 and 7 for the Seagate Exos X16, X18, X20/X22 and
X24 families, documented as “an error rate which consists of a 24-bit error count
and a 32-bit total count”. That name is nominal rather than arithmetic: 24 plus
32 exceeds the 48 bits available, and what smartmontools actually does is take
the top 16 bits as the error count and the low 32 as the total, which is exactly
why selecting bytes 5 and 4 above reproduces the numerator. So the raw value is
two numbers, and the second is the denominator: read errors over sectors read.
201,546,824 is not a count of anything; it is a fraction printed as an
integer.
The database only does this for models it knows. An unlisted, rebranded or recertified model prints the bare 48-bit integer and needs the decode supplied by hand:
smartctl -A -v 1,raw24/raw32 -v 7,raw24/raw32 /dev/sdX
The defaults are not uniform either. Attributes 5 and 196 default to
raw16(raw16) - a 16-bit value plus two optional 16-bit values printed only if
non-zero - and attribute 9 defaults to raw24(raw8). So a figure in parentheses
beside a reallocated sector count is not a decorative flourish, and a bare copy
of the 48-bit integer into a spreadsheet is a different number from the one
smartctl showed you.
Attribute 9 does not always count hours
This is the clean disproof of the “low IDs are standard” belief, and it sits at
the most-quoted attribute on the list. smartmontools’ database records that the
Seagate BarraCuda 3.5 SMR family reports attribute 9 in msec24hour32 format -
milliseconds and hours packed together - rather than plain hours, tested against
an ST4000DM004. Other presets in the same file cover drives counting in minutes,
in seconds, and in half-minutes, the last noted as used by some Samsung disks.
A shingled BarraCuda whose “power-on hours” you read as a plain integer will look
absurdly old, and only ever in that direction: the millisecond half sits above
the hour half, so the integer overstates the age by a factor of billions whenever
that half is non-zero and matches the real hours only when it happens to be zero.
Run smartctl -i first and read the Model Family line, because that line
tells you whether smartmontools recognised the drive and therefore whether the
decoding you are looking at is the drive’s or a fallback.
Attribute 241 has no fixed unit either
Total_LBAs_Written defaults to a plain 48-bit count of 512-byte sectors, and some Western Digital and Seagate drives report it in 1 GiB units instead. The two interpretations differ by a factor of about two million. A lifetime-writes figure computed without checking the model against the database can be wrong by six orders of magnitude in either direction, which is the kind of error that looks like a conclusion. Where the drive offers the Device Statistics log, use that instead - it is unambiguous.
The five attributes the fleet studies actually support
Backblaze publishes drive statistics from a fleet in the hundreds of thousands and watches five: 5, 187, 188, 197 and 198. They count different things, and the differences are the point.
5, Reallocated_Sector_Ct. A sector failed and the drive substituted a spare from a reserved pool set aside at manufacture. Non-zero but stable across 20,000 hours is a drive that lost a few sectors early and got on with its life. A figure that moves while you own it is a drive eating its spare pool, and when the pool is exhausted the remapping stops being invisible to you.
197, Current_Pending_Sector. A sector that failed a read and has not been reallocated, because the drive will not spend a finite spare on the strength of a read failure alone. It resolves on the next write to that LBA: if the write and subsequent verify succeed, the count decrements with nothing remapped; if they fail, the sector is retired and attribute 5 goes up. This is why only a write pass can clear these numbers, why a read pass can only ever add to them, and why a pending sector is unresolved rather than condemned.
198, Offline_Uncorrectable counts sectors the drive could not read or recover during its own background scanning, and usually tracks 197 closely enough that Backblaze’s own analysis says “only SMART 197 and 198 have a good correlation, meaning we could consider them as one indicator versus two”.
187, Reported_Uncorrect counts errors the drive’s error correction could not fix. This is the one where data was lost rather than recovered, and it is the attribute to weight most heavily on a drive you are buying to hold something.
188, Command_Timeout counts commands the drive failed to complete in time. It is the odd one out because it is frequently not the drive: a marginal cable, a failing power supply rail or a flaky controller produces timeouts on a healthy disk. Treat a non-zero 188 the way you treat a non-zero 199 - as a claim about the system the drive was in, pending re-test in yours.
The numbers underneath that list
Backblaze’s October 2016 analysis across 67,814 drives is the clearest published statement of what these five buy you, and it is worth writing out in both directions because only one direction usually gets quoted.
Backblaze, 67,814 drives, published 6 October 2016
failed drives with one or more of 5/187/188/197/198 non-zero 76.7 per cent
operational drives with one or more non-zero 4.2 per cent
failures that gave NO warning on any of the five 23.3 per cent
So the five attributes are an excellent filter and a mediocre alarm. Nearly a quarter of the failures announced nothing. Backblaze also reports attribute 189, High Fly Writes, at 47.0 per cent of failed drives against 16.4 per cent of operational ones - a real signal, weaker separation, and a reasonable thing to look at on a drive you are considering rather than a drive you own.
For dating: Backblaze’s 2025 report, published 12 February 2026, gives an annualised failure rate of 1.36 per cent across the 344,196 drives in the annual cohort, down from 1.55 per cent in 2024, with a lifetime rate of 1.30 per cent.
5, 197 and 198 are three stages of one event
Put plainly, and this is the sentence to carry to a listing:
187 or 198 non-zero is a decline. 197 non-zero is a decline unless you are buying knowingly and cheaply. Attribute 5 small, old and stable is a negotiation; attribute 5 climbing is not.
199, UDMA_CRC_Error_Count, counts errors on the SATA link itself - cable and connector included - so a non-zero value on arrival is most likely the seller’s cabling or the courier’s handling of a drive that was not packed. Google’s study found the same thing at fleet scale: “CRC errors are less indicative of drive failures than that of cables and connectors”. Change the cable, change the port, run a read pass, see whether it moves. If it does not, it is history.
A zero in 197 or 198 is more ambiguous than it looks
Some drives reset attributes 197 and 198 when the underlying uncorrectable
sectors get reallocated, rather than leaving the historical count in place.
smartmontools carries explicit overrides for this, documented as
197,increasing, meaning “Attribute number 197 (Current Pending Sector Count) is
not reset if uncorrectable sectors are reallocated”, with the same wording for
198.
The implication is unpleasant and worth stating: on some drives, a zero in 197 means “resolved and forgotten” rather than “never happened”, and the only record that the event occurred is a corresponding rise in attribute 5 that you have no baseline for. This is one of several reasons to prefer the Device Statistics log, which counts monotonically and does not have presets.
Clearing a pending sector deliberately
If you own the drive and it reports pending sectors, you do not need a full write pass to resolve them. The self-test log reports the LBA of the first error, and writing to that specific block forces the drive to decide:
sudo smartctl -l selftest /dev/sdX # read LBA_of_first_error
sudo dd if=/dev/zero of=/dev/sdX bs=4096 count=1 seek=BLOCK oflag=direct
sync
sudo smartctl -A /dev/sdX # 197 should have decremented
BLOCK is the LBA converted to your block size. A published worked example shows
Current_Pending_Sector falling from 81 to 80 after exactly one such write, which
is also a fair illustration of why this is a diagnostic technique rather than a
repair strategy: a drive with 81 pending sectors is telling you something a
targeted write does not fix. Do this to learn whether the sector is
recoverable, not to make a number look better before resale. And do not do it
at all on a drive still holding data you have not copied off.
Google 2007, dated honestly
The Google study is the most-cited paper in this subject and it is twenty years old. Its population was consumer PATA and SATA drives of 80 to 400 GB, spinning at 5400 to 7200 rpm, with data collected between December 2005 and August 2006. Not one of those drives had perpendicular recording at the densities now normal, none had helium, none had a media cache, and the largest of them held one fiftieth of what a current nearline drive holds. Its conclusions about correlation structure have aged far better than any of its absolute rates, so use it for shape and not for magnitude.
The shape is worth having. After a first event, within a 60-day window:
| First event | Increase in failure probability |
|---|---|
| Scan error | 39x |
| Reallocation | over 14x |
| Offline reallocation | over 21x |
| Probational count | 16x |
In every case the critical threshold is one. The paper’s own framing is that the jump happens at the first occurrence, not at some tolerable count, which is the single most useful thing in it for a buyer: you are not looking for a small number, you are looking for zero or not-zero.
The negative result is stronger than it is usually quoted. The paper reports that “over 56% of them have no count in any of the four strong SMART signals”, and then goes further: “even when we add all remaining SMART parameters (except temperature) we still find that over 36% of all failed drives had zero counts on all variables”. More than a third of failures were invisible to every SMART parameter the drive exposed.
Prevalence matters too, because it sets what a clean report is worth. In that population, fewer than 2 per cent of drives showed scan errors, about 9 per cent had a non-zero reallocation count, about 2 per cent had probational counts and about 2 per cent had CRC errors. A clean report is the normal case, not an achievement.
Two of its findings cut against advice that used drive guides repeat, this one included in its previous version.
Power cycles matter less than the folklore says. The paper found no significant correlation between failures and high power cycle counts for drives aged up to two years. For drives three years and older, higher counts raised the absolute failure rate by over 2 per cent, and the authors attributed even that to population mix rather than to accumulated wear. So the argument below about transitions being the expensive events is a mechanical argument about load cycles and start currents, not an epidemiological one, and it should not be dressed up as the latter.
Seek errors are a single vendor’s artefact. The paper found seek errors “widespread within drives of one manufacturer only”, with no correlation to failure for the others. Attribute 7 is not a cross-brand health signal and should not be read as one.
Splicing the multipliers onto a modern base rate
The Google paper gives multipliers, not probabilities, and the tempting move is to combine them with a current failure rate. That combination is legitimate arithmetic and illegitimate statistics - the populations are twenty years and one technology generation apart - so here it is, clearly labelled as my splice, not either source’s claim:
Backblaze lifetime AFR 1.30 per cent per year
scaled to a 60-day window 1.30 x (60 / 365) = 0.214 per cent
after a first scan error x 39 = 8.3 per cent
after a first reallocation x 14 = 3.0 per cent
after a first offline realloc x 21 = 4.5 per cent
The point is not the numbers. The point is that even a 39x multiplier on a small base rate leaves a drive that is more likely than not to survive the next two months, which is why “it has one reallocated sector” is a price argument rather than a refusal - and why it is a refusal anyway for the only copy of anything.
Device Statistics: the vendor-neutral log nobody reads
There is a standardised alternative to the entire attribute mess, it has existed since ACS-2, and almost nobody asks a seller for it. The ATA Device Statistics log is General Purpose Log address 0x04, its fields are defined by the standard rather than by the vendor, and Seagate’s own openSeaChest documentation recommends it over SMART attributes, describing these as consistent, vendor-neutral health and usage counters.
sudo smartctl -l devstat /dev/sdX
What the pages carry, by the specification’s own field names:
| Page | Fields that matter to a buyer |
|---|---|
| 0x01 General Statistics | Lifetime Power-on Resets, Power-on Hours, Logical Sectors Written, Logical Sectors Read, Number of Write Commands, Number of Read Commands, Pending Error Count |
| 0x03 Rotating Media Statistics | Spindle Motor Power-on Hours, Head Flying Hours, Head Loaded Events, Number of Reallocated Logical Sectors, Read Recovery Attempts, Number of Mechanical Start Failures, Number of Reallocation Candidate Logical Sectors, Number of High Priority Unload Events |
| 0x04 General Errors | Number of Reported Uncorrectable Errors |
| 0x06 Transport | Number of Interface CRC Errors |
Three of those are worth more than their SMART counterparts. Logical Sectors Written and Logical Sectors Read give you lifetime workload in a defined unit, which attribute 241 does not. Number of Mechanical Start Failures is a count of times the drive failed to spin up, which no conventional attribute reports and which is the single most direct statement of bearing and motor health available. And Spindle Motor Power-on Hours against Head Flying Hours is a pair that must obey a physical inequality, which turns out to be a tamper detector.
Self-tests prove less than their name suggests
The long self-test is the standard advice and it is good advice, but it is usually described as “the drive checks its whole surface”, and it is not that.
sudo smartctl -c /dev/sdX # how long the tests will take
sudo smartctl -t long /dev/sdX # start it
sudo smartctl -l selftest /dev/sdX # poll until complete
It stops at the first error it finds
The smartctl manual is explicit: “If a test failure occurs then the device may discontinue the testing and report the result immediately.” So a failed long test proves the drive read cleanly up to the reported LBA and says nothing about the surface beyond it. It never enumerates more than one bad sector. A drive with four hundred unreadable sectors and a drive with one produce the same output, differing only in the LBA printed.
That is why a host-side read pass is not redundant with a long self-test. The
self-test is the drive’s opinion using the drive’s internal retry policy; a
badblocks read pass enumerates every failure to the end of the device.
The log keeps 21 entries and the clock wraps at 7.5 years
The cross-check everyone recommends - compare the power-on hours in the self-test log against the drive’s claimed age - has two limits that break it on exactly the drives it is aimed at.
The ATA self-test log retains only the 21 most recent entries. A seller who runs 21 short tests has erased the log’s entire history, without needing any tool more exotic than smartctl. The SCSI equivalent keeps twenty, with the failing LBA in hexadecimal.
And the timestamp is 16 bits of hours. The manual notes that it “wraps after 2^16 hours, or 2730 days and 16 hours, or about 7.5 years”. Work it through on a genuine high-hours nearline pull:
2^16 hours = 65,536 h
a genuine drive at 70,000 h
logged timestamp 70,000 - 65,536 = 4,464 h
which reads as a drive with six months of use
A 70,000-hour drive logs its self-tests at 4,464 hours. If you use the self-test log as an age check without accounting for the wrap, the oldest drives in the market are the ones that look youngest. Use the FARM log or Device Statistics for age, and use the self-test log for what it is good at: whether tests were run, whether they passed, and whether the timestamps are internally consistent with each other.
Conveyance is the test written for this exact situation
Nearly nobody runs it, and it exists precisely for a drive that arrived in a padded envelope. The manual describes conveyance as an ATA-only routine “intended to identify damage incurred during transporting of the device”, and it runs in minutes rather than hours.
sudo smartctl -t conveyance /dev/sdX
Run it the moment the drive is connected, before the long test, because it answers the cheapest question first: did the courier break it. A conveyance failure on a drive whose seller-supplied report was clean is a packaging dispute rather than a drive dispute, and the distinction matters when you write the case.
Selective self-tests re-test one range
When a long test fails at an LBA, repeating the full multi-hour test to see whether the failure is stable is wasteful. The selective test runs a span:
sudo smartctl -t select,100000000-max /dev/sdX
The manual notes the -t option “can be given up to five times, to test up to
five spans” in one invocation. Re-testing the neighbourhood of a reported error,
plus the outermost and innermost tracks, is fifteen minutes of work that tells
you whether the failure is a single defect or the edge of a region.
The command that starts a test is obsolete
Part of why self-test behaviour varies so much between drives is that the underlying command was deprecated. smartctl’s manual records that SMART EXECUTE OFF-LINE IMMEDIATE “was declared obsolete in ATA ACS-4 Revision 10 (Nov 2015)”. It still works nearly everywhere. It is no longer something the standard requires anyone to implement consistently, and behaviour differences between vendors are therefore expected rather than anomalous.
The Pending Defects log names the sectors
Attribute 197 gives you a count. ACS-4 Revision 01, March 2014, added a log that gives you the addresses.
sudo smartctl -l defects /dev/sdX
The manual describes it as printing “LBA and hours values from the ATA Pending Defects log (General Purpose Log address 0x0c)”, with the first log page’s 31 entries shown by default. Each entry carries the power-on hour at which the defect was found, which is the part that changes a decision.
Consider two drives that both report 6 pending sectors at 40,000 hours. On the first, the Pending Defects log shows all six discovered between hours 3,100 and 3,400 - an early manufacturing-era event that has been stable for the 36,600 powered hours since, more than four years. On the second, it shows them discovered at 39,720, 39,844, 39,901, 39,950, 39,988 and 39,996. That is a drive failing while you read the report. The attribute is identical in both cases.
The SCSI equivalent is better still, because it reflects background scanning rather than host reads: it prints LBAs “that the background scan was unable to read”, with the power-on hours at discovery, and the manual notes that “these pending defects may appear in advance of any application trying to read a defective LBA”.
SMART can be rewritten, and in 2025 it was
Every previous section assumed the counters describe the drive. In early 2025 that assumption failed publicly. Used Seagate nearline drives, many with around 25,000 hours of real service, were sold as new with attribute 9 reset to near zero. They reached buyers across eight countries, eBay among the channels.
There is a documented discrepancy from the affair that shows what detection looks like:
sudo smartctl -a /dev/sdb | grep "Power_On_Hours" # 4083
sudo smartctl -l farm /dev/sdb | grep "Power on Hours" # 28831
4,083 hours is 170 days. 28,831 hours is just over 1,200 days. Same drive, same minute, two logs, a factor of seven.
FARM, and why it survived
Seagate’s Field Access Reliability Metrics log is a separate lifetime record kept independently of the SMART attribute table. It lives at ATA General Purpose Log 0xa6 and, on SAS, at log page 0x3d subpage 0x3. It is Seagate-only - smartctl’s error for anything else is blunt: “FARM log (GP Log 0xa6) not supported for non-Seagate drives” - and the manual warns that “some Seagate drives do not support FARM” even among those that should. smartmontools has decoded it since 7.4, where it is still labelled experimental.
sudo smartctl -l farm /dev/sdX
Page 1, Drive Information, prints the fields that catch a wiped attribute table: Serial Number, World Wide Name, Device Interface, Number of Heads, Device Form Factor, Rotation Rate, Firmware Rev, Model Number, Drive Recording Type, Assembly Date (YYWW), Power on Hours, Head Flight Hours, Head Load Events, Power Cycle Count, Hardware Reset Count, Spin-up Time, Spindle Power on Hours and Depopulation Head Mask.
Three of those deserve calling out separately. Assembly Date in year-and-week form is a manufacture date from the drive itself, independent of the label, which is the check against a relabelled unit. Drive Recording Type is the drive stating whether it is shingled, which is otherwise one of the hardest facts to establish on the used market - see CMR vs SMR for how thin the alternatives are. And Depopulation Head Mask tells you whether the drive has had a head disabled and its capacity reduced in the field.
The reliability and error pages go further than any SMART attribute does, to per-head granularity: Unrecoverable Read Errors, Unrecoverable Write Errors, Number of Reallocated Sectors, Number of Reallocated Sectors by Head, Number of Reallocation Candidate Sectors by Head, MR Head Resistance from Head, Total Flash LED (Assert) Events, Index of the last Flash LED, Cumulative Lifetime Unrecoverable Read errors due to ERC, Helium Pressure Threshold Tripped, Depopulation Head Mask, Has Drive been Depopped, Number of Mechanical Start Failures, Time In Over Temperature (minutes) and Highest Average Long Term Temperature.
A helium drive that has tripped its pressure threshold is a drive whose sealed enclosure is leaking, and no conventional attribute says so in plain language. Reallocations concentrated on one head are a head problem rather than a media problem, which is a different purchase decision.
The tamper test that needs no reference value
Comparing SMART against FARM requires FARM, which requires a Seagate drive. There is a second test that needs neither, and it is the more elegant of the two, because it rests on a physical inequality rather than on a second opinion.
A drive’s spindle runs whenever the drive is powered and spun up. Its heads fly only some of that time. Spindle runtime must therefore exceed head flying hours, always, on every drive, with no exceptions available to firmware that is telling the truth.
FARM page 1: Spindle Power on Hours vs Head Flight Hours
devstat page 0x03: Spindle Motor Power-on Hours vs Head Flying Hours
honest drive: spindle > flying, by the accumulated idle and unload time
tampered drive: spindle == flying, because both were written, not counted
A drive reporting identical large values for the two has had them written rather than accumulated. A published example of a clean drive - four years and eight months old - showed 40,860 hours agreeing across SMART and FARM with head flying hours correctly lower. That is the shape of an honest report.
This check works on the Device Statistics log too, which means it is available on
non-Seagate drives that implement page 0x03. Ask for -l devstat on every
mechanical drive you are considering, and compare those two lines before
anything else.
Hours are not wear; workload is
The used market prices power-on hours like mileage. They are not, and the mistake runs in the direction that costs you money.
| Datacentre pull | Desktop pull | |
|---|---|---|
| Power_On_Hours (9) | 39,820 | 6,100 |
| Elapsed | 39,820 / 24 = 1,659 days = 4.5 years | at about 4 h/day, 4.2 years |
| Power_Cycle_Count (12) | 58 | about 2,400 |
| Load_Cycle_Count (193) | 1,240 | often tens of thousands |
Same age. The datacentre unit spun up 58 times in its life; the desktop unit did it 2,400 times, each one a cold start with the heads unparking and the spindle motor drawing its highest current of the day. Seagate’s Exos X24 product manual specifies 2.6 A peak on the 12 V rail at startup and 25 seconds typical from power-on to ready, 30 seconds maximum - that is the size of the transient, and nothing in normal operation comes near it.
Treat this as a mechanical argument and not a statistical one, because Google’s data does not support the statistical version below three years of age. The transitions are the mechanically expensive events, power-on hours count the opposite of transitions, and neither number is workload.
The workload-rate formula, run on a drive you are looking at
Seagate publishes the formula, which means a buyer can compute a used drive’s lifetime workload from the drive’s own counters and compare it against the rating. From the Exos X24 product manual: Workload Rate equals TB transferred multiplied by 8760 divided by recorded power-on hours. Seagate’s knowledge base states it as annualised workload rate equals lifetime writes plus lifetime reads, times 8760 over lifetime power-on hours, with published limits of less than 550 TB a year for enterprise nearline and less than 180 TB a year for the nearline-lite class.
Here is the calculation end to end. The counter values are illustrative, chosen
to show the shape - they are mine, not a measurement of any particular drive.
Substitute the three numbers from your own -l devstat output.
from smartctl -l devstat, page 0x01:
Logical Sectors Written 41,207,481,958
Logical Sectors Read 138,904,220,113
Power-on Hours 39,820
written 41,207,481,958 x 512 = 21.10e12 B = 21.10 TB
read 138,904,220,113 x 512 = 71.12e12 B = 71.12 TB
transferred 92.22 TB
annualised 92.22 x (8760 / 39,820) = 20.3 TB/year
rating, Exos X24 product manual 550 TB/year
fraction of rating 20.3 / 550 = 3.7 per cent
Four and a half years of continuous operation, and the drive used under four per cent of its workload envelope. That is an ordinary nearline result, and it is the answer to “but it has 40,000 hours on it”. The hours were the cheap part.
Now run it the other way, on a drive from a workload that was not ordinary:
Logical Sectors Written 810,000,000,000
Logical Sectors Read 1,940,000,000,000
Power-on Hours 18,500
transferred (810e9 + 1.94e12) x 512 = 1,408 TB
annualised 1,408 x (8760 / 18,500) = 667 TB/year
fraction of the 550 TB/year rating = 121 per cent
Half the hours, and the drive has been run outside its rated envelope for its
entire life. Workload rate separates these two drives; power-on hours actively
misleads about them. Ask for -l devstat and do the arithmetic. It takes a
minute and it is the single highest-value calculation in this guide.
Note the caveat the same product manual attaches to its reliability numbers, because it applies to everything in this section: the Exos X24 is specified at a 0.35 per cent annualised failure rate, equivalent to 2,500,000 hours MTBF, over a five-year service life at 8760 power-on hours a year - and Seagate writes plainly that “the AFR (MTBF) is a population statistic not relevant to individual units”. It tells you nothing about the drive in front of you. What tells you something about the drive in front of you is the drive’s own log.
Load/unload cycles, and the emergency unload
Every idle head-park and unpark increments attribute 193, and drives are rated
for a finite number - the Exos X24 manual specifies 600,000 load-unload
cycles. SAS drives print their own rating in the report, as
Specified load-unload count over device lifetime, beside the accumulated
figure.
On Western Digital and HGST drives the raw value is not one number. smartmontools decodes it as two 24-bit counters, and the difference between them is the interesting quantity:
smartctl -A -v 193,loadunload /dev/sdX
The manual’s own description: the first value is the number of load cycles, the second the number of unload cycles, and “the difference between these two values is the number of times that the drive was unexpectedly powered off (also called an emergency unload)”. It then supplies the exchange rate: “as a rule of thumb, the mechanical stress created by one emergency unload is equivalent to that created by one hundred normal unloads”.
load cycles 412,904
unload cycles 412,573
emergency unloads 412,904 - 412,573 = 331
stress-equivalent normal unloads
normal 412,573
emergency 331 x 100 = 33,100
total equivalent 445,673 of 600,000 rated
331 unexpected power losses cost this drive more mechanical life than 33,000 ordinary parks would have. That is a statement about the previous owner’s power arrangements, and it is the number to look for on any drive coming out of a home, a workshop or an unspecified “office pull”. A drive from a datacentre with proper power will show the two counters nearly equal.
Why the load cycle count is high in the first place
Enterprise idle timers are aggressive enough to matter, which surprises people who assume a nearline drive sits spinning. The Exos X24 manual gives PowerChoice manufacturer defaults of Idle_a at 100 ms, Idle_b at 2 minutes, Idle_c at 4 minutes and Standby_z at 15 minutes. The firmware refuses anything shorter than two minutes for the deeper states: “a minimum timer value threshold of two minutes ensures the appropriate amount of background drive maintenance activities occur”, and setting a shorter one aborts the command.
A drive left with defaults in a lightly-used desktop can accumulate tens of thousands of load cycles in a year while accruing almost no hours and no workload. That is the case where attribute 193 is the only counter telling you anything, and it is why the NAS guide treats idle behaviour as a configuration decision rather than a drive property.
Counterfeit capacity: the drive tells the OS how big it is
A storage device reports its own size. SATA returns a 48-bit LBA count from
IDENTIFY DEVICE, NVMe returns NSZE from Identify Namespace, SAS answers READ
CAPACITY(16) - and in every case the number comes from firmware and the
operating system believes it. A controller claiming 8,000,000,000 sectors while
sitting on 64 GB of flash appears in Windows, in lsblk, in Disk Utility and in
every partitioning tool as a 4 TB drive. There is no protocol-level verification
step anywhere in the stack, because none was ever thought necessary.
Two ways to build one
A teardown published in December 2025 makes the distinction concrete. The two constructions are:
- Relabelled genuine media. A real 64 GB or 128 GB device with a new label and a firmware string edited to claim more. Everything about it works correctly, including SMART if it supports SMART, up to the real capacity.
- A fraudulent controller. A controller that “tricks the operating system into believing a false drive size, simulates write operations, but only actually stores a fraction of the data”. One examined external “SSD” contained a manipulated controller and a small SD card.
The second is the one that produces the characteristic failure, because it acknowledges writes it has not performed. Past the real capacity the firmware does one of two things. Wrap-around maps high addresses back onto low ones, so new writes destroy the oldest data and the filesystem eats itself. Discard accepts the writes, acknowledges them, and returns zeros or garbage on read - the crueller one, because the copy completes, the progress bar finishes, and the files are hollow.
f3probe classifies exactly this, and its vocabulary is worth borrowing:
good, damaged, limbo, wraparound and chain. Limbo - the drive
accepts writes and returns garbage - is the most common counterfeit type.
Formatting proves nothing
mkfs writes a superblock, allocation metadata and a few backup superblocks at
computed offsets. That is well under one per cent of a 4 TB volume, essentially
all of it inside the first few gigabytes a fake controller keeps real. The volume
mounts, the free-space figure is correct, and the drive is a device timed to fail
at whatever point you exceed its real capacity.
The arithmetic explains the business model:
claimed capacity 4,000,000,000,000 B
real flash 64,000,000,000 B
genuine fraction 64e9 / 4e12 = 1.6 per cent
a 400 GB photo library written to it:
actually stored 64 GB
lost 336 GB = 84 per cent
The counterfeiter bought one flash package and sold four terabytes.
Note where this fraud does and does not live. Faking a 20 TB platter stack is not a firmware edit - the mass, the spin-up current and the seek acoustics are all wrong, and there is no cheap component to hide inside the case. Fake-capacity fraud is overwhelmingly a flash phenomenon: USB sticks, memory cards, and external “SSDs” at capacities that do not exist in the retail channel. On SSDs treat it as the default hypothesis for an extraordinary price; on large hard drives an extraordinary price usually means a genuine decommissioned enterprise pull sold at volume.
Proving the capacity, and the tool that will refuse the job
The only proof is writing the whole device and reading back what you wrote. There is no shortcut that is also a proof, though there is a shortcut that is a good screening test.
badblocks has a 32-bit ceiling, and the usual advice runs into it
badblocks -w writes and verifies at block-device level, and it is the right
tool for a bare disk. Two of its defaults cost you days, and one of its limits
will simply refuse your drive.
The defaults, from the manual: block size is 1024 bytes, the number of blocks
tested at a time is 64, and write mode uses four patterns - 0xaa, 0x55,
0xff and 0x00 - each written to every block and read back. Four patterns is eight
full passes over the device. -t random reduces that to one pattern, two passes.
The limit is that badblocks stores the end block in a 32-bit value. The error
string compiled into the binary is invalid end block (%llu): must be 32-bit value, and it caps the addressable device at 2^32 blocks:
max blocks = 2^32 = 4,294,967,296
-b 1024 4,294,967,296 x 1,024 = 4.398e12 B = 4.40 TB ( 4 TiB)
-b 4096 4,294,967,296 x 4,096 = 17.59e12 B = 17.59 TB (16 TiB)
-b 8192 4,294,967,296 x 8,192 = 35.18e12 B = 35.18 TB (32 TiB)
-b 16384 4,294,967,296 x 16,384 = 70.37e12 B = 70.37 TB (64 TiB)
The commonly-repeated -b 4096 tops out at 17.59 TB. It is fine for a 16 TB
drive and refuses a 20 TB one, which is precisely the size where people are
buying used nearline capacity. A 20 TB or 24 TB drive needs -b 8192. Verified
against e2fsprogs 1.47.2.
Raise -c as well. The default of 64 blocks at a time means a -b 4096 run does
256 KiB of I/O per operation, which is why these runs are slower than the drive
is. A much larger value shortens a multi-day run materially at the cost of RAM
for the buffers:
# 20 TB drive, single random pattern, destructive, log the failures
sudo badblocks -wsv -b 8192 -c 4096 -t random -o /tmp/bb.log /dev/sdX
Non-destructive mode, for a drive that arrived with data on it
The manual’s warning is unambiguous: “never use the -w option on a device
containing an existing file system. This option erases data!” If the drive
arrived with the previous owner’s data on it and you have not yet looked at what
is there, use -n instead, which exercises the write path while restoring the
original contents block by block. It is slower, and it preserves the data.
sudo badblocks -nsv -b 8192 -c 4096 /dev/sdX
Also note -e max_bad_blocks, which stops the run early after a set number of
failures. That is useful for screening and useless for evidence, since the manual
notes it yields “a possibly incomplete list of bad blocks”.
dd is not the same coverage
A read pass with dd and a read pass with badblocks are frequently presented
as interchangeable. They are not, and they differ in exactly the case the test
exists to detect. dd aborts at the first read error unless given
conv=noerror,sync; badblocks records the failing block and continues to the end
of the device. Use dd to measure throughput and confirm the drive is readable
end to end; use badblocks when you want the list.
sudo dd if=/dev/sdX of=/dev/null bs=1M status=progress conv=noerror,sync
The keystream method, with nothing installed
If the machine has OpenSSL and nothing else, this proves capacity in two commands:
openssl enc -aes-256-ctr -pass pass:probe -nosalt </dev/zero 2>/dev/null \
| dd of=/dev/sdX bs=1M status=progress
openssl enc -aes-256-ctr -pass pass:probe -nosalt </dev/zero 2>/dev/null \
| cmp - /dev/sdX
The same passphrase regenerates the same infinite pseudorandom stream, so the
second command checks every byte against what should be there. cmp ending with
EOF on /dev/sdX is the pass condition - the device ran out before the stream
did. Any byte offset reported before that is where the real media ended, and
it is the cleanest evidence you can put in a dispute, because it names a number.
This also defeats the sophisticated version of the fraud. A controller that returns zeros for unwritten regions passes a test that writes zeros. It cannot pass a test that writes an unpredictable keystream, because it would have to store the keystream to reproduce it.
f3, h2testw and ValiDrive, and what each is actually for
| Tool | Works on | Destructive | What it is for |
|---|---|---|---|
f3probe |
Linux only | Yes, with --destructive |
Binary-searches the address space; classifies the device in minutes |
f3write / f3read |
Linux, Windows, macOS | No, fills free space | Filling a mounted filesystem with verifiable 1 GB files |
f3fix |
Linux only | Partitions | Writes a partition covering only the real capacity |
h2testw |
Windows | No, fills free space | The long-standing Windows equivalent of f3write/f3read |
| ValiDrive | Windows | No | Random spot-check across the declared capacity, in minutes |
badblocks |
Linux | -w yes, -n no |
Block-level enumeration of every bad block |
Three caveats the usual advice omits.
f3 is a flash-card tester, framed that way by its own authors: “f3 is a
simple tool that tests flash cards capacity and performance to see if they live
up to claimed specifications”. Its most useful components are Linux-only -
f3write and f3read run on Windows and macOS, but f3probe, f3fix and
f3brew require Linux. It is the right tool for a suspect memory card or USB
stick and largely beside the point for a 20 TB mechanical drive.
A published f3probe example shows the output shape: a device announcing
15.33 GB, 7.86 GB actually usable, detected in 1 minute 13 seconds, after
which f3fix --last-sec=16477878 /dev/sdb writes a partition table covering only
the real space, turning the fake into an honest small drive.
h2testw is a 2008 program. Version 1.4 is the current version, written by Harald Bögeholz for c’t. It is still useful and still recommended, and it is documented as producing false positives from faulty RAM, cables or USB adapters, and as being able to miscalculate on very large drives, potentially crashing with error messages. Use it on the sticks and cards it was designed for.
ValiDrive trades completeness for time. Steve Gibson’s tool, last updated 6 October 2023, “performs a quick, random-sequence spot-check across the drive’s entire declared storage space. At every location it verifies the successful storage and retrieval of random (unspoofable) test data”. It reports access-time statistics as well, which incidentally exposes an SD card hiding inside a case claiming to be an SSD. It is a screening test, not a proof: a device that fails ValiDrive is fake, and a device that passes has passed a sample.
How long it takes, at real speeds
Budget real time, because the return window is the binding constraint and the arithmetic is not close.
Sequential rate on a hard drive falls with radius, because the outer tracks hold more sectors per revolution than the inner ones at constant angular velocity. Seagate’s Exos X24 product manual gives the 24 TB model a maximum sustained transfer rate of 272 MiB/s at the outer diameter, quoted as 285 MB/s max, and the 16 TB model 247 MiB/s, quoted as 260 MB/s. Inner-diameter rates are not published. The assumption below - a linear fall to half the outer rate - is mine, not Seagate’s, and it is the standard rule of thumb rather than a measurement:
Exos X24 24 TB, outer diameter, product manual 285 MB/s
inner diameter, MY assumption of half 142 MB/s
average (285 + 142) / 2 = 214 MB/s
one full pass 24e12 / 214e6 = 112,150 s = 31.2 h
write pass + read pass = 62.3 h = 2.6 days
badblocks -w with default four patterns, x8 = 249 h = 10.4 days
The last line is why -t random matters. Four patterns is not four times the
work; it is eight passes, because each pattern is written and then read back.
| Device | Average sustained | One full pass |
|---|---|---|
| 4 TB HDD | about 150 MB/s | 4e12 / 1.5e8 = 26,700 s = 7.4 h |
| 14 TB HDD | about 190 MB/s | 14e12 / 1.9e8 = 73,700 s = 20.5 h |
| 20 TB HDD | about 200 MB/s | 20e12 / 2.0e8 = 100,000 s = 27.8 h |
| 24 TB HDD | 214 MB/s, derived above | 24e12 / 2.14e8 = 112,150 s = 31.2 h |
| 2 TB SATA SSD | about 480 MB/s | 2e12 / 4.8e8 = 4,200 s = 1.2 h |
| 4 TB QLC SSD, post-cache | about 80 MB/s | 4e12 / 8e7 = 50,000 s = 13.9 h |
The last row is not a mistake. Once an SSD’s SLC write cache is exhausted, a sustained write falls to the native program speed of the flash, and on QLC that can be slower than a hard drive. SSD endurance covers the mechanism.
Add the self-test. A long self-test on a large nearline drive is itself
comparable to a read pass; smartctl -c prints the drive’s own estimate before
you start, and that estimate is generally honest.
What the previous owner left behind
A used drive is not a blank drive, and several of the things left on it are not data.
HPA and DCO: hidden capacity in the other direction
Two mechanisms let a host permanently shrink the capacity a drive reports, and they are different mechanisms with different visibility. Both are common on ex-corporate hardware, where they were used to hide recovery partitions or to normalise drive sizes across a fleet.
A host protected area is visible as a mismatch between the current maximum address and the native one:
sudo hdparm -N /dev/sdX
The output reads max sectors = X/Y, current visible over native true capacity.
If the two differ, the difference is the HPA. At protocol level this is READ
NATIVE MAX ADDRESS compared against IDENTIFY DEVICE.
A device configuration overlay hides itself from that comparison and has to be interrogated separately:
sudo hdparm --dco-identify /dev/sdX
which is DEVICE CONFIGURATION IDENTIFY against IDENTIFY DEVICE. A DCO can also
disable features, not just capacity, which is one explanation for a drive that
inexplicably lacks something its datasheet promises. hdparm --dco-freeze locks
the configuration until the next power-on reset, which is a forensic control
rather than a buyer’s tool, but it is worth knowing the lock exists in case a
seller applied it.
Run hdparm -N before you run the capacity test, because an HPA makes a
genuine drive look short and would otherwise read as fraud.
SCT Error Recovery Control
This is the setting that decides whether a used enterprise drive hangs your desktop, or a used desktop drive gets dropped from your array. When a drive hits a sector it cannot read, it retries. A desktop drive retries for a long time - tens of seconds - because in a single-drive machine there is no other copy and the best outcome is eventual success. An array drive gives up quickly, because the controller has parity and would rather rebuild the sector than wait.
sudo smartctl -l scterc /dev/sdX # read the current values
sudo smartctl -l scterc,70,70 /dev/sdX # set 7.0 s read, 7.0 s write
Values are in deciseconds, so 70 is seven seconds, and NAS drives typically default to that. Settings generally do not persist across a power cycle, which means this belongs in a startup script and not in your memory. Many desktop drives do not implement the feature at all, in which case smartctl says so and the drive is not suitable for an array without host-side timeout tuning. Western Digital calls the feature TLER; Samsung and Hitachi call it CCTL. The NAS guide treats this at length because it is the single most consequential difference between drive classes.
OEM firmware, sector size, and the other surprises
A used enterprise drive can arrive formatted to 520 or 528-byte sectors from a storage array that used the extra bytes for integrity metadata, running vendor firmware that a manufacturer tool will not update, or locked to a specific controller. Those are dealt with in enterprise drives, along with the reformatting that fixes the sector size and the time it takes. If you are buying SAS pulls, read that first.
NVMe publishes its health log; SATA SSDs still do not
NVMe is better served than ATA, because its health log is defined by the specification rather than by the vendor.
sudo nvme smart-log /dev/nvme0
sudo nvme id-ctrl /dev/nvme0 # model, serial, firmware, namespace sizes
The fields that matter, with the specification’s own meanings:
| Field | What it is |
|---|---|
data_units_written, data_units_read |
512-byte units “reported in thousands and rounded up” |
percentage_used |
A vendor estimate of consumed endurance, not a measurement |
avail_spare |
“A normalized percentage (0% to 100%) of the remaining spare capacity available” |
spare_thresh |
The value below which the drive raises an asynchronous event |
media_errors |
“The number of occurrences where the controller detected an unrecovered data integrity error” |
unsafe_shutdowns |
Incremented “when a Shutdown Notification is not received prior to loss of power” |
ctrl_busy_time |
Minutes the controller spent processing I/O |
warning_temp_time, critical_comp_time |
Minutes spent above each temperature threshold |
temperature |
In Kelvin |
The data-unit conversion is exact and worth doing in your head:
1 data unit = 1,000 x 512 bytes = 512,000 B
1 TB = 1e12 / 512,000 = 1,953,125 data units
data_units_written 74,218,750 / 1,953,125 = 38.0 TB written
percentage_used is an estimate, and four things about it are counter-intuitive
The specification is unusually candid. The field “contains a vendor specific estimate of the percentage of NVM subsystem life used based on the actual usage and the manufacturer’s prediction of NVM life”. Four consequences:
- It is a prediction, not a measurement. Two drives with identical host writes can report different values because the manufacturers modelled endurance differently.
- It updates once per power-on hour. A drive you have just written a terabyte to will not reflect it yet.
- 100 does not mean failure. The specification says a value of 100 indicates the estimated endurance has been consumed “but may not indicate an NVM subsystem failure”.
- It can exceed 100, up to 255.
avail_spare against spare_thresh is the harder number and the one to weight.
Spare blocks are physical; when they are gone, the drive’s options narrow to
read-only or worse. media_errors non-zero on a consumer SSD is a decline.
The NVMe self-test nobody runs
NVMe has a real self-test, equivalent in intent to the ATA long test, and almost nobody runs it on a used drive.
sudo nvme device-self-test /dev/nvme0 --self-test-code=0x2 --wait
sudo nvme self-test-log /dev/nvme0
Code 0x1 is the short test, 0x2 the extended one, 0x0 shows the current state and
0xf aborts a running test. smartctl 7.4 and later exposes the same thing as
smartctl -t long /dev/nvme0, documented as an experimental feature that runs
the extended self-test for the current namespace.
Run it before the write test. It is the cheapest question on the list.
SATA SSDs are the minefield
A SATA SSD reports wear through vendor-chosen attributes at vendor-chosen IDs
with vendor-chosen units, and there is no equivalent of the NVMe health log. The
same rules as the mechanical section apply, doubled: check smartctl -i for
whether smartmontools recognised the model, and treat an unrecognised model’s raw
values as uninterpreted bytes. SSD endurance works through
the specific attributes and the arithmetic for backing out write amplification.
SAS answers a different set of questions
SAS drives do not have numbered attributes at all. They have log pages, and the report has a completely different shape.
sudo smartctl -a -d scsi /dev/sgN
sudo smartctl -l background /dev/sgN # background media scan results
sudo smartctl -l defects /dev/sgN # LBAs the background scan could not read
sudo smartctl -l sasphy /dev/sgN # per-phy link error counters
sudo smartctl -l genstats /dev/sgN # general statistics and performance
sudo smartctl -l envrep /dev/sgN # environmental reporting
Two lines carry most of the weight. Elements in grown defect list is the
SAS analogue of attribute 5 - the primary defect list is the factory’s, the grown
list is everything since. Non-zero and stable is tolerable; growing is not. And
the read, write and verify error counter logs each end in a Total uncorrected
errors column, which is the analogue of attribute 187 and should be zero.
The background media scan is the part with no ATA equivalent worth the name. A SAS drive scans itself continuously during idle time and records what it could not read, which means the defects it reports surface before any application has touched the sector. On a used drive that is the most valuable log page available, because it represents the drive’s own unprompted opinion of its own surface, accumulated over the whole of its previous life.
FARM works on SAS Seagates too, at log page 0x3d subpage 0x3.
USB bridges decide whether SMART reaches you at all
Most people testing a used drive put it in an enclosure or a dock, and most
enclosures do not pass SMART commands through by default. The result is a
smartctl -a that returns nothing useful, which readers routinely mistake for a
drive problem.
Start here, because it costs nothing and stops before any test is run:
sudo smartctl -d test /dev/sdX
That prints the device type smartctl guessed, opens the device to confirm it, prints the possibly-corrected type, and then exits without performing any further commands. The autodetection is the point: the type it prints second is the one it would have used.
If it guesses wrong, the odds strongly favour one answer. Counted over the 230
USB bridge entries in the smartmontools 7.5 drive database - revision 5706, the
file shipped with the 7.5 release - 211 carry a device type and 19 are marked
unsupported. Of the 211, 128 are plain -d sat, with the JMicron family a
distant second:
| Device type | Entries in the 7.5 database |
|---|---|
sat |
128 |
usbsunplus |
19 |
usbjmicron |
17 |
usbjmicron,x |
10 |
usbcypress |
10 |
sntasmedia |
9 |
sat,12 |
7 |
sntrealtek |
4 |
Count it again before quoting it. The 7.5 database branch keeps receiving additions between releases, so the figures above are the shipped file rather than a permanent fact:
grep -c '^ { "USB: ' /var/lib/smartmontools/drivedb/drivedb.h
sudo smartctl -d sat -a /dev/sdX
sudo smartctl -d usbjmicron -a /dev/sdX
The NVMe-over-USB types are newer and less known: -d sntjmicron,
-d sntrealtek and -d sntasmedia. Version 7.5 added a /sat suffix to each -
written -d sntrealtek/sat - which falls back to SAT when the NVMe Identify
command fails. That is the right choice for an M.2 adapter that accepts either
kind of stick, because one setting then works for both.
Guessing is not free. The manual carries an explicit caution: “specifying ‘,x’
for a device which does not support it results in I/O errors and may disconnect
the drive”, and the same applies to a port number that does not exist. Find the
bridge’s USB vendor and product ID with lsusb and look it up in the database
rather than cycling through types on a drive you are mid-test on.
One more limitation to expect: 48-bit ATA commands, which -l xerror and several
other logs need, do not work through all bridges and are disabled by default on
JMicron. A drive that reports fewer logs over USB than over SATA is normal,
and the fix is a direct connection rather than a different flag.
Refurbished, recertified, pulled, new old stock
These words carry very different amounts of meaning, and only one of them has a published definition from the party using it.
Recertified, when Seagate says it, means a specific three-step process the company publishes: sanitize, which is to “overwrite data when applicable, following the Data Overwriting Process for Returned Products”; test, to “confirm the product meets Seagate factory performance standards”; and certify, to “apply factory certification (or recertification, as applicable) labeling”. The drive carries a six-month warranty. And the sentence buyers miss: “product documentation including specifications and warranties for Seagate products that include prime drives do not apply to recertified products”. The datasheet you looked up does not describe the drive you are buying.
Recertified, when anyone else says it, means nothing in particular. There is no United States regulation defining “refurbished”, “recertified” or “renewed” for storage devices. The FTC’s only formal guide of this kind - 16 CFR Part 20 - covers the rebuilt, reconditioned and other used automobile parts industry. What actually applies to a drive listing is the general truth-in-advertising prohibition on selling used goods as new, which is the rule the 2025 Seagate affair broke, and which is enforced after the fact rather than before it.
Pulled means removed from working equipment, which may or may not have been working when it was removed. New old stock means unsold retail inventory, which is a real category and a genuinely good buy, and also the exact claim a seller makes about a relabelled drive. Neither term is verified by anyone.
The model number is not necessarily the model number
Recertified drives are frequently rebranded, and smartmontools’ database records
examples. The Seagate Exos X20/X22 family entry includes the match pattern
OOS20000G annotated “20TB refurbished and rebranded” - a model string that
never shipped as a prime product. The Exos X24 entry carries a stranger note
still, recording a ST24000NM000C variant with the annotation “30TB HAMR
recertified to 24TB CMR ?”, question mark included, which is smartmontools’ own
statement that it does not know either.
So: a model number on a used drive is a claim about what the drive is
configured to be, not a fact about what it was manufactured as. The checks that
survive this are the FARM log’s Assembly Date and Drive Recording Type, the
manufacturer’s warranty lookup against the serial, and smartctl -i showing
whether the database recognised the model at all.
Note also what this does to recording technology. smartmontools’ database labels
nine families as SMR - Seagate Archive HDD, Seagate BarraCuda 3.5, Toshiba L200,
Toshiba P300, Toshiba S300, Western Digital Black, Western Digital Blue, Western
Digital Blue Mobile and Western Digital Red - and five as CMR, and smartctl -i
prints the label on the Model Family line. That is a real model-to-technology
mapping for a meaningful set of consumer drives, and it is the best source there
is short of the FARM log. It is also silent about everything not in it, which is
most of the market. CMR vs SMR explains why this site records
“Not stated” rather than guessing, and why the CMR-stated rows in
the hard drive listings are a shorter list than you want.
White-label drives from external enclosures
Some of the drives inside Western Digital Easystore and My Book enclosures are catalogued by smartmontools as Ultrastar-class, under the family “Western Digital Ultrastar (He10/12)”, which matches model strings including WD80EMAZ, WD80EZAZ, WD100EMAZ, WD100EZAZ, WD120EMAZ, WD120EDAZ, WD140EDFZ, WD140EDGZ and WD40EDAZ. Others are not: the same database files the WD80EZZX sold in a My Book, and the WD120EMFZ sold in an Easystore, under “Western Digital Red (CMR)” instead. The enclosure USB IDs are recorded against the individual drives - Easystore at 0x1058:0x25fb, My Book at 0x1058:0x25ee - which means the lookup is per model string, not per enclosure name. An Easystore is not a guarantee of an Ultrastar inside it.
That still turns “what is inside” from a rumour into a lookup, provided you look up the model and not the box. The Ultrastar DC HC560 and HC570 entries then add vendor attributes the database names itself: 82 Head_Health_Score and 90 NAND_Master on both, and 71 Milli_Micro_Actuator on the HC570 only. Attribute 22, Helium_Level, is not one of them - it comes from the database’s DEFAULT block, which applies it to every hard drive, and the line is commented out in both Ultrastar entries. None of those names are documented by the vendor; smartmontools arrived at them by observation, which is worth knowing before you quote one as a specification.
Seagate’s equivalents are in the same position. The Exos family entries define attribute 18 as Head_Health and attribute 200 as Pressure_Limit - and attribute 200 on a non-helium drive is Multi_Zone_Error_Rate, the same ID meaning something entirely unrelated. That is the guide’s central thesis in a single attribute ID.
Condition codes, and the one that defeats a claim
eBay’s condition is a structured numeric field, not a phrase in the title. This site renders the code rather than the seller’s free text, which is why a listing titled “working pull” can still display as for parts.
| Code | Condition | eBay’s own definition, or what it means |
|---|---|---|
| 1000 | New | Sealed and unused |
| 1500 | New (other) | Opened, unused. Common for shucked drives |
| 2000 | Certified refurbished | “Direct from the brand or authorized seller”. Two-year warranty |
| 2010 | Excellent refurbished | Vetted third-party seller. One-year warranty |
| 2020 | Very good refurbished | Vetted third-party seller. One-year warranty |
| 2030 | Good refurbished | Vetted third-party seller. One-year warranty |
| 2500 | Seller refurbished | “An item that has been restored to working order by the eBay seller or a third party” |
| 3000 | Used | “The item may have some signs of cosmetic wear, but is fully operational and functions as intended” |
| 7000 | For parts or not working | “An item that does not function as intended or is not fully operational” |
The four eBay Refurbished grades are gated: sellers must apply and be vetted before they can use any of them, and all four carry a 30-day free return subject to exceptions. That vetting is the actual difference between 2010 and 2500, more than the grading words are.
7000 is a disclosure, and it defeats a not-as-described claim. eBay’s
definition of not-as-described turns on disclosure: a buyer is entitled to a
refund where “the item arrives broken, damaged, or faulty (and was not clearly
described as such)”. A drive listed for parts that arrives not working has
arrived exactly as described. The buyers who want that field are recovery shops
harvesting matched PCBs and head stacks, and this site hides those listings by
default for a measured reason - left in the rankings, they held almost every
cheap position on every page. ?parts=show brings them back if that is what you
are shopping for.
The code that catches people is 3000, because “fully operational and functions as intended” is eBay’s definition and a drive satisfies it by spinning up and enumerating. It asserts nothing about the surface, about SMART, about workload or about any test beyond appearing in Device Manager.
A price far under the floor describes the fraud
Sorting by price per terabyte finds bargains and frauds in the same place, because both surface at the top of the same list.
There is a real floor under solid-state storage and it is the cost of the flash. A price implying that a finished drive - controller, PCB, enclosure, shipping - costs less than the NAND inside it is not a clearance price. It is a statement that the NAND is not there.
This site enforces that as a measured rule rather than an assumed one, and the measurement is publishable. Across every single SSD of 2 TB or more in the catalogue: 175 rows above 60 USD per terabyte, 21 between 20 and 60, nine between 8 and 20, and five below 8. All five of those were junk - four counterfeit 8 TB drives and a carrying case. There is a real gap in the distribution and the floor sits in it, at 8 USD per terabyte, deliberately below every genuine row in the catalogue - the cheapest of those sits between 8 and 20 USD per terabyte - so that a real bargain is never the thing it catches.
Two scoping decisions matter. The rule applies only to single drives, because a lot of twenty 256 GB SSDs can legitimately reach a very low price per terabyte and the arithmetic would then be comparing different things. And it applies only above 2 TB, which was lowered from 8 TB after the first threshold proved useless: setting it above the largest capacity seen counterfeited let the counterfeits sitting exactly at 8 TB through. The methodology removes these rows rather than flagging them, because a ranked table that puts a fraud first is worse than one that omits it.
Mechanical drives differ, for the reason given above: faking a 20 TB platter stack is not a firmware change. A very low price per terabyte on a large hard drive usually means a genuine decommissioned enterprise pull sold at volume, which is the outlier that is real. The rule that generalises: compare a row against the median of its own filtered set, not against your sense of the market. Filter to the capacity and class you are buying - datacentre-class listings, say, or 16 TB drives - and the distribution you are comparing against is the right one.
A bid is not a price, and this site has two filters for it
An auction’s current bid is a lower bound with time left on it, and this site computes price per terabyte from what the listing said when it was last read. Three days out, that figure is fiction in your favour, and it ranks accordingly: measured on the live catalogue, 18 of the 20 cheapest rows in the first full crawl were auctions, and the top one was a penny bid on a 3 TB drive.
This site now starts with those rows out. A listing whose only price is a bid - 1,507 rows of 93,642, and they held the entire front page - is excluded before the ranking is computed, because a bid is not a price: it rises until the auction closes and nobody can pay it today. You arrive on prices somebody can pay, without setting anything.
Two toggles move that line, and they are not the same:
?bids=showputs bid-only listings back, for somebody watching what is about to close rather than shopping. The table marks those rows “at current bid” in the very column being sorted on, which is the deliberate alternative to hiding them.?buy=nowgoes further and drops every auction, including the 348 that also carry a Buy It Now price and therefore do state a price somebody can pay today. It is the right filter when you want a clean comparison and nothing else.
Shipping is the other half, and where auction margins go. If the delivered cost cannot be determined this site leaves the figure empty rather than guessing, and price per terabyte explains why delivered cost is the only honest input. And an auction ending far below the market is usually ending badly - the ones that do not converge are the ones nobody watched. Against that, an auction is also where a seller who does not know what they have gets found out, which is a reading exercise rather than a pricing one.
Lot arithmetic, and the yield to budget for
Lots are where the good prices are and where the arithmetic is least forgiving, because a single dud moves the effective price per terabyte more than any negotiation will.
Start with the headline and then degrade it:
lot of 8 drives, 14 TB each, delivered 560 USD
claimed total 8 x 14 = 112 TB
headline 560 / 112 = 5.00 USD/TB
one dead of eight 7 x 14 = 98 TB
effective 560 / 98 = 5.71 USD/TB (+14 per cent)
two dead 6 x 14 = 84 TB
effective 560 / 84 = 6.67 USD/TB (+33 per cent)
three dead 5 x 14 = 70 TB
effective 560 / 70 = 8.00 USD/TB (+60 per cent)
Every failure costs more than the last one, because the denominator shrinks while the numerator does not. That convexity is the whole reason to test a lot early: the difference between finding two duds inside the window and outside it is the difference between 5.00 and 6.67 per terabyte.
What yield to assume is the part nobody can tell you honestly. One vendor guide suggests budgeting 3 to 5 per cent early failures on lots of fifty or more from untested recycler pulls. I cannot verify that figure against a primary source and you should treat it as a planning assumption rather than a measurement. What can be said with confidence is what it is not: enterprise MTBF ratings of 2,500,000 hours correspond to a 0.35 per cent annualised failure rate, and that is a new-drive population statistic which Seagate itself says is “not relevant to individual units”. Do not use it to predict a lot of untested pulls.
Then the constraint that actually bites:
8 x 14 TB, one write pass each at 190 MB/s average
per drive 14e12 / 1.9e8 = 73,700 s = 20.5 h
serially 8 x 20.5 = 164 h = 6.8 days
four at a time, if the HBA and the PSU allow it
164 / 4 = 41 h = 1.7 days
Seven days of serial testing against a thirty-day window. A lot of eight is not something you start in week three, and a lot of twenty needs parallel testing and a power supply that can start them - recall the 2.6 A peak per drive on the 12 V rail at spin-up, which is where cheap backplanes and daisy-chained SATA power adapters fail. Staggered spin-up exists for this: the Exos X24 pinout notes that P11 grounded means the drive does not use deferred spin-up, which is the opposite of what you want in a twenty-drive test rig.
Filter lots by total capacity rather than per-drive capacity when you shop them.
This site accepts total as a parameter for exactly that - lots of about
100 TB and up - because the delivered total is the number a lot buyer actually
cares about.
Shipping damage, and the head-parking advice that expired
Two pieces of folklore need retiring.
There is nothing you need to do to park the heads before shipping. Modern
drives use ramp load and unload: the heads are lifted off the media onto a plastic
ramp at the outer edge whenever the drive is unpowered, and they have been parked
that way in transit across the industry for many years. A command does exist -
ATA STANDBY IMMEDIATE, reachable as hdparm -y - and it buys you nothing,
because removing the power unloads the heads to the same ramp anyway. The benefit
is already built in, and one of its stated purposes is improved robustness to
mechanical shock under non-operating conditions.
The real shipping risk is shock transmitted to the spindle and actuator, not head-on-platter contact. And here the intuition about enterprise drives is backwards. From the Exos X24 product manual: operational shock 40 G at 2 ms, non-operational shock 200 G at 2 ms. Seagate’s desktop-class non-operating specification is cited at around 300 G at 2 ms, with some high-end products at approximately 150 G. A helium-filled, ten-platter nearline drive - the same manual lists the 24 TB model as 20 heads on 10 disks - is a heavier, more tightly-toleranced mechanism with more mass hanging off the spindle, and it is more fragile in a padded envelope than a two-platter desktop drive, not less.
What follows for a listing is practical. Ask how it will be packed, and ask before you buy, because the answer is also a competence signal. What you want is each drive in an antistatic bag, in a moulded clamshell or at minimum two inches of foam on all six faces, not touching another drive, in a box that does not rattle. Bubble wrap in a padded envelope is the standard failure. A drive that arrives in a jiffy bag and reports CRC errors, command timeouts or a conveyance self-test failure is a packaging dispute, and it is the one dispute where photographs of the packaging matter as much as the logs. Photograph the parcel before you open it.
Run smartctl -t conveyance first, for the reason given above: it is the test
designed for this question and it answers in minutes.
Shucking: what is inside, and the 3.3 V pin
External desktop drives are frequently the lowest genuine price per terabyte available, because they are sold as a finished product rather than as a component. What is inside varies, and the listing will not say:
- A standard 3.5-inch SATA drive, usually with a plain white label and a model number that does not appear in the retail channel. This is the case worth doing, and for the Western Digital enclosures the database lookup above tells you what that white-label model string means once you can read it.
- A native-USB drive, the bridge soldered to the drive’s own PCB with no SATA connector at all. Near-universal on 2.5-inch portables. It cannot move to a SATA port, and bridge-implemented encryption can make the data unreadable if the bridge dies, which is an argument against using one as a backup target.
- A drive whose recording technology is stated nowhere. Shucked drives are often the same platform as a retail model, and no vendor publishes a complete mapping from model number to whether a drive is shingled. See CMR vs SMR.
Then the pin. SATA revision 3.3 redefined pin 3 of the 15-pin power connector, one of its three 3.3 V pins, as PWDIS - power disable. Assert 3.3 V on it and a drive implementing the feature stays unpowered. Enterprise-derived drives implement it; an enclosure’s power board never supplied 3.3 V at all, so the drive worked in the box and appears dead on a PC power supply.
The Exos X24 pinout confirms both halves of the standard fix. P1 is listed as “V33 - Not Used (P1 and P2 tied internally)”, P2 as “V33 - Not Used”, and P3 as “PWRDIS - Enter/Exit Power Disable (option)”. All three pins are 3.3 V and none of them is needed, which is why covering all three costs nothing.
- Polyimide tape over the first three pins at the narrow end of the L. Use Kapton, not electrical tape, which creeps when warm and leaves adhesive.
- A Molex-to-SATA power adapter, since 4-pin Molex carries only 5 V and 12 V and physically cannot deliver 3.3 V.
Many modular PSU cables omit the orange 3.3 V wire anyway, in which case there is nothing to fix and nothing to tape.
You can check whether a drive implements the feature before plugging it in,
because it is reported in the drive’s own IDENTIFY data. In the SATA Capabilities
word, bit 30 is “power disable feature supported” and bit 31 is “power disable
feature always enabled”; in SATA Features Enabled, bit 11 is “power disable
feature enabled”. Both smartctl -x and hdparm -I print these. Not every
enterprise drive implements it - the ST24000NM00xH sample published in the Exos
X24 manual reports zero for all three bits - so the tape is a precaution rather
than a certainty.
The cost of shucking is warranty. Cover is on the product you bought, enclosure and drive as a unit, and opening it ends that; the bare drive generally returns no independent entitlement against its serial, which you can confirm with the warranty lookup before you commit. Right trade on a drive going into an array with redundancy, wrong one on a drive holding the only copy - and either way it is a used drive the moment it leaves the box, so it gets the same treatment as everything else.
The clock has two ends
The eBay Money Back Guarantee is usually summarised as thirty days. It has both a longer floor and a hard ceiling, and the ceiling is the one that loses cases.
Opening the request. For an item not as described, you may request a return “within 30 calendar days after the estimated or actual delivery date or within the seller’s stated returns window, whichever is longer”. That second clause is worth reading twice: a seller offering a 60-day return policy has given you a 60-day not-as-described window, not thirty. For an item not received, the window is 30 calendar days from the estimated delivery date.
Escalating it. You may ask eBay to step in no earlier than 3 business days after opening the request and no later than 21 business days after the request date. Miss the ceiling and the case closes on its own, in the seller’s favour, regardless of how good your evidence was. Put a calendar entry on the day you open a case.
What counts. Not-as-described turns on disclosure: “if the buyer receives the wrong item, or the item arrives broken, damaged, or faulty (and was not clearly described as such), they are entitled to return it for a refund”. The seller pays return postage when that is the reason. A “no returns” listing does not remove the right to return something defective or misdescribed; change of mind falls under the seller’s own policy instead.
And the clause that makes testing safe. Buyers must return an item “in the same condition in which it was received”, with an explicit exception where “the use was necessary to determine the quality or functioning of the item, or the damage was the result of that use”. That is the clause that lets you run a destructive full-device write test inside the return window. Erasing a used drive to establish whether it holds its claimed capacity is necessary to determine its functioning, by the policy’s own language.
Now put those against the timing table. A 20 TB drive needs about 28 hours for a read pass, another 28 for a write pass, and a long self-test before either. Qualifying one properly is a four to five day job. Order the test before you need the answer.
If it fails, save the evidence with the serial visible. The diff between two
timestamped smartctl -x captures is the ideal artefact, and a cmp output
naming the byte offset where a claimed 4 TB device stopped being real is about as
unambiguous as evidence gets. Message the seller first, since the earliest
escalation is three business days anyway, then open the case and escalate inside
the ceiling.
You may also find the previous owner’s data on the drive. That is a fact about the seller rather than about the drive, and the next section is what to do with it.
Testing on arrival, in order
Do this the week it lands. The order matters, because each step changes what the next one can tell you, and two of the steps are irreversible.
0. Photograph the packaging before opening it, and the label before connecting it. Both are free and both are evidence.
1. Capture SMART before touching anything. Check the serial against the listing photograph and the seller’s report now, while a mismatch is still a clean case rather than an argument.
sudo smartctl -x /dev/sdX > ~/drive-before.txt
sudo smartctl -l devstat /dev/sdX >> ~/drive-before.txt
sudo smartctl -l farm /dev/sdX >> ~/drive-before.txt # Seagate only
sudo smartctl -i /dev/sdX | grep -E 'Serial|Model|Family|Rotation|Firmware'
2. Run the three cross-checks that cost nothing.
- SMART attribute 9 against FARM Power on Hours, or against Device Statistics page 0x01 Power-on Hours.
- Spindle Motor Power-on Hours against Head Flying Hours. Spindle must be the larger.
sudo hdparm -N /dev/sdXfor a host protected area, before any capacity test makes an HPA look like fraud.
3. Conveyance self-test. Minutes, and it answers the courier question first.
sudo smartctl -t conveyance /dev/sdX
4. Long self-test. smartctl -c gives the estimate before you commit the
hours. Remember it stops at the first error.
sudo smartctl -c /dev/sdX
sudo smartctl -t long /dev/sdX
sudo smartctl -l selftest /dev/sdX # poll until it completes
5. Full read pass from the host, which enumerates every failure rather than stopping at the first.
sudo badblocks -sv -b 8192 -c 4096 -o /tmp/bb-read.log /dev/sdX
6. Full write pass, if the drive is yours to erase. This resolves pending
sectors and it is the only step that proves capacity. Use badblocks -w with a
block size that fits the device, or the keystream method, or f3probe on flash.
Do not run it on data you have not copied off, and use -n instead if you need
the contents preserved.
7. Capture SMART again and diff it.
sudo smartctl -x /dev/sdX > ~/drive-after.txt
sudo smartctl -l devstat /dev/sdX >> ~/drive-after.txt
diff ~/drive-before.txt ~/drive-after.txt
Hours and temperature will have moved. Anything moving in 5, 187, 188, 197 or 198 is your answer, and you now hold two timestamped files saying so, which is better evidence than any description of the problem.
8. If the drive is going into an array, set error recovery control and put the command in a startup script, since it does not survive a power cycle.
Erasing it properly, which the capacity test already did
You bought a used drive, which means it may still hold someone else’s data, and you are going to hand it on eventually yourself. The vocabulary for this is NIST Special Publication 800-88 Revision 1, and it distinguishes three things: Clear “applies logical techniques to sanitize data in all user-addressable storage locations for protection against simple non-invasive data recovery techniques”; Purge “applies physical or logical techniques that render Target Data recovery infeasible using state of the art laboratory techniques”; and Destroy renders the media unusable.
The useful finding, which contradicts a great deal of folklore about multi-pass wiping, is that one pass is enough. NIST states that “for storage devices containing magnetic media, a single overwrite pass with a fixed pattern such as binary zeros typically hinders recovery of data even if state of the art laboratory techniques are applied to attempt to retrieve the data”. For Purge on magnetic media, “a single write pass should suffice”, via the SANITIZE DEVICE OVERWRITE EXT command - and NIST notes that the ATA Sanitize commands “are preferred over the ATA Security feature set SECURITY ERASE UNIT command”.
So the destructive capacity test you ran in step 6 is also a correct Clear. The two jobs are the same write, and you only have to do it once.
Two caveats. NIST notes that overwriting does not reach “areas not currently mapped to active Logical Block Addressing (LBA) addresses (e.g., defect areas and currently unallocated space)” - reallocated sectors keep whatever was in them. And flash is different: an SSD needs a block erase or a cryptographic key scramble rather than an overwrite, because the controller’s indirection means a write to an LBA does not necessarily touch the cells that held the old contents.
What to do with this on a listing page
- Start where the page puts you: bid-only rows are already out, so the
cheapest rows are prices somebody can pay today rather than bids with three
days left. Switch to
?buy=nowwhen you want a clean comparison,?bids=showwhen you are watching an auction close, and?parts=showonly when you genuinely want broken drives. - Filter to the class and capacity you are actually buying before you judge
any price.
?class=DATACENTER,?class=NAS,?cap=for per-drive capacity and?total=for lot totals. Compare a row against the median of its own filtered set, never against your sense of the market. - Message the seller before bidding, and paste the commands.
smartctl -x,-l devstat,-l selftest, and-l farmon any Seagate. A refusal is a price, not a verdict; a substitute like “tested, no bad sectors” is neither. - Do the workload arithmetic on the spot. Logical Sectors Written plus Read, times 512, times 8760 over power-on hours, against 550 TB a year for nearline or 180 for the lighter class. It takes a minute and it separates two drives that power-on hours cannot.
- Run the two tamper checks before you bid. SMART attribute 9 against FARM Power on Hours, and Spindle Motor Power-on Hours against Head Flying Hours. The second needs no reference value and works on any drive with Device Statistics page 0x03.
- Price a used drive as an option, not a purchase. Budget the return shipping, the test time and one failure per lot, and buy at a price where all three are survivable.
- Book the test before the drive arrives. A 20 TB qualification is four to five days; a lot of eight is a week of serial testing against a window that starts the day it is delivered and has a 21-business-day escalation ceiling on top of it.
- Take the capacity test seriously on anything solid state above 2 TB, and treat a price below the floor as a description of the fraud rather than a bargain. On large mechanical drives the same outlier is usually a genuine decommissioned pull, which is the one place the cheap row is the right buy.
A high-hours enterprise drive with a clean surface record, a workload rate at four per cent of its rating, matching SMART and FARM hours, and spindle hours correctly above head flying hours is a better purchase than a low-hours consumer drive with anything at all moving in 5, 197 or 198. That trade is only visible if you read the log, and only safe if you verify the log was not written by hand.