NVMe caching for mechanical drives: how to set it up
By Harry Saarinen ·
More RAM first, then metadata on mirrored flash, and a block cache in front of the disks only third - a step most readers never need to reach. When you do reach it, decide which of three things you need - a read cache, a write cache or a tier - before choosing the software: a read cache that dies costs nothing but warmth, while a write cache holds the only copy of acknowledged writes and needs power-loss protection and a mirror. For a media library or a backup target, a cache does nothing at all.
Three things are called caching here, and only one is safe to lose
A 7200 rpm drive’s average rotational latency is 4.16 ms, and it was 4.16 ms twenty years ago. Seagate quotes that figure on the IronWolf Pro and Exos product manuals today because it is not an engineering choice, it is arithmetic on the spindle speed, and the spindle speed has not moved.
7200 rpm -> one revolution = 60 / 7200 s = 8.33 ms
average rotational wait = half a turn = 4.16 ms
plus an average seek, high single-digit ms on a
3.5 inch drive
-> about 12 ms to serve one random read at QD1
Exos X24 datasheet: 168 random read IOPS at 4K, QD16
-> 1 / 168 = 5.95 ms per completed I/O with 16 in
flight, which beats QD1 on throughput because a queue
that deep lets the drive reorder the seeks. Each of
those 16 I/Os still waits 16 / 168 = 95 ms
Capacity over the same period went from under a terabyte to twenty-plus
terabytes per spindle. The seeks per second did not. A modern hard drive is
therefore a device that holds enormously more data than it can reach
quickly, and every year that ratio gets worse. That gap, and only that
gap, is what a cache buys back. Sequential throughput is not the problem:
an IronWolf Pro 20 TB is rated at up to 285 MB/s at the outer diameter, which is
faster than
most home networks and most backup jobs can consume. Random access is the
problem. A consumer TLC NVMe serves a 4K read in roughly 80 to 120 µs -
that range is an approximation, because vendors publish IOPS rather than
QD1 latency, so measure your own at queue depth 1 with fio rather than
trusting it - against the 12 ms derived above, which is arithmetic on the
spindle and the seek, not a published number. That is two orders of
magnitude, and it is the entire premise.
Read cache, write cache, tier
Three mechanisms get the same word, and conflating them is where most setups go wrong. They differ in exactly one respect, which happens to be the one that matters at three in the morning: where the only copy of your data lives.
| What the fast device holds | Losing it costs | |
|---|---|---|
| Read cache | A second copy of hot blocks | Performance only |
| Write cache | The only copy, until destaged | The un-destaged writes |
| Tiering | The only copy, permanently | That data, until restored |
A read cache is a bet with no downside except the money, the wear and the
RAM. ZFS’s l2arc_meta_percent exists because L2ARC headers sit in ARC and
are not evicted under memory pressure, so an oversized cache device buys
flash by spending the memory that was already serving hits.
ZFS’s L2ARC is the clean example: the OpenZFS manual pages state that a
read error on a cache device is simply reissued to the original pool
device, and zpool remove tank nvme0n1 is a no-drill operation. dm-cache
in writethrough mode and bcache in writethrough or writearound make
the same promise. lvmcache(7) states it plainly: “the loss of a device
associated with the cache in this case would not mean the loss of any
data.”
A write cache is a different risk class entirely. In writeback, the fast
device holds acknowledged writes that exist nowhere else until a background
thread destages them, meaning copies them down to the hard drive and clears
the dirty bit. lvmcache(7) is
equally plain in the other direction: “the loss of a cache device can
result in lost data.” dm-writecache by design acknowledges a write the
moment it is in the fast device. A writeback cache in front of a
redundant array converts that array into a non-redundant one, because the
NVMe is now a single point of failure sitting upstream of all the
redundancy you paid for. Everything later in this guide about power-loss
protection, mirrored cache volumes and drain procedures follows from this
one sentence.
Tiering is not caching at all, and the projects that implement it say so in
their own words. OpenZFS describes the special allocation class as storage
dedicated to specific block types - metadata, the indirect blocks of user
data, optionally small file blocks - and then gives the instruction you
give about a vdev rather than about a cache: “The redundancy of this device
should match the redundancy of the other normal devices in the pool.” The
blocks routed there exist only there. Windows Storage Spaces heat-based
tiering moves sub-file extents between an SSD tier and an HDD tier
overnight; bcachefs background_target moves replicas off the foreground
device rather than copying them. A tier is a placement decision with the
permanence of a top-level vdev. On a pool that contains a raidz or draid vdev a
special vdev can never be removed at all - adding one is a decision you do
not get to revisit.
Find out which workload you have before buying anything
Here is the uncomfortable part. For a media library or a backup target, a
cache does nothing, and adding one makes things worse. A film server
reads sequentially at a rate the disk already sustains; bcache defends
against exactly this with sequential_cutoff, which sends a sequential
stream past the cache once it crosses a threshold, 4 MB by default, and
Microsoft’s own wording is narrower: the Storage Spaces write-back cache
“buffers small random writes to solid-state drives”, which is not the
shape of a film server’s traffic. Storage Spaces can also pin a file to a
chosen tier with Set-FileStorageTier, which is the standard remedy for a
media volume that keeps being promoted onto the flash. A backup target is
written once and read never, so every byte through the cache is wear with
no hit to show for it - and wear is a budget, spent in TBW, that
TBW, DWPD and how worried to be puts numbers to.
Which workloads pay for the flash, which ones only wear it out, and how to tell which you have is the whole of the next section.
The rest of this guide settles the mechanisms in the order they are built: dm-cache and LVM, bcache, dm-writecache and Open CAS, then the full ZFS stack of ARC, L2ARC, special vdev and separate log device, then btrfs, bcachefs and the dead ends; Storage Spaces tiering, the storage bus cache and the third-party options on Windows; the NAS appliances; the RAID controller’s own cache, the clustered and scale-out stacks and the enterprise arrays; then sizing, tuning, measurement, safeguards, what the role costs the NVMe, four worked builds end to end, and what a small business should actually deploy. Device selection - capacity, endurance rating and whether the drive has power-loss protection at all - follows the mode rather than leading it, and the NVMe listings are where to check those figures against price.
The access pattern decides everything
A cache is a bet that your reuse distance is shorter than your cache is big. If the bet is wrong you have not bought performance, you have bought a wear-out mechanism sitting in the write path. The bet is testable before you spend anything, and the test is arithmetic rather than opinion.
Working set, not total capacity
The number that decides the outcome is the count of distinct blocks touched in a window, not the volume of IO passing through that window. A workload reading 2 TB a day across an 80 GB footprint caches almost perfectly. One reading 200 GB a day scattered over a 6 TB footprint does not, and no NVMe device repairs that. Capacity of the array is irrelevant to the decision; footprint is the whole of it.
Two independent properties settle it. The first is footprint size relative to cache size. The second is drift rate, meaning how fast the footprint turns over. Microsoft states the second one plainly in its Storage Spaces Direct cache documentation: “If the active working set exceeds the size of the cache, or if the active working set drifts too quickly, read cache misses increase and writes need to be destaged more aggressively, hurting overall performance.”
The degradation is not linear, and this is the part most sizing advice gets
wrong. For uniformly random re-reference over a footprint F with a cache of
size C, hit ratio tracks roughly C/F, which degrades gracefully. For a
footprint traversed in a repeating cycle under an LRU-ish policy, every block
is evicted exactly before it is next needed, and hit ratio goes to zero the
moment F exceeds C. Real workloads sit between those two, which is why a
cache that was fine at 400 GB of footprint can collapse at 550 GB rather than
losing a proportional slice. Size for the footprint you measure plus headroom
for drift, not for the number that looks affordable.
Hit ratio is the only number that matters
Everything else in the stack - chunk size, policy, cache mode, device model -
is instrumentally interesting. It matters exactly as far as it moves h.
L_eff = h * L_fast + (1 - h) * L_slow
S = L_slow / L_eff
L_slow = 12 ms 7200 rpm, QD1 random 4K. DERIVED, not quoted. The 4.16 ms
average rotational latency is spec'd on the Seagate
Exos X24 product manual, and the Exos X24 datasheet
gives 168 random read IOPS at 4K QD16. QD16 lets the drive
reorder, so QD1 throughput is well below that figure; at
roughly half of it, 84 IOPS, one read costs about 12 ms. The
halving is an assumption. Substitute your own measurement.
L_fast = 0.1 ms consumer TLC NVMe, 4K read, QD1. Vendors publish IOPS,
not QD1 latency; this guide will not claim a spec'd one.
h = 0.50 L_eff = 0.050 + 6.000 = 6.050 ms S = 2.0x
h = 0.80 L_eff = 0.080 + 2.400 = 2.480 ms S = 4.8x
h = 0.90 L_eff = 0.090 + 1.200 = 1.290 ms S = 9.3x
h = 0.95 L_eff = 0.095 + 0.600 = 0.695 ms S = 17.3x
h = 0.99 L_eff = 0.099 + 0.120 = 0.219 ms S = 54.8x
h = 0.999 L_eff = 0.100 + 0.012 = 0.112 ms S = 107.2x
Three things fall straight out of that block. The curve is hyperbolic. Moving 50 to 90 per cent buys 4.7 times; moving 90 to 99 per cent buys another 5.9 times on top. Almost all the value lives in the last few per cent, which is precisely where people stop tuning.
Think in misses, not hits. Above about 90 per cent the h * L_fast term
stops mattering and L_eff is approximately (1 - h) * L_slow. Halving the
miss rate halves the latency. “95 per cent hit ratio” is a flattering way of
saying “5 per cent miss rate”, and 5 per cent of 12 ms is 600 µs, still six
times slower than the flash it is running on.
Amdahl’s law sets the floor. If 2 per cent of your IO can never be cached, no device on earth takes the mean below 240 µs.
Two corrections to how the number is usually read. The mean is the wrong
summary statistic, because the miss tail is what a user feels: at h = 0.99 one
request in a hundred still goes to the platter, so everything above the 99th
percentile is hard disk latency however good the mean looks. Report clat
percentiles from fio; PerfMon on Windows gives averages rather than
percentiles, though its Cluster Storage Hybrid Disk counter set does expose
Cache Miss Reads/sec, which Microsoft suggests comparing against your overall
read IOPS. And partial hits count as misses - bcache documents this explicitly,
so hit ratio is sensitive to your request size against the cache block size.
dm-cache keeps read hits, read misses, write hits and write misses as four
separate counters for the same reason; collapsing them hides the case where
reads are fine and writes are thrashing.
A hit ratio quoted without the IO rate beside it is decoration. Ninety-nine per cent of 20 IOPS is not a result.
What pays for the flash, and what does not
| Benefits | Mechanism |
|---|---|
| Filesystem metadata | Tiny, hot, random, on the critical path of everything. find, ls -lR, rsync traversals and backup scans collapse. The highest-leverage target there is. |
| VM images and VDI | A small hot region inside an enormous file. Microsoft’s Windows Server 2012 R2 Storage Tiers documentation: “if only 30 percent of the data on a virtual hard disk is ‘hot’, only that 30 percent moves to your solid-state drives”. |
| OLTP databases | Random 4K to 16K reads over an index working set far smaller than the table. Caveat: the buffer pool already holds the hottest data, so the block cache sees a larger, colder second tier. |
| Compile trees, package caches, CI workspaces | Vast numbers of small files, near-total reuse, a footprint of a few GB. |
| Mail spools, container layers, Git object stores | The same shape. |
| Random write bursts | A write cache derandomises before destaging, converting seeks into streams. |
| Does not benefit | Mechanism |
|---|---|
| Sequential streaming, media serving | An IronWolf Pro 20 TB is spec’d at 285 MB/s maximum sustained transfer on the outer tracks, and less further in. The cache adds nothing and evicts what mattered. |
| Backup targets, archive | Write once, read never. Every byte is wear for zero hits. See SSD endurance: TBW, DWPD, how worried to be. |
Full-table scans, zfs send, whole-tree rsync, virus scans, weekly indexers |
Single-touch streams that wipe the cache. The commonest cause of a cache that “used to work”. |
| Anything the page cache or ARC already serves | Check before buying. |
| Any footprint larger than the cache | You pay lookup cost and promotion writes for a hit ratio trending to zero. |
Characterise your own workload first
Start above the block layer, because a miss the page cache already absorbed is not your problem:
cachestat 1 # page cache hits/misses, from bcc
bitesize # IO size histogram: 4K random or 1 MiB streams?
biolatency -D 1 # per-device latency histogram
iostat -x 1 # r_await, w_await, rareq-sz, wareq-sz, aqu-sz
Ignore %util on any NVMe device. The iostat man page says so itself: “for
devices serving requests in parallel, such as RAID arrays and modern SSDs, this
number does not reflect their performance limits.”
Then count distinct chunks rather than bytes:
blktrace -d /dev/sdb -a queue -w 3600 -D /var/tmp/bt
blkparse -i /var/tmp/bt/sdb -f "%d %S %n\n" \
| awk '$1 ~ /R/ && $3 ~ /^[0-9]+$/ { for (s = $2; s < $2 + $3; s += 8) print int(s/128) }' \
| sort -u -S 2G | wc -l
%d is RWBS, %S the start sector, %n the sector count; 128 sectors is
64 KiB, LVM’s compiled-in default cache chunk size, which is worth confirming
with lvmconfig --type default allocation/cache_pool_chunk_size because LVM may
choose a larger chunk for a large cache pool. Multiply the count by 64 KiB for
the read footprint. The numeric test on the third field keeps blkparse’s
end-of-run summary out of the count. Verify the format string against your own
blkparse
before trusting it, and bound the sort, because sort -u on a busy device will
eat the machine. Repeat at 5 minutes, 1 hour and 24 hours to sketch a
miss-ratio curve. The rigorous version is Mattson stack-distance analysis;
there is no packaged, production-ready tool for it in mainstream distributions,
and this guide will not pretend otherwise.
The better advice is cheaper: build the smallest cache you can, run the real workload for a week, read the counters, then grow it. Synthetic estimation is a planning tool. The deployed cache’s own counters are the measurement. Size the device from that result before shopping the NVMe listings, and pick on write budget rather than headline read IOPS.
Sequential IO is bypassed on purpose
Most implementations refuse to cache streams, because one cp of a
10 GB file would otherwise evict the entire hot set. bcache’s
sequential_cutoff defaults to 4 MiB (set in super.c) and decides on the
larger of a task’s current and average sequential run, tracked over a
128-entry table of recent IOs with a 5-second window, so the rolling per-task
average makes a backup bypass wholesale rather than caching 4 MiB after every
seek. ZFS ships l2arc_noprefetch=1, so buffers that were prefetched but never
used by an application are not written to the cache device. dm-writecache
offers metadata_only, which promotes metadata alone and leaves bulk data on
the slow device. Storage Spaces is quieter about thresholds: Microsoft
documents a default cache page size of 16 KB and an algorithm that
derandomises writes before destaging, but the 256 KB large-IO bypass figure
that circulates in write-ups is documented only for the older tiered Storage
Spaces write-back cache, whose Windows Server 2012 R2 page states that writes
larger than 256 KB are not written to the cache. For the current storage pool
cache, treat it
as unconfirmed.
dm-cache is the exception, and the exception is instructive. The old mq
policy’s sequential_threshold still appears in every tutorial on the web; the
kernel documentation says mq “is now an alias for smq” and lists
sequential_threshold among the tunables that “are accepted, but have no
effect”, dm-cache-policy-mq.c was deleted in Linux 4.6, and no evidence of a
replacement bypass path turned up. The absence is the evidence. If you
run dm-cache over a workload with large sequential passes, assume nothing
protects the hot set beyond smq’s hotspot queue, which the kernel documentation
describes only as deciding which blocks to promote at a coarser granularity than
a cache block, and keep the streams on a separate volume. That
architectural point matters more for a general-purpose box than any tuning
knob, and it is the same reasoning behind keeping bulk media off flash entirely
in what differs from a desktop.
Linux: dm-cache and LVM cache
Every guide that tells you to set sequential_threshold 512 and
random_threshold 4 is tuning a kernel that stopped existing in 2016.
dm-cache-policy-mq.c was deleted in Linux 4.6, mq has been a pure alias
for smq ever since, and the kernel’s cache-policies document lists both of
those names among the tunables that are “accepted, but have no effect”.
lvmcache(7) says the same settings are “silently ignored”. smq is documented
as not having “any cumbersome tuning knobs”, and for once the documentation
is not being modest.
The second uncomfortable fact belongs beside it. The raw dm-cache target
defaults to writeback; LVM defaults to writethrough. DEFAULT_CACHE_MODE "writethrough" sits in LVM2’s lib/config/defaults.h, while the kernel’s
constructor treats writeback as the no-feature-argument case. Same target,
opposite defaults, and the difference is whether a dead NVMe costs you
performance or costs you data.
Origin, cache, metadata
The target takes three sub-devices, in this order: the origin (the hard drive), the cache (the NVMe), and a metadata device holding the block map, the dirty bits and the policy hints. The kernel doc notes that a metadata device “may only be used by a single cache device”; it is separable precisely so a volume manager can mirror it.
cache <metadata dev> <cache dev> <origin dev> <block size>
<#feature args> [<feature arg>]*
<policy> <#policy args> [policy args]*
Feature arguments are exactly writethrough, passthrough, metadata2 and
no_discard_passdown. Ask for smq by name rather than default: the doc
warns that “as the default policy could vary between kernels, if you are
relying on the characteristics of a specific policy, always request it by
name”. metadata2 stores dirty bits in a separate btree and shortens cache
shutdown; LVM selects it through --cachemetadataformat auto|1|2, which is
fixed at creation and cannot be changed later.
Metadata sizing is arithmetic, not guesswork. _cache_min_metadata_size() in
cache_manip.c uses 16 + 20 + 8 bytes per chunk over a 4 MiB transaction
overhead:
min_meta = 4 MiB + 44 bytes per cache chunk
1 TiB cache at 1 MiB chunks = 1,048,576 chunks
1,048,576 x 44 B = 46,137,344 B = 44 MiB
+ 4 MiB = 48 MiB of metadata
The kernel ceiling is DM_SM_METADATA_MAX_BLOCKS = 255 * ((1 << 14) - 64) =
4,161,600 blocks of 4096 bytes = 15.875 GiB, so you cannot overshoot by
accident. The metadata block size is fixed at 4096 bytes, which is why field
four of dmsetup status is always 8.
Chunk size, and the cap that overrides your choice
The kernel permits a cache block “between 64 sectors (32KB) and 2097152
sectors (1GB) and a multiple of 64 sectors (32KB)”. LVM restates this as
--chunksize, takes its 64 KiB default from
allocation/cache_pool_chunk_size, and warns that chunks bigger than 512 KiB
should be used only when necessary.
You probably will not get 64 KiB. allocation/cache_pool_max_chunks
defaults to 1,000,000. LVM divides the cache size by that cap and rounds up to
a multiple of 32 KiB to get a minimum chunk size. Leave --chunksize off and
it raises the default to meet that minimum, printing a line to say so; name a
smaller size yourself and the command fails instead of rounding it up for you.
A terabyte-class cachevol therefore lands in the megabyte range, and that
decides your write amplification:
4 KiB write, promotion writes one whole chunk
64 KiB chunk -> 16x amplification
1 MiB chunk -> 256x
2 MiB chunk -> 512x
When you name no size at all, that minimum is rounded up again to a power of
two. The exact figure is not worth predicting. Run lvs -o+chunksize after
conversion and read it.
One hard incompatibility, verbatim from lvmcache(7): a cache pool created on a device with a 4096-byte logical block size cannot serve a main LV with logical block size 512, and “if the cache pool is attached, the main LV will likely fail to mount”. Check the filesystem’s sector size before attaching anything.
The three modes and what each survives
| Mode | Kernel guarantee |
|---|---|
| writeback | writes land in the cache only, block marked dirty |
| writethrough | write completes only after both devices have it |
| passthrough | reads and writes go to the origin, a write invalidates a hit |
Power loss is documented plainly: metadata is committed on every FLUSH or FUA bio and otherwise once a second, so “the cache behaves like a physical disk that has a volatile write cache. If power is lost you may lose some recent writes.” The sharper clause is about dirty bits, which are treated as a hint: “If the system crashes all cache blocks will be assumed dirty when restarted.” After an unclean shutdown of a writeback cache, the next detach drains the entire cache, not the genuinely dirty part of it. Rehearse that drain before you need it.
Writethrough is often described as costing performance but never data. Half of
that is documented and half of it is folklore. lvmcache(7) is right that
losing the fast device loses no data, but dm-cache-target.c remaps a read
hit straight to the cache device and propagates the bio status upward with no
fallback to the origin. No data loss is not the same as still working: the LV
stays unusable until the cache is detached. That behaviour is a source read,
not a documented promise.
So: mirror the fast device if it matters. lvmcache(7) endorses it directly, and the same reasoning that governs arrays with NVMe drives applies here. For writeback, the fast device holds acknowledged writes that exist nowhere else, which makes power-loss protection a requirement rather than a preference - that is the line dividing enterprise SSDs from the rest.
Tunables that still do something
migration_threshold is a core argument, not a policy argument, default
2048 sectors (1 MiB). It is not a bandwidth cap. spare_migration_bandwidth()
compares in-flight migration volume against it and requires the device to have
been idle for a full second first:
concurrent migrations ~= migration_threshold / sectors_per_block
64 KiB chunks (128 sectors), threshold 2048 -> 16
1 MiB chunks (2048 sectors), threshold 2048 -> 1
LVM papers over the second row by enforcing a floor of “2048 sectors (1 MiB) or 8 cache chunks whichever of those two values is larger”. The kernel default is flatly 2048; the eight-chunk floor is LVM policy.
The rest of smq is fixed: FREE_TARGET 25u is consulted only when the cache is
already full - free_target_met() is reached from queue_promotion() on the
allocator-empty path - so it paces demotion to make room rather than holding a
quarter of the cache
blocks free, and clean_target_met() returns true while the
device is busy, on the reasoning that “if we’re busy we don’t worry about
cleaning at all”. On a server that is never idle, a writeback cache never
drains by itself. Dirty grows to the cache size, and the only real drains
are --cachepolicy cleaner, converting to writethrough, or splitting.
smq_max_background_work (default 4096, under
/sys/module/dm_cache_smq/parameters/) exists only from Linux 7.2 onward, and
the comment above it in the source says it applies to newly created caches
rather than running ones.
The worked sequence
Choose the fast device first. An M.2 slot on the board is the cheap path - NVMe in M.2 - and the endurance column matters more than the sequential-read column, because a cache device’s job is absorbing writes forever.
# origin on the HDD, fast LV on the NVMe, same VG
lvcreate -n main -L 8T vg /dev/sda
lvcreate -n fast -L 400G vg /dev/nvme0n1
# mirrored fast LV instead, if the mode demands it
lvcreate --type raid1 -m 1 -n fast -L 400G vg /dev/nvme0n1 /dev/nvme1n1
cachevol puts data and metadata in different parts of one fast LV and works
for dm-cache and dm-writecache. cachepool splits them across separate LVs,
allows explicit placement, and is dm-cache only; lvmcache(7) calls that the
“preferred way of using dm-cache”.
# simplest: one fast LV, cache data and metadata inside it
lvconvert --type cache --cachevol fast vg/main
# cache pool with explicit data and metadata volumes
lvcreate -n fast -L 400G vg /dev/nvme0n1
lvcreate -n fastmeta -L 48M vg /dev/nvme1n1
lvconvert --type cache-pool --poolmetadata fastmeta vg/fast
lvconvert --type cache --cachemode writethrough \
--cachepool fast vg/main
Mode, policy and settings after the fact, then the two exits:
lvchange --cachemode writeback vg/main
lvchange --cachepolicy smq \
--cachesettings 'migration_threshold=8192' vg/main
lvchange --cachesettings 'migration_threshold=default' vg/main
lvconvert --splitcache vg/main # flush, detach, keep the fast LV
lvconvert --uncache vg/main # flush, detach, delete it
Both flush before separating. lvconvert(8) accepts
writethrough|writeback|passthrough for --cachemode while lvmcache(7)
documents only the first two; passthrough through LVM is real and
under-documented, and the kernel requires a clean cache to enter it.
Reading the counters
dmsetup status on a named device prepends start, length and target type, so
the hit and miss counters sit at fields 8 to 11: read hits, read misses, write
hits, write misses, then demotions, promotions and dirty blocks.
dmsetup status --target cache --noflush vg-main | awk '{
split($7, c, "/")
r = $8 + $9; w = $10 + $11
printf "used %.1f%% rhit %.2f%% whit %.2f%% dirty %d\n",
100*c[1]/c[2], r ? 100*$8/r : 0, w ? 100*$10/w : 0, $14
}'
Two traps. cache_status() calls commit() unless the no-flush flag is set,
“to ensure statistics aren’t out-of-date” - a one-second monitoring poll
drives a metadata write to the NVMe every second, forever. Pass --noflush,
even though dmsetup(8) explains that flag for thin pools alone. And the
counters are 32-bit, read with atomic_read() and printed as %u: they wrap,
and they reset to zero on any table reload or deactivation. Graph deltas with
wrap detection, never raw totals.
LVM exposes the same numbers as report fields, which is the saner interface:
lvs -o lv_name,cache_total_blocks,cache_used_blocks,cache_dirty_blocks,\
cache_read_hits,cache_read_misses,cache_write_hits,cache_write_misses vg/main
lvs -o+cache_mode,cache_policy,kernel_cache_settings,chunksize vg/main
The operating mode is readable from either side: emit_flags() prints
exactly one of writethrough, passthrough or writeback in the status
line, followed by the core arguments, the policy name, rw or ro, and
needs_check or -. A cache whose entire status line reads Fail has hit
CM_FAIL, which the source marks as never leaving fail mode; one showing
needs_check comes up read-only until cache_check and cache_repair from
thin-provisioning-tools have been near it.
bcache: buckets, not chunks
bcache is not dm-cache with a different spelling. dm-cache is a block-remapping layer bolted under a volume manager; bcache is a small log-structured allocator with a btree index that happens to present a block device. That difference decides almost everything else about it - how it writes, what it wears, and which of its failure modes you can survive.
Architecture
Three objects, and the vocabulary matters because the sysfs tree is organised
around it. A backing device is the slow one. A cache device is the fast
one. A cache set is the group a cache device belongs to, and a backing
device attaches to the set, not to a device. The format reserves
MAX_CACHES_PER_SET 8, but bcache.rst is blunt: “Cache devices are managed
as sets; multiple caches per set isn’t supported yet but will allow for
mirroring of metadata and dirty data in the future.” There is no native
mirroring of a bcache cache. Putting the cache on an md RAID1 and
registering the md device works, and is the usual community answer, but no
kernel document endorses it - treat it as convention, not policy.
Allocation is in buckets, and this is the design’s whole thesis. bcache
allocates only in erase-block-sized buckets and writes into them sequentially,
which is why bcache.rst describes it as “designed to avoid random writes at all
costs”. A 4 KiB write does not trigger a bucket-sized read-modify-write the way
a 4 KiB promotion into a 2 MiB dm-cache chunk writes 2 MiB. bcache pays its
amplification later, in garbage collection and btree writes, instead of
immediately on every miss.
Bucket size should match the SSD’s erase block. make-bcache validates -b
through hatoi_validate(), which enforces a power of two, and the on-disk
superblock carries bucket size as a 16-bit sector count, so the ceiling is
16 MiB whatever the tool accepts. The default in the bcache-tools source is
1024 sectors, 512 KiB; the man page shipped with 1.0.8 still says
“Defaults to 128k”. The man page is stale, and a guide that repeats it is
repeating a bug. Vendors do not publish erase-block size, so this is a choice
between 256 KiB and 1 MiB made on the drive’s class rather than its datasheet,
and the wear consequences are the ones set out in
SSD endurance: TBW, DWPD, how worried to be.
The superblock sits at sector 8, byte offset 4096, and backing-device data
starts at BDEV_DATA_START_DEFAULT, 16 sectors, 8 KiB in. On RAID, align
that: the documentation’s worked example takes a 64k stripe and uses
bcache make --data-offset with 64k * 2*2*2*3*3*5*7 = 161280k, wasting
157.5 MB to pre-align for the data-spindle counts it lists - 3 to 10, then 12,
14, 15, 18, 20 and 21.
Making and attaching
bcache make -B /dev/sdb # backing device
bcache make -C /dev/nvme0n1p1 # cache device
bcache make -B /dev/sdb -C /dev/nvme0n1p1 # both, auto-attached
echo /dev/sdb > /sys/fs/bcache/register # manual registration, no udev
Those are the unified bcache tool’s spellings, the ones the kernel
documentation now uses; upstream bcache-tools lives in Coly Li’s tree, not
Overstreet’s. A distribution still packaging bcache-tools 1.0.8, as Debian
bookworm does, has no bcache binary at all: there the three calls above are
make-bcache -B, make-bcache -C and make-bcache -B ... -C ..., same flags
and same devices. Attach and detach are sysfs writes against the resulting
/dev/bcacheN:
echo <CSET-UUID> > /sys/block/bcache0/bcache/attach
echo 1 > /sys/block/bcache0/bcache/detach
echo 1 > /sys/block/sdb/bcache/running # force up, no cache
Attachment persists across reboots. For a partition the sysfs path nests:
/sys/block/sdb/sdb2/bcache, never /sys/block/sdb2/bcache. The
documentation’s own recommendation is to format every slow device as a backing
device from the start, with no cache, and add flash later - the one decision
that makes bcache retrofittable at all.
The cache modes
echo writethrough > /sys/block/bcache0/bcache/cache_mode, with
writeback, writearound and none the other values.
| Mode | Cache device lost |
|---|---|
writethrough (default) |
no data loss; discard the device |
writearound |
no data loss; reads-only caching |
writeback |
dirty data gone, backing device stale |
Writeback is off by default, and the documentation says why in one sentence: “not due to a lack of maturity, but simply because in writeback mode you’ll lose data if something happens to your SSD.” bcache has no clean-shutdown concept because it never acknowledges a write before it is on stable storage, so an unclean shutdown is a handled case in every mode. A write error in writethrough invalidates that LBA; the same error in writeback is passed up to the filesystem. Any writeback cache in front of spinning disks wants power-loss protection, which is a filter you can apply directly on enterprise SSDs.
sequential_cutoff
Default 4 << 20, 4 MiB, set in super.c. The rationale is stated
plainly: “if you copy a 10 gigabyte file you probably don’t want that pushing
10 gigabytes of randomly accessed data out of your cache.”
The mechanism is better than the name suggests. bcache keeps a 128-entry hash
of recent I/Os per device (RECENT_IO_BITS 7), matching on the sector where
the previous request ended within a 5-second window, and accumulates a
per-task running total. The bypass test uses
max(task->sequential_io, task->sequential_io_avg), a rolling average, so a
backup that seeks between files still bypasses wholesale instead of caching the
first 4 MiB after every seek. Several conditions bypass regardless of the
cutoff: a detaching device, discards, CACHE_MODE_NONE, writes in
writearound, unaligned I/O, and in_use > CUTOFF_CACHE_ADD, which is 95 -
above 95 per cent cache utilisation everything bypasses.
writeback_percent and writeback_rate
writeback_percent defaults to 10. It is a dirty-data target, fed to a
controller the source calls a PI controller (the documentation’s “PD” is
wrong; there is no derivative term). Defaults from
bch_cached_dev_writeback_init(): writeback_rate 1024 sectors/s initial,
writeback_rate_minimum 8, writeback_rate_update_seconds 5,
writeback_rate_p_term_inverse 40, writeback_rate_i_term_inverse 10000. The
40 means the proportional term aims to retire all dirty blocks in 40 seconds;
the 10000 accumulates one ten-thousandth of the error per second.
The target is shared, and this surprises people:
target = cache_sectors * writeback_percent / 100
per-device = target * (this_dev_sectors / c->cached_dev_sectors)
5 equal backing devices, writeback_percent = 10
-> each device's dirty target is 10% / 5 = 2% of the cache
By default you cannot set writeback_percent above 40. sysfs.c clamps it
to bch_cutoff_writeback, whose default is CUTOFF_WRITEBACK 40; that is a
module parameter, fixed at module load and itself capped at
CUTOFF_WRITEBACK_MAX, 70. The same two numbers are then reused against a
different quantity. should_writeback() tests them against gc_stats.in_use,
the cache set’s own utilisation: above bch_cutoff_writeback_sync, 70 per cent
used, writeback stops entirely, and between 40 and 70 per cent only sync writes,
REQ_META, REQ_PRIO and dirty partial stripes still go to writeback -
everything else silently degrades to writethrough. That is the mechanism behind “my
writeback cache stopped absorbing writes”, and it is visible read-only at
/sys/fs/bcache/<cset-uuid>/internal/cutoff_writeback. Against it,
idle_max_writeback_rate (default on) slams the rate to INT_MAX whenever the
set is idle, so a reader of writeback_rate during an idle drain sees a figure
that means nothing.
Congestion
/sys/fs/bcache/<cset-uuid>/congested_read_threshold_us = 2000
/sys/fs/bcache/<cset-uuid>/congested_write_threshold_us = 20000
2 ms and 20 ms; zero disables. bcache tracks latency to the cache device and
“gradually throttles traffic if the latency exceeds a threshold (it does this
by cranking down the sequential bypass)”. bch_get_congested() returns a
sector threshold compared against the same per-task counter that
sequential_cutoff uses, so congestion literally lowers the effective cutoff,
with i -= hweight32(get_random_u32()) fuzzing the boundary. A cache set that
looks like it has stopped caching under load has usually not failed; it has
decided the cache device is slower than the threshold says it should be.
Reading the hit ratio
Four stat directories exist under both the device and the set, one running total
and three that decay:
stats_total, stats_five_minute, stats_hour, stats_day, each with
cache_hits, cache_misses, cache_hit_ratio, cache_bypass_hits,
cache_bypass_misses, cache_miss_collisions and bypassed.
cat /sys/block/bcache0/bcache/stats_hour/cache_hit_ratio
cat /sys/block/bcache0/bcache/dirty_data
cat /sys/block/nvme0n1/nvme0n1p1/bcache/priority_stats
Two caveats. Hits are counted per individual I/O and a partial hit counts as
a miss, so the ratio is sensitive to request size against block_size. And
bypassed, in bytes, is the number that tells you the sequential cutoff is
working rather than silently eating your cache. priority_stats is the real
working-set instrument: unused percentage, metadata overhead, and priority
quantiles. Benchmark honestly - a full btree node cannot take the key a cache
miss would insert, so a read-only warm-up under-fills the cache. The
documentation’s own answer: “warm the cache by doing writes.”
The sharp edges
Three, stated without softening.
Formatting is required. The backing superblock plus the 8 KiB data offset
means an existing filesystem cannot be converted in place in the general case.
You back up, format with bcache make -B, and restore. The reverse direction
is easier than it looks - the untouched filesystem is still sitting at an 8 KiB
offset, reachable through a loop device (losetup -o 8192) - but that is a
recovery path, not a migration plan.
Dirty data on a lost cache device is lost. echo 1 > .../running will
force the backing device up without its cache, and if it was in writeback the
documentation promises “massive filesystem corruption, though ext4’s fsck does
work miracles”. The device state then reads inconsistent. If the old cache
reappears later, its contents are invalidated; do not try to be clever.
stop_when_cache_set_failed governs the automatic behaviour - default auto
stops the bcache device only when it has dirty data, always stops it
regardless.
Detach must finish before the cache goes. detach flushes first, and the
documentation admits “it currently doesn’t do anything intelligent if it fails
to read some of the dirty data.” Confirm dirty_data reads 0. Then remember
that detaching does not unregister: “Ooops, it’s disabled, but not
unregistered, so it’s still protected.” The full sequence is detach, wait,
echo 1 > /sys/fs/bcache/<cset-uuid>/stop, then wipefs.
bcache or LVM cache
| bcache | LVM cache (dm-cache) | |
|---|---|---|
| Retrofit onto live data | no, format required | yes, lvconvert in place |
| Mirrored cache device | only via md underneath |
documented, --type raid1 |
| Small random write wear | lower, log-structured | chunk-sized promotions |
| Sequential bypass | yes, sequential_cutoff |
no tunable on smq kernels |
| Load-sensitive throttling | yes, congestion thresholds | no equivalent |
| Tunables that still work | many, documented | few; smq ignores most |
| Default write mode | writethrough | writethrough under LVM |
Choose bcache when the volume is being built from scratch, when the workload is small random writes that would otherwise pay chunk-sized promotion costs, and when you want the cache to stand down under load rather than fight the backing array. Choose LVM cache when the data already exists and cannot be recreated, when the cache device must be mirrored without a second layer, or when the volume already lives in a VG and adding a second stack is its own risk. Neither is faster than the other in any way this guide can evidence - no vendor-neutral head-to-head measurement of the two is cited here, and one will not be invented. The deciding factors are retrofit, mirroring and wear, in that order.
dm-writecache: the one that does not cache reads
Every other layer in this guide is trying to keep your hot data on flash. dm-writecache is not. It exists to make a write return faster, and it has no opinion at all about what you read.
The kernel documentation says it in one sentence, and it is the sentence most readers have never seen:
“The writecache target caches writes on persistent memory or on SSD. It doesn’t cache reads because reads are supposed to be cached in page cache in normal RAM.”
No read ever pulls a block into the cache. There is no promotion path, no
policy, no hint metadata, no hotspot tracking. A read is served from the fast
device only in the narrow case where that block is still sitting there from a
recent write and has not been written back yet - the target counts those as
read hits in its status line - and every other read goes to the hard drive.
lvmcache(7) restates it from the other side: “Data read from the main
LV is not stored in the cache, only newly written data.” If you attach a
writecache expecting your find runs or your VM boot storms to get quicker,
nothing will happen, and nothing is broken. You picked the wrong target.
What it is actually for
A synchronous write workload in front of slow media. A write arrives, lands on
the fast device, and is acknowledged - the origin has not been touched yet.
That is the whole mechanism. It flattens the fsync and commit latency of a
database, a mail spool, an NFS server, a metadata-heavy filesystem journal,
anything whose user-visible speed is governed by how long fdatasync() blocks.
It is not adaptive and there is no policy to tune, because there is nothing for
a policy to decide. Writes land in the cache, the high watermark starts
writeback and the low watermark stops it, and a write goes straight to the
origin when cleaner mode is on, when metadata_only is set and the write is not
flagged REQ_META, or, on the SSD path, when no free cache block is available.
Contrast dm-cache in writeback mode, which is a hotspot cache that
happens to absorb writes: it runs the smq policy, promotes and demotes on read
as well as write, and carries per-block metadata and hints. dm-writecache
carries neither policy nor hints, which is why its crash behaviour is a
question of what had been committed, rather than the rebuild dm-cache falls
back to - the dm-cache documentation is blunt about that one: “If the system
crashes all cache blocks will be assumed dirty when restarted.”
It is pointless when your writes are already asynchronous. Buffered writes are absorbed by the page cache and flushed by writeback threads; they were never waiting on the disk in the first place. Confirm before you build: if the workload issues few FLUSH or FUA requests, the latency win is small. What is left is that writeback walks the cache in origin-sector order and merges neighbouring blocks into larger writes, which a random write stream landing on a hard drive can still feel.
Two modes, and the RAM bill nobody budgets for
<type> in the constructor is p for persistent memory or s for SSD. The
pmem path writes into the cache directly, and its fua setting governs the
writes it later sends back to the origin rather than the incoming ones; the SSD
path batches and commits.
Their autocommit defaults differ by three orders of magnitude, from the kernel
doc: autocommit_blocks is 64 for pmem and 65536 for ssd, alongside
autocommit_time of 1000 ms, high_watermark 50, low_watermark 45 and
pause_writeback 3000 ms. That 65536 is a durability parameter wearing a
performance parameter’s clothes - it is how many blocks may sit uncommitted
when no FLUSH arrives.
Then the figure that kills most designs. lvmcache(7), verbatim: at 4096-byte blocks, “each 100 GiB of writecache cachevol uses slightly over 2 GiB of system memory”; at 512-byte blocks, “a little over 16 GiB”. The man page states the rate and stops there, so scaling it is your own arithmetic - 10.24 times the per-100-GiB figure for a 1 TiB cachevol, and “slightly over” means treat the results as floors:
1 TiB cachevol, block_size=4096 -> about 21 GiB of RAM
1 TiB cachevol, block_size=512 -> about 164 GiB of RAM
Size the writecache against your RAM, not against your NVMe. And the block
size is not free to choose: if an existing xfs was made with 512-byte sectors,
lvmcache(7) warns the filesystem “will likely fail to mount” behind a 4096
writecache. Check xfs_info for sectsz= and match it.
Setup
# sizes below are illustrative, not a recommendation
lvcreate -n main -L 8T vg /dev/slow_hdd
lvcreate -n fast -L 200G vg /dev/fast_nvme
lvconvert --type writecache --cachevol fast \
--cachesettings 'block_size=4096' vg/main
lvs -o lv_name,writecache_block_size,writecache_total_blocks,\
writecache_free_blocks,writecache_writeback_blocks,writecache_error vg/main
lvchange --cachesettings 'high_watermark=40 writeback_jobs=2048' vg/main
Mirror the fast LV if the data matters, exactly as for dm-cache:
lvcreate --type raid1 -m 1 -n fast -L 200G vg /dev/nvme0n1 /dev/nvme1n1.
Detaching is a two-step: set cleaner=1, wait for
writecache_writeback_blocks to reach zero, then lvconvert --splitcache,
which is near-instant once drained. Skip the first step and splitcache turns on
cleaner mode itself and waits for the flush before it detaches, which is the
same work in a less convenient place. cleaner and max_age need Linux 5.7;
metadata_only and pause_writeback need 5.14; the target itself arrived in
4.18.
metadata_only deserves a look before you size anything: it promotes only
metadata, which “improves performance for heavier REQ_META workloads” and cuts
the flash written per day by whatever fraction your bulk data represents. See
SSD endurance: TBW, DWPD, how worried to be for what that
saving is worth in drive-years.
Where power-loss protection stops being optional
dm-cache in writethrough can lose its cache device and lose nothing. A writecache cannot. It acknowledged writes that exist on exactly one device, and that device is a single point of failure sitting in front of whatever redundancy the origin has. The same is true of the drive’s own volatile buffer: the acknowledged write may be in DRAM on the NVMe and not yet in NAND, so the guarantee you are relying on is the capacitor bank, not the filesystem. This is the one stack in this guide where a cache device without power-loss protection is a defect rather than a trade-off, so shop the enterprise SSDs and confirm power-loss protection (PLP) on the datasheet rather than inferring it from the price.
The published gap makes the point twice over. One third-party set of fio results, measuring fsync after a 16 KB write, reports 891 µs on a Crucial T500 and 2,974 µs on a Samsung 990 Pro, both without PLP, against 12.4 µs on a Solidigm D7-P5520 and 1.6 µs on a Samsung PM9A3, both with it. Those are one person’s drives on one rig rather than vendor figures, so read the ranking and not the decimal places. A 7200 rpm drive’s average rotational latency is 4.16 ms. A consumer NVMe at 2.97 ms per fsync is not buying you much in front of a drive that averages 4.16 ms of rotational latency on its own. Correctness follows the same line: the drive must honour FLUSH, and power-cut testing has repeatedly found drives that do not.
Removal without LVM, from the kernel doc, is a fixed sequence worth rehearsing
on a scratch volume: send flush_on_suspend, load an inactive linear table
pointing at the underlying device, suspend, read the status and confirm field 1
is zero, resume onto the linear table, then delete the cache device. Skipping
the status check is how a drained cache turns out not to have been drained.
Open CAS: the out-of-tree caching layer, and what it does with DRAM
For most readers the honest answer is still LVM cache. Open CAS is a BSD-licensed, out-of-tree kernel module with its own userspace tool, its own config file and its own kernel-version ceiling, and adopting it means accepting that a kernel upgrade can leave your cache module unbuildable. Read the rest of this section as a description of what the extra operational cost buys, not as a recommendation to pay it.
It earns a place here for one reason the alternatives cannot
match: Open CAS documents DRAM as a first-class cache medium rather than a hack.
The tool reference defines the cache device as “an SSD or any NVM block device
or RAM disk shown in the /dev/disk/by-id directory”, and the project’s own
description of OCF lists the media it caches onto as HDDs on SSDs, QLC on TLC,
Optane, “RAM memory, or any combination of above including all kinds of
multilevel configurations”.
Architecture, and why it is not in the tree
Open CAS Linux is a thin Linux shim over OCF, the Open CAS Framework: by its
own documentation a “high performance block storage caching meta-library written
in C… entirely platform and system independent, accessing system API through
user provided environment wrappers layer”. The same library is the caching core
of SPDK. That portability is the architectural difference from dm-cache, a
device-mapper target managed through LVM, and from bcache, a block-layer cache
in drivers/md/bcache/ carrying MAINTAINERS status S: Maintained under Coly Li
and Kent Overstreet. Both of those are GPL and in mainline. Open CAS is
BSD-3-Clause and bundles safeclib under MIT, and it has stayed outside the tree:
it ships as DKMS source and builds against whatever kernel is on the target.
It is not an Intel project any more. Intel open-sourced the Linux version in 2018; the README, read on 2026-09-25, states that Open CAS Linux “is developed and maintained by Unvertical, founded by the project’s core maintainers”, with Robert Baldyga as lead and commercial support sold by Unvertical. The repository is not archived and was last pushed on 2026-09-24, so “abandoned” is the wrong word. “Trailing” is the right one.
The kernel ceiling is the sharp edge and it is dated. As of 2026-09-25, release v26.03.4 was published 2026-08-24; v26.03.3, on 2026-07-07, added support for kernel versions up to 7.0. A pull request adding kernel 7.2 support was merged on 2026-09-21, so master is ahead of every release, while mainline sits at 7.3-rc4. The README adds that “bugfix releases are guaranteed only for the latest major release line”. On a current distribution kernel the newest release is at or past its ceiling, and those numbers move - check the release page before planning around them.
Accelerated devices appear as /dev/cas<cache ID>-<core #>, for example
/dev/cas1-1; partitions on an accelerated parent are auto-mapped, so
/dev/sdc1 becomes /dev/cas1-1p1, and the originals are not meant to be
addressed directly afterwards. Persistence across boot comes from
/etc/opencas/opencas.conf, read by udev, with [caches] and [cores]
sections.
The RAM arithmetic, which is the whole point here
Open CAS is unusual in publishing its metadata cost as a formula:
RAM = 1 GiB + (2% * 4 KiB / cache_line_size + 0.05%) * cache_capacity
# 2.05% at 4 KiB is the project's own example; the rest is arithmetic
4 KiB line (default): 1 GiB + 2.05% of cache capacity
64 KiB line: 1 GiB + 0.175% of cache capacity
1 TB cache device @ 4 KiB -> ~21.6 GB of DRAM in metadata
1 TB cache device @ 64 KiB -> ~2.8 GB of DRAM in metadata
A terabyte of cache at the default line size works out at roughly twenty-one and a half gigabytes of host memory before a single block is cached, and the line size is settable only at cache start and immutable afterwards. The permitted values are 4, 8, 16, 32 and 64 KiB. Sizing this wrong is a rebuild, and on a memory-constrained host it is a DIMM purchase rather than a tuning exercise. The other published requirements: CPU overhead “less than 10% in most cases”, supported filesystems ext3 (up to 16 TiB), ext4 and xfs, S3 suspend and S4 hibernate must be disabled, and the OS must not live on the cached core device.
Five modes, and what each promises on power loss
| Mode | Flag | Guarantee |
|---|---|---|
| Write-through | wt (default) |
Writes land in cache and core together; the core is fully in sync; reads accelerated only |
| Write-back | wb |
Acknowledged from cache, written back later; a cache device failure “may lead to the loss of data that has not yet been flushed to the core device” |
| Write-around | wa |
Writes go to cache only if that block is already cached, and always through to core; avoids write-once pollution |
| Write-only | wo |
Writes as in wb, but reads are never promoted; a read hits only data previously written there; same dirty-loss exposure as wb |
| Pass-through | pt |
Caching disabled; lets you attach every core first and switch modes live |
Write-only is the mode the in-tree options do not have, and the one worth understanding. It separates the two jobs a cache normally does at once: it buffers writes without letting a read-heavy scan evict the write buffer. The documented behaviour is that reads are never promoted, so on a workload whose reads are already served by the page cache, the DRAM tier does not end up holding a second copy of them.
Eviction is LRU, and it is the only policy. The admin guide names LRU as the
cache’s eviction policy, and no casadm command sets another;
the published tool reference lists seq-cutoff, cleaning, cleaning-acp,
cleaning-alru, promotion and promotion-nhit. Any source offering a menu of
eviction policies is describing something else. Cleaning, a separate axis
governing write-back flush behaviour, does have choices: alru (approximate LRU,
periodic, the default), acp (aggressive, intended “to stabilize the timing of
write-back mode data flushing”), and nop. Promotion is always by default or
nhit, which promotes only on the n-th access, threshold 2 to 1000.
A multi-level DRAM tier, built on the project’s own example
modprobe brd rd_size=102400 # ramdisk; /dev/ram0 is reserved
casadm -S -d /dev/disk/by-id/nvme-SSD # level 1: SSD over HDD
casadm -A -i 1 -d /dev/disk/by-id/wwn-0x50014ee0aed22393
ln -s /dev/ram1 /dev/disk/by-id/ram-disk-part1
casadm -S -d /dev/disk/by-id/ram-disk-part1 -i 2 # level 2: DRAM over the cas device
casadm -A -i 2 -d /dev/cas1-1
mkfs.xfs -f -i size=2048 -b size=4096 -s size=4096 /dev/cas2-1
casadm --set-cache-mode --cache-mode wb --cache-id 1 --flush-cache yes
casadm -X -n cleaning -i 1 -p acp
casadm -X -n seq-cutoff -i 1 -j 1 -p always -t 4096
casadm -P -i 1 --filter conf,req --output-format csv # hit ratio, machine readable
casadm -F -i 1 # flush
casadm -T -i 1 # stop
The documentation is explicit that the ramdisk tier “is volatile and will need
to be set up again following a system restart”, that data survives in the SSD
across a clean restart, and that “if the RAMdisk cache tier is started in
write-back mode and there is a dirty shutdown, data loss may occur”. Because the
SSD tier underneath stays fully inclusive, a wt or wa DRAM tier carries
little durability risk of its own; wb on DRAM is the configuration that can
cost you data. That inclusive design is also what separates this from stacking
two independent caches yourself.
One trap worth memorising: in --set-cache-mode, -f is --flush-cache and
takes yes or no; in --start-cache the same -f is --force, creating the
cache even if a filesystem is already on the cache device, and in
--remove-core it is --force again, meaning removal without flushing dirty
data. Same letter, opposite consequences.
What it buys, and what it costs
Set against dm-cache and bcache, Open CAS buys five cache modes where those two
have three and four, an IO-class policy engine neither has an equivalent of,
documented multi-level stacking, and machine-readable statistics from a single
tool. It costs mainline membership, a kernel ceiling that trails the tree, a
documentation gap (target_failover_state=standby appears in the config
reference while casadm itself carries standby, attach-cache and
detach-cache commands that the tool reference never lists, the attach/detach
pair having been announced in the v25.03 release notes of 2025-04-28), and a
warning of its own that lazy_startup on cores under a write-back cache “might
cause data corruption”.
No vendor-published, reproducible throughput or latency figure for Open CAS was found, and this guide will not claim one. The only quantitative claims the project makes are the RAM formula above and that CPU overhead is under 10 per cent in most cases. If your workload is a general-purpose fileserver on ext4 or xfs and your team already runs LVM, the in-tree option is the correct one. Open CAS earns its keep when you need write-only mode, IO classes, or a DRAM tier over a flash tier over spinning disks - and when someone will own the DKMS rebuild on every kernel upgrade.
ZFS: ARC and L2ARC
On ZFS the cache is the RAM, and the NVMe device is the consolation prize.
Every other stack in this guide treats flash as the cache and memory as an
accident of the page cache. ZFS inverts that: the Adaptive Replacement Cache
lives in RAM, and the device you add with zpool add tank cache /dev/nvme0n1
is a second-level spill for blocks the ARC was about to throw away. OpenZFS
says so itself, in its own pool-structure documentation: “RAM is by far the
most effective ZFS ‘tuning knob’. Before adding any cache device, check
whether the ARC is simply too small.”
The ARC, and why it gets the money first
The ARC keeps an MRU list and an MFU list, each shadowed by a ghost list of
headers for blocks already evicted. Hits on the ghost lists are what move the
target balance, which is how the cache adapts to a workload instead of
behaving like a fixed LRU. In OpenZFS 2.2 and later the old single arc_p
balance variable, counted in bytes, is gone; the commit that replaced it
(“More adaptive ARC eviction”, a8d83e2) splits the job across three
fixed-point fractions in module/zfs/arc.c - arc_meta for the
metadata-versus-data balance, arc_pd and arc_pm for the MRU-versus-MFU
balance of data and of metadata separately. How hard a ghost hit on metadata
pulls the first of those is zfs_arc_meta_balance, default 500;
zfs_arc_meta_limit and zfs_arc_meta_min no longer exist and appear
nowhere in zfs.4 at 2.4.4.
The default ceiling changed, and most of the advice on the web predates it.
In module/os/linux/zfs/arc_os.c, arc_default_max() returned
MAX(allmem / 2, min) through 2.1.16 and 2.2.11 - the familiar “half your
RAM”, and what zfs.4 for 2.2 still describes. From 2.3 onward Linux uses the
formula FreeBSD already had, MAX(allmem * 5 / 8, allmem - 1 GiB). On a
64 GiB Linux box running 2.3 or newer the default ARC cap is 63 GiB, not 32.
zfs_arc_max=0 means auto; the floor is 64 MiB and it cannot be reset to 0 at
runtime. zfs_arc_min defaults to the larger of 32 MiB and allmem/32.
Observability is the arcstats kstat file, /proc/spl/kstat/zfs/arcstats on
Linux; note that OpenZFS 2.4.0 renamed the wrappers, arcstat to zarcstat
and arc_summary to zarcsummary.
So the first purchase for a pool of spinning disks is memory, and on most second-hand server boards that means registered ECC DDR4. Only after the ARC has been given everything the chassis will hold does a cache vdev become an interesting question.
The “1 GB of RAM per TB” rule
It is not an OpenZFS rule. It appears nowhere in OpenZFS documentation or source. It is FreeNAS-era hardware guidance that the community carried forward, and the current TrueNAS hardware guide does not state it. What that guide states instead is a base of 8 GB, plus roughly 5 GB per TB where deduplication is enabled, plus approximately 1 GB for every 50 GB of L2ARC.
Read those last two carefully. The figure that genuinely scales with pool capacity is the dedup table, and it is five times the folk rule. The figure that scales with cache-device capacity is the L2ARC header cost, which is the subject of the next section. The 1 GB per TB rule conflates the two and applies neither correctly. A 200 TB archive pool with no dedup and a cold working set does not want 200 GB of RAM; a 20 TB VM pool with dedup on wants far more than 20 GB. This guide will not offer a replacement single number, because the two costs have different drivers and one number cannot carry both.
What L2ARC holds, and what it never sees
L2ARC is a read cache only. The arc.c comment is unambiguous: “The
L2ARC does not store dirty content. It never needs to flush write buffers back
to disk based storage.” If a buffer resident in L2ARC is written and
dirtied, the stale L2ARC copy is dropped immediately. No write ever lands
there first, so no write is ever at risk on it, and the device needs no
power-loss protection. A read error against a cache device is simply
reissued against the pool. It cannot be mirrored and cannot be part of a
raidz, and it does not need to be - if you are assembling several NVMe
devices into something redundant for other roles in the same box,
Arrays with NVMe drives covers that separately.
The feed thread scans near the tail of the MRU and MFU lists and copies
buffers that are about to be evicted, in their ARC form, which means
compressed when compressed ARC is on. Two consequences follow directly.
Blocks never read cannot be in L2ARC, so a scrub, a resilver, a backup walk
or a find over cold metadata gets nothing from it. And prefetched blocks
that were never consumed are skipped, because l2arc_noprefetch defaults to
1. Per-dataset gating is secondarycache=all|metadata|none, with the ARC’s
own gate primarycache above it; setting primarycache=none on a dataset
effectively disables its L2ARC too, since L2ARC is fed only from ARC
evictions.
The header tax, and the cap that already exists
Every block resident in L2ARC but not in ARC keeps a stub header in ARC.
arc.c charges HDR_L2ONLY_SIZE, defined as
offsetof(arc_buf_hdr_t, b_l1hdr), which works out to roughly 96 bytes
on x86-64 for the 2.2 to 2.4 layout of struct arc_buf_hdr in
include/sys/arc_impl.h. The TrueNAS hardware guide puts the same number in prose,
“for every data block in the L2ARC, the primary ARC needs a 96-byte
entry”, so size with it and then check the real one:
the running total is published as the l2_hdr_size kstat, which makes this
measurable rather than estimated. The 70-byte figure still carried in Klara
Systems’ L2ARC sizing formula is Solaris-era and understates current headers.
Header cost per TiB of L2ARC content, at 96 B per record.
Record size here is the PHYSICAL (post-compression) block size,
because that is what lands on the device.
128 KiB blocks: 1 TiB / 131072 B = 8,388,608 recs x 96 = 768 MiB
32 KiB blocks: 1 TiB / 32768 B = 33,554,432 recs x 96 = 3 GiB
16 KiB blocks: 1 TiB / 16384 B = 67,108,864 recs x 96 = 6 GiB
8 KiB blocks: 1 TiB / 8192 B = 134,217,728 recs x 96 = 12 GiB
4 KiB blocks: 1 TiB / 4096 B = 268,435,456 recs x 96 = 24 GiB
TrueNAS's "1 GB per 50 GB of L2ARC" is 2%, which back-solves to an
assumed average physical block of about 4,800 B (4.7 KiB). It is a worst-case
small-block figure, not a universal one.
Those headers are not evictable under memory pressure, unlike ordinary
ARC buffers. Left unbounded, a large cache device would convert an
increasing share of a fixed ARC into bookkeeping for a slower cache - paying
RAM to make RAM smaller. ZFS already bounds it. l2arc_meta_percent
defaults to 33, documented as “Percent of ARC size allowed for L2ARC-only
headers. Since L2ARC buffers are not evicted on memory pressure, too many
headers on a system with an irrationally large L2ARC can render it slow or
unusable.” The enforcement lives in l2arc_hdr_limit_reached(), which
weighs the l2_hdr_size total against that percentage of the ARC target -
arc_c once the ARC is warm, arc_c_max before it is - and stops both
feeding and persistent rebuild above the line.
Where a 16 GiB ARC stops filling its L2ARC:
16 GiB = 17,179,869,184 B
x 33 / 100 = 5,669,356,830 B of header budget
/ 96 B per record = 59,055,800 resident records
at 16 KiB physical blocks: 59,055,800 x 16384 B = 901 GiB
at 128 KiB physical blocks: 59,055,800 x 131072 B = 7.04 TiB
The honest reading: an oversized L2ARC does not silently melt the ARC, it stops growing, and you have bought flash you cannot use while still paying for every header it did manage to install. A 2 TB cache device in front of a 16 GiB ARC on a database pool at 16 KiB records will never fill past about 900 GiB. Buy the ARC first; the cache device’s usable size is a function of it.
Filling it: the four knobs that matter
l2arc_write_max caps what one feed writes, and its default of 33554432
is 32 MiB. Feeds are l2arc_feed_secs apart, default 1 second, which reads
as 32 MiB/s - but that is the quiet-list interval, not a hard ceiling.
l2arc_feed_again defaults to 1, and whenever a feed wrote more than half
of what it asked for, l2arc_write_interval() shortens the gap to
l2arc_feed_min_ms, default 200, so five feeds a second and roughly
160 MiB/s are reachable while the lists stay busy. While arc_warm is still
B_FALSE - it flips to B_TRUE in
arc_reap_cb_check() the first time free memory goes negative, that is, once
the ARC has grown into memory - l2arc_write_boost (also 32 MiB) is added on
top.
What a feed asks for also carries the persistent-L2ARC log block overhead, and
l2arc_write_size() clamps the result to a quarter of the cache device, so a
very small device never sees the full write size. Both defaults were 8 MiB
through 2.2.6 and became 32 MiB in 2.2.7, backported
into the 2.2 branch rather than held for 2.3 - zfs.4 prints 8 MiB at 2.2.6
and 32 MiB at 2.2.11 - which is why so much older advice says an L2ARC takes
days to warm.
1 TiB of L2ARC content at the 2.2.7+ write size of 32 MiB per feed,
feeds 1 s apart (the quiet-list interval, not a hard ceiling):
1 TiB / 32 MiB/s = 32,768 s = 9.1 hours of CONTINUOUS feeding
cold, with boost: 64 MiB/s = 4.6 hours
on 2.2.6 and earlier (8 MiB): = 36.4 hours
l2arc_feed_again (default 1) drops the gap to l2arc_feed_min_ms
= 200 ms while the lists are busy, so up to 5 feeds a second:
1 TiB / 160 MiB/s = 6,554 s = 1.8 hours
The 1 s interval read as a wear budget (the 200 ms interval can
multiply it by up to five):
32 MiB/s x 86,400 s = 2.64 TiB written per day
on a 1 TiB cache device that is about 2.6 DWPD, indefinitely
Real fill tracks neither figure closely, because the feeder writes only what
is eligible near the list tails. The wear number is the one people miss: a
consumer TLC drive rated at 0.33 drive writes per day (DWPD) does not survive a
cache role at
default settings, which is the argument for a
write-rated NVMe device rather than the fastest one on the shelf, or
for lowering l2arc_write_max. l2arc_headroom (default 8 since 2.2.7, 2
before) sets how far through the ARC lists the feeder looks per cycle, as a
multiplier of the effective write size; 0 disables that per-cycle limit.
l2arc_noprefetch=1 and l2arc_mfuonly=1 both trade hit ratio for less
flash traffic.
Persistent L2ARC
Introduced in OpenZFS 2.0 - l2arc_rebuild_enabled is present in
zfs-2.0.7/module/zfs/arc.c and absent in zfs-0.8.6 - and enabled by
default. L2ARC writes occasional log blocks describing what was written,
anchored by a pair of log block pointers in the device header, from which the
arc_buf_hdr_t structures are reconstructed after import. What is restored
is the headers, not the data: the blocks are still served from the SSD,
not from RAM. The rebuild is asynchronous, consumes ARC as it proceeds, and is
subject to the same l2arc_meta_percent cap above. Devices below
l2arc_rebuild_blocks_min_l2size (1 GiB) get no log blocks written at all.
One operational trap from zpoolconcepts(7): a device re-added with
zpool add has its label and header overwritten and its contents are not
restored, while a device brought back with zpool online is restored.
The verdict
An L2ARC earns its slot when the re-read working set is larger than any ARC
the chassis can hold, the blocks are big enough that the header tax stays
small, and the device can absorb at least 2.6 TiB a day without wearing out. It earns
nothing on a pool whose pain is first-access metadata, cold scans or
sequential streaming, and it is a net loss when it is large relative to the
ARC that must index it. OpenZFS master - unreleased at the time of writing,
and absent from the 2.4 source - carries the first upstream sizing ratio this
guide has found, in the comment on L2ARC_PERSIST_THRESHOLD (arc_c_max * 2):
an L2ARC smaller than twice arc_c_max either cyclically overwrites itself or
merely duplicates what the ARC already holds. Minimum twice your ARC, and a
workload that re-reads more than the ARC holds. Short of both, spend the
money on memory.
ZFS: the special vdev
This is the highest-value trick in the article, and it is one of the few devices here that takes the pool with it when it dies. Both halves of that sentence are load-bearing, and the second one is why this section is longer than it looks like it needs to be.
OpenZFS states the distinction before anything else: “It is not a cache. The blocks routed there exist only there.” An L2ARC holds a second copy of something the pool already has. A special allocation class vdev holds the only copy. Everything below follows from that.
What the allocation class holds
From zpoolconcepts(7), verbatim: “Allocations in the special class are
dedicated to specific block types. By default, this includes all metadata, the
indirect blocks of user data, intent log (in absence of separate log device),
and deduplication tables.”
The allocation order documented upstream is: dedup tables first (to a dedup
vdev if one exists, otherwise here, unless zfs_ddt_data_is_special is set to
0), then all metadata, then the indirect blocks of user data
(zfs_user_indirect_is_special, on by default), then small user data blocks if
a dataset opts in, then everything else to the normal class. ZIL blocks follow
their own ladder: a dedicated log vdev, then log space reserved on the
special vdev, then the special class proper, then log space on the normal
vdevs, then the normal class. A special vdev absorbs the ZIL when there is no
separate log device, and a dedicated log vdev - the SLOG of the next section -
always wins.
Two structural rules: “A pool must always have at least one normal (non-dedup/-special) vdev before other devices can be assigned to the special class”, and “If the special class becomes full, then allocations intended for it will spill back into the normal class.” A full special vdev degrades quietly rather than failing. New metadata simply lands on the spindles again.
special_small_blocks, and the trap inside it
The property is per-dataset and defaults to 0, meaning off. Metadata goes to the special vdev with or without it set; this property only adds small user data. The valid range moved in 2.4: through 2.3 the man page says “zero or a power of two from 512 up to 1048576 (1 MiB)”; 2.4 says “small file or zvol blocks … after compression and encryption” with a range up to the maximum block size, 16 MiB.
The words “after compression” are the trap. The comparison is against the
physical size, so “a 128 KiB record that compresses to 20 KiB counts as a
20 KiB block.” Set the threshold anywhere near your recordsize and a
well-compressing dataset will route essentially all of its data onto the
special vdev, which is not what you bought it for.
The second trap is a floor you cannot see from the property.
zfs_special_class_metadata_reserve_pct defaults to 25: small blocks stop
being accepted once the class is 75 per cent full, reserving the last quarter
for metadata. That is a floor for metadata, not a ceiling for the vdev.
Metadata alone can still fill it.
Practical choice: leave it at 0 first, measure, then raise it deliberately per dataset - 16K on a source tree or a maildir, 0 on a media dataset. Never set it globally because a number looked reasonable.
Why metadata on flash beats an L2ARC
Five mechanisms, none of them a matter of taste:
- Permanent placement, no warm-up. L2ARC fills at
l2arc_write_max, 32 MiB (33554432 B) per feed, and the feed interval defaults to one second; before 2.2.7 the write ceiling was 8 MiB. Filling 1 TiB at that ceiling takes 32,768 intervals - about 9.1 hours - of continuous feeding, and the feeder only writes what is eligible near the list tails, so the real figure is worse. - Almost no ARC headers. Every L2ARC-resident block keeps a stub header
in ARC:
HDR_L2ONLY_SIZE, defined inarc.cas the offset of the L1 header insidearc_buf_hdr_t, so of the order of a hundred bytes on a 64-bit build. Those headers are not released by memory pressure the way cached data is, which is whyl2arc_meta_percent, 33 per cent of ARC by default, exists to bound them. A special vdev costs none of this. - First-access service. The special vdev serves metadata blocks that have
never been read before. Scrub, resilver,
find,zfs list, a backup walk, aduacross ten million inodes. L2ARC by construction holds only what was already read and then evicted, so a traversal of metadata that has never been read cannot hit it. - Prefetch blindness.
l2arc_noprefetch=1is the default, so prefetched buffers that an application never used are kept out of the L2ARC, so a sequential metadata walk largely bypasses the feeder. - Scrub time. A scrub traverses the whole block tree. Move the tree off 7200 rpm media and the seek component of that traversal disappears.
By how much? Upstream says only that this is “usually a larger and more
reliable win than adding an L2ARC, because it is a permanent placement rather
than a cache that must warm up and that costs ARC headers.” There is no
first-party quantified comparison, and this guide will not invent one. The
measurement to take is wall-clock on a find and on a scrub, before and after.
That is the shape of the hybrid-pool argument: bulk capacity from hard drives at the lowest price per terabyte, with the entire random-access tax paid once, by a mirrored pair of small NVMe devices.
Sizing one
No credible fixed percentage exists. Upstream calls sizing “workload-dependent and hard to predict from first principles”, and the commonly quoted “0.3% of pool” has no first-party source anywhere in the OpenZFS documentation; treat it as folklore. Measure instead.
Measure: zdb -bbb tank -> block-type breakdown, metadata total
Two corrections to that output, both community-sourced, not documented:
PSIZE under-counts - it ignores redundant metadata copies (~2x)
ASIZE over-counts - it includes raidz parity overhead
Worked shape, not a prediction. Suppose zdb reports:
metadata, PSIZE 180 GiB
x2 for the redundant metadata copies 360 GiB
/ 0.75 for the 25% metadata reserve floor 480 GiB
+ small blocks, if you enable them + your own figure
+ headroom, because you cannot shrink it
-> a mirrored pair of 1 TB NVMe, not a pair of 480 GB
Upstream’s own advice ends the same way: leave “generous headroom, because the cost of guessing low is a vdev you cannot shrink.” Round up.
The part that has to be said plainly
“A special vdev must be at least as redundant as the pool’s normal vdevs.
Losing it loses the pool, exactly as losing any other top-level vdev would.
Never add one as a single device to a redundant pool.”
Read literally, that puts a three-way mirror under RAIDZ2 data and a four-way under RAIDZ3, though upstream states the principle rather than printing that mapping. The same power-loss and endurance reasoning that applies to arrays of NVMe drives applies here, with the added point that this vdev is not optional to the pool’s survival.
Removal is worse than you expect. zpool remove of a top-level vdev requires
no raidz or draid top-level vdev anywhere in the pool, the same ashift on
every top-level vdev, the keys for encrypted datasets loaded and the
device_removal feature enabled. Upstream is blunt about the consequence:
“on a raidz pool, a special vdev can never be removed.” Adding one is a
permanent decision. If zpool add demands -f because the special vdev is
less redundant than the data vdevs, that is the pool telling you not to.
The migration reality
“Existing blocks are never migrated, in either direction.” The special vdev
takes newly written data only. An existing pool gets nothing retroactively, the
same way recordsize “affects only files created afterward”. A scrub does
not move blocks. A resilver does not move blocks.
So retrofitting means rewriting: zfs send into a new dataset and swap, or
rewrite in place dataset by dataset, or accept that the benefit arrives
gradually as data turns over - which on an archive pool means never.
The commands
# add a mirrored special vdev, matched to the pool's redundancy
zpool add tank special mirror /dev/nvme0n1 /dev/nvme1n1
# metadata only; all user data stays on the spindles
zfs set special_small_blocks=0 tank
# opt one dataset into small-block placement
zfs set special_small_blocks=16K tank/src
# what is in there, and how close to the 75% cliff
zpool list -v tank
zdb -bbb tank
The ZFS SLOG, and what it is not
A separate log device is not a write cache. It absorbs nothing the way dm-writecache does, it holds no hot data, and on a pool whose workload issues no synchronous writes it changes nothing measurable. Far more SLOG devices have been bought on that misunderstanding than on any measurement.
Every pool already has a ZIL
From zpoolconcepts(7): “The ZFS Intent Log (ZIL) satisfies POSIX
requirements for synchronous transactions … By default, the intent log is
allocated from blocks within the main pool.” A log vdev does not add the
ZIL. It relocates it, off the spinning disks that are also serving pool data,
onto something with lower commit latency. It is a placement change, not a new
layer, and the thing being placed is only ever read once - after a crash, to
replay what had not yet reached a transaction group. In normal operation the
ZIL is written and never read back.
Asynchronous writes never touch it
This is the sentence that should settle most purchases. A write that is not
synchronous is aggregated in memory and issued with the next transaction
group; the ZIL is not consulted unless an fsync() lands on that file first,
which makes those writes synchronous. That window is bounded by
zfs_txg_timeout
(default 5 seconds) and by the dirty-data limits, zfs_dirty_data_max
defaulting to one tenth of physical RAM but capped by zfs_dirty_data_max_max
at min(RAM/4, 4 GiB). A large asynchronous write stream
gets nothing from a log device at all.
The workloads that genuinely issue synchronous writes are a short list: NFS exports, which issue them as a matter of course; databases committing transactions; and VM datastores configured for sync. Most home and media pools, most backup targets and most desktop use are asynchronous throughout. If that describes the pool, the money belongs in RAM or in a metadata special vdev, not here.
Find out before you buy
zilstat ships with OpenZFS from 2.2 onward (cmd/zilstat.in is present at
tag zfs-2.2.0 and absent from the 2.1 series). It reads the ZIL kstats -
global at /proc/spl/kstat/zfs/zil on Linux, per-dataset through the objset
kstats, with -a for every dataset, -p for a pool and -d for one
dataset.
# Run the real workload for an hour, then read the two fields that decide it:
grep -E 'zil_itx_metaslab_(normal|slog)_bytes' /proc/spl/kstat/zfs/zil
zil_itx_metaslab_normal_bytes ZIL traffic landing in the main pool
zil_itx_metaslab_slog_bytes ZIL traffic landing on a log vdev
# Both near zero under production load -> a SLOG will do nothing for you.
# normal_bytes large, no log vdev present -> a SLOG is the right purchase.
Do this before ordering anything. It costs an hour and it is the only honest answer to the question.
Sizing, which is far smaller than people expect
The device has to hold the synchronous writes that could still be lost, which is a couple of transaction groups’ worth and nothing more. The arithmetic, with the assumptions stated so you can re-run it with your own measured rate:
Sustained synchronous write rate (your figure) 500 MB/s
zfs_txg_timeout 5 s
Bytes accumulated before one commit 500 x 5 = 2.5 GB
Two transaction groups outstanding 2.5 x 2 = 5 GB
A saturated 10 GbE link carrying nothing but synchronous writes, call the
link 1.1 GB/s, lands near 11 GB by the same method: 1.1 x 5 x 2. The TrueNAS
hardware guide puts it at “a high-endurance, low-latency device between 8 GB
and 32 GB is adequate for most modern networks”, which matches the arithmetic
rather than contradicting it. Buying a 2 TB drive changes no figure above.
What surplus capacity is good for is over-provisioning: partition 8 to 32 GB,
leave the rest unpartitioned and never written after a secure erase or
blkdiscard, and the controller keeps it as spare area. The write
pattern here is small, relentless and synchronous, so the endurance column is
the one that decides the purchase: see
SSD endurance: TBW, DWPD, how worried to be.
Two dataset knobs interact with all of this: sync=standard|always|disabled
(default standard) and logbias=latency|throughput (default latency,
where throughput means ZFS will not use configured log devices to store
that dataset’s written data). The zil_slog_bulk default also moved
upstream, from 786432 bytes to 67108864 bytes, in a change merged in October
2023; writes above it are issued at asynchronous priority to stop a single
ZIL writer monopolising the device. Which of the two your own build ships
with depends on its vintage, so read it rather than assume it:
cat /sys/module/zfs/parameters/zil_slog_bulk.
Power-loss protection is the product, not a nicety
Elsewhere in this guide power-loss protection is a recommendation. Here it is
the entire function. The SLOG exists to make one promise - this write is
durable - and a drive that acknowledges a FLUSH before the data is safe
breaks precisely that promise, silently, and reveals it only during a power
cut. Published fsync measurements (Small Datum, fio with O_DIRECT, fsync
after a 16 KB write) put a Samsung 990 Pro near 2970 µs and a Crucial T500
near 890 µs, against roughly 12 µs for a Solidigm D7-P5520 and under 2 µs for
a Samsung PM9A3, each figure from a different host. A 7200 rpm drive’s specified
average rotational latency is
4.16 ms. A consumer NVMe SLOG at 3 ms is competing with the disks it was
bought to rescue, and losing on correctness as well. Published power-pull
testing has found consumer drives losing writes they had already
acknowledged. Filter for the protection rather than the brand:
enterprise SSDs, the power-loss-protection filter.
Losing one
On current OpenZFS a dead log device is an inconvenience. The pool continues
using its in-pool ZIL, zpool remove tank <logdev> takes it back out, and
synchronous write performance collapses back to where it was before you
bought anything. Data is lost only if the device dies and the system crashes
in the same moment, taking the sync writes it still held. Log devices can be
mirrored and are worth mirroring for exactly that overlap; raidz is not
supported for the intent log.
Very old implementations behaved far worse. Log device removal arrived with ZFS pool version 19; on a pool predating that, a missing log device could refuse to import and take the pool with it. That is the origin of the mirror-your-SLOG-or-lose-everything advice still repeated on forums. It was correct before log device removal landed and has been wrong ever since.
The rest of the Linux landscape: btrfs, bcachefs and the dead ends
The filesystem with the best native tiering design on Linux is not in Linux. bcachefs was removed from the mainline tree in the 6.18 merge window, and btrfs - the copy-on-write filesystem most readers already have - has never had a cache of its own and still does not. Everything below is either a layering exercise or an out-of-tree bet.
btrfs has no cache, so it borrows one
btrfs has no equivalent of ZFS’s L2ARC or special vdev. To put NVMe in front of a btrfs volume you build dm-cache, dm-writecache or bcache underneath it and make btrfs on the composite device. That works, and it costs three things.
The cache cannot see what btrfs knows. Below the filesystem there are no
file boundaries, no trees and no distinction between a metadata block and a
data block, except for the flags on the bio. One lever survives that gap: btrfs
tags metadata I/O REQ_META, and dm-writecache’s metadata_only option keys
off exactly that - the kernel documentation says it means “only metadata is
promoted to the cache” and that it “improves performance for heavier REQ_META
workloads”. That is the nearest a layered stack gets to a ZFS special vdev, and
it is a write path only, because dm-writecache “doesn’t cache reads because
reads are supposed to be cached in page cache in normal RAM”.
Multi-device btrfs multiplies the cache. dm-cache takes exactly one origin device. A btrfs raid1 across four hard drives is four independent cache stacks and four splits of the flash budget, or one stack over an md array - in which case btrfs sees a single device, so a raid1 data profile is not available at all, dup is the only duplication one device will carry, and a checksum failure becomes EIO rather than a repair from the sibling copy. The ordering decides whether btrfs keeps its best property. Cache each drive and mirror above; do not mirror below and cache once, unless you have accepted losing the repair path. The same reasoning applies to the array shapes in Arrays with NVMe drives.
Writeback removes the guarantee btrfs was bought for. Lose a writeback cache device and the btrfs trees reference extents whose contents never reached the disk. Checksums turn that into loud errors instead of silent corruption, which is better than the alternative and is not a recovery.
bcachefs: the right model, in the wrong place
bcachefs expresses tiering as targets rather than modes, which is a cleaner
design than anything else here. foreground_target takes new writes,
background_target is where reconcile moves them afterwards,
promote_target receives a cached copy when data is read, and
metadata_target places btree nodes. All four are unset by default.
Writeback caching is foreground and promote on the flash, background on the
disks; writearound is foreground on the disks, promote on the flash.
The trap is durability. Per the Principles of Operation, with
data_replicas=2 and foreground_target=ssd, foreground writes place both
durable copies on the SSDs, assuming the SSDs carry enough durability between
them, and reconcile later moves both to the disks - they are not split one per
tier. Set durability=0 on the cache devices and they become a pure cache
whose copies never count towards replication, after which “losing the cache
device never causes data loss”. There is also a per-device rotational option,
default false, which has reconcile additionally track pending work in device
LBA order so it can be processed sequentially rather than in key order; the
v1.33.0 notes say it cannot be changed once set and the current documentation
lists it as a runtime option, and this guide will not decide between them for
you.
Now the standing, with dates, because it has moved twice. Merged for Linux 6.7,
released 7 January 2024. MAINTAINERS changed to “Externally maintained” in
commit ebf2bfec412a on 28 August 2025, shipping in 6.17. The code was removed
in commit f2c61db29f27 on 29 September 2025, about 117,000 lines, with the
stated reason “It’s now a DKMS module, making the in-kernel code stale, so
remove it to avoid any version confusion.” Linux 6.18, released 30 November
2025, was the first mainline kernel without it, and fs/bcachefs was still
absent from mainline when this section was last checked, in September 2026.
Development continues out of tree - v1.39.6 is dated September 2026 - shipped
as DKMS from the project’s own packages. Its own Kconfig prompt still reads
EXPERIMENTAL. The project itself no longer calls it that: Kent
Overstreet took the label off the bcachefs site and made it official with the
v1.38.6 release in June 2026, writing “Consider this the belated official
announcement”. A separate v1.37.0 changelog line had already dropped the
experimental tag from erasure coding, which is a different claim about a
different thing.
One more reason old recipes fail: bcachefs-tools v1.33.0 (4 December 2025)
renamed rebalance to reconcile, gated behind an incompatible metadata upgrade.
rebalance_enabled is now reconcile_enabled. Any recipe written before that
date may use option names that no longer exist.
flashcache and EnhanceIO, named so you can recognise them
Search for SSD caching on Linux and the 2012 to 2015 forum threads still rank. They recommend flashcache, Facebook’s out-of-tree module, or EnhanceIO, the fork of it. Neither was ever merged. dm-cache arrived in 3.9 and dm-writecache in 4.18, and those, with bcache, are what a current kernel offers you. This guide will not print a last-known-good kernel version for either dead end, because there is no maintained tree against which to check one - and that absence is the whole finding. A page that tells you to clone a cache module and build it against your running kernel is telling you its code was never reviewed by the people who maintain the block layer.
The comparison that decides it
| Option | Caches reads | Caches writes | Survives cache loss | Reformat backing | Maintained |
|---|---|---|---|---|---|
| dm-cache writethrough | yes | no | data yes, volume not automatic | no | mainline |
| dm-cache writeback | yes | yes | no | no | mainline |
| dm-cache passthrough | no | no | yes | no | mainline |
| dm-writecache | no | yes | no | no | mainline |
| bcache writethrough | yes | no | yes | yes, superblock | mainline |
| bcache writearound | yes | no | yes | yes, superblock | mainline |
| bcache writeback | yes | yes | no | yes, superblock | mainline |
bcachefs, durability=0 |
yes | yes | yes | yes, mkfs | out of tree |
| bcachefs, durability 1+ | yes | yes | no | yes, mkfs | out of tree |
| ZFS L2ARC | yes | no | yes | no | yes |
| btrfs native | none | none | n/a | n/a | n/a |
| flashcache, EnhanceIO | - | - | - | - | no |
Every row that cannot survive cache loss makes the fast device the only copy of acknowledged data, which turns the device choice into a durability decision before it is a speed one: power-loss protection first, then a write budget you can defend against the figures in SSD endurance, and only then throughput. Those are the filters to apply when choosing the cache device itself.
Windows: Storage Spaces tiering
Storage Spaces tiering is not a cache. It moves data on a timer, and the timer is the whole story. A dm-cache promotion decision is made on the I/O that misses; a Storage Spaces placement decision is made by a scheduled task that runs at 1:00 a.m. Everything else in this section follows from that one difference, including the safeguards, which are milder than the Linux writeback stacks need precisely because nothing here acknowledges a write that exists in only one place.
The mechanism is a heat map. Storage Spaces records access counts through the day and a scheduled task moves data between an SSD tier and an HDD tier afterwards. Movement is sub-file: Microsoft’s own description is that “the data is mapped and moved at a sub-file level. So if only 30 percent of the data on a virtual hard disk is ‘hot’, only that 30 percent moves to your solid-state drives.” The slab size of that movement is widely quoted as 1 MB and appears in no Microsoft document this guide could find, so no figure is given here.
Building a tiered space
Two StorageTier objects, one per MediaType, then one virtual disk drawing
from both. Microsoft’s own New-Volume example:
New-Volume -StoragePoolFriendlyName "CompanyData" -FriendlyName "UserData" `
-AccessPath "M:" -ResiliencySettingName "Mirror" -ProvisioningType "Fixed" `
-StorageTiers (Get-StorageTier -FriendlyName "*SSD*"),
(Get-StorageTier -FriendlyName "*HDD*") `
-StorageTierSizes 20GB, 80GB -FileSystem NTFS
Three constraints bite at purchase time, not at deployment time.
Provisioning must be Fixed - Microsoft’s requirement is that a tiered
virtual disk uses fixed provisioning, so thin provisioning and heat tiering
are not a combination. Parity is not a documented tiering layout either: the
Microsoft material works through simple and mirror throughout, and vendor
guidance states the restriction outright. And the column count is identical
on both tiers, which sets the drive count arithmetic:
four-column two-way mirror = 4 columns x 2 copies = 8 devices per tier
= 8 SSDs AND 8 HDDs
two-column two-way mirror = 2 x 2 = 4 devices per tier
= 4 SSDs and 4 HDDs
The four-column line is Microsoft’s own worked example, and it is the reason a tiered space is a poor fit for the single-NVMe-in-front-of-a-pair-of-HDDs shape this guide opened with. You do not buy one fast device; you buy as many fast devices as you have slow ones, in matched sets, which is a different budget and a different shopping list of SSDs entirely. The cmdlets are present on client Windows and a tiered space can be built there from PowerShell, but the documentation, the reports and the worked examples are all written for Server, and Microsoft describes no supported client scenario.
Resizing the fast tier later has an ordering trap - tier, then disk, then partition:
$SSDtier = Get-VirtualDisk -FriendlyName $Name | Get-StorageTier -MediaType SSD
Get-StorageTierSupportedSize -InputObject $SSDtier # TierSizeMin/Max/Divisor
Resize-StorageTier -InputObject $SSDtier -Size 34GB
Get-VirtualDisk $Name | Get-Disk | Update-Disk
Resize-Partition -DriveLetter D -Size 65.75GB
Tier sizes snap to a TierSizeDivisor that is configuration-dependent;
Microsoft’s sample output shows 2147483648 bytes, 2 GiB, for that particular
space. Read it from the cmdlet rather than assuming it. The 65.75 GB above is
not a typo: Microsoft recommends leaving 256 MB unallocated for GPT, so a
66 GB space gets a 65.75 GB partition.
The optimisation task, and the lag it imposes
Task path: \Microsoft\Windows\Storage Tiers Management\
Storage Tiers Optimization
Default trigger: 1:00 a.m., daily
Get-ScheduledTask -TaskName "Storage Tiers Optimization" | Start-ScheduledTask
Optimize-Volume -DriveLetter H -TierOptimize
defrag H: /g /h /# is the same thing from cmd: /g optimises the storage
tiers on the volume, /h runs the job at normal rather than the default low
priority, and the command returns a Post Defragmentation Report and a Storage
Tier Optimization Report. Optimize-Volume with no parameters picks its
operation by drive type, and for a tiered storage space that default is
-TierOptimize, described as placing “file data on the optimal storage tier
according to heat or desired placement.”
Run it more often by changing the task’s repetition interval, six hours being Microsoft’s stated recommended maximum frequency: “There’s nothing to gain by running the task any more frequently.”
$Task = "\Microsoft\Windows\Storage Tiers Management\Storage Tiers Optimization"
schtasks /change /tn $Task /ri 360
Now do the arithmetic on what that buys:
Default schedule, one run at 01:00:
working set turns over at 09:00 -> served from HDD until 01:00 next day
lag in this example = 16 hours (worst case nears 24)
Every 6 hours (/ri 360):
worst-case lag = 6 hours, mean ~3 hours
dm-cache smq, bcache:
promotion considered on the missing I/O itself
A working set that changes during the working day is served from spinning disks for the rest of that working day. That is not a defect to tune out; it is the architecture. Heat tiering suits a footprint that is stable over days - a VDI parent image, a departmental share, a media library with a reliable hot corner. It does not suit a build agent, a database whose index hot spots move with the month-end, or anything where yesterday’s heat map is a poor prediction of this afternoon.
The report earns its keep here. It gives the percentage of I/O currently served by the SSD tier, the tier size that would be needed to hit a target percentage, and the share of tier I/O attributable to pinned files. Expected far above actual means optimise more often. That table is a measured miss-ratio curve, which is more than the Linux stacks hand you.
Pinning
Set-FileStorageTier -FilePath "D:\VMs\Parent.vhdx" `
-DesiredStorageTierFriendlyName "Space01_SSDtier"
Get-FileStorageTier -VolumeDriveLetter D # PlacementStatus, State
Clear-FileStorageTier -FilePath "D:\VMs\Parent.vhdx"
Three behaviours to hold in mind. Pinning does not move anything now -
Storage Spaces “will attempt to move the pinned files to the desired volumes
during its next Storage Tiers Optimization run”, so force the task if you
want it today. The whole file moves, never part of it. And a pinned file is
excluded from heat optimisation thereafter, so its capacity leaves the pool
of space the optimiser has to work with. Microsoft’s own uses are the parent
VHDX in pooled VDI pinned to SSD, and streaming-media VHDs pinned to the HDD
tier to keep sequential traffic off the flash - the same instinct as bcache’s
4 MB sequential_cutoff, expressed by hand.
The write-back cache
Layered under all of this is a genuine cache, and a small one. Default size
1 GB, created automatically when the pool holds enough devices with
MediaType = SSD or Usage = Journal: one for simple, two for two-way mirror
and single parity, three for three-way mirror and dual parity. Short of that
it is set to 0, except on parity spaces where it is set to 32 MB. The pool
property WriteCacheSizeDefault is Auto; New-Volume -WriteCacheSize
overrides per volume, and the cmdlet refuses combinations that would force a
cache into a configuration where it would be slower.
Writes larger than 256 KB are not written to the cache. That single
sentence from Microsoft’s storage-tiers monitoring page explains most confused
benchmark results on tiered spaces: the write-back cache absorbs small random writes and
large sequential writes go straight to the HDD tier by design. The diagnostic
follows directly - SSD-tier and HDD-tier write latency should look similar,
because the cache is absorbing HDD-tier writes. If HDD write latency is
higher, either the cache is full (check Bytes Used on the Cache object) or
the writes are bypassing it (compare Cache Writes/sec against
Tier Writes/sec on the Tier object). There is no read-cache equivalent:
-ReadCacheSize is documented as “no longer supported.”
Microsoft describes the WBC as tolerant of power failures. That guarantee rests on the devices honouring FLUSH, which is a property of the drive and not of Windows, so enterprise SSDs with power-loss protection remain the right class of part for the fast tier even here.
Device failure in a tiered space
This is where tiering pays back what the timer cost. Data lives on one tier or the other, both tiers are built from the same virtual disk’s resiliency setting, and a device failure in either tier is an ordinary Storage Spaces device failure handled by the mirror. Losing an SSD does not lose data that exists nowhere else, because nothing here holds data that exists nowhere else beyond the 1 GB write-back cache, and that cache inherits the space’s resiliency too. Compare a dm-cache in writeback mode, where the fast device holds the only copy of acknowledged writes and its loss is a restore from backup. Tiering trades responsiveness for a failure domain you already own.
The corollary is that the fast tier must be as redundant as the slow one. A two-way mirror needs two SSDs per column on both sides; you cannot save money with a single fast device and expect the space to survive it, and the column rule will not let you build it that way in the first place.
Storage Spaces Direct, in server editions
S2D is a different engine wearing similar cmdlets, and it is a real cache rather than a tier. Where the drive types differ, cache devices are chosen automatically by the hierarchy PMem, NVMe, SSD, HDD; every device of the fastest type present becomes cache and contributes zero usable capacity. With NVMe or SSD in front of HDD the cache mode is read and write; with flash in front of flash it is write-only. It sits below the rest of the stack, at drive level - Microsoft’s framing is that it creates “‘hybrid’ (part flash, part disk) drives which are then presented to the operating system” - and it derandomises writes before destaging “to emulate an IO pattern to disk that seems sequential even when the actual I/O coming from the workload (such as virtual machines) is random.” Microsoft puts the resulting hybrid latency at “often ~10x better”, which is a vendor figure and the only one of its kind found for this guide.
Get-ClusterStorageSpacesDirect # CacheModeHDD : ReadWrite
Set-ClusterStorageSpacesDirect -CacheModeSSD ReadWrite
Get-PhysicalDisk # cache drives show Usage = Journal
Sizing is measured, not guessed: the Cluster Storage Hybrid Disk PerfMon
object exposes Cache Miss Reads/sec, one instance per capacity drive,
compared against total read IOPS. Microsoft requires at least two cache
drives per server to preserve performance and recommends a capacity-drive
count that is a multiple of the cache count. Cache-drive loss costs only
un-destaged writes on that one server, and the other copies survive on the
others.
Two honest boundaries. S2D is a clustered, multi-server product: hybrid configurations are not supported in a single-server configuration, and all-HDD deployments are not supported at all. A home reader is not deploying it.
The standalone descendant they might deploy is the Windows Server 2022
storage bus cache, built on the same storage bus layer. It wants exactly two
media types, one of which must be HDD, and the Failover Clustering feature
installed on a server that is not itself in a cluster. Enable-StorageBusCache
creates the pool, binds fast media to slow and turns the cache on;
Get-StorageBusCache shows the settings, none of which can be changed once it
is enabled. Read its resiliency table before buying, because this is the one
Windows feature that comes closest to a read-and-write cache in front of hard drives
without a cluster, and only on one side of that table: a simple space with no
resiliency gets read and write caching, while the resilient option,
mirror-accelerated parity, is listed as read caching only, though Microsoft adds
that the mirror tier still provides write caching in the default Shared
provision mode. By default the cache takes
15 per cent of the fast media (SharedCachePercent, settable from 5 to 90
before enabling, and Microsoft advises against going above 50 with
mirror-accelerated parity). It is server-edition territory rather than a
Windows 11 feature.
Windows, the rest: Intel’s caching stack, ReadyBoost and PrimoCache
Intel shipped two consumer answers to this question and has discontinued both, and the driver versions that ended them are published numbers. Smart Response Technology stopped at Intel RST 15.9. Optane Memory support was dropped from Intel RST 19.0 onward. Nothing has replaced Optane Memory. On a desktop Windows machine the maintained choices are a Storage Spaces configuration Microsoft documents for Server editions, or a paid third-party filter driver.
Intel Smart Response Technology
SRT was the SATA-era equivalent of dm-cache: a small SSD placed in front of a hard disk or a RAID volume, caching at the LBA level rather than per file. Third-party documentation is consistent on the LBA point and Intel’s own pages do not spell it out, so treat the mechanism as well attested rather than first-party.
The setup constraints matter to anyone reviving an old board. The SATA controller has to be in RAID On mode in firmware, and moving an installed Windows between AHCI and RAID On without preparing the driver first is the classic route to an unbootable machine. The cache device needs at least 18.6 GB of unallocated space, and at most 64 GB of it is ever used - surplus capacity on a larger SSD is presented as a separate independent disk. Two modes were offered: Enhanced, which is write-through, and Maximized, which is write-back and carries the failure semantics any write-back cache carries.
Intel’s support article 000059331, last reviewed 12 August 2021, is unambiguous: “Intel will no longer support Intel Smart Response Technology which was used on legacy platforms for acceleration.” It names the boundary exactly - “The latest Intel RST version that supports Intel SRT is Intel RST 15.9. There will be no further releases or updates to the product” - and adds the sentence that catches people out: “Intel SRT feature capabilities are not supported with Intel RST 16.0 or later driver. Once the Intel SRT volume is disabled, the user can no longer enable it with an Intel RST 16.x driver or later.”
Disabling an SRT volume under a 16.x or newer driver is a one-way door. If you inherit a machine with a live SRT volume and a modern driver stack, decide before you touch it, because there is no re-enable path afterwards.
The stated replacement was Optane Memory, “starting with Intel RST version 16.x”. That replacement is now itself dead, which tells you most of what you need about planning around either.
Optane Memory, and what a second-hand module actually buys
The product dates, from Intel’s own end-of-life notices:
| Product | End of life | End of interactive support |
|---|---|---|
| Optane Memory Series | 5 August 2019 | 29 May 2024 |
| Optane Memory M10 Series | 13 January 2021 | 15 March 2024 |
| Optane Memory H10 / H20 | 8 September 2022 | not stated |
After the interactive-support dates Intel stops ticket, phone, chat, forum and email support; the five-year warranty from date of purchase is unaffected, and by now that has run out for most units in circulation. Intel wound down the whole Optane business with its Q2 2022 results, taking a 559 million dollar impairment charge - a figure from trade coverage of that release rather than from an Intel page read for this guide.
The driver side is separate and more decisive. Intel support article 000088727, last reviewed 12 March 2025, states that from RST driver 19.0 and greater the Optane Memory Series and M10 Series are no longer supported, that the H10 and H20 SKUs are discontinued, and that “The Pinning feature will no longer be supported” - previously pinned data keeps being accelerated by the ordinary algorithm only. The same article notes the RST driver supports 11th and 12th generation platforms only on a VMD-capable BIOS.
So, plainly. Buying an Optane Memory Series or M10 module today gets you an M.2 NVMe device whose media is 3D XPoint, with the latency floor that implies. The H10 and H20 are a different animal, as Intel’s own product names admit - “Intel Optane Memory H10 with Solid State Storage” and “Intel Optane Memory H20 with Solid State Storage” are combined modules, not pure 3D XPoint devices. Neither kind of module gets you Windows caching: the driver that drove them has dropped them, the management application it shipped with is end-of-life, and the platform requirements are frozen at a generation that is itself now old. It does get you a perfectly ordinary NVMe namespace, which dm-cache, dm-writecache or bcache will accept like any other - a legitimate use for a module bought cheaply, bounded by the tens of gigabytes the Series and M10 modules carry. Set against what the same money buys in current NVMe stock, the case is thin unless the QD1 latency floor is specifically the thing you want and the small capacity is enough for your metadata working set.
ReadyBoost is not an answer to this question
ReadyBoost is a read cache on removable flash, driven by the SysMain service, designed for machines short of RAM. It does not accelerate writes. Its backing store is a USB stick or a memory card, slower than every device discussed in this guide, including the SATA SSDs. Its premise - that RAM is the scarce resource and flash is the abundant one - has been false for a decade. Microsoft is reported to have removed the management UI in Windows 11 22H2; no Microsoft announcement confirming that version was found, so take the version with caution and the direction as settled. It is a footnote, not an option.
PrimoCache, and where a paid tool earns its money
PrimoCache from Romex Software is the serious third-party answer on Windows. Version 4.4.1 dates from 26 January 2025. It is a two-level cache: Level 1 is RAM, including the invisible memory Romex describes as usually existing on 32-bit Windows with 4 GB or more installed, and Level 2 is an SSD or NVMe device, whose contents the vendor documents as persistent across restarts. Cache block size runs from 4 KB to 512 KB, with the vendor’s guidance being a value equal to or less than the file system’s cluster size; smaller blocks raise hit rate for the same capacity and cost more memory and CPU. Level-1 cache space is allocated either unified across reads and writes, or split by a ratio. A Volatile Cache Contents setting clears the level-2 cache at every restart instead of keeping it, and that is the right choice on a multi-boot machine, where another OS can change the volume underneath a stale cache.
Defer-Write is the feature that earns the money and the feature that carries the risk. Its latency is the interval, in seconds, at which deferred data is flushed to disk, and a latency of 0 disables the feature outright. Five modes are documented - Native, Intelligent, Idle-Flush, Buffer and Average - and they differ in what triggers a flush ahead of that interval: a proportion of the cache filling up, Windows going idle, or a smoothed trickle rather than a burst. Holding writes back that way lets repeated writes to the same block collapse into one and lets a scattered stream leave the cache as a batch, which is exactly the derandomisation a mechanical array wants.
Romex states the risk itself - “because of the risk that data may be lost on system crashes or power failures, usually latency is set to a low value” - and recommends enabling Defer-Write only on logical disks that store “temporary, unimportant, or reproducible data, such as download disks, temporary file disks, and logical disks with good backup mechanisms”. The vendor’s forum is blunter, and this is the sentence to read twice: deferred writes can “write out of order data to disk, which is a huge difference from filesystems that write data in order.”
The deferred-write window does not merely lose the last N seconds of writes - it loses their ordering. A journalling file system’s consistency guarantee rests on barriers being honoured in sequence. Break the sequence and what survives a power cut is not an older consistent state, it is an arbitrary mix of old and new blocks. NTFS metadata recovery assumes the journal reached the platter before the data it describes did. Size your Defer-Write latency as the window of arbitrary corruption you are willing to accept, and put a UPS behind it - Romex’s own page says the same, recommending Defer-Write for disks with good backup mechanisms and noting that a stable system with a UPS configured can also be considered.
Where does a paid tool earn its keep? Precisely here: Windows client editions have no supported HDD-plus-NVMe caching mechanism at all. Heat-based Storage Spaces tiering is PowerShell-only and undocumented for client, storage bus cache is documented for Windows Server 2022 and needs the Failover Clustering feature installed on a server that is not itself in a cluster, and Intel’s consumer answers are both retired. PrimoCache is the best known maintained product filling that hole, it has a 30-day full-function trial that costs nothing to evaluate against your own counters, and its RAM tier does something no block cache on the NVMe can. Its headline marketing claim of benchmark scores “increased more than 70 times in sequential read/write” on a mechanical drive is a RAM-cache artefact measured against a cache large enough to hold the benchmark file, and should be read as such - see SSD endurance for the other side of that ledger, because every deferred write eventually lands on flash.
NAS appliances: Synology, QNAP, unRAID, TrueNAS
On an appliance you do not choose a caching architecture. You choose a vendor, and the architecture comes with it. Three of the four systems below are sold in the same shop, in boxes that look alike, and not everything they call a cache is one in the sense this guide has used the word. One vendor also ships a tiering engine that moves the only copy of your data onto flash. One ships a landing zone with a scheduled mover and no promotion path at all. Getting that distinction wrong is the mistake that no amount of tuning fixes.
Synology: two caches, and only one of them needs a mirror
Synology’s own white paper is unusually precise, and it is the only source for most of this. A read-only cache can be built from a single SSD; a read-write cache “must be created with at least two SSDs of the same type (i.e., NVMe or SATA)”. The permitted RAID types differ accordingly: Basic, RAID 0, RAID 1 or RAID 10 for read-only, and RAID 1, RAID 5, RAID 6 or RAID 10 for read-write. The mirror is not advice. It is the product refusing to put acknowledged writes on a single device, which is the conclusion the dm-cache and bcache sections arrive at by a longer route.
Write handling is write-back with LRU replacement, and caching is at block level: reading 400 KB out of a 4 GB file promotes 400 KB. On RAID degradation an “automatic protection mechanism” halts writes to the cache, routes new writes straight to the drives, and drains what is already there in sequential order.
The sizing guidance is the honest kind. “The recommended SSD cache size should be just enough to cover the size of frequently accessed data”, and for 100 GB of hot data, raising the cache from roughly 100 GB to 500 GB “will not result in significant performance increase. Excess cache space will only be used to store cold data.” Storage Manager’s SSD Cache Advisor measures the hot set for you: at least 7 days of analysis, automatically ending after 30, reporting the average rather than the peak day.
The RAM cost is where Synology differs sharply from ZFS, and the arithmetic is worth doing:
Synology mapping table: ~400 KB of RAM per 1 GB of SSD cache
1 TB of cache: 1024 x 400 KB = 409,600 KB ~ 400 MB
DSM spends at most 25% of installed RAM on that table, so
1 GB of RAM allows: 1,048,576 KB / 400 KB per GB ~ 2,600 GB of cache
a 4 GB NAS allows: (4 GB / 4) = 1 GB of table -> ~2.6 TB of cache
ZFS L2ARC, for contrast: 96 B per resident record; 1 TiB of cache
at 16 KiB physical blocks = 67.1 M records ~ 6 GiB of ARC
On Synology, RAM almost never limits the cache. On ZFS it usually does. Same job, header costs an order of magnitude apart, because the block sizes are. The exception Synology names is Alpine-CPU models, capped at 930 GB of cache in total.
Two more first-party details that matter. Enabling “pin all metadata” keeps the Btrfs metadata of the mounted volume on the cache - the appliance version of a ZFS special vdev, without the permanence. And Synology states plainly where the cache will do nothing: “SSD cache will not improve performance in scenarios involving sequential access patterns”, naming single-channel HD video streaming and surveillance recording. A camera NAS does not want a cache.
QNAP: Qtier is tiering, and tiering is not caching
QNAP ships both, and its documentation draws the line for you. With SSD cache, “QTS copies data to the cache as it is accessed”. With Qtier, “QTS writes incoming data to the SSD tier and moves data to different tiers based on access frequency”. Copies against moves is the whole difference:
- Qtier’s SSD capacity is usable capacity, added to the pool. QNAP’s comparison lists Qtier storage space as HDDs plus SSDs, and SSD cache as HDDs only.
- Because the SSD tier holds the only copy of what lives there, it must carry its own redundancy. A Qtier SSD is a pool member, with a pool member’s consequences on failure.
- Migration is scheduled or runs during low-load periods, rather than continuously on access. Qtier reacts in hours; a cache reacts in milliseconds.
- QTS caps SSD cache by installed memory and places no such cap on Qtier. The cache acceleration requirements for QTS 5.0.0 and later list 2 GB of RAM for 1 TB of cache and 8 GB for the documented maximum, 8 TB on ARM models and 16 TB on x86 ones; the 1 TB ARM ceiling belongs to QTS 4.5.x and earlier, and QNAP’s Qtier comparison table still prints an older 4 TB figure.
- One SSD cannot do both: an SSD already used for Qtier is unavailable to SSD cache.
Qtier can be turned on in an existing storage pool - QTS documents “Enabling Qtier in an Existing Storage Pool” as a supported operation, which is more than Windows heat tiering allows.
unRAID: the cache pool is not a cache
An unRAID pool is a landing zone with a scheduled mover, and this trips up more people than any other item in this guide. New files are written to the pool; the mover later transfers them to the array on the schedule you set in Settings, Scheduler. Nothing is left behind on the flash, and reading a file that already lives on the array never promotes it. There is no hit ratio to measure because there are no hits. What you get is write absorption and a place to keep things that should never migrate - appdata, Docker images, VM disks.
Two consequences. Array parity does not extend to pools, so anything sitting on a single-device pool between mover runs exists in exactly one place; use at least two devices if the window matters. And since 6.12 the share settings are Primary storage, Secondary storage and Mover action, replacing the old Use Cache values Yes, No, Only and Prefer. Any tutorial still using those four words predates the rename.
TrueNAS: ZFS with a management layer
Everything in the ZFS material applies unchanged; the UI renames vdev roles and adds guard rails. Cache (L2ARC) and log (SLOG) vdevs can be attached and detached later whatever the pool topology, because ZFS permits it. A special vdev cannot be removed while raidz data vdevs are present, which makes it a permanent decision taken through a web form. TrueNAS’s own hardware guidance asks for roughly 1 GB of RAM per 50 GB of L2ARC, and its documentation states that releases before 24.04 shipped with persistent L2ARC disabled by default, because the rebuild slowed restarts on large pools; 24.04 and later enable it. That is vendor policy rather than upstream behaviour, so check the default on the release you are running.
The question to ask before you buy
Ask whether the cache can be added, and removed, without destroying the array. The answers are not alike. Synology and QNAP SSD caches attach to an existing volume and detach again. Qtier enables in place. ZFS cache and log vdevs come and go; a special vdev on a raidz pool never leaves. Windows heat tiering is the outlier - a tiered virtual disk has to be created as one, with fixed provisioning and equal column counts on both tiers, so retrofitting means rebuilding.
Then ask what the cache costs you in bays. A read-write Synology cache consumes two M.2 slots or two drive bays, permanently, and mirrored pairs are the whole reason to check the slot count before the drive price. Power-loss protection matters more here than on a desktop, because an appliance write-back cache acknowledges data that exists nowhere else: see enterprise SSDs for the parts that have it, and NVMe in M.2 for what will physically fit. The sizing rules above tell you what capacity to aim for, and what differs from a desktop covers the drives underneath. Beyond that, the vendors document the behaviour rather than the implementation, so their documentation is what these rules rest on.
Caching on the RAID controller
A small business almost never decides to deploy a RAM cache. It buys a server, the server arrives with a RAID card, and the card carries a gigabyte or four of DDR4 running in write-back mode behind a battery or a bank of supercapacitors. That is the RAM cache. On a mechanical array it is the largest single performance lever anyone in the building has, it costs nothing extra, and nobody looks at it until it stops.
A dead cache battery is one of the commonest causes of an unexplained slowdown in an SMB server. The array does not report an error anyone reads. No disk fails. The CPU looks busier, because it is sitting in iowait. The server “went slow overnight” and stayed slow, and the ticket gets closed as “ageing hardware”.
Half the entry-level cards have no cache at all
Before diagnosing a cache, confirm one exists. Broadcom’s 94xx MegaRAID and HBA Tri-Mode Storage Adapters User Guide (pub-005851) is explicit: “The MegaRAID 9440-8i Tri-Mode storage adapter and the HBAs do not support CacheVault data protection.”
| Adapter | On-board cache | Energy backup |
|---|---|---|
| MegaRAID 9460-16i | 4 GB DDR4-2133 | CVPM05 |
| MegaRAID 9480-8i8e | 4 GB DDR4-2133 | CVPM05 |
| MegaRAID 9460-8i | 2 GB DDR4-2133 | CVPM05 |
| MegaRAID 9440-8i | none | not supported |
The pattern repeats across vendors: the cheapest tier of a family is usually cacheless, and Broadcom positions the 9440-8i for “non-business critical applications”. Nothing in this section applies to those cards. In the current 9600 generation (24G SAS, PCIe Gen4), only the 9660 and 9670 RAID members take energy backup at all, via CVPM05 / FBU345, part 05-50039-00, per the 9600 Series 24G PCIe 4.0 Tri-Mode product brief, 9600-Series-PB111, dated 16 September 2024. That brief does not print the DRAM size in its comparison table, and this guide will not quote one from a review site as if it were Broadcom’s figure.
What the capacitor protects, and what it does not
Broadcom’s CVPM02, CVPM05 Power Modules, CVFM04 Cache Module User Guide (CVPM02-05-CVFM04-UG101, version 1.1, 8 June 2021) states the risk and the remedy in its own words. The risk: “the cached data can be lost if the AC power fails before the data is written to the storage device.” The remedy is “an ultracapacitor-based energy source that powers the critical adapter components in the event of a power loss, long enough to offload cached data to the nonvolatile flash memory on the CVFM.”
Read that mechanism carefully, because it decides the honest answer to “how long is my data safe”. A CacheVault is not a scheme for holding DRAM alive on a timer. It is a one-shot dump of DRAM to NAND, after which retention is flash retention. Broadcom publishes no hold-up time and no retention figure in that guide, and this guide will not invent one.
What the same guide does publish is the environmental envelope, and that is what actually kills these modules: operating temperature 0 °C to 55 °C ambient, storage minus 40 °C to 70 °C, humidity 5 % to 95 % non-condensing, with repeated instructions to “design your system with an airflow that allows the CacheVault device to operate within the operating temperature range”. A remote-mounted supercapacitor cable-tied into the dead air of a tower case is the classic SMB failure, and heat is what shortens its life. HPE centralises the same function in the Smart Storage Battery, one lithium-ion pack backing several controllers. Microchip integrates it on most SmartRAID 3200 adapters as an onboard capacitor module (ASCM) with a stated five-year lifetime.
The write-through collapse
Write-back acknowledges a host write once it is in controller DRAM. Write-through acknowledges only once the write has reached the media. On a RAID 5 or RAID 6 array of mechanical disks, a small random write then costs a full read-modify-write across spindles at seek latency. That is the whole collapse: same hardware, same software, one policy bit, and small random writes drop to disk latency.
Two mechanisms flip that bit without anyone touching it.
The first is the periodic relearn, on cards old enough to lack a transparent one. Dell documents a Transparent Learn Cycle that runs “once every 90 days”, and says plainly that it “causes no impact to the system or controller performance” and that “virtual disks stay in Write-Back mode, if enabled, during transparent learn cycle” (Dell KB 000141687). TLC arrived with PERC 8. On PERC H700 and earlier, Dell states, “virtual disks automatically switch to Write-Through mode when the battery charge is low because of a learn cycle” - and the same is true of a plain LSI card with a real battery under the default policy that disables write cache on a bad BBU. The relearn is a discharge and a recharge, so the degraded window is hours. A 60-day relearn interval is widely repeated for older LSI cards; it appears in third-party writeups only, and is not Broadcom’s published figure.
The second is a battery or capacitor that has genuinely failed, at which point the array stays in write-through permanently.
Check it in one command:
storcli /c0/cv show status # CacheVault: state, temperature, design capacity
storcli /c0/cv show all
storcli /c0/bbu show status # older cards with a real battery
storcli /c0/v0 show all # the virtual drive's current cache policy
Dell ships its own build of the same tool as perccli; HPE’s tool is ssacli,
whose flag
spellings for cache ratio and no-battery write cache circulate widely in
community posts and should be checked against the SR Gen10 configuration guide
before being typed into a production server.
The fix that circulates online is not a fix. Always Write Back means
write-back with no working energy backup, and the MegaCli recipe passed around
in forums for forcing write-back on a failed BBU sets exactly that. It restores
the speed by removing the protection, and converts a temporary latency problem
into permanent silent exposure to a power cut. Replace the module.
One more state worth knowing before it appears at two in the morning: preserved, or pinned, cache. Dell’s PERC 9 guide states that if a virtual disk goes offline or is deleted because of missing physical disks, “the controller preserves the dirty cache from the virtual disk”. That preserved cache blocks the creation of new virtual disks until it is imported or discarded, and Dell warns that any foreign configuration should be imported first, “otherwise, you might lose data.”
The third RAM tier nobody accounts for
Below the controller sits a third volatile cache: the 64 MB to 512 MB of DRAM on
each mechanical drive. Whether it is used is a controller policy, and the
defaults are asymmetric. Dell’s OpenManage documentation gives the default Disk
Cache Policy as Enabled for virtual disks built on SATA drives and Disabled for
those built on SAS drives. The reasoning is durability, not speed - that cache
is not battery backed, so a power cut loses whatever sits in it. StorCLI exposes
it as pdcache=on|off|default.
So a mechanical array has three RAM tiers: host memory, controller DDR, drive DRAM. Only the middle one is both fast and safe. The other two will acknowledge a write they cannot guarantee. It is one of the sharper differences between a consumer disk and the SAS parts covered in enterprise drives, where the platform expects the host to make that decision explicitly.
SSD caching as a controller feature, with dates
Every card vendor also sold a flash-caching add-on, and a second-hand buyer will be offered one. Their status is the perishable part.
| Feature | Modes | Licence | Status |
|---|---|---|---|
| MegaRAID CacheCade 1.1 | read only | key or electronic licence | superseded |
| CacheCade Pro 2.0 | read and write | key or electronic licence | absent from current docs |
| Microchip maxCache 4.0 | read and write | not stated as chargeable | current, SmartRAID 3100 and 325x |
| Dell PERC CacheCade | read and write | PERC 9 era | absent from PERC 11 docs |
| HPE SmartCache | read and write | per-server licence | unresolved |
CacheCade’s original FAQ (LSI doc 12350215, April 2010) capped the pool at
512 GB; Pro 2.0 adds read-write caching and a 32-SSD limit per controller.
Broadcom has published no dated end-of-life notice for it that this guide could
find. The evidence is negative and should be read as negative: pub-005851
contains zero occurrences of “CacheCade”, and the September 2024 9600-series
brief lists the full RAID feature set while mentioning neither CacheCade nor
FastPath. The absence is the evidence. StorCLI still accepts
add vd cachecade, because the 12Gb/s StorCLI binary spans several controller
generations, while the 9600 series is driven by a separate StorCLI2, and a
command surviving in a tool is not a statement that the firmware behind it does.
maxCache 4.0 is the one with current, dated documentation: Microchip white paper DS90003224B (2024) gives up to 32 cache pools, an aggregate ceiling of 6.8 TB, write-through on a non-redundant RAID 0 pool, and write-back only on a redundant cache pool, which the paper gives as RAID 1 or RAID 5 - because in write-back the pool briefly holds the only copy. Its published figures are 8× read and 36× write IOPS at a 100 % hit rate against a twenty-drive 7200 rpm array, falling to roughly 2× on both at 50 %. That collapse between the two rows is the real lesson. Write-back pools should be built from enterprise SSDs, whose endurance ratings are covered in SSD endurance.
On Dell, the generational break matters more than the feature: CacheCade appears in the PERC 9 documentation, where its presence “disables FastPath for all eligible HDD virtual disks”, and is absent from the PERC 11 controller-cache documentation entirely. Dell’s own community answers contradict each other on H730 against H730P support, so this guide will not claim one. HPE SmartCache is documented for Gen10 Smart Array controllers, licensed per server, and disables a long list of other features while enabled - I/O performance mode, cache ratio selection, write cache bypass threshold, no-battery write cache, drive write cache control and more. Whether the licence is still sold now that current ProLiants ship MegaRAID-based MR controllers could not be established from an HPE document, and this guide will not claim a status it cannot date.
Two caches that cannot see each other
A controller cache under a host filesystem cache is two caches with no shared
state. For reads that is mostly wasted: 4 GB of controller DDR under 64 GB of
page cache adds almost nothing to the hit rate, because the host already holds
everything the controller would have held. Leave MegaRAID’s iopolicy at its
direct default rather than cached, which double-buffers reads the host has
already cached, and skew HPE’s cache ratio towards write.
Keep the controller cache for writes and the host RAM for reads. The
controller’s entire value on a parity array is coalescing writes and absorbing
the read-modify-write penalty, and it is the only one of the two that can
acknowledge a write safely, because it has the capacitor and the NAND behind it.
Host page cache is free and safe for reads, free and unsafe for writes:
fsync() still has to reach stable storage, which is why a 4 GB backed
controller cache is worth more to a database than 64 GB of extra page cache.
The exception is a filesystem that implements its own durable write acknowledgement. A copy-on-write filesystem with its own intent log already does what the battery does, and a write-back controller cache underneath it inserts a reordering layer it cannot see or verify. The defensible split is the conventional one: HBA or eHBA mode under such a filesystem, write-back-with-battery under NTFS, ext4 or XFS on hardware RAID.
Hyperconverged and virtualised clusters
Four platforms, one rule, and it inverts everything a workstation build teaches. In every clustered system here the DRAM cache is node-local, not shared between nodes, and read-only. Microsoft says it in one sentence in Use the CSV in-memory read cache: “Writes cannot be cached in memory.” vSAN’s client cache reads only. The Nutanix Bible calls the Unified Cache “a read cache which is used for data, metadata and deduplication and stored in the CVM’s memory”. The durable write buffer is always a device - NVMe, SSD, NVDIMM - and the reason is not squeamishness about volatile media. A write acknowledged from the memory of one node is a write that dies with that node, and a cluster exists precisely because a node is a thing you have already decided you can lose.
The sizing question therefore changes shape. On a desktop the failure domain is a drive. Here it is a host, and the arithmetic that matters is how much RAM each node spends on caching, multiplied by the node count, and multiplied again by the nodes you are paying to be able to lose.
| Platform | RAM cache | Scope | Caches writes |
|---|---|---|---|
| vSAN (since 6.2) | client cache, 1 GB per host cap | host running the VM | no |
| Horizon CBRC | 100 MB to 2,048 MB, default 1,024 MB | ESXi host | no |
| S2D / Azure Local | CSV in-memory read cache | server-local | no, stated outright |
| Nutanix AOS | Unified Cache in CVM memory | node | no |
vSAN: the 70/30 split, and the disk group that dies with its cache
For hybrid Original Storage Architecture, Characteristics of a vSAN Cluster (still present in the VMware Cloud Foundation 9.0 documentation) is exact: “vSAN allocates 70 percent of all available cache for read cache and 30 percent of available cache for the write buffer.” All-flash OSA inverts that - the cache tier is write buffer only, and “all read requests come directly from the flash pool capacity”.
Hybrid sizing comes from Design Considerations for Flash Caching Devices in vSAN, where the rule sits under a heading scoped to hybrid configurations: the caching device “must provide at least 10 percent of the anticipated storage that virtual machines are expected to consume, not including replicas such as mirrors”. The exclusion is where readers go wrong. It is 10 per cent of consumed capacity before FTT, not of raw capacity after it.
vSAN hybrid OSA, one disk group, 7 TB consumed per host
cache floor 0.10 x 7 TB = 700 GB
read cache 0.70 x 700 GB = 490 GB
write buffer 0.30 x 700 GB = 210 GB
write buffer cap 600 GB in 7.0 and earlier; hybrid stays at 600 GB
Broadcom KB 326411 raises that ceiling to 1.6 TB on cache devices of 1.6 TB and
larger, but only on all-flash OSA, only on vCenter and ESXi 8.0 or newer, only
at on-disk format v17 or later, and only after setting
esxcfg-advcfg -s 1 /LSOM/enableLargeWb and recreating the disk groups.
“Hybrid clusters are not supported.” The same KB prices the change in this
guide’s currency: every disk group running the large write buffer costs an extra
5 GB of host memory, up to 25 GB on a host carrying the supported maximum of
five disk groups.
In OSA the cache device is not one drive among several; it is the front door to its entire disk group. Every capacity device behind it is reachable only through it, so its loss is a disk-group event rather than a drive event. That is the design lever: more disk groups per host means a smaller blast radius per cache device, and it is why the per-host cache budget above is really a per-disk-group budget.
Date the deprecation, because it is recent. The vSAN Design Guide for VMware Cloud Foundation 9.1, dated 5 May 2026, states that “the hybrid configuration in vSAN Original Storage Architecture feature will be discontinued in a future VCF release. This was announced with vSAN 9.0.” Express Storage Architecture, introduced with vSAN 8, has no cache tier and no disk groups at all, and its smallest ESA ReadyNode profile, called AF-0 in one part of the VCF 9.1 design guide and vSAN-HCI-SM in the profile table of the same document, lists a minimum of 128 GB of RAM per node.
The DRAM layer is the client cache, on by default since vSAN 6.2, sitting in the memory of the host running the VM. The figure repeated across secondary sources is 0.4 per cent of host memory capped at 1 GB per host; others say 4 per cent, and no Broadcom document stating either could be found. The 1 GB cap is solid, this guide will not claim a percentage, and the vSAN Design Guide for VMware Cloud Foundation 9.1 does not mention the client cache at all.
vSphere Flash Read Cache, and why it is still in your search results
vFRC used host-local flash to cache reads for individual VMDKs. Deprecation was
announced with vSphere 6.7 Update 2. In vSphere 7.0 its actions were
removed from the vSphere Client and the feature is unsupported; at least one
credible write-up argues the underlying vFlash plumbing survived in 7.0 even
with the interface gone, so the accurate phrasing is removed from the UI and
unsupported from 7.0, rather than deleted from the product. The replacements
Broadcom points at are vSAN caching and third-party VAIO filters. The fossil
is visible in current storage policies, where the vSAN design guide still
lists FlashReadCacheReservation: 0% and annotates it as a legacy hybrid OSA
policy. Anything written before 2019 that describes vFRC as a tuning option is
describing a product you can no longer configure.
The RAM cache in this stack that survived is CBRC, switched on by the View Storage Accelerator setting in Horizon, a product that moved from VMware to Omnissa in 2024. CBRC caches virtual machine disk blocks in ESXi host memory and is configurable from 100 MB to 2,048 MB with a 1,024 MB default.
Storage Spaces Direct: two caches, and only one is RAM
Microsoft now calls the device layer the storage pool cache. It claims drives on its own: “in deployments with multiple types of drives, Storage Spaces Direct automatically uses all drives of the fastest type for caching”, ranking NVMe above SSD above HDD. NVMe plus HDD and SSD plus HDD cache reads and writes; all-flash pairings cache writes only, because “reads don’t significantly affect the lifespan of flash”. Cache drives “don’t contribute usable storage capacity to the cluster”, and all-HDD deployments are not supported.
This is the opposite of the desktop Storage Spaces tiering covered earlier. That is a tier: the fast drives add capacity and hold the only copy. This is a cache: the fast drives add nothing to usable capacity, bindings to capacity drives are dynamic “from 1:1 up to 1:12 and beyond” and rebalance automatically, and the engine “derandomizes writes before de-staging them” so a random workload reaches the spindles looking sequential.
The hardware requirements page carries the number that decides the build: 4 GB of RAM per terabyte of cache drive capacity on each server, for Storage Spaces Direct metadata, on top of the OS and every workload.
8 nodes, 2 x 3.2 TB NVMe cache per node
cache per node 6.4 TB
metadata RAM 6.4 x 4 GB = 25.6 GB per node
cluster total 8 x 25.6 GB = 204.8 GB no tenant will ever see
CSV read cache 8 x 1 GiB = 8 GiB at the Azure Local default
Cache devices must be 32 GB or larger, and Microsoft recommends high write endurance of at least 3 DWPD or at least 4 TB written per day; the minimum counts are two cache plus four capacity drives per server, and at least two cache drives per server are required to preserve performance after one fails. Persistent memory is still listed as a cache device, though Intel announced the wind-down of its Optane business in July 2022, so this is a note about installed hardware more than a buying option. The restriction attached to it makes the point of this whole tier: “when using persistent memory devices as cache devices, you must use NVMe or SSD capacity devices (you can’t use HDDs)”.
The RAM layer is separate and Microsoft documents it separately. The CSV in-memory read cache defaults to 0 on Windows Server 2016 and 1 GiB on Windows Server 2019 and Azure Local, and may consume “up to 80% of total physical memory”. It exists because Hyper-V uses unbuffered I/O to VHD and VHDX files, which the Windows Cache Manager will not cache.
(Get-Cluster).BlockCacheSize # MiB per server; 1024 = 1 GiB
(Get-Cluster).BlockCacheSize = 2048
Microsoft also warns that DISKSPD and VM Fleet “may produce worse results with the CSV in-memory read cache enabled than without it”, because uniformly random synthetic reads have no reuse. That is a vendor stating that a cache benchmark measures the benchmark’s access pattern, not the cache. Measure with the PerfMon counter set Cluster Storage Hybrid Disk, counter Cache Miss Reads/sec, which reports one instance per capacity drive.
When a cache drive fails, “any writes which haven’t yet been destaged are lost to the local server”, and the bound capacity drives report unhealthy until rebinding and repair complete. Node-scoped again.
Nutanix, briefly
The Unified Cache is host DRAM inside the Controller VM, sized by the Nutanix
Bible’s formula ((CVM memory - 12 GB) x 0.45). A 32 GB CVM yields roughly
9 GB. A first read lands in a single-touch pool; any subsequent read promotes
the 4 KB entry to a multi-touch pool, and both pools evict by LRU. Writes go to
the OpLog, “a persistent write buffer” on the CVM’s SSD tier, replicated
synchronously to the OpLogs of other CVMs before the write is acknowledged, and
bypassed for sequential streams once more than 1.5 MB of write I/O to a vDisk
is outstanding.
8 nodes, 64 GB CVMs
per node (64 - 12) x 0.45 = 23.4 GB cache
cluster 8 x 23.4 = 187 GB, none of it shared between nodes
CVM RAM 8 x 64 = 512 GB reserved before a single VM boots
That last block is the tier’s whole argument. The cache is bought per node, in DIMMs, and duplicated by the replication factor rather than pooled by it, so the planning question is how many of these node-sized caches you can afford to have cold at once - which is the same number as the nodes you can lose. The DIMM pricing that decision rests on sits with the modules themselves (DDR4 RDIMMs), not with the cache device.
Scale-out: Ceph, and the feature that was withdrawn
Everything above this point has been a cache in the ordinary sense: a copy of something that also lives somewhere slower, and losing it costs you time rather than data. Ceph breaks that assumption twice. In a modern Ceph OSD the thing that behaves like a cache is the RAM, and the flash device people put in front of a spinning OSD is not a cache at all. The feature that genuinely was a cache, cache tiering, has been deprecated and has lacked a maintainer for years. That is three inversions in one product, and it is why a reader who arrives from a home bcache setup mis-sizes an entire cluster.
BlueStore caches in DRAM because it cannot use the page cache
BlueStore writes to raw block devices from userspace, so the kernel page cache
that carries every stack earlier in this guide is not in the path. The OSD
keeps its own onode, buffer and RocksDB caches inside its own address space, and
the knob that sizes all of them together is osd_memory_target, default
4Gi, with bluestore_cache_autotune on by default. Around that:
osd_memory_base 768Mi, osd_memory_cache_min 128Mi, and the manual fallbacks
bluestore_cache_size_hdd 1Gi and bluestore_cache_size_ssd 3Gi with
bluestore_cache_meta_ratio and bluestore_cache_kv_ratio both 0.45.
Reads are buffered, writes are not: bluestore_default_buffered_read is true
and bluestore_default_buffered_write is false, the latter to avoid paying
eviction overhead for data nobody is going to read back. Red Hat’s Ceph
administration guide advises contacting support before enabling buffered write.
Treat that as a warning, not a tunable.
The sentence in Ceph’s hardware recommendations that matters most to a guide
about RAM is this: an effective osd_memory_target of at least 6 GiB can
help mitigate slow requests on HDD OSDs. Ceph’s documented answer to a slow
spinning OSD is to buy memory. The same page says under 2 GB is not
recommended, 2 to 4 GB may degrade performance, total system memory should
exceed OSD count multiplied by the target multiplied by two, a 1U chassis with 8
to 10 OSDs is well provisioned with 128 GB, and you should budget at least 20
per cent extra to keep OSDs out of the OOM killer. Swap is not advised.
A 12-bay HDD node, tuned per Ceph's own HDD advice:
12 OSDs x 6 GiB target = 72 GiB of targets
Ceph's rule: total RAM > OSDs x target x 2 = 144 GiB
plus the documented 20 % headroom ~ 173 GiB
round up to a populated, sane DIMM count = 192 GiB
Per usable terabyte this is small. Per node it is a second
memory purchase most first-cluster budgets did not contain.
Under cephadm, osd_memory_target_autotune is on by default: it takes a
fraction of total RAM, subtracts the memory consumed by non-autotuned daemons,
and divides what is left by the remaining OSDs. The
ratio is mgr/cephadm/autotune_memory_target_ratio, default 0.7, dropping
to 0.2 on hyperconverged nodes in Red Hat’s and IBM’s builds. Ceph’s own
documentation warns that autotuning is usually not appropriate where a node runs
Ceph and compute together, which is exactly the configuration a small business
reaches for first. Opt a single OSD out with
ceph config set osd.123 osd_memory_target_autotune false.
Monitors and managers want 32 GB on a very small cluster, 64 GB up to roughly
300 OSDs and 128 GB beyond that, and CephFS metadata servers default to
mds_cache_memory_limit 4Gi with mds_cache_reservation 0.05. The MDS
documentation states plainly that the cache size is not a hard limit. Size the
node for the overshoot, not the setting.
The DB device is not a cache, and the failure model proves it
block.db and block.wal on flash in front of an HDD OSD look like the
caching pattern described earlier and are not it. They hold RocksDB metadata
and the write-ahead log, and those are the only copies. A cache device can be
pulled and the array still reads; a lost DB device takes every OSD bound to
it, and those OSDs are reprovisioned, not repaired. No upstream sentence
states that consequence verbatim, and this guide will not quote one; it follows
from an OSD being unable to start without its block.db. The practical reading
is that Ceph’s published fan-out figures, 4 to 5 HDD OSDs per DB/WAL SATA
SSD and up to 15 per NVMe, are a failure-domain decision wearing a
performance costume.
Sizing, from Ceph’s BlueStore configuration reference: when offloading WAL and
DB, make block.db at least 2.5 per cent of the slow device. That figure
goes with the RocksDB compression the documentation ties to Squid; for releases
older than Squid, where RocksDB compression is off, it gives shares by workload
instead: at least 4 per cent for RGW, whose omap keys live there, and 1 to 2 per
cent for RBD. The folklore figure of 30 GB
per OSD comes from the level sizes of older releases, where usable sizes stepped
roughly 3 GB, 30 GB, 300 GB and anything in between was wasted.
When the DB outgrows its device the metadata spills back onto the spinner and
Ceph raises BLUEFS_SPILLOVER, described as metadata having “spilled over” onto
the slow device. The cluster keeps working and quietly returns to HDD metadata
latency, which is the worst kind of regression.
ceph health detail # BLUEFS_SPILLOVER
ceph daemon osd.123 bluestore bluefs device info # BDEV_DB / BDEV_SLOW free
ceph-bluestore-tool bluefs-bdev-expand --path /var/lib/ceph/osd/ceph-123
ceph config set osd bluestore_warn_on_bluefs_spillover false # silence only
Expanding needs the OSD stopped and the backing logical volume already grown. The other remedy is to destroy and reprovision the OSD. Silencing the warning is not a remedy.
Cache tiering: deprecated, and the project said why
Older guides describe a hot replicated pool in front of a cold erasure-coded one. The Ceph documentation now carries this banner: “Cache tiering was deprecated in the Reef release and has lacked a maintainer for a long time. It will be removed in a future release without notice. Do not deploy new cache tiers. Migrate existing cache tier deployments as soon as possible.”
The deprecation names a release, Reef (v18), not a date, and this guide will not attach one. It was still shipping through the Squid series while its tests were taken out around it: the v19.2.4 changelog contains “test/rbd: remove unit tests about cache tiering”. The documentation still describes removal as something that happens without notice, so read the notes of the release you actually run rather than this paragraph.
The reasons the project gives are unusually candid: benefit is highly workload dependent, warm-up costs make it difficult to benchmark honestly, it is usually slower for non-optimal workloads, the librados object enumeration API “is not meant to be coherent” across a tier, and the whole thing is complex. The one known-good case is RGW with time-skewed access. The known-bad cases are RBD with a replicated cache over an erasure-coded base, and RBD with a replicated cache over a replicated base. That is most of what people built.
What survives, for block, is client-side caching, in RAM by default. librbd
cannot use the
Linux page cache either, so it carries its own: rbd_cache true,
rbd_cache_policy writearound, rbd_cache_size 32Mi per image,
rbd_cache_max_dirty 24Mi, rbd_cache_target_dirty 16Mi,
rbd_cache_max_dirty_age 1.0 s. It is per image and per client, with no
coherency between clients, and the documentation states that GFS or OCFS on top
of RBD will not work with caching enabled.
The threshold
The honest cost line is memory, and this guide will not put a core count on an OSD without a citation to hang it on. A twelve-spindle Ceph node needs roughly 192 GB of server DIMMs before it stores a byte, a DB device whose failure domain you have chosen deliberately, and a replication factor that multiplies every disk you buy. Under three nodes none of it is available in the way it promises to be. Ceph earns its place when you need a second and third chassis for availability, not when you need a single array to be faster. For one box of spinning disks, everything earlier in this guide is cheaper, simpler and, per pound, quicker.
Enterprise arrays and auto-tiering
On each array quoted below, the vendor documents the host write being acknowledged out of a nonvolatile buffer before it reaches the data drives. NetApp states it as the reason its flash caching is not what acknowledges a write: “The NetApp Data ONTAP operating system is already write-optimized through the use of write cache and nonvolatile memory (NVRAM or NVMEM).” HPE’s CASL architecture guide (a50002410enw) is blunter - “All incoming write requests, regardless of I/O size, must land in NVRAM/NVDIMM” - and describes that NVDIMM-N as a DDR4 module with a supercapacitor and onboard flash, mirrored by PCIe DMA into the standby controller’s NVDIMM before the active controller acknowledges the host. Dell’s PowerStore platform guide (H18149.13, May 2024) says “All host data is written to the NVMe NVRAM drives from DRAM before the host is acknowledged”, and records that the PowerStore 500 ships with no NVRAM drives at all and uses internal DRAM for write caching. The flash an array vendor sells you sits behind that acknowledgement. Memory, or a small nonvolatile buffer in front of the data drives, is the write path.
A cache holds a copy; a tier holds the only copy
Dell publishes the cleanest statement of the layering in Dell EMC Unity: FAST Technology Overview (H15086.3, February 2021), Table 6:
| System Cache (DRAM) | FAST Cache (SSD) | |
|---|---|---|
| Position | Closest to the CPU | Between System Cache and pool drives |
| Best for | Sequential I/O, I/O larger than 64 KB, zero-fill, high-frequency access | Random I/O, I/O under 64 KB, high locality |
| Response | Nanosecond to microsecond | Microsecond to millisecond |
| Granularity | Fixed 8 KB page | Fixed 64 KB page |
| Power loss | Volatile; BBUs hold the system up while dirty pages vault to the mSATA drive | Non-volatile |
Read that first column again. A storage vendor is documenting that DRAM is the correct medium for sequential and large-block I/O and that flash is for small random I/O with locality - the inverse of the reflex to solve every slow array by adding SSDs.
The cache/tier line is not always clean. NetApp’s Flash Pool read caching policies are write-through, so cached read data also exists on the HDDs. Its write caching policies do not: “Data inserted into the cache by using the write caching policy exists only in cache; there is no copy in HDDs.” That is why a Flash Pool’s SSDs are RAID protected and a Flash Cache module’s are not. A cache that holds the only copy of anything is a tier wearing a cache’s name.
The named implementations
| Product | Copy or only copy | Reads / writes | Granularity | Cadence |
|---|---|---|---|---|
| NetApp Flash Cache | Copy | Random reads only | WAFL buffers | Reactive |
| NetApp Flash Pool | Both, by policy | Read and write policies, default auto |
WAFL buffers | Reactive |
| Dell FAST Cache | Copy | Promotion on access from HDD; writes ack from System Cache | 64 KB | Reactive |
| Dell FAST VP | Only copy | Reads and writes counted | 256 MB slice | Hourly analysis, scheduled window |
| HPE adaptive flash (Nimble Adaptive Flash, Alletra 5000) | Copy | Populated on write and on read miss | Variable block | Reactive, ABE eviction |
| IBM Easy Tier | Only copy | Both | 1 GiB (DS8000) or pool extent | 24-hour migration plan, moved in 5-minute batches |
| Hitachi Dynamic Tiering | Only copy | Both | 42 MB page | 0.5, 1, 2, 4 or 8 h auto cycles |
Details that decide behaviour. Flash Cache is a per-node PCIe or NVMe module driven by WAFL External Cache, and data in a Flash Pool aggregate is not cached by it. FAST Cache promotes a 64 KB chunk when an eligible block is accessed from spinning media three times “within a short period of time”; H15086.3 does not quantify that period, and this guide will not invent a figure. FAST Cache is supported only on Unity Hybrid and Unity XT hybrid systems. Nimble’s flash cache is populated on the write path - “cache worthy blocks are immediately cached in the SSD flash cache layer” - but “by default, sequential read/write streams are not cached”, and the cache “holds copies of blocks”, so an SSD failure costs no data. Easy Tier monitors continuously, creates a cross-tier migration plan at least once every 24 hours, then moves a limited number of extents every five minutes until that plan is complete, and can ship that heat map to a replica so the target pre-warms itself. Hitachi’s automatic cycles are anchored at midnight, and relocation begins only at the end of a cycle.
Standing: FAST Cache and FAST VP are Unity and VNX-era features. PowerStore is all-flash and supports no HDDs, with NVMe in the base enclosure, SAS SSD only in SAS expansion shelves, and NVMe SCM dedicated to metadata; the platform introduction describes no FAST equivalent, which is absence of evidence rather than a Dell statement that they were dropped. Flash Pool has a live future - FAS70 and FAS90 were announced in autumn 2024 and the FAS50 followed in early 2025. HPE’s hybrid line continues as Alletra 5000.
Granularity and lag
Cache granularity and tier granularity differ by three to five orders of magnitude, and that gap is the whole argument.
Hitachi Dynamic Tiering, fixed 42 MB page:
region actually hot 1 MB
amount relocated 42 MB
cold data dragged to fast media 41 MB = 97.6% of the move
Dell EMC Unity FAST VP, 256 MB slice:
analysis runs hourly
default relocation window 17:00 - 01:00 local
workload turns hot 09:00
ranked as a candidate 10:00
earliest possible movement 17:00
lag 8 h, and only if the window
has the bandwidth to finish
A cache is fine-grained and reactive; a tier is coarse-grained and scheduled. Neither mechanism is broken. They are answering different questions.
The heat map is a map of cache misses
Dell documents the interaction that most capacity planners get backwards: “FAST VP only monitors I/O which reaches the drives within the Pool. I/O handled by FAST Cache does not affect the analysis done by FAST VP”, and I/O serviced by System Cache is likewise not factored into promotion. FAST Cache page-cleaning I/O, by contrast, is counted and weighted as normal traffic.
So the array’s temperature map is built from what the DRAM and the flash cache failed to absorb. Adding controller memory makes the tiering engine blinder, and a tiering report can look nothing like the application owner’s model of their own workload without either of them being wrong.
The month-end report that is always on the slow tier
Unity’s temperature is a weighted average with recent activity weighted higher and the weight decaying over time, and FAST VP demotes cold slices to hold 10% free capacity in each tier. A batch set that is idle for twenty-nine days is cold by construction, is demoted, and is served from the capacity tier on the one day it matters - with promotion landing after the run has finished. That is reasoning from documented mechanics, not a vendor claim.
Hitachi names the failure in its own manual. Period mode, the default, decides “based solely on the monitoring data from the last complete cycle”, and “if the I/O loads vary greatly, relocation may not finish in one cycle”. Continuous mode exists to weight recent monitoring against earlier cycles “so that even if a temporary decrease or an increase of the I/O load occurs, unnecessary relocation can be avoided”. A vendor shipping a second mode to stop the first one thrashing is the strongest available evidence that cyclic workloads and scheduled tiering fight each other.
Note where the guards live. FAST Cache refuses to promote small-block sequential I/O, zero-fill requests, and I/O larger than the RAID stripe length - above 512 KB in an 8+1 RAID 5 layout - so a backup pass does not flatten it. Nimble excludes sequential streams by heuristic. Those are cache-layer defences; the tiering documentation reviewed here describes no equivalent, which is a reading of the documents rather than a stated fact.
The case for tiering, where it holds, is skew. IBM measured five real customer systems (WP102295, May 2013): 5% of allocated capacity took between 50% and 87% of small-I/O accesses, and after initial promotion a 140 TB system moved “typically around 200 GB” a day, under 0.2% of the data. The skew curve carries forward. The drive tables in that paper, which run from 146 GB 15K SAS through 900 GB 10K SAS to 3 TB nearline, do not.
What this means for the second-hand shelf
IDC’s Worldwide Quarterly Enterprise Storage Systems Tracker for Q3 2025, released 11 December 2025, puts all-flash arrays at +17.6% year on year, hybrid flash at -9.8% and HDD arrays at -6.3% in the same quarter. Hybrid is being relegated, not abolished, and the shelves come off support contracts and onto the used market with their SAS drives still in the carriers. That is where a great deal of the enterprise-class stock in this catalogue originates.
Check the caching feature before the capacity. Unity XT accepts only 400 GB SAS Flash 2 drives for FAST Cache, in RAID 1 pairs, so an XT 880 needs thirty drives of a single legacy SKU to reach its 6 TB maximum, and the Create FAST Cache wizard in Unisphere offers only the drive sizes the system supports for FAST Cache. A NetApp RAID group cannot be removed from an aggregate after creation. Nimble’s NVDIMM and its supercapacitor are monitored, serviceable parts, and they sit on the write path rather than the read path.
The layer that ages best is the memory. Registered DDR4 for a decade-old controller is a commodity with a live second-hand market of its own, priced at ram-index.net, while the flash SKU the array insists on is not.
Persistent memory as the cache device: the top of the ladder, and a dead end
The last and fastest persistent-memory module Intel designed for this job never reached volume production, and the company that made it put that in writing. Intel’s support note on the Optane PMem 300 Series (Crow Pass) says Intel “does not intend to ramp production, take future orders, take last time buys, provide warranty coverage or technical support” for it, effective 31 January 2023, and Intel’s July 2022 Optane business update says “Intel plans to cease future development of our Optane products”, citing the position of being “a single-source supplier of Optane technology”. Everything below is therefore a second-hand buying question, not a product recommendation. It is still worth understanding, because the mechanism explains what a cache tier can and cannot buy you.
What byte addressability changed
An NVMe drive is a block device behind a submission queue and a completion interrupt. A persistent-memory DIMM sits on the memory bus and is reached with load and store instructions, and its persistence comes from ADR (Asynchronous DRAM Refresh), which drains the memory controller’s write-pending queues into the DIMM when power fails. A durable commit costs a cache-line flush and a fence rather than a descriptor, a doorbell and an interrupt.
The kernel’s own defaults price that difference. dm-writecache sets
autocommit_blocks to 64 on persistent memory and 65536 on SSD:
dm-writecache autocommit_blocks, at a 4096 B block size
pmem : 64 blocks x 4096 B = 256 KiB written before an auto commit
ssd : 65536 blocks x 4096 B = 262144 KiB (256 MiB) before an auto commit
ratio 1024 : 1
source: Documentation/admin-guide/device-mapper/writecache.rst
The same document opens with the most useful sentence in the kernel tree for
anyone putting flash in front of a hard drive: “The writecache target caches
writes on persistent
memory or on SSD. It doesn’t cache reads because reads are supposed to be
cached in page cache in normal RAM.” dm-cache and bcache do cache reads;
this target never does. One correction to a claim in circulation:
lvmcache(7) does not
require persistent memory for --type writecache. Only the fua and nofua
options are pmem-only.
Two device-mapper targets, one of them very new
# dm-writecache: type p = persistent memory, s = SSD
lvconvert --type writecache --cachevol fast vg/main
# dm-pcache, merged in Linux 6.18, DAX device only
dmsetup create pcache_sdb --table \
"0 $(blockdev --getsz /dev/sdb) pcache /dev/pmem0 /dev/sdb 4 \
cache_mode writeback data_crc true"
dm-pcache (maintainers Dongsheng Yang and Zheng Gu, GPL-2.0,
drivers/md/dm-pcache/) is log-structured, uses 16 MiB segments, duplicates
its metadata with CRC and sequence numbers, and takes a pure DAX path with no
extra BIO round trips. It accepts write-back as its only cache mode,
invalidates FIFO today with LRU and ARC listed as planned, and does not yet
support discard. Its Kconfig depends on DEV_DAX and says plainly that the
feature “is experimental and should be tested thoroughly before use in
production environments”. Treat that as the maintainers’ warning it is.
Memory Mode against App Direct
| Memory Mode (2LM) | App Direct | |
|---|---|---|
| What the OS sees | volatile main memory | a persistent DAX device |
| Role of the DDR DIMMs | invisible L4 write-back cache, run by the memory controller | ordinary main memory |
| Persistent? | no | yes, via ADR |
| Software control | none | ipmctl, then ndctl |
| Useful as a disk cache | no | yes |
Intel publishes a support article titled “Why Is the Intel Optane Persistent Memory in Memory Mode Not Persistent?” - the answer being that in Memory Mode the data is volatile by design. That mode is the purest hardware example of caching as a placement decision made below the operating system: DRAM used as a cache for slower memory, with the policy in silicon and no knob at all.
App Direct capacity, provisioned alone or as the App Direct share of Mixed Mode,
is the only thing on which dm-writecache type p or dm-pcache
make any sense. Provisioning is a goal written to the DIMMs and read by the
BIOS on the next boot:
ipmctl create -goal PersistentMemoryType=AppDirect
ipmctl create -goal MemoryMode=25 PersistentMemoryType=AppDirect Reserved=25
ndctl create-namespace -m fsdax # /dev/pmemX, filesystem DAX on xfs/ext4
ndctl create-namespace -m sector # adds the sector atomicity raw pmem lacks
What a used DIMM commits you to
The 100 and 200 Series run only on specific Cascade Lake, Cooper Lake and Ice Lake Xeon generations with matching firmware, and population rules are per socket, so the DIMMs are not the purchase - the platform is. Intel’s wind-down note says “the 5-year warranty terms from date of sale remain unchanged”, which on a pulled DIMM means a clock that started for somebody else. Intel’s own list of discontinued Optane products gives the 100 Series an end-of-life date of 29 June 2023 and an end of interactive support date of 30 June 2025, and gives the 200 Series an end-of-life date of 26 June 2024 with no interactive support date at all. Broadcom’s KB 326910, “vSphere/vSAN Support update for Intel’s Optane Persistent Memory (PMem)”, states that Optane PMem and both of its modes are no longer supported from vSphere 9.0, and that the certified server SKUs no longer appear in the 9.0 compatibility guide, which strands the technology on the commonest virtualisation platform.
The decisive fact is smaller and better sourced. Microsoft’s Storage Spaces Direct hardware requirements state: “When using persistent memory devices as cache devices, you must use NVMe or SSD capacity devices (you can’t use HDDs).” Microsoft’s supported configuration forbids putting a persistent-memory cache in front of spinning disks, which is the one arrangement that would make the DIMMs worth buying for a home array.
CXL is the direction, not the answer
CXL 4.0 was released on 18 November 2025. It doubles the rate to 128 GT/s “with
zero added latency”, adds bundled ports, memory RAS enhancements, native x2 and
support for up to four retimers, keeps the 256B flit format and claims full
backward compatibility with 3.x, 2.0, 1.1 and 1.0. Persistence is not the
missing piece it is often called: CXL 2.0 already defined persistent-memory Type
3 devices and a Global Persistent Flush to drive cached and buffered data to a
durable medium on power loss. What the 4.0 feature list does not do is advance
it. CXL’s centre of gravity remains memory expansion and pooling, with Type 3
devices most often presenting DRAM over CXL.mem. The clearest place the
kernel joins CXL to persistent caching today is dm-pcache’s Kconfig text,
which names “CXL persistent memory” as an example device class. Shipping
silicon and kernel plumbing are not a part number you can buy for a file
server.
The verdict
In front of mechanical drives, no. Reads belong to the page cache, on DRAM already installed. The only thing a persistent-memory tier adds over an ordinary power-loss-protected enterprise NVMe drive is commit latency measured in hundreds of nanoseconds rather than tens of microseconds - and behind either of them sits an array whose floor is seek time plus the RAID read-modify-write, two to four orders of magnitude higher. The advantage dissolves into the backing store. A flash-backed or battery-backed controller DRAM cache delivers the same durable acknowledgement on hardware many servers already carry.
Where persistent memory still earns its place is inside a product, not bolted onto one: HPE’s Nimble and Alletra 5000/6000 write path lands every incoming write in NVDIMM-N - DDR4 with onboard flash and a supercapacitor - and mirrors it to the standby controller before acknowledging the host. That is persistent memory bought as part of an enterprise array, with a vendor on the hook for it. Buy the DIMMs only if the platform they lock you to is a platform you were buying anyway.
Sizing the cache, arithmetically
A cache is sized by the working set, not by a percentage of the array. Every ratio in circulation - 10 per cent of capacity, 1 GB of RAM per TB of pool, five times RAM for L2ARC - is a guess that acquired authority by repetition. The RAM-per-TB rule is not an OpenZFS requirement and does not appear in its documentation; it is FreeNAS-era hardware guidance, and the current TrueNAS hardware guide does not state it. What that guide states instead scales with different things entirely: about 5 GB per TB for a dedup table that lives in RAM, and roughly 1 GB of RAM for every 50 GB of L2ARC. Those are two unrelated quantities that the folk rule welded together.
What the working set actually is
Distinct blocks touched in a window, and how fast that set drifts: the two
properties the opening sections rested the argument on. Sizing turns them into
a number, and counting distinct chunks is a blktrace job. Capture queue events on the
backing device, fold each request into chunk indices at your cache’s chunk
size, and count the unique ones:
blktrace -d /dev/sdb -a queue -w 3600 -D /var/tmp/bt
blkparse -i sdb -D /var/tmp/bt -f "%d %S %n\n" \
| awk '$1 ~ /R/ && $3 ~ /^[0-9]+$/ {
for (s = $2; s < $2 + $3; s += 8) print int(s/128) }' \
| sort -u -S 2G | wc -l
128 sectors is 64 KiB, lvm2’s default cache chunk size;
lvm2 raises it on large cache volumes to hold the chunk count near the
recommended ceiling of a million, so read the real one off
lvs -o+chunksize VG/LV. Multiply the count by
64 KiB for the read footprint. Repeat at 5 minutes, 1 hour and 24 hours and
you have a crude miss-ratio curve. blktrace wants root and a mounted
debugfs; the numeric test on the third field keeps blkparse’s end-of-run
summary out of the count. Check the format specifiers against your version
first, and bound sort memory on a busy device.
What a cache one tenth the working set buys
Take a footprint of one million 64 KiB chunks and a cache holding a hundred thousand of them. Use 12 ms for the mean read service time of a 7200 rpm drive at queue depth 1. That is an assumption for the arithmetic, not a quoted spec: Seagate gives the Exos X24 an average latency of 4.16 ms, half a rotation, and 168 random 4K read IOPS at QD16, which is one completion every 5.95 ms with sixteen requests in flight and free rein to reorder them. A single outstanding request gets none of that reordering, so 12 ms is the same drive seen from its worst angle. Use 0.1 ms for the NVMe.
L_eff = h*L_fast + (1-h)*L_slow S = L_slow / L_eff
uniform access, h = 0.10:
L_eff = 0.10*0.1 + 0.90*12 = 10.81 ms S = 1.11x
pure Zipf s=1, top 1e5 of 1e6:
h ~= ln(1e5)/ln(1e6) = 11.51/13.82 = 0.833
L_eff = 0.833*0.1 + 0.167*12 = 2.09 ms S = 5.7x
The same ratio buys 1.11x or 5.7x, and the ratio does not tell you which.
Skew is the entire bet. The second line is arithmetic on an assumed
distribution rather than a measurement - a pure s=1 law and the harmonic
approximation, which your workload owes nothing to. Under uniform access a
tenth-sized cache is a rounding error, and each promotion it does make writes
a whole chunk of NAND for a 4 KiB read - 64 KiB at lvm2’s default. dm-cache’s
default smq policy exists to refuse most of those promotions, and
migration_threshold, 2048 sectors or eight chunks whichever is larger, caps
what it will move at once.
The cliff is at the top, not at the bottom
Latency is linear in the hit rate; the speedup S = L_slow / L_eff is the
hyperbola. While the miss term dominates - here up to about h = 0.95, where
h*L_fast is still only a seventh of the total - L_eff ≈ (1-h)*L_slow and
latency tracks the miss rate:
h = 0.50 6.05 ms 2.0x 500 misses per 1000 I/Os
h = 0.90 1.29 ms 9.3x 100
h = 0.95 0.695 ms 17.3x 50
h = 0.99 0.219 ms 54.8x 10
h = 0.999 0.112 ms 107x 1
Halving the miss rate halves the latency only while the miss term dominates.
The bottom of the table shows where that stops: a tenfold cut in misses from
h = 0.99 to h = 0.999 buys barely 2x, because by then h*L_fast is 0.0999 ms
of the 0.112 ms total. Going from 50 to
90 per cent buys 4.7x; going from 90 to 99 buys another 5.9x. Think in
misses, not hits. And report p99, not the mean: at h = 0.99 the misses are
exactly the top 1 per cent,
so everything above the 99th percentile is a disk access and the tail sits
near 12 ms however good the average looks. At h = 0.98, p99 itself is one.
Metadata is the cheap case
Metadata is small, hot, random and on the critical path of everything, which
is why a ZFS special vdev is the highest-leverage flash in a pool of spinning
disks. Work the floor. ZFS addresses each record with a 128-byte block
pointer and, under the default redundant_metadata=all, keeps two copies of
most metadata:
10 TiB of data at recordsize=128K
= 10 * 2^40 / 131072 = 83,886,080 records
* 128 B block pointer = 10 GiB
* 2 default metadata copies = 20 GiB
same pool at recordsize=1M = 2.5 GiB
Add dnodes, directory ZAPs and space maps and you are still in tens of GiB
for a 10 TiB pool, not hundreds. That is the floor, not the answer - read the
real figure off zdb -bbb pool, noting that PSIZE undercounts multiple
metadata copies and ASIZE overcounts raidz. The commonly quoted “0.3 per cent
of pool” has no first-party source and this guide will not endorse it; the
arithmetic above shows why it cannot be one number, since an eightfold change
in recordsize moves it eightfold.
An L2ARC has the opposite economics. Every resident block costs a header in
ARC that memory pressure cannot evict, and l2arc_meta_percent, documented
as the “percent of ARC size allowed for L2ARC-only headers” and defaulting to
33, stops the feed once those headers pass a third of the ARC. Third-party
guidance sizes that header at about 70 bytes; the 2.2 to 2.4 structure gives 96, the
true cost varies by build, and the only number that binds is your own
l2_hdr_size from /proc/spl/kstat/zfs/arcstats. Take 96 and a 16 GiB ARC:
headers stop at 5.28 GiB, about 59 million of them, which at 16 KiB physical
blocks is roughly 900 GiB of content and at 4 KiB blocks roughly 225 GiB -
close enough to TrueNAS’s 1 GB per 50 GB to show where that rule came from.
Your RAM sizes your L2ARC, and your block size decides how far it stretches,
whatever the SSD says on the label.
A write cache is sized by burst and drain
A write cache absorbs the difference between arrival and drain for as long as the burst lasts. Working set does not enter into it. The three inputs below are stand-ins to show the shape; only your own three matter.
arrival (10 GbE payload; the wire
itself carries 1,250 MB/s) 900 MB/s
drain (array's derandomised write
throughput - measure it, do not
use the sequential spec) 250 MB/s
burst duration 480 s
accumulation = (900 - 250) * 480 = 312,000 MB = 312 GB
Double it, because dm-writecache starts writeback at high_watermark
(default 50) and you want the burst absorbed before throttling starts:
624 GB, so a 1 TB device. Then check the RAM bill, because LVM’s is severe -
“each 100 GiB of writecache cachevol uses slightly over 2 GiB of system
memory” at block size 4096, and “a little over 16 GiB” at 512. That 624 GB is
581 GiB, so about 12 GiB of RAM at 4096 and something near 95 GiB at 512. If
drain equals or exceeds arrival, you need no write cache at all, only enough
flash to flatten fsync latency.
Three starting points
Starting points to be measured and then moved. No vendor publishes a table like this; this one is judgement, not doctrine.
| Build | HDD raw | Cache | Notes |
|---|---|---|---|
| Desktop, 4-bay NAS | 4 to 16 TB | 250 to 500 GB | writethrough, unmirrored |
| 8 to 24 bay server | 50 to 200 TB | 1 to 2 TB | mirror anything writeback |
| Hypervisor node | 200 TB and up | 2 to 4 TB | PLP mandatory, 1 DWPD or better |
On ZFS, split the middle tier: a mirrored special vdev sized from zdb -bbb,
plus a separate unmirrored L2ARC that costs nothing to lose. Choose the
device by write budget rather than read IOPS, since the job description is
“absorb writes forever” - the reasoning is in
SSD endurance: TBW, DWPD, how worried to be. Capacity
candidates are ranked in the NVMe listings. And run the comparison
honestly against the alternative: if the footprint will not fit whatever cache
you can afford, the money buys more spindles instead, and
the full drive listing ranks every drive by price per terabyte.
Tuning
Most of the dm-cache tuning advice on the internet is a no-op. Tutorials
still tell readers to set sequential_threshold 512 and random_threshold 4.
Those belonged to the mq policy, and dm-cache-policy-mq.c was deleted in
Linux 4.6; since then mq is an alias for smq. The kernel’s
cache-policies documentation is explicit that the tunables are “accepted, but
have no effect”, and lvmcache(7) says they are “silently ignored”. The same
kernel text says “smq also does not have any cumbersome tuning knobs”, and
that is close to true.
What follows is organised by symptom. A knob with no symptom attached is a knob nobody should turn.
dm-cache
Chunk size is the one you must check rather than assume. The kernel accepts “between 64 sectors (32KB) and 2097152 sectors (1GB) and a multiple of 64 sectors (32KB)”. LVM’s default is 64 KiB, but LVM also caps chunk count:
allocation/cache_pool_chunk_size = 64 KiB (DEFAULT_CACHE_POOL_CHUNK_SIZE)
allocation/cache_pool_max_chunks = 1000000 (DEFAULT_CACHE_POOL_MAX_CHUNKS)
min_chunk_size = ceil(cache_size / cache_pool_max_chunks), rounded up
to a multiple of 32 KiB, then rounded up to a
power of two when no --chunksize was given
update_cache_pool_params() raises your chunk size to that floor, so a large
cachevol does not get 64 KiB chunks. The symptom is write amplification you
cannot explain: a 4 KiB promotion into a 1 MiB chunk writes 1 MiB of NAND.
The check is one command, and it is not optional.
lvs -o+chunksize,cache_mode,cache_policy vg/main
migration_threshold is not a bandwidth cap, despite being described as
one almost everywhere. Kernel default 2048 sectors (1 MiB); lvmcache(7) puts
the LVM default at “2048 sectors (1 MiB) or 8 cache chunks whichever of those
two values is larger”. It is a core arg, not a policy arg, and
spare_migration_bandwidth() compares it against in-flight migration volume
after checking dm_iot_idle_for(&cache->tracker, HZ) - a full second of no
I/O:
concurrent migrations ~= floor(migration_threshold / sectors_per_block)
64 KiB chunks (128 sectors), threshold 2048 -> 16 in flight
1 MiB chunks (2048 sectors), threshold 2048 -> 1 in flight
Symptom: the cache warms at a crawl, or CacheDirtyBlocks climbs and never
falls. Raise it with lvchange --cachesettings 'migration_threshold=8192' vg/main, and put it back with --cachesettings 'migration_threshold=default'.
A busy writeback cache never drains on its own. In dm-cache-policy-smq.c
the CLEAN_TARGET constant is defined but no longer gates anything;
clean_target_met() returns true unconditionally when the device is neither
idle nor running the cleaner, with the comment “If we’re busy we don’t worry
about cleaning at all.” On a server with no idle window, dirty grows to the
cache size. The three reliable drains are lvchange --cachepolicy cleaner vg/main, lvchange --cachemode writethrough vg/main, and lvconvert --splitcache vg/main. Rehearse one before you need it, because “if the
system crashes all cache blocks will be assumed dirty when restarted” and the
drain is cache-sized.
Mainline has since made the background-work limit a module parameter,
/sys/module/dm_cache_smq/parameters/smq_max_background_work, declared in
the policy as smq_max_background_work = 4096. It is writable at runtime but
read only when a policy is created, so it applies to newly created cache
devices; kernels without the file hardcode the same 4096. Check the file is
there before planning around it.
For dm-writecache the durability knobs are autocommit_blocks (default 64
for pmem, 65536 for SSD) and autocommit_time (default 1000 ms). Those
two bound your exposure window, and the SSD default is large enough to
deserve a decision rather than a shrug. LVM’s writecache block_size takes
4096 or 512, set at attach and never after: match sectsz= from xfs_info
or the volume will likely fail to mount.
bcache
Everything lives in sysfs, and the paths differ for a partition: it is
/sys/block/sdb/sdb2/bcache, not /sys/block/sdb2/bcache.
/sys/block/bcache0/bcache/sequential_cutoff 4M (4 << 20)
/sys/block/bcache0/bcache/writeback_percent 10
/sys/block/bcache0/bcache/writeback_delay 30 (seconds)
/sys/block/bcache0/bcache/writeback_rate 1024 (sectors/s initial)
/sys/block/bcache0/bcache/writeback_rate_minimum 8
/sys/block/bcache0/bcache/writeback_rate_p_term_inverse 40
/sys/block/bcache0/bcache/writeback_rate_i_term_inverse 10000
/sys/block/bcache0/bcache/writeback_rate_update_seconds 5
/sys/fs/bcache/<uuid>/congested_read_threshold_us 2000
/sys/fs/bcache/<uuid>/congested_write_threshold_us 20000
sequential_cutoff exists because, in the documentation’s words, “if you
copy a 10 gigabyte file you probably don’t want that pushing 10 gigabytes of
randomly accessed data out of your cache”. The decision uses
max(task->sequential_io, task->sequential_io_avg), whose second term is an
exponentially weighted moving average of that task’s sequential run lengths
(bcache keeps 128 recent IOs per cached device to spot the runs), so a backup
bypasses wholesale
rather than caching 4 MiB after every seek. Symptom for lowering it:
bypassed stays near zero while a nightly rsync flattens your hit ratio.
Symptom for raising or zeroing it: a benchmark that never warms.
writeback_percent stops at 40 on a default build. sysfs clamps it to
bch_cutoff_writeback, which defaults to CUTOFF_WRITEBACK = 40 and can
itself be raised as far as 70 with the cutoff_writeback module parameter at
load time. should_writeback()
refuses writeback entirely above bch_cutoff_writeback_sync (70 per cent in
use), and between 40 and 70 only sync, REQ_META and REQ_PRIO writes are
still written back - everything else quietly degrades to writethrough. That
is the mechanism behind “my writeback cache stopped absorbing writes”, and no
amount of turning writeback_percent up will address it.
The dirty target is also shared. __calc_target_rate() takes
cache_sectors * writeback_percent / 100 and scales it by each backing
device’s share of c->cached_dev_sectors:
5 equal backing devices, writeback_percent = 10
-> each device's dirty target ~= 2% of the cache, not 10%
The writeback controller is a PI controller, not the PD the documentation
claims; there is no derivative term. p_term_inverse = 40 targets retiring
all dirty blocks in 40 seconds, i_term_inverse = 10000 accumulates
1/10000th of the error per second. Leave both alone unless dirty sits pinned
at the ceiling under a workload that does have idle gaps.
Congestion is the subtle one. bcache tracks cache-device latency and “does
this by cranking down the sequential bypass” - bch_get_congested() returns
a sector threshold compared against that same per-task counter, so congestion
literally lowers your effective sequential_cutoff. Setting both thresholds
to 0 disables it. Read hit ratio from the decaying windows rather than the
cumulative total, and remember that “a partial hit is counted as a miss”:
cat /sys/block/bcache0/bcache/stats_hour/cache_hit_ratio
cat /sys/block/bcache0/bcache/dirty_data
ZFS
The per-dataset knobs are the ones that pay. primarycache and
secondarycache both default to all and both take all|metadata|none.
Setting secondarycache=metadata on a media dataset stops streaming content
consuming the L2ARC; setting primarycache=none also disables L2ARC for that
dataset in effect, since L2ARC is fed from ARC evictions.
recordsize defaults to 128 KiB and sets the block count, which is what
the cache actually costs you. Every block resident in L2ARC but not ARC keeps
a header in ARC, offsetof(arc_buf_hdr_t, b_l1hdr) = 96 bytes on LP64 - not
the 70 or 180 bytes older articles quote:
1 TiB of L2ARC content, 96 B per record:
128 KiB blocks -> 8,388,608 records -> 768 MiB of ARC
16 KiB blocks -> 67,108,864 records -> 6 GiB of ARC
Those headers are not evictable under memory pressure, which is why
l2arc_meta_percent (default 33) stops both feeding and persistent rebuild
once they pass a third of the ARC target. Symptom: l2_size stops growing
well short of the device size. The answer is a larger recordsize or a
smaller cache, not a bigger one.
special_small_blocks defaults to 0 and compares against the physical,
post-compression size, so a threshold near your recordsize can pull an
entire well-compressing dataset onto a vdev you cannot shrink. Small blocks
also stop being accepted once the class passes 75 per cent full, the
remaining quarter being held for metadata by
zfs_special_class_metadata_reserve_pct, which zfs(4) documents with a
default of 25 per cent.
Module parameters live under /sys/module/zfs/parameters/. The fill rate is
the endurance knob:
l2arc_write_max 33554432 B (32 MiB) since 2.2.7, 8 MiB before it
l2arc_headroom 8 since 2.2.7, 2 before it
l2arc_noprefetch 1 (prefetched-but-unused buffers are not cached)
l2arc_mfuonly 0 (1 = MFU only; 2 = all metadata, MFU data only)
32 MiB/s sustained = 2.64 TiB/day -> ~2.6 DWPD on a 1 TiB cache SSD
1 TiB of L2ARC / 32 MiB/s = 32,768 s ~= 9.1 h of continuous feeding
That DWPD figure is the reason to choose the cache device on its write budget
rather than its read IOPS, and the reason SSD endurance: TBW, DWPD, how
worried to be is the companion read to this section. If
the arithmetic says a consumer drive lasts eighteen months in the role, the
fix is either l2arc_write_max or a drive with the endurance rating to
survive it, which in practice means the
enterprise SSDs with power-loss protection.
On ARC itself, one default changed and most articles are stale. The zfs(4)
man page for 2.2 gives the Linux limit as half of system memory; 2.3
documents the larger of all_system_memory - 1 GiB and five eighths of
all_system_memory, for every platform rather than FreeBSD alone. A 64 GiB
box now caps ARC at 63 GiB by default, not 32. zfs_arc_meta_limit and
zfs_arc_meta_min no longer exist; the balance is adaptive, steered by
zfs_arc_meta_balance (default 500). Note also that 2.4.0 renamed the tools:
arcstat is zarcstat, arc_summary is zarcsummary. Scrape
/proc/spl/kstat/zfs/arcstats and parse neither.
Underneath all of them: scheduler, queue depth, readahead
/sys/block/sdb/queue/scheduler # the HDD sees all the destage traffic
/sys/block/nvme0n1/queue/scheduler # usually none, and should stay none
/sys/block/sdb/queue/nr_requests
/sys/block/sdb/queue/read_ahead_kb
The distinction that matters: the NVMe wants no scheduler at all, because
sorting requests for a device with no seek penalty mostly adds latency; what a
scheduler still offers there is starvation control and I/O priority, not
throughput. The HDD underneath still wants a merging, sorting scheduler, because
destage traffic is exactly the workload elevator sorting was invented for.
Distro udev rules set both and they differ between distros, so read the files
rather than assume. Raising read_ahead_kb on the backing device helps
sequential misses and costs nothing on a cache that already bypasses
sequential streams; raising it on the NVMe is pointless.
Queue depth is a measurement question rather than a tuning one. Seagate’s
Exos X24 datasheet quotes about 168 random read IOPS at 4K QD16 for that
7200 rpm drive, so deeper queues buy throughput at the cost of per-request
latency, and the cache exists precisely so that queue never gets deep. Watch
aqu-sz, r_await and w_await in iostat -x 1 on all three devices at
once. Ignore %util on the NVMe: iostat(1) says outright that for “devices
serving requests in parallel, such as RAID arrays and modern SSDs, this
number does not reflect their performance limits”.
The filesystem: turn off atime first
noatime is the cheapest win available on a spinning pool, and it costs
you nothing unless something on the box reads access times, such as a Maildir
client or a temporary-file reaper. Linux defaults to relatime, which still
writes an inode update
once per day per file that is read. On a plain HDD that is a stray metadata
write. On a writeback cache it is a dirty block created by a read: promoted,
written back, and never read again - pure wear and pure destage traffic,
generated by workloads that only looked at the data. A nightly backup that reads a few
million files is a few million metadata writes under relatime and zero
under noatime. A find only calls stat, which does not update a file’s
access time, so it moves the directory timestamps instead. ZFS spells the same
property atime=off, with relatime=on
as its middle setting. On a box that mostly serves files, this is the first
mount option to change and the last one you will regret. The same reasoning
applies to any array of spinning disks, cached or not, which is why it turns
up again in what differs from a desktop.
The other filesystem interaction is a hard compatibility rule rather than a
tuning choice, and it will stop a volume mounting. From lvmcache(7): a cache
pool created on a device with a 4096-byte logical block size cannot be
attached to a main LV whose filesystem uses 512, and “the main LV will likely
fail to mount”. Check xfs_info /dev/vg/main for sectsz= before attaching
anything, and set block_size=512 on a writecache if that is what the
filesystem was made with. A cache that will not mount is a worse outcome than
a cache that is slightly misconfigured, and this one has no runtime fix - the
block size is fixed at attach.
Measuring, because a cache nobody measures is an expensive assumption
A cache you have not instrumented is not a cache, it is a second failure domain with a marketing story. Every stack described above ships counters. Nothing collects or graphs them for you, and all of them mean something slightly different.
Reading the hit ratio in each stack
For dm-cache, the counters live in the target’s status line. Read it with
--noflush. The flag sits in the dmsetup status synopsis, and although
dmsetup(8) describes it against the thin target, cache_status() in
dm-cache-target.c tests the same flag and commits metadata on every status
call without it, so a one-second poll drives a metadata write to the NVMe once
a second, forever:
dmsetup status --target cache --noflush vg-main | \
awk '{split($7,c,"/"); printf "used %.1f%% read hit %.2f%% write hit %.2f%% dirty %d\n", \
100*c[1]/c[2], 100*$8/($8+$9), 100*$10/($10+$11), $14}'
Fields 8 to 11 are read hits, read misses, write hits, write misses. On a cache
that has only just been activated those denominators are still zero, so let it
take some traffic before you trust the line. The counters come from
atomic_read() and are emitted as unsigned 32-bit values, so they wrap, and
they reset to zero on any table reload or deactivation. Graph deltas with wrap
detection, never raw totals. LVM exposes the same numbers without the awk:
lvs -o lv_name,cache_used_blocks,cache_dirty_blocks,cache_read_hits,\
cache_read_misses,cache_write_hits,cache_write_misses vg/main
bcache adds decaying windows on top of the running totals, which is better instrumentation than dm-cache offers:
cat /sys/block/bcache0/bcache/stats_hour/cache_hit_ratio
cat /sys/block/bcache0/bcache/stats_five_minute/bypassed
Read bypassed beside the ratio. The kernel documentation defines it as the
I/O, reads and writes both, that has gone round the cache entirely, most of
which will be the long runs past the 4 MiB sequential_cutoff, and a healthy
media server shows a large number there. The same documentation is explicit
that hits are counted per individual I/O and that a partial hit is counted as
a miss, so a mismatch between request size and cache block size depresses the
ratio without depressing the benefit.
ZFS separates the two caches. Read the kstat file, not a tool’s output:
awk '/^l2_(hits|misses|size|hdr_size)/ {print $1, $3}' /proc/spl/kstat/zfs/arcstats
arcstat 1 gives the live view, renamed to zarcstat in OpenZFS 2.4.0 along
with arc_summary becoming zarcsummary. Check the ARC hit ratio first.
Anything that reaches L2ARC at all is something the ARC missed, and if the board
still has empty slots, the cheaper fix is
RAM that fits it.
On Windows the equivalent is PerfMon. For storage bus cache and Storage Spaces
Direct, compare Cache Miss Reads/sec on the Cluster Storage Hybrid Disk
object against total read IOPS, one instance per capacity drive. For heat-based
tiering, Microsoft’s storage tiers guidance has you run defrag D: /g /h /#,
which returns a Storage Tier Optimization Report giving the percentage of I/O
the SSD tier actually served, the percentage expected for that tier size, and
what tier size would reach a target. That is a measured miss-ratio curve, which
is more than Linux hands you out of the box. PerfMon’s counters are rates and
averages rather than percentiles, so pair it with DiskSpd when the tail matters.
Eighty per cent can be a failure and ninety-nine is transformative
Put the ratio you have just read back into the model from the top of the guide: 2.48 ms at 80 per cent, 1.29 ms at 90, 0.695 ms at 95, 0.219 ms at 99.
Above about 90 per cent, h * L_fast is noise and L_eff is simply the miss
rate times the HDD. Think in misses, not hits. Eighty per cent means one
request in five still costs 12 ms, and the mean stays in milliseconds; the
cache is working exactly as designed and the user still feels a hard disk.
Ninety-nine per cent is nearly 55 times faster than the same array without it.
Almost all the value sits in the last few per cent, which is where most people
stop tuning.
Two corrections follow. Mean latency is the wrong summary: at h = 0.99 the 99th percentile sits on the boundary between a hit and a miss and everything above it is a miss, so p99.9 is about 12 ms however good the average looks. And a ratio without an I/O rate beside it is meaningless - 99 per cent of 20 IOPS is not a result.
Has the spindle actually been spared
iostat -x 1 on the backing device, the cache device and the composite mapping
at once. The healthy pattern is high r/s on the composite and the NVMe, low
r/s on the HDD, and a rareq-sz on the HDD that has grown, because what
reaches it should now be migrations and destages rather than scattered 4 KiB
reads. Watch r_await and w_await, which include queue time. Ignore %util
on the NVMe entirely: iostat(1) states that for devices serving requests in
parallel “this number does not reflect their performance limits”. It is routinely
misread as a saturation gauge.
fio profiles that measure your workload, not your cache
Three rules separate a useful run from a worthless one. The file must be larger than the cache, or you are testing the cache’s ability to hold everything. The distribution must be skewed, because uniform random over a 4 TB file returns a hit ratio of about cache size over file size, which measures nothing. And the warm-up must be discarded.
fio --name=hotset --filename=/mnt/cached/testfile \
--size=2T --io_size=200G \
--rw=randread --bs=4k --iodepth=32 --numjobs=4 \
--ioengine=io_uring --direct=1 \
--random_distribution=zipf:1.2 --norandommap --randrepeat=0 \
--ramp_time=300 --runtime=1800 --time_based \
--percentile_list=50:90:99:99.9:99.99 --group_reporting
--ioengine=io_uring wants a recent fio and kernel; libaio is the safe
substitute. Report clat percentiles, not bandwidth. Then run the same profile
as --rw=randwrite with --fdatasync=1 to see what power-loss protection is or
is not doing for your commit latency, and --refill_buffers so compression on
the SSD does not flatter the result.
A sequential benchmark will show the cache doing nothing, and that is the
cache behaving correctly. bcache’s sequential_cutoff bypasses runs past
4 MiB; Storage Spaces documents that writes larger than 256 KB are not written
to the write-back cache; ZFS ships l2arc_noprefetch=1, which the OpenZFS
manual page defines as not writing buffers to L2ARC if they were prefetched but
not used by applications. A 1 MiB sequential read test is a measurement of the HDD, by
design, and a modern high-capacity HDD already does 250 MB/s or better on its
outer tracks. Reporting that number as a cache failure is reporting a safeguard
as a bug.
Every first measurement is a lie
Caches start empty. L2ARC fills at l2arc_write_max per feed interval, 32 MiB
against a one-second l2arc_feed_secs since OpenZFS 2.2.7, so a 1 TiB cache needs
1 TiB / 32 MiB/s = 32,768 s, about 9.1 hours of continuous feeding at the
ceiling, and the feeder only takes what is near the list tails. On the 2.0 and
2.1 default of 8 MiB the same terabyte took around 36 hours, which is why so
much older advice says days. bcache warns that if a btree node is full a cache
miss cannot insert a key for the new data, and its own documentation’s remedy is
to “warm the cache by doing writes”. Windows heat-based tiering does not move
anything until the 01:00 optimisation task runs, so a benchmark at noon measures
last night’s heat map.
Prove it earned its place by difference, not by belief: record the counters,
nvme smart-log /dev/nvme0 and a p99 figure before, run the real workload for
a week, and read all three again. data_units_written, reported in thousands of
512-byte units and so 512,000 bytes apiece, tells you what the cache cost in
flash, which is the other half of the ledger and the subject of
SSD endurance: TBW, DWPD, how worried to be.
If the hit ratio is 60 per cent and percentage_used moved two points in a
week, the honest conclusion is that you bought wear. Devices sized for that
duty cycle are filtered under enterprise SSDs.
Safeguards
The three shapes from the top of this guide - read cache, write cache, tier - are three different insurance claims, and the separation only earns its keep here. Losing a read cache costs you time. Losing a writeback cache costs you data. Losing a tier costs you the data that lived only there. Everything below follows from which of the three you built.
The uncomfortable part first. “No data loss” and “keeps working” are not the
same claim, and the documentation for the most popular Linux stack only makes
the first one. lvmcache(7) says of writethrough that “the loss of a device
associated with the cache in this case would not mean the loss of any data” -
true, and it is the origin that holds it. But in dm-cache-target.c, ordinary
bios remapped to the cache device have their bi_status propagated straight up;
there is no fallback-to-origin path for a read hit. A dead cache device under a
writethrough dm-cache would therefore return EIO on every read hit until you
detach it. The data is intact and the volume is unusable. That behaviour is a
reading of the source and not a documented guarantee - the man page does not
state it - so treat it as one, and rehearse the detach.
Power-loss protection and what the capacitors actually buy
A capacitor bank on an SSD guarantees one thing: that data the controller has already acknowledged, sitting in the drive’s own DRAM, reaches NAND after mains power goes away. It guarantees nothing about data still in the host page cache, nothing about a filesystem that never issued a FLUSH, and nothing about ordering above the drive.
What it buys in practice is two separate things that are routinely conflated. The first is latency: with PLP the drive can complete a FLUSH as soon as the data is in DRAM. One published set of fio measurements of fsync after a 16 KiB write puts a Samsung PM9A3 at 1.6 µs and a Solidigm D7-P5520 at 12.4 µs, against 891 µs for a Crucial T500 and 2,974 µs for a Samsung 990 Pro, both consumer drives. A writeback cache issues that operation constantly. A 3 ms fsync on the NVMe is the same order of magnitude as the 4.16 ms average rotational latency of a 7200 rpm disk (half of one 8.33 ms revolution) - the delay the cache was bought to hide - which is close enough to make the whole build pointless.
The second is correctness insurance, and it matters because you cannot verify FLUSH honesty from a datasheet. A 2022 power-pull test of four consumer NVMe drives found two of them - a Sabrent Rocket and an SK hynix Gold P31 - losing data that had already been flushed; the Samsung and WD parts did not. The older Ohio State and HP Labs fault-injection study (FAST ’13) tested 15 SSDs and found 13 of them failing in ways their interface said they would not, including shorn writes and unserialisable writes. If a cache device will ever hold the only copy of an acknowledged write, buy one with capacitors on the datasheet. The enterprise SSDs listing is where those drives sit, though the class marks duty rating rather than capacitors, so the datasheet for the exact model is still the thing to read.
Do not assume a read cache is exempt. Ahmadian, Taheri and Asadi’s fault-injection work on SSD I/O caches (arXiv:1912.01555) found that “read accesses to the I/O cache are subjected to failure in presence of sudden power outage”, because a promotion is itself a write, and reported the failure rate growing “by more than 14X” as request size fell.
Mirroring, and which stacks let you
| Stack | Mirrored cache | Mechanism |
|---|---|---|
| LVM dm-cache | Yes | lvmcache(7) documents lvcreate --type raid1 -m 1 for the fast LV; with a cache pool the data and metadata sub-LVs can each be raid1 |
| LVM dm-writecache | Yes | the same raid1 cachevol |
| bcache | Not natively | MAX_CACHES_PER_SET is 8 in the source, but the admin guide says “multiple caches per set isn’t supported yet”; an md RAID1 underneath is common practice, not documented policy |
| ZFS L2ARC | No, and unnecessary | “Cache devices cannot be mirrored”; a read error is reissued to the pool device |
| ZFS SLOG | Yes | log vdevs can be mirrored; raidz is not supported for the intent log |
| ZFS special vdev | Mandatory | not a cache; its redundancy “should match the redundancy of the other normal devices in the pool” |
| S2D cache | By architecture | cache sits below the rest of the storage stack; Microsoft asks for two cache drives per server to hold performance when one dies |
Two copies of the same model, bought together and written identically, wear identically and reach end of life together. Stagger the pair, or at least watch both wear indicators and replace one early; the arithmetic for that is in SSD endurance: TBW, DWPD, how worried to be.
Unclean shutdown, per stack
dm-cache is the sharp one. The kernel documentation is explicit that the dirty bit is a hint - it “changes far too frequently for us to keep updating it on the fly” - and that “if the system crashes all cache blocks will be assumed dirty when restarted”. After a crash there is no partial drain. Every block in the cache must be written back before the cache can be detached.
1 TiB writeback cache, all blocks presumed dirty after a crash
HDD sustained sequential write (IronWolf Pro 20 TB spec): 285 MB/s
1,099,511,627,776 B / 285,000,000 B/s = 3,858 s = 64 min
That is a floor, not an estimate: a destage is not one long sequential write, so the real figure is worse, and it lands precisely when you are least able to wait. Rehearse it before you need it.
dm-writecache bounds the window explicitly. It commits on FLUSH or FUA, or
after autocommit_blocks (default 65536 in SSD mode) or autocommit_time
(default 1000 ms). Those two defaults are the durability parameter; lower them
if the exposure matters more than the throughput.
bcache treats unclean shutdown as the normal case: “bcache simply doesn’t return writes as completed until they’re on stable storage.” Its dangerous state is a missing cache device, not a crash.
ZFS has no writeback block cache to recover, which is its strongest safety argument in this comparison. S2D loses only un-destaged writes, and only on the local server, because the other copies live elsewhere.
Removal, done correctly
Pulling a cache device without detaching it is how people lose pools.
# LVM dm-cache: flush and keep the fast LV, or flush and delete it
lvconvert --splitcache vg/main
lvconvert --uncache vg/main
# LVM dm-writecache: pre-flush first so --splitcache does not block for hours
lvchange --cachesettings 'cleaner=1' vg/main # wait for it to drain
lvconvert --splitcache vg/main
# bcache: detach flushes; confirm it drained before stopping the set
echo 1 > /sys/block/sdb/sdb1/bcache/detach # the BACKING device's dir
cat /sys/block/bcache0/bcache/dirty_data # wait for 0
echo 1 > /sys/fs/bcache/CSET-UUID/stop
# ZFS: a cache vdev leaves with no ceremony
zpool remove tank nvme0n1
Raw dm-writecache has its own six-step sequence in the kernel doc: send
flush_on_suspend, load an inactive table with a linear target that maps to the
underlying device, suspend, check the status for errors, resume, then delete the
cache device. Microsoft’s storage bus cache documentation gives no rebind
cmdlet - Remove-StorageBusBinding then New-StorageBusBinding, and the
existing read cache is lost.
When the device dies anyway
- dm-cache, writethrough or passthrough. Data is intact; LVM will still
fight you, because
--uncacherefuses while a PV is missing. The community path isvgreduce --removemissing --forcethenlvconvert --uncache --force. This is forum and issue-tracker guidance, not vendor guidance, and it is version-sensitive. Test it on your LVM version first. - dm-cache, writeback. Assume loss. Writeback loses ordering as well as content, so the filesystem is in an arbitrary state, not merely a stale one. Restore from backup. The drill is short because there is nothing else to do.
- bcache, writethrough or writearound.
echo 1 > /sys/block/sdb/bcache/runningbrings the backing device up alone, or/sys/block/sdb/sdb1/bcache/runningwhere the backing device is a partition; register a new cache and attach. - bcache, writeback. The same command works, and the documentation warns that with dirty data in the cache you “will have massive filesystem corruption”. Mount read-only, assess, restore. If the old cache device reappears its contents are invalidated; do not be clever.
- ZFS L2ARC.
zpool remove,zpool add, no downtime, no risk. - ZFS special vdev. Replace the failed member now. If it is lost, the pool is lost; and on a pool whose primary storage holds a raidz or draid vdev it cannot be removed at all, mirrored or not. The real drill was not adding it that way.
- bcachefs. Set
durability=0on the cache devices, which is what the documentation describes for writethrough caching: copies there do not count towards the replica target, and the device can leave without taking data with it. Without it, a write can be durable only on the fast tier until the background job moves it, so a cache failure inside that window is a data failure.
A cache is not a backup, and a mirror is not a backup either. A mirrored writeback cache survives one dead NVMe; it does not survive a deleted file, a bad write the filesystem accepted, or the controller that killed both halves. Every configuration in this guide assumes a restorable copy somewhere else.
What caching does to the NVMe drive
The cache device’s wear rate has very little to do with how much the workload writes. A read-only job that issues no writes at all can still push terabytes a day onto the flash, because every promotion the policy decides on is a write. The hard drives behind the cache are indifferent to this. The NVMe is not, and it is the only part of the stack that is being consumed.
Every promotion is a write, and promotions are chunk-sized
dm-cache promotes in whole chunks. LVM’s default chunk size is 64 KiB
(DEFAULT_CACHE_POOL_CHUNK_SIZE in lib/config/defaults.h), so a 4 KiB
read miss that the policy elects to promote writes 64 KiB to flash -
sixteen times the bytes the application asked for. It then gets worse
quietly. LVM caps chunk count at allocation/cache_pool_max_chunks,
default 1,000,000, and raises the chunk size to stay under that cap, so a
large cachevol does not get 64 KiB chunks. Check rather than assume:
lvs -o+chunksize vg/main
At 1 MiB chunks, that same 4 KiB promotion costs 256 times its size.
bcache is log-structured and packs writes into buckets instead of doing read-modify-write, so small writes do not pay bucket-sized amplification; it pays in garbage collection and btree traffic instead. No published head-to-head measurement of NAND writes across dm-cache, bcache and dm-writecache is easy to point at, and this guide will not invent a ranking.
Two designs bound the damage on purpose. dm-writecache never promotes on
read - the kernel documentation is explicit that it “doesn’t cache reads
because reads are supposed to be cached in page cache in normal RAM” -
so read traffic costs the flash nothing at all. ZFS caps L2ARC’s fill
rate outright: l2arc_write_max defaults to 33554432 bytes per
l2arc_feed_secs interval, and that interval defaults to one second, so
32 MiB/s since OpenZFS 2.2. Sustained, that is 2.64 TiB per day, and once the
cache is warm that is
the ceiling on cache wear, the best argument for L2ARC that nobody
makes. While the device is still cold, l2arc_write_boost (32 MiB by
default) is added to l2arc_write_max and l2arc_feed_again shortens
the feed interval to l2arc_feed_min_ms, so warm-up writes faster than
the steady-state figure. OpenZFS master documents a further l2arc_dwpd_limit, default
100, meaning 1.0 drive write per day with unspent budget carried over.
That is unreleased at the time of writing; do not plan around it yet.
Mirroring the cache doubles the figure. A raid1 cachevol writes every
promotion twice, to two drives that wear at the same rate and were
probably bought at the same time.
The arithmetic, with your numbers substituted
Promotions accepted by the policy 25 /s (measure: the
promotions counter
in dmsetup status)
Chunk size from lvs -o+chunksize 1 MiB
Bytes written to flash per second 25 MiB/s = 26.2 MB/s
Per day 26.2 MB/s x 86400 = 2.26 TB
Samsung 990 Pro 2 TB, 1200 TBW rating:
1200 TB / 2.26 TB/day = 531 days ~ 17 months
Solidigm D7-P5810 800 GB, 50 DWPD x 5 y = 73 PB:
73000 TB / 2.26 TB/day = 32301 days ~ 88 years
Twenty-five promotions a second is a quiet cache. The workload in that example may be writing nothing at all.
One subtlety that makes the consumer figure worse than it looks: TBW is normally quoted in host writes under an assumed JEDEC client workload. The cache-layer amplification above eats that budget directly, byte for byte. The amplification inside the drive’s own flash translation layer does not show up in the TBW arithmetic at all - it erodes the assumption the rating was built on, because a cache’s write pattern is more random and more sustained than the pattern the rating assumed. SSD endurance: TBW, DWPD, how worried to be works through what those ratings do and do not promise.
The drive that belongs in this slot
Pick by write budget, not by read IOPS. A cache device’s job description is “absorb writes indefinitely”, which is the DWPD column.
Power-loss protection is the second filter, and the safeguards section above gives both halves of it at length: the published fault-injection work, in which most of the drives tested failed in ways their interface said they could not, and the fsync gap, which puts protected drives in the tens of microseconds against hundreds or thousands for consumer parts. You cannot verify a drive’s flush honesty from its datasheet, and a writeback cache performs that operation constantly. Those consumer figures sit in the same millisecond range as the 4.16 ms average rotational latency of the 7200 rpm drive being accelerated.
Which is why a used enterprise drive is usually the better purchase than a new consumer one at the same price. The right ones have the capacitors and a write-intensive endurance rating, and a drive that has spent 20 per cent of its rated life with a 50 DWPD budget has more writes left in it than a new 0.33 DWPD consumer drive has in total. Enterprise drive pulls covers reading the wear on a second-hand one before you buy, and the power-loss-protection filter on the enterprise SSD listing is the shortlist.
Then over-provision it. Leave 10 to 20 per cent of the NVMe unpartitioned and never written; the FTL uses it as spare and the drive’s own write amplification factor (WAF) falls. Capacity for lifetime is the right trade on a cache.
Watching the wear counters
nvme smart-log /dev/nvme0 -H
data_units_written is in units of 1000 x 512 bytes, so multiply by
512,000 for bytes. percentage_used is a vendor estimate updated once
per power-on hour and can exceed 100 per cent without the drive failing.
Watch available_spare against available_spare_threshold, and
critical_warning bit 3 - the bit that reports the media has been placed in
read-only mode.
Record data_units_written weekly and difference it. A single reading
tells you nothing; the slope tells you the replacement date. Divide by
seven, project against the TBW rating, and compare against the hard
drive’s own SATA attribute 241 where the drive exposes it, to get the
destage ratio. For true device WAF you need NAND-level writes, which base
NVMe does not expose: nvme ocp smart-add-log (log page C0h) gives
physical media units written on OCP-compliant datacentre drives, and
nvme intel smart-log-add gives nand_bytes_written against
host_bytes_written on Intel and Solidigm parts.
Budget for the replacement at build time. The hard drives are storage and may outlive the machine. The cache device is a wear part with a calculable service life, and the correct question is not whether it will be replaced but in which year.
Four builds, end to end
Four complete recipes follow. Each names the hardware, the commands in order, the check that proves it worked, and the way out. Read the removal procedure before you run the creation commands. Three of these four can be undone, though recipe 3 must flush its dirty data first and recipe 4 destroys the volume on the way out. The second cannot be undone at all once the pool uses raidz, and that asymmetry is the most important fact on this page.
Recipe 1: Linux home server, LVM cache in writethrough
For: a general file server whose pain is read latency - photo libraries, media metadata, package mirrors, a Samba share that stalls on directory walks. Not for: sync-heavy workloads. Writethrough completes nothing until the origin has it, so every write still costs a seek. If your complaint is fsync latency, this recipe changes nothing and you want recipe 3 or a dm-writecache.
Hardware: two CMR hard drives of equal size, mirrored, and one NVMe SSD in M.2. Writethrough never holds the only copy of anything, so a single consumer drive from the M.2 NVMe listing is defensible here in a way it is not anywhere else in this article. Partition only 80 to 90 per cent of it and leave the rest untouched as spare area for the controller.
pvcreate /dev/sda /dev/sdb /dev/nvme0n1p1
vgcreate vg0 /dev/sda /dev/sdb /dev/nvme0n1p1
# mirrored origin across the two spindles; sized by hand on 4 TB disks so the
# volume group keeps a little slack for later
lvcreate --type raid1 -m 1 -n main -L 3.5T vg0 /dev/sda /dev/sdb
# cache pool: separate data and metadata sub-LVs, which lvmcache(7) credits
# with slightly better dm-cache performance than a single cachevol
lvcreate --type cache-pool -L 180G -n fast --chunksize 256K vg0 /dev/nvme0n1p1
lvconvert --type cache --cachepool vg0/fast \
--cachemode writethrough --cachepolicy smq vg0/main
mkfs.xfs /dev/vg0/main
Writethrough is already lvm2’s default cache mode. The flag is on the attach step so the mode is stated in the command rather than assumed by the reader.
Verification, and one of these two lines is mandatory rather than decorative:
lvs -o lv_name,cache_mode,cache_policy,chunksize,\
cache_used_blocks,cache_read_hits,cache_read_misses vg0/main
dmsetup status --target cache --noflush /dev/mapper/vg0-main
Check chunksize against what you asked for. LVM caps the chunk count at
allocation/cache_pool_max_chunks, which lvm2 treats as 1,000,000 when the
setting is left unset, and raises the chunk size to stay under that cap - a large
cache pool does not get the compiled-in default, and the default itself is
worth reading back with lvmconfig --type default rather than assumed. The hit
counters are 32-bit and they wrap, but they do not reset when the LV is
reactivated: dm-cache saves them into the cache metadata on suspend and reads
them back on activation, so sample twice and difference. Use --noflush:
cache_status() in
dm-cache-target.c commits metadata on every status read that omits it, and a
one-second poll therefore writes to the NVMe once a second forever.
Removal:
lvconvert --splitcache vg0/main # flush, detach, keep the fast LV
lvconvert --uncache vg0/main # flush, detach, delete the fast LV
If the NVMe dies first, no data is lost - but the target has no
fall-back-to-origin path, so the LV errors rather than quietly serving from the
spindles, and it stays unusable until the cache is detached. lvconvert --uncache refuses while the PV is missing; the community route is to run
vgreduce --removemissing --force first. That sequence is not in any man page
and its behaviour varies by lvm2 version, so rehearse it on a loopback pair
before you need it.
Recipe 2: ZFS pool with a mirrored special vdev
For: a pool of many spindles where the slow operations are metadata
operations - find, rsync traversals, snapshot listing, scrub and resilver,
backup scans. This is the highest-value build in this article, because a special
vdev is permanent placement rather than a cache: no warm-up, no ARC header cost,
and it serves blocks on first access, which an L2ARC by construction never can.
Not for: anyone who cannot mirror it. It is not a cache. OpenZFS is blunt
that the blocks routed there exist only there, and losing it loses the pool
exactly as losing any other top-level vdev would. The same reasoning drives
drive selection generally, covered in
what differs from a desktop.
Hardware: six hard drives in raidz2, plus two NVMe SSDs with power-loss protection, mirrored. Size the special vdev from measurement, not from the 0.3-per-cent figure that circulates without a first-party source:
zdb -bbb tank | less # space by block type; read off the metadata total
Note the caveat before you trust the output: PSIZE does not account for multiple metadata copies, and ASIZE includes raidz overhead. Leave generous headroom, because the vdev cannot be shrunk.
zpool create -o ashift=12 tank raidz2 \
/dev/disk/by-id/ata-HDD1 /dev/disk/by-id/ata-HDD2 \
/dev/disk/by-id/ata-HDD3 /dev/disk/by-id/ata-HDD4 \
/dev/disk/by-id/ata-HDD5 /dev/disk/by-id/ata-HDD6
zpool add tank special mirror \
/dev/disk/by-id/nvme-SSD_A-part1 /dev/disk/by-id/nvme-SSD_B-part1
zfs create tank/media
zfs set recordsize=1M tank/media
zfs create tank/vms
zfs set special_small_blocks=32K tank/vms
special_small_blocks defaults to 0, meaning metadata only. Raise it per
dataset, never pool-wide by reflex: the threshold is compared against the
physical size after compression, so a 128 KiB record that compresses to
20 KiB counts as 20 KiB, and setting the threshold near your recordsize on a
well-compressing dataset can pull essentially all of it onto the flash.
zfs_special_class_metadata_reserve_pct defaults to 25, so small blocks stop
being accepted once the class is 75 per cent full; metadata alone can still fill
it.
Verification:
zpool list -v tank # per-vdev capacity; watch the special vdev fill
zpool status tank # both mirror members ONLINE
Removal takes the vdev name as zpool status prints it, so a mirrored special
vdev is removed as zpool remove tank mirror-1, and that works only if no
top-level vdev in the pool is raidz or draid, ashift is uniform, and
device_removal is enabled. On the raidz pool above it can never be removed:
the attempt fails with a message about all top-level vdevs needing the same
sector size and none being raidz. Adding one is a permanent decision, which is
why the mirror is not optional.
Recipe 3: Linux workstation, bcache in writeback
For: a single-user machine doing heavy small random writes - compile trees, CI workspaces, container image layers, a scratch dataset you can rebuild. Not for: anything whose loss would matter. In writeback the cache device holds the only copy of acknowledged data, and the kernel documentation is explicit that writeback defaults to off because “in writeback mode you’ll lose data if something happens to your SSD”.
Hardware: one hard drive from the hard drive listings, and two NVMe drives with
power-loss protection, mirrored with md beneath bcache. bcache has no native
cache-set mirroring - MAX_CACHES_PER_SET is 8 in the format but multiple
caches per set is not supported - so md RAID1 is the only route. PLP is the
thing to filter on rather than a luxury, and it is an
enterprise-class feature: published fsync
measurements put PLP drives in the tens of microseconds against roughly 0.9 ms
and 3 ms for two consumer drives, and a 3 ms fsync is no better than a 7200 rpm
disk’s 4.16 ms average rotational latency by any margin that justifies the
build.
mdadm --create /dev/md10 --level=1 --raid-devices=2 \
/dev/nvme0n1p1 /dev/nvme1n1p1
make-bcache -B /dev/sdc -C /dev/md10 -b 512k
# (bcache-tools 1.0.8 defaults the bucket to 512 KiB; the shipped man page
# still says 128k. Set -b explicitly rather than trusting either. Newer
# packages ship the unified front end instead: bcache make -B ... -C ...)
mkfs.ext4 /dev/bcache0
echo writeback > /sys/block/bcache0/bcache/cache_mode
echo always > /sys/block/sdc/bcache/stop_when_cache_set_failed
The safeguards this mode demands, all readable and settable in sysfs:
cat /sys/block/bcache0/bcache/dirty_data # bytes at risk, right now
cat /sys/block/bcache0/bcache/stats_hour/cache_hit_ratio
cat /sys/fs/bcache/CSET-UUID/internal/cutoff_writeback # 40
cat /sys/fs/bcache/CSET-UUID/internal/cutoff_writeback_sync # 70
writeback_percent defaults to 10 and is clamped to bch_cutoff_writeback,
default 40 - a larger value is silently clamped down to 40, not rejected. Past
40 per cent in use,
only sync and metadata writes are still written back and everything else
degrades to writethrough; past 70 per cent, writeback is refused outright. That
clamp, not a bug, is the mechanism behind “my writeback cache stopped absorbing
writes”. sequential_cutoff defaults to 4 MiB and keeps large streams off the
flash; set it to 0 only while benchmarking.
Removal, in this order:
echo 1 > /sys/block/sdc/bcache/detach # flushes dirty data first
cat /sys/block/bcache0/bcache/dirty_data # must read 0 before continuing
echo 1 > /sys/fs/bcache/CSET-UUID/stop
wipefs -a /dev/md10
Detach alone leaves the cache registered and protected; the stop is what frees
it. If the cache is already dead, echo 1 > /sys/block/sdc/bcache/running forces
the backing device up alone - mount it read-only and assess, because the kernel
documentation’s warning applies exactly here: with dirty data in the cache, do
not expect the filesystem to be recoverable.
Recipe 4: Windows tiered Storage Space, in PowerShell
For: an NTFS volume on Windows Server whose hot region is a stable minority of a large dataset. Not for: real-time acceleration. Movement is scheduled, not live - a heat map built through the day and a task that moves sub-file extents at 01:00. Nothing you write at noon is on the SSD tier before tonight.
Note two hard constraints before ordering drives: the virtual disk must use fixed provisioning, and the column count is identical on both tiers, so a four-column two-way mirror wants eight SSDs and eight hard drives.
$sub = (Get-StorageSubSystem -FriendlyName "Windows Storage*").FriendlyName
New-StoragePool -FriendlyName Pool1 -StorageSubSystemFriendlyName $sub `
-PhysicalDisks (Get-PhysicalDisk -CanPool $true)
# set MediaType by hand on any disk that reports Unspecified
Get-PhysicalDisk | Where-Object MediaType -eq Unspecified |
Set-PhysicalDisk -MediaType HDD
New-StorageTier -StoragePoolFriendlyName Pool1 -FriendlyName SSDTier -MediaType SSD
New-StorageTier -StoragePoolFriendlyName Pool1 -FriendlyName HDDTier -MediaType HDD
New-Volume -StoragePoolFriendlyName Pool1 -FriendlyName Data -DriveLetter D `
-ResiliencySettingName Mirror -ProvisioningType Fixed `
-StorageTiers (Get-StorageTier -FriendlyName SSDTier),
(Get-StorageTier -FriendlyName HDDTier) `
-StorageTierSizes 200GB, 4TB -FileSystem NTFS
Verification, and a manual first run so you are not waiting until 01:00:
Get-StorageTier | Format-Table FriendlyName, MediaType, Size
Optimize-Volume -DriveLetter D -TierOptimize
defrag D: /g /h /# # /h runs at normal priority, /# writes the report
Get-FileStorageTier -VolumeDriveLetter D
schtasks /change /tn "\Microsoft\Windows\Storage Tiers Management\Storage Tiers Optimization" /ri 360
The report is the measurement: it states the share of I/O the SSD tier actually
served against what was expected for that tier size, and what tier size would
reach a target share - effectively a measured miss-ratio curve, which Linux
gives you nothing comparable to. Microsoft puts the recommended maximum
frequency at every 6 hours and says there is nothing to gain from running the
task more often than that. Two behaviours to hold on to: writes larger than
256 KB are not written to the 1 GB write-back cache, by design, so large
sequential writes land on the spindles; and Set-FileStorageTier does not move
a file immediately and excludes it from heat optimisation entirely thereafter.
Removal is ordinary and destructive to the volume, so back up first:
Remove-VirtualDisk -FriendlyName Data
Remove-StorageTier -FriendlyName SSDTier
Remove-StorageTier -FriendlyName HDDTier
Remove-StoragePool -FriendlyName Pool1
On Windows 11 and 10 the Control Panel does not expose tiers at all; the PowerShell path above is the only one, and Microsoft documents no supported client-edition tiering scenario.
What a small business should actually deploy
The largest caching lever at this tier is already in the server, already paid for, and usually left on a default nobody chose. It is the RAID controller’s own DDR4, and the second largest is the DIMM slots that were left empty to hit a quote. Neither is a purchase decision about a cache device. Both are configuration decisions, and both are reversible on a Tuesday morning.
The uncomfortable part comes first. The entry controller that turns up in SMB server quotes often cannot do any of this. Broadcom’s 94xx MegaRAID and HBA Tri-Mode Storage Adapters User Guide (pub-005851) lists the 9440-8i with cache “N/A” and states that “the MegaRAID 9440-8i Tri-Mode storage adapter and the HBAs do not support CacheVault data protection”. The same guide gives the 9460-8i 2 GB of DDR4-2133 and the 9460-16i 4 GB, both backed by the CVPM05 supercapacitor module. A cacheless controller in front of a mechanical array is a write-through array whatever the datasheet implies, and on RAID 5 or 6 a small write that does not fill a stripe then pays a read-modify-write at seek latency.
The five shapes, and what each one actually costs
| Shape | Money | Operational load | What it buys | The failure that happens |
|---|---|---|---|---|
| One server, cached hardware RAID | Controller step-up plus energy pack | Low, one-off | Write coalescing, RAID 5/6 parity absorption | Supercap cooks, or a relearn drops the array to write-through |
| One server, Linux, RAM instead of a controller | DIMMs | Low, but needs a DBA’s attention | Read hits at DRAM latency, bigger buffer pools | fsync still lands on spindles; the write path is untouched |
| ZFS box, ARC sized properly | DIMMs, ECC | Low | The best return per DIMM at this tier | A 02:00 backup scan evicts Monday’s working set |
| Two-node cluster | Licences, two cache drives per node, RAM for metadata | High | Node-local read acceleration, survivable writes | Failover is cold; the RAM cache does not move |
| NAS appliance, read-write cache | Two SSDs, a UPS | Low to buy, high to recover | Appliance-level acceleration | A single cache SSD holding dirty data dies |
One server with a cached RAID controller
Three tiers of RAM sit under an SMB array, and only the middle one is both
fast and safe: host DRAM, the controller’s backed DDR, and the drives’ own
unbacked DRAM, commonly 64 MB to 256 MB. Dell’s Server Administrator
storage guide documents an asymmetry that integrators rarely look at: the
default disk cache policy is Enabled for virtual disks built on SATA drives
and Disabled for SAS. StorCLI exposes the same switch as pdcache.
storcli /cx/vx set wrcache=wt|wb|awb # write through | back | ALWAYS back
storcli /cx/vx set rdcache=ra|nora
storcli /cx/vx set iopolicy=cached|direct # direct is the usual default
storcli /cx/vx set pdcache=on|off|default # the drives' own unbacked DRAM
storcli /cx/cv show all # CacheVault state
awb is the landmine. It keeps write-back switched on when the energy
backup is missing, failed or relearning, and it is exactly what an
integrator reaches for when the battery warning appears. It converts a
temporary latency complaint into permanent silent exposure.
The failure that will actually happen is not the capacitor’s age. Broadcom gives the CVPM operating range as 0 °C to 55 °C ambient and repeats that the system has to be designed with airflow that keeps the module inside it. A remote-mounted supercapacitor cable-tied behind a tower’s blanked-off slot is the classic SMB death. The second failure is procedural: Dell’s PERC 9 guide describes preserved dirty cache from a virtual disk that went offline, and that preserved cache blocks the creation of new virtual disks until it is imported or discarded, with a data-loss warning attached to discarding it in the wrong order.
On relearns, be precise. Dell KB 000141687 documents a Transparent Learn
Cycle every 90 days and says it was first introduced with PERC 8; Dell’s
PERC user guides add that “virtual disks stay in Write-Back mode, if
enabled, during transparent learn cycle”. On PERC H700 and earlier, and on plain
MegaRAID cards with a genuine battery under the default
No Write Cache if Bad BBU policy, the array silently drops to
write-through for as long as a discharge and recharge takes. The commonly
repeated 60-day LSI relearn interval turns up in third-party writeups
rather than in the vendor documentation consulted here, and this guide will
not claim it as Broadcom’s figure.
One Linux server where the money goes into DIMMs
Free memory is already a read cache. The write-side knobs are not a cache
in any durable sense: vm.dirty_ratio and vm.dirty_background_ratio
buffer dirty pages in volatile DRAM, and fsync() still has to reach
stable storage. That is the whole reason 2 GB of backed controller cache
can beat 64 GB of page cache on anything transactional.
The biggest return here is usually one line in a config file. MySQL ships
innodb_buffer_pool_size at 134217728 bytes, which is 128 MB; the MySQL
manual’s own advice for a
dedicated database server is to set the buffer pool to 80 per cent of
physical memory, which is a recommendation and not a default. PostgreSQL ships
shared_buffers at
128 MB and its documentation treats 25 per cent of RAM as a starting point,
with the caveat that beyond roughly 40 per cent it rarely beats leaving the
memory to the kernel. A server running a production database on stock
128 MB buffers does not have a disk problem.
If a DRAM tier is built explicitly with Open CAS, price the metadata before the DIMMs:
Open CAS RAM requirement, default 4 KiB cache line:
1 GiB + 2.05% of cache device capacity (the documented formula)
a 100 GiB RAM tier -> ~3.1 GiB of further RAM just for metadata
a 1 TB cache device -> ~21.6 GB
raising the cache line to 64 KiB divides the 2% term by 16
the two worked figures are arithmetic from the formula, not vendor tables
The Open CAS documentation is candid that a RAMdisk tier is volatile, must
be rebuilt after every restart, and that “if the RAMdisk cache tier is
started in write-back mode and there is a dirty shutdown, data loss may
occur”. Run it in wt or wa and it carries little durability risk of its
own. Run it in wb and it is a fast way to lose a morning.
A ZFS box, where memory does the most work
ARC holds most blocks in their compressed on-disk form, so it caches more
logical data than its size suggests, and it sizes itself from installed
memory: zfs_arc_max defaults to 0, which means half of RAM on Linux up to
OpenZFS 2.2, and from OpenZFS 2.3 the larger of RAM minus 1 GiB and 5/8 of RAM.
zfs_arc_min
defaults to the larger of 32 MiB or RAM divided by 32. The
per-dataset primarycache and secondarycache properties are the lever
that stops a nightly backup walk evicting the working set, which is the
failure that actually happens on these boxes: nothing breaks, Monday is
just slow. Filling the empty slots with registered ECC modules, on a board
that takes them, is the simplest upgrade in this article, and
what the board accepts is a
five-minute check before any cache device is priced.
Two nodes, where the failure domain becomes the design
Microsoft states the constraint that governs every clustered design in one
sentence: for the CSV in-memory read cache, “writes cannot be cached in
memory”. Microsoft’s documentation, as it stood in September 2026, gives
the default as 1 GiB per server on Windows Server 2019 and Azure Local and
0 on 2016, raised to as much as 80 per cent of physical memory with
(Get-Cluster).BlockCacheSize, in MiB. It is server-local, so a failover
starts cold.
The durable half of the cache must be on devices, and it bills in RAM:
S2D storage pool cache metadata (Microsoft hardware requirements,
as documented at the time of writing):
4 GB of RAM per TB of cache drive capacity, per server
2 x 1.6 TB NVMe = 3.2 TB -> 12.8 GB of RAM before a VM starts
Minimums: two cache drives per server; 2 cache + 4 capacity drives;
cache devices 32 GB or larger, 3 DWPD or 4 TB written per day; HDD capacity requires cache devices.
Two cache drives per server is not a performance recommendation. On a cache-drive failure the capacity drives bound to it go unhealthy until rebinding and repair finish, and what protects the data in the meantime is the cluster’s resiliency across servers rather than anything on that node. Buy the devices with power-loss protection from the enterprise class and check the endurance rating against that 3 DWPD line before the quote is signed; the enterprise drive guide covers how those ratings are derived.
The rule, stated plainly
A five-person outage usually costs more than the hardware it saved. A write cache without power-loss protection and without a mirror is not a saving. It is a deferred incident with a date on it. The mirror covers the device failure, the capacitor covers the power failure, and a cache that has neither is holding the only copy of acknowledged writes in something that forgets.
Choosing across the whole range
Start with the uncomfortable part. From one desktop to a rack of storage nodes, the order of the three moves that actually pay does not change: more DRAM first, then metadata on flash, then everything else. The elaborate layer - the block cache, the tiering engine, the licensed accelerator - is the third move, and most readers never need to reach it.
The second uncomfortable part is commercial. The flash-caching features
vendors sold for a decade are largely gone, while the DRAM write-back cache on
the same controller is still shipping and needs no licence key. “CacheCade” appears
nowhere in Broadcom’s 94xx user guide (pub-005851, version 1.6, 28 May 2021)
or in the 9600 series product brief (9600-Series-PB111), although StorCLI’s
reference guide still lists storcli /cx add vd cc because one binary serves
every generation. A command in the syntax is not a feature on the card.
Ceph’s cache tiering page carries a deprecation banner from the Reef release
and tells you not to deploy new cache tiers. Gluster’s tier xlator was listed
as deprecated in the 6.0 release notes and moved out of tree. Intel cancelled
Optane PMem 300 on 31 January 2023 and stated it would not take last-time buys
or provide warranty coverage. Meanwhile vm.dirty_ratio, ZFS ARC,
osd_memory_target and the CacheVault-backed DRAM on a MegaRAID are all
exactly where they were.
The ladder
Rung 0 - one machine, one or two disks. The page cache is already the
cache. Raise read_ahead_kb from its 128 KiB default on sequential HDD work
and stop. Move up when the hot data no longer fits in RAM you could simply
buy: more DIMMs is the cheapest rung on
this ladder and the one with no data to lose when it fails.
Rung 1 - workstation or small NAS. ZFS ARC, capped by default at the larger of
system memory minus 1 GiB and 5/8 of system memory since OpenZFS 2.3.0, and at
half of system memory on Linux before that (zfs_arc_max default 0), with
primarycache and
secondarycache set per dataset so a nightly scan cannot evict the working
set. Move up when metadata walks, not data reads, are what hurt - then a
special vdev, mirrored, before any L2ARC. L2ARC headers live in ARC; the
classic wrong purchase is an L2ARC on a box that should have bought DIMMs.
Rung 2 - one SMB server on hardware RAID. The controller’s DDR cache in write-back with working energy backup. A 9460-8i carries 2 GB of DDR4-2133; a 9460-16i carries 4 GB; the 9440-8i carries none, and Broadcom’s user guide says it and the HBAs “do not support CacheVault data protection”. Check which card you are being sold before believing anything about write latency. Move up when the read working set exceeds host RAM and the array is still seek-bound.
Rung 3 - host-side block cache on the same server. LVM cache, bcache, dm-writecache, or Open CAS where its IO-class policy engine earns its keep. Open CAS documents a multi-level SSD-then-RAMdisk configuration in which the RAMdisk is the upper cache level and the SSD below it stays fully inclusive, so the DRAM copy is always also on the SSD. That tier is volatile across reboots. Move up when one server is no longer the unit of failure.
Rung 4 - the application, not the block layer. MySQL ships
innodb_buffer_pool_size at 134217728 bytes and PostgreSQL ships
shared_buffers at 128 MB. A database on stock buffers does not have a disk
problem. This rung is out of order on purpose: check it before rungs 2 and 3.
Rung 5 - a two-to-sixteen-node hyperconverged cluster. Storage Spaces Direct claims the fastest media automatically and runs read plus write caching for HDD capacity drives, derandomising writes before destaging; the CSV in-memory read cache sits above it, defaults to 1 GiB on Windows Server 2019 and Azure Local, and may take up to 80 per cent of physical memory. Microsoft states the boundary plainly: “Writes cannot be cached in memory.” Move up when you need more capacity than a cluster of servers should hold.
Rung 6 - the rack. Ceph, where osd_memory_target defaults to 4 GiB and
the hardware guidance says an effective target of at least 6 GiB helps
mitigate slow requests on HDD OSDs, or a dual-controller array whose DRAM
system cache does the sequential and large-block work while a flash cache or
tiering engine handles small random I/O.
The DRAM bill for the rungs above is arithmetic, not opinion:
Open CAS metadata, 1 TB cache device, default 4 KiB cache line:
1 GiB + (2% * 4KiB/4KiB + 0.05%) * 1 TB = 1 GiB + ~20.5 GB = ~21.5 GB
At a 64 KiB cache line the 2% term falls by 16x:
1 GiB + (0.125% + 0.05%) * 1 TB = ~2.8 GB
S2D, per server: 4 GB of RAM per TB of cache drive capacity, for metadata
2 x 1.6 TB NVMe cache = 3.2 TB -> 12.8 GB before any workload runs
The comparison
| Layer | Caches reads | Caches writes | Survives losing the fast device | Fast device needs redundancy | Operational cost | Who it is for |
|---|---|---|---|---|---|---|
| Linux page cache | yes | volatile only | n/a | n/a | none | everyone |
| ZFS ARC | yes | no | n/a | n/a | none | rung 1 and up |
| Controller DRAM write-back | yes | yes | only with energy backup | n/a | firmware, plus a consumable capacitor | one SMB server |
| Drive’s own DRAM | yes | yes, unprotected | no | n/a | one policy flag | nobody, deliberately |
| LVM cache (dm-cache) | yes | writethrough by default | yes in writethrough | in writeback | low | Linux single host |
| bcache | yes | optional | yes in writethrough | in writeback | low | Linux single host |
| dm-writecache | no, by design | yes | no | yes | medium | fsync-bound hosts |
Open CAS wt / wa |
yes | no | yes | no | out-of-tree module | tuned Linux hosts |
Open CAS wb |
yes | yes | no | yes | out-of-tree module | tuned Linux hosts |
Open CAS wo |
no, writes only | yes | no | yes | out-of-tree module | tuned Linux hosts |
| ZFS special vdev | metadata | metadata | no, it is a tier | yes, mirrored | medium | ZFS pools |
| ZFS L2ARC | yes | no | yes | no | low, plus ARC headers | large read sets |
| CSV in-memory read cache | yes | never | n/a | n/a | one cluster property | Hyper-V and VDI |
| S2D storage pool cache | yes on HDD | yes | undestaged writes lost on that server | two cache drives per server | cluster-level | 2 to 16 nodes |
| vSAN OSA cache tier | hybrid only | yes | no | per-host design | cluster-level | vSphere estates |
| Ceph BlueStore DB/WAL | metadata | metadata | no, the bound OSDs go with it | a failure-domain choice | high | rack scale |
| Array tiering (FAST VP, Easy Tier, HDT) | yes | yes | no, it is a tier | array RAID | licensed | arrays |
The three questions
What is the working set? Measure it. ONTAP ships the Automated Workload
Analyzer (storage automated-working-set-analyzer start -node nodename, at the
advanced privilege level), which analyses up to
one rolling week and estimates from the highest loads seen, not the average.
S2D exposes the PerfMon counter set “Cluster Storage Hybrid Disk” with a “Cache
Miss Reads/sec” counter per capacity drive. IBM’s WP102295 (May 2013) charts
five real workloads in which 5 per cent of allocated capacity took between 50
and 87 per cent of small-I/O accesses. Skew that steep is what makes any of
this work, and its absence is what makes all of it fail. Microsoft’s own
warning about DISKSPD and VM Fleet producing worse results with the CSV cache
enabled is the cleanest published statement of the trap: uniformly random
synthetic reads have no reuse, so the benchmark measures the benchmark.
Can you lose the fast device without losing data? A one-bit question with a
large bill attached. Open CAS says of write-back that a cache device failure
“may lead to the loss of data that has not yet been flushed to the core
device”. Microchip’s maxCache 4.0 permits write-through on a non-redundant
RAID 0 cache pool but requires RAID 1 or RAID 5 for write-back. S2D wants at
least two cache drives per server, rated at 3 DWPD or better, which is a narrow
slice of enterprise flash. A ZFS special vdev and a
Ceph block.db are not caches at all: they hold the only copy, and an OSD does
not start without its DB device.
Who is on call when it fails? That answer decides more than the benchmarks do. Always-write-back on a MegaRAID means write-back with no working energy backup, and it is the setting an integrator reaches for the moment a battery warning appears. Preserved cache from an offline virtual disk blocks the creation of new virtual disks until it is imported or discarded, and Dell warns that discarding before importing a foreign configuration can lose data. Open CAS is out of tree: release 26.03.4, published 24 August 2026, supports kernels up to 7.0, and a kernel upgrade can leave the module behind.
If nobody owns those three answers, buy DIMMs, put metadata on mirrored flash, and leave the rest alone. For a second-hand hybrid array the same test applies before the hardware question does.
When caching is the wrong answer
A cache is a bet that your reuse distance is shorter than your cache is big. Lose that bet and you have not bought performance. You have bought a wear-out mechanism, an extra failure domain, and a second metadata format to repair at three in the morning. Five cases lose that bet reliably, and some of them have a better answer sitting one purchase order away.
The working set is larger than any cache you would buy
Run the arithmetic from the opening model with a uniform access distribution, which is what a workload with no hot region behaves like:
h ~= C / F uniform random reads, cache C over footprint F
F = 6 TB, C = 500 GB -> h = 0.083
L_eff = 0.083 * 0.1 ms + 0.917 * 12 ms = 11.0 ms
Speedup = 12 / 11.0 = 1.09x
Nine per cent, and that is the optimistic reading, because every miss that a
read-promoting cache chooses to promote costs a chunk-sized write to the NVMe. At LVM’s
64 KiB default chunk a 4 KiB miss writes 64 KiB of flash. Size the cachevol
large enough and LVM raises the chunk size to hold the chunk count near the
million that the cache_pool_max_chunks note in lvm.conf calls the recommended
maximum for the cache target, and the write amplification rises with it. You
are paying endurance for a rounding error of latency.
Footprint is only half of it. Microsoft states the other half plainly for Storage Spaces Direct: “If the active working set exceeds the size of the cache, or if the active working set drifts too quickly, read cache misses increase and writes need to be destaged more aggressively, hurting overall performance.” A stable 400 GB footprint caches. A 400 GB footprint that turns over completely every six hours does not.
The workload is sequential, and the spindles were never the problem
A 20 TB IronWolf Pro is specified at 285 MB/s maximum sustained transfer. If
you are streaming media, restoring a backup or feeding a transcoder, the disk
is not the constraint and a cache in front of it is a machine for evicting
things that mattered. Most stacks know this and defend against it: bcache
bypasses above a sequential_cutoff of 4 MB by default, OpenZFS ships
l2arc_noprefetch at 1 so that blocks the prefetcher pulled in and nobody read
stay out of the L2ARC, and Microsoft’s Windows Server 2012 R2 tiering guidance
suggests pinning the VHDs that hold video and audio to the HDD tier, on the
grounds that such media are read sequentially and a hard disk handles that
fine. When the
vendor’s tuning advice is a mechanism for keeping your workload out of the
cache, believe it.
The money belongs in RAM first
OpenZFS says it without hedging: “RAM is by far the most effective ZFS ‘tuning knob’. Before adding any cache device, check whether the ARC is simply too small.” The kernel’s own write cache says the same from the other direction - the dm-writecache documentation states that “It doesn’t cache reads because reads are supposed to be cached in page cache in normal RAM.”
Check cachestat or cachetop before you check a price list. A read that the
page cache already served is not a problem you can buy flash for, and an L2ARC
makes RAM scarcer rather than less so: every block held on the flash keeps a
header in the ARC, tens of bytes apiece, a cost that scales with block count
rather than with capacity, and l2arc_meta_percent caps those headers at
33 per cent of the ARC target. On OpenZFS 2.3 and later the Linux default ARC
maximum is the larger of all_system_memory - 1 GiB and
5/8 * all_system_memory, so on a 64 GiB box that is 63 GiB of cache you
already own.
Buy flash for the data that needs it, and leave the rest spinning
The case that beats every caching layer in this guide is the boring one. If you know which 800 GB is hot - the database, the VM images, the build tree - put it on an SSD and be done. A cache achieves a hit ratio you must measure and defend; a filesystem on flash achieves 1.0 by construction, with no promotion writes, no dirty-block drain after an unclean shutdown, and no writeback device holding the only copy of an acknowledged write. The comparison to make before buying a cache device is the price per terabyte across every drive we list against the price of an SSD big enough to hold the working set outright, and if the hot set is small the second column usually wins. Where each still wins covers the split in detail. The spinning disks keep the bulk, which is what they are good at.
The filesystem, or the access pattern, was at fault
Some problems that look like slow media are not media problems, and a cache
will bury the evidence rather than fix it. A 512-byte-sector XFS that will not
mount over a 4096-byte cache pool is a geometry mismatch, and lvmcache(7) warns
about it explicitly. A ZFS media pool left at the default 128 KiB recordsize
carries eight times the block count, and therefore eight times the L2ARC header
cost and the special-vdev metadata, of the same bytes at 1 MiB. A workload that
calls fsync() per file is bounded by flush latency, not seek latency, and on
a consumer NVMe without power-loss protection that flush can be slower than the
4.16 ms rotational latency of the disk you were trying to escape. And if the
drive underneath is shingled, the problem is the rewrite penalty described in
shingled recording, which a write cache postpones and never
removes.
A read cache, a write cache and a tier
The three things this guide opened by separating stay separate all the way
down, and the separation is the decision. A read cache - L2ARC, bcache in
writearound, dm-cache in writethrough - costs you flash wear and buys latency,
and losing it costs nothing but warmth. A write cache - dm-writecache, bcache
in writeback, the Storage Spaces write-back cache - holds acknowledged data
that exists nowhere else, which makes it part of the data path and demands
power-loss protection and a mirror. A tier is not a cache at all: ZFS’s special
vdev, Storage Spaces heat tiering and bcachefs background_target place data
rather than copy it, so the fast device is storage, with the redundancy
requirements of storage.
Decide which of the three your problem needs before you decide which stack implements it. If the honest answer is none of them, the cheapest correct configuration is the one you already have, with more RAM in it.