Single host
The R0–R3 comparison runs the same stock QEMU, the same guest, and the same NVMe device, and varies only the storage behind the device.
Hypothesis 1 predicts a tie on capture with ZFS fast dedup. The curve beneath the tie, capture against index memory across three chunk sizes, is this page's result.
Configurations
R0. Raw file on XFS.
QEMU's raw driver on the dedicated NVMe.
The control, with no deduplication anywhere in the path.
R1. Zvol on pinned OpenZFS fast dedup.
Its own pool on the same device, created and destroyed per run, opened by QEMU as a block device. The current Nixpkgs lock selects OpenZFS 2.4.4. Freeze the release, host kernel and pool settings for each measurement cohort; a version change reruns the affected controls.
| Setting | Value | Why |
|---|---|---|
feature@fast_dedup | enabled | DDT update log |
dedup | blake3 | SHA-256 with dedup=on |
volblocksize | 16K primary arm, 4K second arm | the zvol's dedup granularity |
compression | zle | all-zero blocks become holes |
dedup_table_quota | none | uncapped DDT |
zpool ddtprune | never run during a measurement | no entries dropped |
primarycache | all | explicit ARC policy |
| DDT memory | zpool status -DD, dedupcached | resident DDT bytes |
OpenZFS 2.4.4 does not support its ARC-bypassing direct-IO path for zvols, so R1 remains ARC-backed. Filesystem direct writes instead skip deduplication; replacing the zvol with a direct-IO file would change the comparator. OpenZFS properties
The R1 profile verifies nonzero deduplication and zero handling before measurement. Record pool features, quota, sync settings, ARC residency and total host memory alongside the DDT measures. Report dedup_table_size as disk footprint and zero-run compression savings separately. Pool properties, DDT statistics
R2. Raw file on XFS over dm-vdo (optional).
Fixed 4 KiB deduplication in the kernel, with its own XFS instance on the vdo device.
Index memory from vdostats.
R3. The backend on one host.
Local store only, so k does not apply.
Three chunk-size arms, below.
R0 against R3 is the cost of the daemon with everything else held constant.
R1 is the deployed comparator and differs in kernel boundary, caching, and allocation, so it is a case study beside the controlled pair, and deltas are attributed accordingly.
Chunk-size arms
Fixed 4 KiB chunks cost one index entry per 4 KiB:
- about 250 million entries per TB
- about 10 GB of memory per TB at 40 bytes per entry, a 32-byte hash and an 8-byte offset
The alignment argument on page 00 predicts that they capture nearly every duplicate a Linux guest holds. The census measures the remainder.
FastCDC with a 16 KiB mean cuts the index by four and loses an aligned 4 KiB match whenever the rest of its chunk differs.
The one prior curve on VM images is Liquid's: 77% of bytes removed at 4 KiB, falling to 59% at 256 KiB on 183 images, with 256 KiB chosen for HDD seek cost. On NVMe the seek term is gone and the trade is index memory against capture.
Three arms: fixed 4 KiB, fixed 16 KiB, FastCDC 8 to 64 KiB with a 16 KiB mean.
CDC boundaries snap to 4 KiB, so no guest block straddles two chunks and a 4 KiB overwrite invalidates one chunk, not two.
Reported per arm: bytes stored, index bytes per TB, guest p99, write amplification, compactor CPU per GB.
The census below predicts the capture column for each arm before any run.
Workloads
- fio: 4 KiB random write and read at QD1 and QD32; 128 KiB sequential.
- Boot storm: n clones of one image booted together, n = 4, 16, 32. A clone is a copy of the manifest with its own staging log.
- Fleet replay: the synthetic fleet below written onto n guests, at two points on its timeline.
- Overwrite: a small SQLite database rewriting its pages in place for an hour, with guest discard on.
- Pressure: sustained unique writes and overwrite bursts beside a reading guest; run through admission pressure and idle drain. Repeat reads with shared and disjoint working sets and a scanning neighbor.
Metrics
- Guest p50 and p99 write and read latency against R0, compactor active and idle. Reported first.
- Bytes stored after compaction completes and the sweep has run, against the census prediction at the configuration's chunk or block size. Bytes the sweep reclaimed reported beside it as the leak.
- Index or DDT bytes per stored TB.
- Write amplification: device bytes written per guest block-device byte, from NVMe counters. Record application bytes, virtio payload, staging records and fences, chunk writes, and metadata/GC traffic separately. Guest filesystem journaling can make application bytes differ from block-device bytes.
- Sustainable ingest, staging allocations, live and dead bytes, and compaction progress. Report the point where admission slows, each guest's latency, and idle drain time.
- Chunk traffic against the settle window: chunks produced per guest byte written, on the overwrite workload.
- Compactor CPU per GB ingested, per chunk-size arm.
- Host payload bytes copied, peak append/fetch-buffer bytes, cache occupancy and total resident memory, with shared mappings counted once.
- Recovery: the page 01 tests pass before any number is reported.
Controls
Pinned vCPUs, performance governor, discarded warm-up, fresh filesystem or pool per repetition, at least five repetitions, variance beside every number.
With cache=none, R0 and R2 bypass the host file-data cache. R1 retains the ARC. Choose zfs_arc_max and the R3 clean-cache limit within the equal total memory budget below. Report actual ARC data/metadata residency, DDT residency, ZFS dirty memory and daemon buffers/indexes separately; equal cache caps do not establish equal total memory use.
All configurations are observed at the guest boundary (fio's histograms, guest-side blktrace for the boot storm) plus host device counters.
The daemon adds per-request stage timestamps drained to ndjson, cross-checked once against bpftrace with the delta reported.
zpool and vdostats figures are supplementary.
Guest filesystem workloads retain normal caching. Direct guest fio runs isolate block-device costs and are labeled separately. Host IO mode, guest IO mode and FLUSH frequency are recorded independently.
The host buffered/direct staging comparison holds format, queue depth, durability and total memory budget constant. Record packing and submission concurrency are varied separately. A one-block write followed by FLUSH exposes fence overhead.
Memory comparisons use equal total host budgets, including guest RAM, daemon buffers, indexes, caches and kernel file-data cache. Shared mappings are counted once. Slowed-compactor and owner-outage runs test the reserve and failure paths from page 01.
Guest memory sharing proposed
A separate experiment would compare the same immutable image through virtio-blk and virtio-pmem/DAX, with identical private OverlayFS uppers. First verify DAX use, cross-guest isolation, and recovery of synchronized upper-layer writes. Then measure aggregate resident memory without double-counting shared pages, guest memory use, read latency, CPU, and copy-up bytes as guest count grows. Hold contents, memory budgets, and workloads constant; separate cold host cache, warm host cache with cold guest cache, and repeated accesses.
This changes the guest storage stack, so it has its own results table alongside R0–R3. Its first question is whether mapped reads reduce memory duplication for shared files. No performance threshold is set; further CAS integration depends on the feasibility result.
Payload placement proposed
The baseline copies surviving staging data into a separate chunk store. An alternative at fixed 4 KiB would hash surviving records and publish their existing payload locations. Segment cleaning could still move live data to reclaim space. Larger chunks assembled from scattered writes may require copying.
Both layouts must recover acknowledged FLUSHes before comparison. Hold workload, durability, memory and available disk space constant. Measure unique writes, overwrites and deletion at increasing disk occupancy; report all payload writes, retained segment bytes, cleaning traffic and guest latency. This experiment leaves the guest contract unchanged.
Census
A small census supplies the numbers the rest of the study is measured against: how many unique bytes the fleet holds under each arm's chunker, and how many bytes copy-on-write would already have shared.
Phase 0.
zdb -S on a ZFS pool holding the cloned fleet.
Pool traversal starts each dataset at its origin snapshot's transaction group, so blocks a clone inherited are counted once and the simulated ratio is duplicates beyond what clones already share. This reading of dmu_traverse.c is confirmed with a two-clone test before the number is cited.
The fleet.
Ubuntu publishes dated cloud images and a dated package archive.
The controlled Ubuntu fleet uses that archive; Debian's snapshot.debian.org is an alternative for Debian guests.
An image installed as of T0 and upgraded monthly against the archive as of T1, T2, and on replays a real update history.
n such clones with scripted drift (hostnames, logs, a few packages each) form the fleet.
It is rebuilt by one command, dated, and is also the replay workload above.
The split.
Per byte range: zero or unallocated (from the guest allocation map, excluded), unique, shared with the T0 base in place, duplicate at an aligned 4 KiB or 16 KiB boundary elsewhere in the fleet, or duplicate only at a shifted offset.
The aligned columns predict R1 and the fixed arms.
The CDC arm is predicted by running FastCDC with the arm's parameters over the images, because a 16 KiB mean chunk captures fewer aligned matches than fixed 4 KiB chunks and more shifted ones, and the two effects do not add.
There are no donors, no real fleets, and no claims about time.