← index
page 02 / 06

Single host

The R0–R3 comparison runs the same stock QEMU, the same guest, and the same NVMe device, and varies only the storage behind the device.
Hypothesis 1 predicts a tie on capture with ZFS fast dedup. The curve beneath the tie, capture against index memory across three chunk sizes, is this page's result.

Configurations

R0. Raw file on XFS.
QEMU's raw driver on the dedicated NVMe.
The control, with no deduplication anywhere in the path.

R1. Zvol on pinned OpenZFS fast dedup.
Its own pool on the same device, created and destroyed per run, opened by QEMU as a block device. The current Nixpkgs lock selects OpenZFS 2.4.4. Freeze the release, host kernel and pool settings for each measurement cohort; a version change reruns the affected controls.

SettingValueWhy
feature@fast_dedupenabledDDT update log
dedupblake3SHA-256 with dedup=on
volblocksize16K primary arm, 4K second armthe zvol's dedup granularity
compressionzleall-zero blocks become holes
dedup_table_quotanoneuncapped DDT
zpool ddtprunenever run during a measurementno entries dropped
primarycacheallexplicit ARC policy
DDT memoryzpool status -DD, dedupcachedresident DDT bytes

OpenZFS 2.4.4 does not support its ARC-bypassing direct-IO path for zvols, so R1 remains ARC-backed. Filesystem direct writes instead skip deduplication; replacing the zvol with a direct-IO file would change the comparator. OpenZFS properties

The R1 profile verifies nonzero deduplication and zero handling before measurement. Record pool features, quota, sync settings, ARC residency and total host memory alongside the DDT measures. Report dedup_table_size as disk footprint and zero-run compression savings separately. Pool properties, DDT statistics

R2. Raw file on XFS over dm-vdo (optional).
Fixed 4 KiB deduplication in the kernel, with its own XFS instance on the vdo device.
Index memory from vdostats.

R3. The backend on one host.
Local store only, so k does not apply.
Three chunk-size arms, below.

R0 against R3 is the cost of the daemon with everything else held constant.
R1 is the deployed comparator and differs in kernel boundary, caching, and allocation, so it is a case study beside the controlled pair, and deltas are attributed accordingly.

Chunk-size arms

Fixed 4 KiB chunks cost one index entry per 4 KiB:

The alignment argument on page 00 predicts that they capture nearly every duplicate a Linux guest holds. The census measures the remainder.
FastCDC with a 16 KiB mean cuts the index by four and loses an aligned 4 KiB match whenever the rest of its chunk differs.
The one prior curve on VM images is Liquid's: 77% of bytes removed at 4 KiB, falling to 59% at 256 KiB on 183 images, with 256 KiB chosen for HDD seek cost. On NVMe the seek term is gone and the trade is index memory against capture.

Three arms: fixed 4 KiB, fixed 16 KiB, FastCDC 8 to 64 KiB with a 16 KiB mean.
CDC boundaries snap to 4 KiB, so no guest block straddles two chunks and a 4 KiB overwrite invalidates one chunk, not two.
Reported per arm: bytes stored, index bytes per TB, guest p99, write amplification, compactor CPU per GB.
The census below predicts the capture column for each arm before any run.

Workloads

Metrics

Controls

Pinned vCPUs, performance governor, discarded warm-up, fresh filesystem or pool per repetition, at least five repetitions, variance beside every number.
With cache=none, R0 and R2 bypass the host file-data cache. R1 retains the ARC. Choose zfs_arc_max and the R3 clean-cache limit within the equal total memory budget below. Report actual ARC data/metadata residency, DDT residency, ZFS dirty memory and daemon buffers/indexes separately; equal cache caps do not establish equal total memory use.

All configurations are observed at the guest boundary (fio's histograms, guest-side blktrace for the boot storm) plus host device counters.
The daemon adds per-request stage timestamps drained to ndjson, cross-checked once against bpftrace with the delta reported.
zpool and vdostats figures are supplementary.

Guest filesystem workloads retain normal caching. Direct guest fio runs isolate block-device costs and are labeled separately. Host IO mode, guest IO mode and FLUSH frequency are recorded independently.
The host buffered/direct staging comparison holds format, queue depth, durability and total memory budget constant. Record packing and submission concurrency are varied separately. A one-block write followed by FLUSH exposes fence overhead.

Memory comparisons use equal total host budgets, including guest RAM, daemon buffers, indexes, caches and kernel file-data cache. Shared mappings are counted once. Slowed-compactor and owner-outage runs test the reserve and failure paths from page 01.

Guest memory sharing proposed

A separate experiment would compare the same immutable image through virtio-blk and virtio-pmem/DAX, with identical private OverlayFS uppers. First verify DAX use, cross-guest isolation, and recovery of synchronized upper-layer writes. Then measure aggregate resident memory without double-counting shared pages, guest memory use, read latency, CPU, and copy-up bytes as guest count grows. Hold contents, memory budgets, and workloads constant; separate cold host cache, warm host cache with cold guest cache, and repeated accesses.

This changes the guest storage stack, so it has its own results table alongside R0–R3. Its first question is whether mapped reads reduce memory duplication for shared files. No performance threshold is set; further CAS integration depends on the feasibility result.

Payload placement proposed

The baseline copies surviving staging data into a separate chunk store. An alternative at fixed 4 KiB would hash surviving records and publish their existing payload locations. Segment cleaning could still move live data to reclaim space. Larger chunks assembled from scattered writes may require copying.

Both layouts must recover acknowledged FLUSHes before comparison. Hold workload, durability, memory and available disk space constant. Measure unique writes, overwrites and deletion at increasing disk occupancy; report all payload writes, retained segment bytes, cleaning traffic and guest latency. This experiment leaves the guest contract unchanged.

Census

A small census supplies the numbers the rest of the study is measured against: how many unique bytes the fleet holds under each arm's chunker, and how many bytes copy-on-write would already have shared.

Phase 0.
zdb -S on a ZFS pool holding the cloned fleet.
Pool traversal starts each dataset at its origin snapshot's transaction group, so blocks a clone inherited are counted once and the simulated ratio is duplicates beyond what clones already share. This reading of dmu_traverse.c is confirmed with a two-clone test before the number is cited.

The fleet.
Ubuntu publishes dated cloud images and a dated package archive. The controlled Ubuntu fleet uses that archive; Debian's snapshot.debian.org is an alternative for Debian guests.
An image installed as of T0 and upgraded monthly against the archive as of T1, T2, and on replays a real update history.
n such clones with scripted drift (hostnames, logs, a few packages each) form the fleet.
It is rebuilt by one command, dated, and is also the replay workload above.

The split.
Per byte range: zero or unallocated (from the guest allocation map, excluded), unique, shared with the T0 base in place, duplicate at an aligned 4 KiB or 16 KiB boundary elsewhere in the fleet, or duplicate only at a shifted offset.
The aligned columns predict R1 and the fixed arms.
The CDC arm is predicted by running FastCDC with the arm's parameters over the images, because a 16 KiB mean chunk captures fewer aligned matches than fixed 4 KiB chunks and more shifted ones, and the two effects do not add.
There are no donors, no real fleets, and no claims about time.