Thesis
Addressing a virtual machine's data chunks by their content lets a fleet of hosts deduplicate, provision, and cache across host boundaries.
At what scale does content-addressing your data become economically viable?
This study measures those advantages and their cost on a multi-host cluster, and projects from them the fleet size at which a content-addressed store becomes advantageous.
Existing solutions such as ZFS and dm-vdo also hash blocks for deduplication, but the hash serves only as a key into a table scoped to one pool while the block remains addressed by its location on disk, so no identity the table records is visible beyond a single host.
Our system has three pre-requisites
A guest is provisioned or migrated by moving its manifest
Each unique chunk is stored k times across the fleet instead of once per host
A chunk many guests read is served from its owner's memory. Pages 03 and 04 measure each individually.
A cold read of a chunk another host holds costs one round trip on the network,
the latency for this is measured in this study over NVMe-tcp as well as RDMA
Durability before acknowledgment becomes a choice between this host's disk alone and a peer's disk as well.
The system is a content-addressed block backend under unmodified QEMU.
Its scope, called the testbed from here on:
- Two hosts with static membership
- Linux guests
- Assumed single-digit terabytes scale
Deduplication within a host
Two guests that install the same package hold equal bytes that no clone or snapshot can share, because neither copy descends from the other.
A deduplication table shares them. OpenZFS keeps one per pool, the DDT, and Linux has had dm-vdo in the mainline kernel since 6.9.
Both hash every block and share equal blocks at one fixed, aligned size: 4 KiB for dm-vdo, and the volblocksize of a ZFS zvol, 16 KiB by default.
This study calls the guest's 4 KiB unit a block and the store's unit a chunk, whether the chunk's boundaries are fixed or content-defined.
Nearly everything a Linux guest writes is 4 KiB aligned. ext4 uses 4 KiB blocks, partitions start at 1 MiB, and package managers write whole files.
Jin and Miller found on VM disk images that fixed-size chunks reach nearly the same deduplication ratio as content-defined chunking (CDC), which places chunk boundaries by the bytes themselves.
We therefore predict that on one host the backend stores within 10% of what ZFS fast dedup stores when its chunk size equals the zvol's block size.
Page 02 tests this as hypothesis 1.
None of this reaches across hosts.
The DDT is per pool, dm-vdo has no replication, and zfs send has not carried a deduplicated stream since OpenZFS 2.0.
A fleet of N hosts, each with its own table, stores a chunk shared by all of them N times. Moving a guest to another host sends every block of its image, whether or not the destination holds an equal one.
Shared-storage systems deduplicate across hosts by placing every write on the network before it is acknowledged. Ceph RBD with TiDedup is the open example.
Page 06 lists the systems on either side.
What is gained across hosts
Each consequence below is measured on pages 03 and 04.
Transfer.
Provisioning a guest from an image whose chunks exist in the fleet moves the manifest and no chunk data.
Migrating a guest moves the manifest plus the bytes written since the last compaction.
Capacity.
The fleet stores each unique chunk k times rather than once per host, and each host's index holds entries only for the chunks it owns, plus surplus copies until their owner acknowledges them.
Cache.
Every host sends its reads of a chunk to the same k owners. We predict that a chunk many guests read is therefore hot at its owner, and page 04 reports the owner's hit rate.
On the published figures page 04 stacks, a chunk in a peer's memory arrives before a chunk on the local disk. Hypothesis 3 tests this.
What is paid
Three costs come with any post-process deduplicating store, and each is measured in this study:
- write amplification, because every surviving byte is written to the staging log and again to the store
- compactor interference, because compaction shares the guest's disk
- index memory, one entry per chunk in RAM
One cost belongs to distribution alone. For a chunk this host does not hold, the network is on the read path.
Page 04 measures that read over TCP and over RDMA, from the peer's memory and from the peer's NVMe.
It then measures how much prefetch, described on page 01, removes.
One cost is a tradeoff the design makes on purpose: durability before acknowledgment.
In local class, the default, a FLUSH is acknowledged after fdatasync on this host, the contract a local disk gives, and a host lost before compaction has shipped its bytes loses them.
In fleet class the FLUSH also waits for a peer's fdatasync, as the hyperconverged products on page 06 do, so fsynced bytes survive the loss of this host.
Page 01 defines the classes and page 04 measures the round trip fleet class adds.
Hypotheses
Each hypothesis states a metric with its conditions, a comparator, a threshold, the source of the threshold, and what a miss would show.
Thresholds are frozen at the end of week 2, after R0 has measured the testbed's fdatasync and media times, and do not move after that.
1. Single-host parity.
Bytes stored by the backend after compaction and sweep, under the fleet replay at fixed 4 KiB and 16 KiB, are within 10% of the bytes ZFS fast dedup stores at the same volblocksize. Bytes stored in each chunk-size arm are within 10% of the census prediction for that arm.
Guest write and read p99 at 4 KiB QD1, with the compactor idle and again with it active, are within 20% of a raw file on XFS.
The 10% is the alignment argument above plus record headers, since the sweep runs before every capacity number. The 20% is the passthrough bound of gate G1 plus an equal allowance for the log append and compactor interference.
A miss on capture would show that fixed aligned chunks lose duplicates a Linux guest produces. A miss on p99 would show that the host daemon (not deduplication) is the cost paid.
2. Transfer and capacity across hosts.
Bytes on the wire to provision a guest are within 10% of its manifest size, and to migrate one within 10% of the manifest plus the staging tail, against the allocated image size that zfs send or rsync moves.
Bytes sent to synchronize two drifted guests are within 10% of the census's unique-byte count for the pair.
Bytes stored on both hosts in partitioned mode, after the sweep, are at most 55% of the bytes two per-host ZFS pools hold for the same guests.
The 10% on transfer covers framing and HAS replies, since no chunk an owner holds is sent by design. The 55% is one copy of the unique set instead of two, because guests cloned from one image give the two pools nearly the same unique set, plus five points for record headers and manifests.
A miss on transfer would mean chunks were sent that an owner already held, a HAS or fence defect. A miss on capacity would mean the two pools shared less than the census predicted, which the census would show first.
3. Reads over the wire.
For a 4 KiB read at QD1 whose chunk is not in the local cache, over the daemon on kernel TCP:
Served from the owner's memory, guest-visible latency is lower than the same read served by the daemon from its own NVMe.
Served from the owner's NVMe, it is at most 40% over the local read.
With reads in flight at or above the bandwidth-delay point, remote sequential throughput is within 10% of local.
In a partitioned boot storm of 16 guests with profile prefetch, guest p99 is within 25% of the same storm in replicated mode.
The kernel nvme-rdma probe, which stands in for a daemon over RDMA, serves the owner's NVMe at most 15% over the local read. The probe is not the architecture, and its number is reported as the floor the ibverbs arm could approach.
The thresholds are the literature stack on page 04: about 80 µs of media, plus 20 to 30 µs for a userspace daemon over kernel TCP and about 12 µs for kernel nvme-rdma.
A miss on the first part would show that the kernel stack or the daemon's wakeup costs more than the media. A miss on throughput would show that the fabric bounds it. A miss on the boot storm would show that prefetch does not hide the remote read under a real access pattern.
4. Durability before acknowledgment.
Write p99 at 4 KiB QD1 in fleet class is within 3x of local class over TCP, and within 2x over RDMA if the ibverbs arm lands.
The window of local class, the seconds between a FLUSH acknowledgment and the durability of those bytes at their owners, is reported as a distribution under the fleet replay.
The 3x is one round trip plus one peer fdatasync alongside the local fdatasync, on the page 04 figures and an fdatasync near 40 µs (NEED DATA; measured on the testbed drive in week 1).
A miss would show that the journal path rather than the transport is the cost. The peer's fdatasync time is reported separately so the two can be told apart.
Results
The system.
A content-addressed block backend for VMs under unmodified QEMU on a stock Linux kernel, over kernel TCP, with source, configuration, and the scripts that produce every table.
The single-host table.
The backend against ZFS fast dedup and a raw file on XFS: bytes stored, guest p99, write amplification, and index memory, at three chunk sizes.
Hypothesis 1 is decided here. The capture against index memory curve across the three arms is reported without a threshold, because no prior curve on NVMe exists to bound it.
The multi-host table.
Bytes moved to provision and to migrate a guest, bytes sent to synchronize two drifted guests, fleet bytes stored with one copy per chunk, and index bytes per host, each against what zfs send or rsync moves and what two per-host ZFS pools hold.
The remote-read measurement.
A content-addressed chunk fetched from a peer under a VM block device, at microsecond resolution, over the daemon on kernel TCP and over NVMe-oF on TCP and RDMA, from the peer's memory and from its NVMe, with and without prefetch.
The durability trade.
Local class against fleet class on the same hardware: the write latency fleet class costs per transport, and the seconds of acknowledged data local class puts at risk.
Scope
The study covers hosts that serve guests from local flash, from a small homelab setup up to rack scale (storage arrays and hyperscale economics are out of scope).
Hosts hold each other's chunks, which couples the failure domains of compute and storage that shared-storage designs keep apart.
This study measures what that costs on the read path and does not model its availability.
The testbed is two hosts with static membership. Membership changes, failure detection, rebalancing, authentication and encryption on the wire, measurement on more than two hosts, and concurrent garbage collection are out of scope, and none affects a number reported here.
Each image has only one writer.
Ownership state, the root record that names the writer and carries a generation number, is held on both hosts. Two hosts form no quorum, so failover of a lost writer is a scripted decision and not automatic.
The study migrates disks only. Memory migration is QEMU's own live migration.
The guest contract is virtio-blk with a volatile write cache. An acknowledged FLUSH is durable; ordinary write completion does not promise power-loss durability.
Equal BLAKE3 hashes are taken to mean equal bytes. A sample of matches is verified byte for byte and the sample size is reported.
The store is trusted infrastructure, so deduplication side channels are documented and excluded.
Experiments run at single-digit TB, and larger figures are projections from measured constants, labeled as such.