← index
page 03 / 06

Multiple hosts

Page 03 runs the backend on two hosts with one parameter, k, and measures bytes moved and bytes stored against what zfs send, rsync, and two per-host ZFS pools would move and hold.

Two placement modes

k, the number of owners per chunk (page 01), takes two values on two hosts, and they are two different experiments.
Surviving a dark host at two hosts costs a full mirror of chunks (k = 2) plus fleet class for the staging tail.

REPLICATED · K = 2 host Astaging log and manifestsevery chunk host Bstaging log and manifestsevery chunk PUT each new chunk each unique chunk is transferred onceevery read is localcapacity: the full store on each host PARTITIONED · K = 1 host Astaging log and manifestschunks whose hash selects A host Bstaging log and manifestschunks whose hash selects B PUT and GET by hash each chunk is stored on one hostone half of cold reads are remotecapacity: the store split across hosts
k = 2 provides transfer savings and keeps every read local. k = 1 provides capacity savings at the cost of remote reads. Two hosts with k = 1 send one half of cold reads to the peer in expectation, the largest share this testbed can produce.

Provisioning

A new guest on host B from an image whose chunks exist anywhere costs a copy of the manifest: at least 32 bytes per chunk, about 80 MB for a 40 GB image at 16 KiB chunks. Every chunk it names already exists at its owner.
In replicated mode no other data is transferred.
In partitioned mode no other data is transferred either, because chunks are fetched on first read.

The baseline is qemu-img convert or scp of the raw file and zfs send | zfs recv of the zvol, each moving the allocated size of the image.
Liquid cloned an image by copying its metadata file, in milliseconds. Its distribution benchmark moved an 8 GB image to seven nodes on 1 GbE in 35 s against 730 s by scp, and still moved every unique block to every node. Here provisioning is bytes on the wire at 100 GbE.

Migration

To move a guest from A to B, the daemon freezes the device on A and takes E, hands the image to B by one fenced swap of its root record, ships the manifest and the staging extents in (D, E], and resumes on B.
The root record names the writer and carries a generation number, and the swap is written durably on both hosts before B resumes. A accepts no write after the swap, and B resumes only after the swap names it.
On resume the log is reconciled by evidence, the replayed E against what is durable on disk, never by who claims to own it. In a prior implementation by the author, a refusal keyed on writer identity kept healthy guests from restarting.
A 40 GB image that compacted recently moves its manifest, about 80 MB at 16 KiB chunks, plus the staging tail, which was under 9 MB for an idle guest in that implementation and is workload-bound for a busy one.

We predict that bytes are the small part of a migration.
The disk cut measured 3 to 6 ms in that implementation and the rest of the blackout was orchestration, so the blackout is reported decomposed into freeze, swap, transfer, and resume, beside the bytes. Governor pacing is disabled while the guest is paused.

The baseline is rsync of the raw file and zfs send of the zvol, which since 2.0 emits no deduplicated stream (page 00), so both move the allocated size.

Synchronization after drift

Two guests, one on each host, are cloned from the same image and each updated independently to the same package set.
Compaction on each host sends only the chunks the owner lacks, packed in sealed segments.
Bytes on the wire are read against the census's unique-byte count for the pair.
Chunks per second is reported beside bytes per second, the compactor's ship rate, because we predict per-chunk cost caps the path before the link does.
This is the apt upgrade case from page 00, measured.

Capacity

Partitioned mode stores each chunk once across the fleet.
Bytes on both stores after the fleet replay and the sweep are measured against two per-host ZFS pools holding the same guests, with index bytes on each host. Hypothesis 2 predicts at most 55% of the pools' bytes and about half the index per host.
The fraction of a guest's cold reads served by the other host is measured too.
On two hosts with k = 1 that fraction is one half in expectation. In general it is 1 − k/N, so a larger fleet at fixed k sends a larger share of its cold reads over the network, and the two-host number is a lower bound on that share.

The local-class window

In local class, between a FLUSH acknowledgment and the chunk being durable at its owner sits the compaction lag, (O, E] in the watermark's terms.
It is reported in seconds under the fleet replay, as a distribution, with the segment size as the parameter.
This window measures owner-confirmation lag. Host-loss recovery also requires surviving chunk copies and mapping metadata, as specified on page 01.

Fleet class protects the staging tail through the journal peer and pays one round trip and one remote fdatasync per FLUSH. Page 04 measures that cost.

Measurements

FlowDaemonBaselineRead against
provisionbytes transferred, both modesscp of raw file; zfs sendmanifest size
migratebytes transferred, both modes; blackout decomposedrsync; zfs sendmanifest size + staging tail; milliseconds for the cut
sync after driftbytes and chunks per second sent by compactionrsync; zfs sendcensus unique bytes
capacitybytes stored, partitioned, after the sweeptwo per-host ZFS poolsat most 55% of the pools' bytes (hypothesis 2); census prediction
index per hostindex bytes on each host, both modesDDT bytes per poolk/N of the fleet index
remote fractioncold reads served by the peerone half in expectation
local-class windowseconds from ack to owner-durableowner-confirmation lag

The locality objection

Dong et al. (page 06) rejected per-chunk hash placement for backup streams on locality grounds. This is primary storage with a local cache, so page 04 measures the fragmentation cost directly, and super-chunk placement is the knob if it is large.