The write path: a log, a sequence number, and a fence.
The guest sees a normal virtio-blk disk. Behind it is the daemon, my process, which QEMU talks to over a vhost-user socket.
Write
Appended to one file, the staging log, stamped with a sequence number. Nothing is overwritten in place.
FLUSH
The guest saying: everything before this must survive a power cut.
The daemon appends a fence record and calls fdatasync. The same contract a physical disk gives.
Not on this path
Hashing. It happens later, in the compactor, which is not built yet.
Where the daemon sits: guest driver, QEMU, vhost-user-backend, my code.
The guest kernel's virtio-blk driver puts requests into two rings in guest memory.
QEMU sets the device up, hands the daemon the ring addresses and a file descriptor for guest memory over a unix socket, then leaves the data path.
vhost-user-backend, from rust-vmm, speaks that socket protocol: it maps guest memory into the daemon, runs the event loop, and calls my code when the guest kicks a queue.
My code reads the request out of guest memory, runs it against the storage worker, writes the status byte, and publishes the used entry.
O_DIRECT: writes go to the NVMe, not to host RAM.
Normally a write lands in the host kernel's page cache and returns at once; the kernel writes it to the NVMe later. If the host dies in between, the only copy was in RAM. O_DIRECT bypasses that cache: a write returns once the NVMe has accepted it. fdatasync still flushes the device's own cache, so FLUSH is unchanged.
Tradeoffs
Every buffer must be 4 KiB aligned in memory or the kernel rejects it. A Rust type with a compile-time check makes a misaligned buffer impossible to construct.
Every 4 KiB guest write is one 8 KiB device write today, synchronously. The fix is batching in the daemon, not in the kernel: the daemon knows where the FLUSH boundaries are. Measured before G1.
The host page cache was a free read cache, and it is gone. Metadata-heavy work, many small reads, renames, stats, is where a second-level cache pays.
The read cache is not decided
A chunk cache inside the daemon, keyed by hash, size a parameter.
Buffered reads on the host, so the page cache is a second level under the guest's own.
virtio-pmem with DAX for the immutable base image: the one place a shared mapping is a read cache.
Recovery: finding where the log ends.
A fence is the record FLUSH appends. It names the last sequence number that is now durable. Recovery trusts everything up to the last fence and nothing after it.
Problem
After a crash the tail is garbage: a half-written record, or zeros from preallocation. And the guest's own data can contain bytes that look exactly like a fence.
Format
Fixed 8 KiB slots, a 4 KiB metadata block and a 4 KiB data block, each with a CRC32 checksum. A fence can only be a metadata block, so guest bytes are never mistaken for one.
Recovery
Scan back to the last fence with a valid checksum, replay forward checking every record, cut everything after. A bad record before that fence is an error, never a shorter replay.
Live recovery: the daemon dies mid-request.
Requests live in two rings in memory shared between guest and daemon: the guest posts to one, the daemon posts completions to the other. A replacement daemon sees the rings, but not which requests the old one had started.
vhost-user has a mechanism for this, inflight I/O tracking: QEMU hands the backend a shared file to record per-request state in, and a replacement reads it back. vhost-user-backend 0.23 returns unsupported for both of those messages.
Workaround today: one request in flight, made durable before it is marked used, so the used counter alone is the resume point. Killed at four points inside a request; the same guest keeps running. Too slow for any latency number; lifting it is the next two weeks.
Implementing virtio-blk by hand.
Addressing
The guest is told blocks are 4 KiB, but a virtio-blk request counts in 512-byte sectors, a rule from the original spec. Sector 8 means byte 4096.
Anything that is not whole 4 KiB blocks is rejected.
Framing
A request arrives as a chain of memory descriptors; the 16-byte header can be split across them or share one with the data.
Parsing works on bytes, not descriptors. Tested at all 17 split positions.
Negotiation
A guest that will not negotiate FLUSH is refused before any IO, instead of acknowledging unsynced writes as stable.
Ordering
io_uring does not order an fsync behind earlier writes unless asked; the drain flag does that.
An empty discard once took a sequence number nothing wrote, so the next FLUSH waited forever. Now it takes none.
Chunk size: the decision the compactor gets built around.
The compactor is not built. It will take settled data out of the log, cut it into chunks, hash each with BLAKE3, and store every distinct chunk once. Chunk size is the one choice everything downstream depends on.
Against 4 KiB
Memory. Every distinct chunk needs an index entry, a 32-byte hash and an 8-byte offset: about 10 GB of RAM per TB at 4 KiB, a quarter of that at 16 KiB.
For 4 KiB
Change one 4 KiB block inside a 16 KiB chunk and the whole chunk is new. The other 12 KiB of sharing is lost every time.
Third option, untested
Content-defined chunking: cut where the bytes say to, using a rolling hash, so an insert moves one boundary instead of every chunk after it. Cuts are rounded to 4 KiB so chunks still line up with the guest's blocks.
The census: measuring before committing.
Two dated Ubuntu images share 38% of their bytes at 4 KiB, 19% at 16 KiB.
Three clones of one image, each upgraded on its own. T0 is the control: everything shared, nothing new. By T2, 0.67 GB of new duplicates that only hashing can see.
At T2 the three guests need 2.11 GB at 4 KiB and 2.84 GB at 16 KiB: 16 KiB costs 35% more for the same guests.
Proves 4 KiB gets the index. Proves nothing about speed.
4 KiB chunks
16 KiB chunks
Three guests, ancestor excluded. Bar length is what three separate stores would hold; only the unique part is stored when they share.
virtio-pmem for reads: examined, parked.
What it does
Maps one host file into every guest as memory. With DAX the guest reads the host's pages directly, no copy in its own cache: one resident copy of a base image per host.
What it costs
Fixed at boot, immutable while mapped. The host never sees the writes, so it cannot hash. Needs a flat file per image per host: a second copy of the bytes dedup collapses.
Where it still fits
As the read cache for the immutable base image, one of the three candidates on the O_DIRECT slide. It shares what is already known to be equal; the amber on the census slide is what it cannot see.
Next
Write path first: no number is reported before recovery and ordering pass.
Write path
- concurrent inflight recovery: the two vhost-user inflight handlers, or the lower-level vhost crate
- multiqueue: FLUSH covers the highest completed write on any queue
- batched IO in the daemon: one group of device writes per FLUSH
- bounded staging: a governor that paces compaction
Compactor
- settle window, then fixed 4 KiB chunks hashed with BLAKE3
- manifest: image offset to chunk hash, journaled
- trim the log below D once chunks and manifest are durable
Storage backend
- chunk store: append-only records, hash inline
- index: hash to offset, in memory, rebuilt by scanning the store
- GC sweep from the live set
Read path
- staging log, then chunk cache, then local store
- manifest and index lookup for settled data
- remote GET by hash, with the two-host work
Testbed
- finalize the Nix guest image and the XFS, ZFS and CAS host profiles
- CloudLab pair; fallback two OVHcloud servers on 25 GbE
Measurements
- G1: raw XFS against the passthrough guest, p99 within 10%
- write path: per-write direct vs buffered + fdatasync vs daemon-batched direct
- read cache: daemon chunk cache vs host page cache vs pmem for the base image
- ZFS two-clone control; content-defined chunking arm of the census