← index
page 01 / 06

Architecture

In local class the network is on the read path only, and only for a chunk this host does not hold.
In fleet class it is also on the FLUSH (fsync) path, once per FLUSH, to only one fixed peer.

Agreed baseline and scope

The C0–C5 implementation is one host daemon serving multiple private images: local durability, fixed 4 KiB chunks, a separate copying store, COW manifests, a bounded clean read cache and quiescent GC. Stock QEMU and normal guest caching remain in place. On this single host, O = D.

The diagrams also show the later distributed design: peer GET/PUT, placement, surplus repair, migration and fleet-class journaling. Those paths are outside C0–C5. The 16 KiB/CDC compactor arms, prefetch, DAX/shared guest memory, guest-buffer zero-copy and in-place payload placement require their own experiments; the baseline starts with prefetch off.

This page summarizes the architecture agreed in PR #26. The C0 design defines the full local contract and failure traces; storage-format.md fixes the exact bytes. Linked implementation notes refine that contract. The 16 MiB manifest-page cache below is one such later implementation choice.

Implementation status · 14 September 2026

CheckpointValidatedRemaining
C0–C3Agreed design; build-bound validator; one-copy packed append; four queues, FLUSH ordering and retained live recovery.Accepted development checkpoints.
C4Store, manifests, compaction, snapshots/clones and GC. All 40 functional scenarios passed and independently verified, including six compaction crash cuts.Final allocation audit before aggregate closure.
C5Two private ext4/SQLite guests, retained restart and fresh boots; shared caches, scheduling, all 11 pressure stages and the consolidated 41-scenario inventory passed independent verification.Final host/guest memory accounting and dedicated-media repetitions.

The initial implementation merged in PR #27. Subsequent atomic increments validate scheduling, telemetry and competing workloads, including staging backpressure and a burst above measured drain. The latest-main integration passed all 41 C5 scenarios and independent verification on b8399ad, plus Rust and Nix/guest CI. The final Markdown receipt retains earlier failures and their corrections. Final memory accounting remains open. Since that run, PRs #48, #49 and #53 changed capacity waiting, compaction preparation and read admission; the C5 suite has not been rerun on those changes. TODO.md tracks completion; the implementation status lists the remaining work and evidence.

Spark/KVM checks establish development correctness under recorded failure models. Dedicated-media repetitions, physical power-loss experiments and research gates G1–G6 remain pending. This dated snapshot does not establish CI, merge or live deployment.

Components on one host

The guest runs a normal Linux filesystem on a virtio-blk device under stock QEMU.
Guest caching remains enabled. Applications can choose guest O_DIRECT independently of the host IO mode.

QEMU configures the device and passes its queues and guest-memory mappings to one daemon per host over vhost-user-blk. Mapping guest memory adds no payload copy. The daemon submits storage IO through Linux io_uring.
The device advertises a 4 KiB logical block and validates whole-block requests. Virtio addresses remain in 512-byte sectors. Host buffer alignment is checked separately.

QEMU reconnects after a daemon restart. Recovery preserves completed writes and replays requests still in flight without overwriting newer data. Concurrent retained recovery has passed the development checkpoint; the implementation plan distinguishes it from the preserved serial reference.

Watermark

sequence numbers of one image, increasing O D E owner-durableevery owner has acknowledged compactedin a store, manifest committed durable prefixFLUSH waits here, a snapshot cuts here TRIMMED FROM THE STAGING LOG STAGING TAIL, REPLAYED MAY DEPEND SOLELY ON THIS HOST IN LOCAL CLASS (O, E] (D, E]
O tracks owner durability, D compaction and E the durable prefix. Data in (O, E] may depend solely on this host in local class.

Three durability watermarks describe each image.
E is the highest sequence number with no unconfirmed append before it. In local class confirmed means on local NVMe. In fleet class it means on the journal peer too.
D is the highest sequence number whose chunks are durable in a store, at their owners or as surplus copies on this host, and whose manifest entries are committed.
O ≤ D is the highest sequence number whose chunks are durable at every owner. O equals D except while a surplus copy stands in for an unreachable owner.
FLUSH waits until E covers its captured boundary. A snapshot cuts at E. The staging log is trimmed below D; discarded regions can be reclaimed by the drive. Durable recovery replays (D, E]. A daemon-only restart also preserves completed writes beyond E while the host and device remain live. Owner confirmation alone does not protect a chunk whose only owner is the lost host.
E never skips a hole, because a maximum over confirmations forgets the append still in flight, and that is the answer that loses acknowledged data.

The committed manifest and staging must agree after a crash. A valid durable root supplies data through D; staging supplies newer mappings. Reclaimed payload below D is not required for replay. A selected complete root with missing or corrupt durable chunks is an error.
Re-running compaction over the replayed extents yields a manifest whose every offset maps to the same bytes. It need not yield the same chunk boundaries under CDC, and the sweep reclaims the orphans of the first run.
kill -9 at any point, then this replay, must pass fio --verify before any number from the daemon is reported. The log's torn tail is tested in both shapes, a shortened file and a partial record followed by preallocated zeros.
Three more cases have tests because each is a defect the author met in a prior implementation:

Retained recovery and ownership

The local runtime also records P, the ordered publication frontier, in QEMU's retained INFLIGHT_SHMFD trailer before guest completion. A completed WRITE can be beyond the last durable FLUSH. On daemon replacement, the same live guest and retained FD preserve request identities, PREPARED/ACTIVE transitions and the prefix through P. Cold recovery instead relies on synchronized durable state; retaining P is not a power-loss guarantee.

Replay reuses original identities and mutation order. File locks prevent repair while old kernel IO owns a segment. Reads pin their captured staging/manifest view; reset, cancellation and timeout retain buffers and mappings until kernel ownership ends. A terminal storage failure publishes FAILED before IOERR and prevents later success from crossing that failure. The C0 design specifies each transition and interruption case.

Write path

guest filesystempage cache and dirty pagesvirtio-blk queues stock QEMUqueue and memory setup DAEMON ON HOST A admissionper-image and host byte/request limitsspace and IO reserved for reclamation one copy final append buffersaligned; daemon-ownedencoding and IO use the same allocation vhost-user setup O_DIRECT staging log on local NVMeone log per imagefdatasync covers the FLUSH boundary buffers released after IO completion fleet class journal peerappend, fdatasync, acknowledgein parallel with local fdatasync WRITE: ordered visibilityFLUSH: covered writes durablechunking follows in background host file-data cache bypassedfor staging and chunk-store IO
Target write path: one host payload copy into a bounded append buffer. WRITE completes after staging IO and ordered visibility; FLUSH waits for durability.

Each WRITE copies guest payload once into its final append buffer. Record encoding and storage submission use the same allocation. The buffer remains owned until its IO completes.
The daemon bounds the append-buffer pool by bytes and request count, with admission limits per image. The copy count describes this host write path; guest application copies, peer transfers and later storage writes are counted separately.

An aligned guest buffer could instead feed vectored direct IO without a daemon payload copy. That path would need stable guest bytes and valid mappings until IO completion, a compatible record layout, and recovery across resets. The baseline uses one owned buffer for a stable, aligned record. Linux io_uring(7)

Guest writes append at block granularity to a staging log on local NVMe, one log per image.
Sequence numbers and log positions are assigned together. Mutation publication follows sequence order even when storage IO completes out of order; later writes wait for earlier mutations before index visibility and guest completion. The index holds block-to-log offsets and ordering metadata, with no retained payload.
A WRITE completes after its staging IO and ordered index update. A later read observes that write or a newer write to the same block.

The device requires FLUSH support. FLUSH captures a per-image sequence boundary covering writes completed on every queue, waits for the covered appends, and calls fdatasync before acknowledgment. Local class waits only for the local log. Fleet class also waits for the journal peer, as defined below.
Hashing and chunking run after log durability, outside the FLUSH path.

Staging and chunk-store payload IO use O_DIRECT to bypass the host file-data page cache. Direct IO still requires synchronization for persistence. Buffer addresses, file offsets and lengths satisfy the backing filesystem's direct-IO alignment requirements. Linux open(2)
A waiting FLUSH starts fdatasync as soon as its covered appends finish, without a batching delay. Later writes wait to submit until the active sync finishes; a FLUSH with a higher boundary follows those writes in the next cohort. An idle sync every 50 ms advances durability for writes without a guest FLUSH.

The governor limits allocated staging space per image and across the host. It reserves output space and minimum IO service for compaction and manifest commits before admitting more writes.
When append exceeds safe reclamation, admission slows before consuming that reserve. Progress deadlines and the recovery or failure action are recorded before each run. Storage errors return IOERR; lack of space does not authorize admission beyond the reserve.
The experiment records the point where pressure engages, per-guest latency and the time to drain after writes stop.

Admission, scheduling and resource limits

The agreed scheduler uses per-image FIFO admission and byte deficit round robin across images with a 1 MiB quantum. At the host, one of every four ready bulk submission opportunities is reserved for compaction; three favor demand IO, and unused opportunities can be borrowed. FLUSH/control bypass bulk work and have their own reserve. This is a service policy; guest fairness and latency still require measurement.

BoundDevelopment default
Requests / queues1 MiB payload; 4 queues × 256 entries; 128 outstanding requests per image, 1,024 per host.
Read requestsSeparate pool: 8 per image, 64 per host.
Descriptor snapshot quota6,297,600 bytes per frontend, within foreground metadata; actual spans allocate on discovery.
Control reserve8 requests / 64 KiB per image; 32 requests / 256 KiB per host.
Append / read buffersEach pool: 8 MiB per image, 64 MiB per host, including retained allocations.
Clean chunk cache256 MiB per host; read-fill LRU, no automatic write admission.
Metadata256 MiB per host: 128 MiB foreground, 128 MiB reserved for compaction. The 16 MiB manifest-page limit is inside foreground metadata.
Segments / staging64 MiB segments; allocated staging capped at 256 MiB per image, 1 GiB per host.
CompactionOne worker; at most 1 MiB input payload and 128 MiB new manifest pages per transaction; 100 ms settle, 1 s maximum durable-version age.
DeadlinesCapacity waits do not expire; 30 s submitted IO; 60 s recovery.

Reserve R = 3S + M + 16 MiB for background progress, where S is segment size and M the manifest transaction cap: 336 MiB at these defaults. Account allocated bytes plus accepted promises. Start pressure compaction at 75% of a staging cap; stop foreground admission at the cap or before invading the disk reserve; resume below 60% when all budgets permit. Unique live data can exhaust usable capacity and return ENOSPC/IOERR. These are development settings, not measured optima; record every override.

Compactor

trigger: settle window, maximum age or staging pressure staging logdurable, immutable version chunk4/16 KiB or FastCDC BLAKE3one hash per chunk owners = rendezvous(hash)first k hosts; HAS asks what they lack this host owns itappend to the local storefdatasync another host owns itPUT a sealed segmentowner appends, fdatasyncs, acks owner unreachablepinned surplus in local storea repair queue retries the PUT extent compactedchunks durable at owners or as local surplusmanifest committed before staging is reclaimed new unique payload is written again in this baselinethe governor reserves output space and IO for reclamation
Background compaction in the copying-store baseline. New unique payload is written to a chunk store before staging is reclaimed; this is separate from the ingress copy count.

A background pass reads settled extents from the staging log, cuts them into chunks, hashes each with BLAKE3, and skips any hash that every current owner already holds and has fenced.
A copy in a cache does not count as held.
Chunking is fixed 4 KiB, fixed 16 KiB, or FastCDC with boundaries snapped to 4 KiB (one per measurement arm on page 02).
Normally an extent is compacted after a settle window without writes. A maximum age or staging pressure forces a pass over an immutable version even if the guest keeps overwriting that extent.
The window is a parameter, and its effect on chunk traffic is measured.
A discarded or zero-filled range names no chunk and consumes no store payload. Range updates still require index and manifest work proportional to the affected mappings. A read of such a range returns zeros.

Rendezvous order of the hash names the k owners (Placement, below).
If this host is an owner, the chunk is appended to the local store and made durable with fdatasync.
Otherwise it goes to each owner in a sealed segment of many chunks, which the owner appends, fdatasyncs once, and acknowledges.
An extent counts as compacted at D, when its chunks are durable in a store, at their owners or as surplus copies here, and its manifest entry is committed. It counts as owner-durable at O, after every owner's acknowledgment.
If an owner is unreachable, the compactor appends the chunk to the local store as a surplus copy, pinned until that owner acknowledges it later. A repair queue retries the send.
In local class, durable surplus copies and the manifest commit allow log reclamation during an owner outage while local capacity remains. Surplus copies consume the same disk budget. The sweep reclaims them after owner acknowledgment.
A chunk the compactor has produced stays pinned, in the staging log or in a store, until the manifest commit that references it is durable. An owner never reclaims a chunk it acknowledged before that fence.
The staging log is therefore the write-ahead log for every chunk this host produces, wherever the chunk might end up.

The compactor releases append and FLUSH locks before chunk IO, manifest IO or owner RPC. A test slows the store to one second per append while the log remains healthy and checks FLUSH against its recorded deadline.

CDC over a dirty extent re-chunks from the last settled boundary before it to the first boundary after it that agrees with the existing cut.
Two published properties make the rule exact (LBFS locality; Xet's boundary reset), and it is why CDC never runs on the hot path. One aligned write can move every boundary in its neighborhood.

Read path

guest cache miss or guest O_DIRECT in the staging log? no in the chunk cache? no in the local store? no GET(hash)first reachable ownercache first, then store yes yes yes one NVMe readstaging index: offsets only clean cache hithash-keyed; byte limit one NVMe readindex lookup first one round tripreply hashed before it is served one host cache; read fills populate itsame-hash misses share a fetch; fetch and prefetch buffers are bounded prefetch: on sequential reads the daemon asks for the next configured number of hashes in one GET; the guest's own readahead adds to the prefetch depth
Reads after a guest-cache miss or guest direct IO. The host cache is clean and bounded; outstanding fetches have separate byte and request limits.

A guest-cache miss or guest direct read reaches the daemon. The staging index resolves fresh data first. For settled data, the manifest supplies a hash for the chunk cache, then the local store if this host holds the chunk. Otherwise the daemon sends GET to the first reachable owner in rendezvous order.
The owner answers from its cache if the chunk is hot and from its store otherwise.
Every chunk that arrives over the network is hashed before it is used, so a wrong or corrupt reply is detected and never served. A record read from a store is checked against its inline checksum.
Fresh data is served without indirection. Settled data incurs the manifest lookup, the index lookup, and, if the chunk is remote, one round trip.
GET uses separate connections and has priority over bulk PUT. Disk scheduling preserves the compactor service reserved by the governor.

One daemon-owned LRU chunk cache serves the host, keyed by hash and bounded by bytes. Read fills populate it; writes do not automatically enter it. Eviction drops a cached copy without changing durable storage.
Evicted bytes remain charged until the final reader releases them. Coalesced reads retain the original fetch credit through the final shared reader; leaders and waiters are bounded.
A fetched chunk this host does not own lives in that memory cache only. Concurrent misses for one hash share a fetch. Fetches and prefetch have bounded buffers and request counts, with demand reads served first.
A disk tier for fetched chunks, as Liquid had, is a knob measured only if time remains, since page 04 predicts a refetch from a peer's memory costs less than a local disk hit.

Prefetch is the daemon issuing the next configured number of hashes from the manifest in one GET when it sees sequential reads, and optionally replaying a recorded boot profile.
The guest's own readahead is left at its default and adds to that depth.
Page 04 sweeps the prefetch depth; its parameter is separate from the publication frontier P above.

Shared guest read memory proposed

An optional virtio-pmem frontend would expose one shared, immutable filesystem image through direct access (DAX). File-backed virtio-pmem uses the host page cache and bypasses the guest file-data cache. An OverlayFS upper would retain normal guest caching and the same private virtio-blk write contract. Copied-up files would read from that upper. The mounted lower must remain unchanged.

The first experiment uses a prepared local image. Mapping CAS chunks into guest address ranges, publishing new content, and fetching remote misses remain separate work. Sharing an existing image establishes frontend sharing; sharing independently produced equal content would establish the CAS-specific benefit. Page 02 defines the comparison.

Store, index, and manifest

The baseline local store is an append-only log of records (length, hash, checksum, bytes), separate from staging, and is authoritative for the chunks this host owns. Page 02 compares this layout with retaining and indexing surviving staging payloads in place; the guest contract is the same in both.
The index maps hash to store offset, lives in memory, and is rebuilt by scanning the store without re-hashing, because the hash is inline.
Its bytes per TB is the constant the chunk-size arms measure.
In partitioned mode a host indexes the chunks it owns plus any surplus copies awaiting an owner, so per-host index memory is k/N of the fleet's once the repair queue is empty.
An index entry is added only after the data it points to is durable, at every fence.
The manifest, one per image, is a copy-on-write tree from disk offset to chunk hash, packed in offset order. A root commit becomes durable after its referenced chunks and before staging reclamation.
It lives with the guest's host and moves when the guest does.

The local manifest uses checksummed 4 KiB COW B+tree pages, height at most eight, and fixed COMMIT pages identifying generation, root and D. Chunk data is synced before the manifest commit, and D publishes before covered staging can be reclaimed behind all read/replay pins. A shared manifest-page cache keys verified pages by incarnation, COMMIT end and offset. Cached bytes retain no root/file pins; every hit passes the checked lookup descent.

A local snapshot pauses mutation admission, drains and FLUSHes, and compacts through the captured cut. XFS FICLONE creates an immutable standalone manifest; file and directory sync precede publication. A writable clone gets a private identity/log and a new COMMIT with D = E = O = 0 in its own sequence namespace. Catalog changes use temp-file, fsync, rename and directory fsync. Snapshot membership changes are excluded from host-wide quiescent GC; measure the pause.

Protocol after C5

MessageReplyUsed by
GET(hashes)bytes per hashcold read, prefetch
PUT(segment)ack after one fdatasynccompactor sending a sealed segment of chunks to an owner
HAS(hashes)bitmap of hashes the owner lacks or has not fencedcompactor before PUT, so only missing chunks are sent; provisioning verification
LIVE(epoch, hashes)ackgarbage collection
JOURNAL(image, range)ack after fdatasyncfleet class: the appends since the last FLUSH, sent to the journal peer

Messages are length-prefixed over kernel TCP with TCP_NODELAY, driven by io_uring.
GET and JOURNAL have their own connections and priority. PUT is bulk.
Every message is idempotent and named by hash or sequence number, so any of them can be retried.
The daemon runs busy-polling or blocking. Page 04 measures both, because the scheduler wakeup is part of the cost.

Placement and the parameter k

Chunks are placed over N hosts by rendezvous hashing. Every host scores each (chunk, host) pair with one hash function, and the k highest-scoring hosts own the chunk.
Every host computes the same owner set without shared state, a ring, or a lookup, at N hash evaluations per chunk.
CRUSH's straw2 bucket is the same computation with per-host weights (NEED CITE).
When a host joins or leaves, only the chunks whose top-k set changes move.
The journal peer for fleet class is not chosen this way. A journal needs a fixed home with ordered replay, so each image names one peer at creation and keeps it.
If a migration lands the guest on its own journal peer, the image names a new peer in the same fenced swap. On two hosts the journal peer is always the other host.
k is the one multi-host parameter.
With N hosts, k = N places every chunk on every host (replicated) and k = 1 places each chunk on exactly one (partitioned). On the two-host testbed these are k = 2 and k = 1.
Page 03 measures both, and a deployment would run k ≥ 2 on N ≥ 3 hosts.

Durability classes

Durability is a per-image class on one pipeline. The class changes who waits at FLUSH and for how long. Chunks reach the same owners either way; fleet class adds a copy of the staging tail at the journal peer until compaction catches up.
Local class, the default: FLUSH returns after fdatasync of the staging log on this host.
Fleet class: the appends since the last FLUSH are sent to the image's journal peer, which appends it to its own log and fdatasyncs. The send proceeds in parallel with this host's fdatasync, FLUSH returns after both, and FLUSHes from several images to the same peer share one round trip and one fdatasync.
Fleet class is what Nutanix AOS and HPE SimpliVity do before they acknowledge, and page 04 measures what it costs.

FailureLocal classFleet class
daemon crashcompleted writes remain readable; recover the log and inflight requests, then re-run compactionsame
host crash, power lossFLUSH-covered writes survive; ordinary completed writes may be lostsame
host lostthe tail (O, E] can be lost; compacted chunks survive only if another host retains a copythe staging tail survives: the journal peer replays (D, E] onto a new host; chunks the lost host owned survive only if k ≥ 2, as in the row below
peer lost, k = 1chunks it owned are unreadable until it returns, and lost if its disk is; a read that needs one waits or fails with an error, never returns stale bytes; writes use surplus copies while local capacity remainsthe same for reads; a FLUSH waits for journal durability within its deadline or fails. The image does not silently change durability class

Reclamation preserves the selected durability class. In fleet class, the journal peer retains recovery data until durable chunks and recoverable mapping metadata on surviving hosts can replace it. A local surplus copy alone cannot release that remote journal data.

Two rules hold in both classes:

Garbage collection (GC)

A chunk is live if any manifest on any host references it, or if an in-flight compaction has pinned it. A copy in a cache is never a reference.
The baseline sweep pauses admission, drains IO and compaction, and marks active images and snapshots. Each host sends owners the complete live set for an epoch with LIVE before reclamation. Delete wholly dead store segments; copy live chunks from selected mixed segments into durable destinations before deleting the originals. Dead manifest pages can be hole-punched.
Refcounting is not a concept in this architecture.
ZFS frees an overwritten block when its reference count drops. This design does not, so space can leak between sweeps.
The sweep therefore runs before every capacity measurement and when disk pressure requires reclamation. Report reclaimed bytes, copied live bytes and pause duration beside the capacity number.

Provenance

ComponentSourceLicense
hypervisorstock QEMU, unmodified, vhost-user-blk front endGPL-2.0
vhost-user protocolrust-vmm vhost-user-backend, vm-memory, virtio-queue; Cloud Hypervisor's vhost_user_block read as referenceApache-2.0 / BSD-3-Clause
hashingblake3 crateCC0 / Apache-2.0
chunkingfastcdc crateMIT
host filesystemXFS on the dedicated NVMe, O_DIRECT, hole punching; ZFS never sits under the daemon
staging, watermark, governor, compactor, store, index, manifests, cache, protocol, journal peer, garbage collectionthis studynew code