← index Update 04 · 16 September 2026

The whole system, one diagram

Every part of the project at 7a555f3, top down: guest and QEMU, the cas-host process, the cas-core library, the kernel, the bytes on disk and the tooling. Zoom and drag; a +N box opens; choose any box for what it does and where its source is; an operation button traces one request.

100%
CAS-HOSTone process · crate cas-daemon−CAS-CORElibrary · formats, indexes, accounting−descriptor chains · kick eventfdvhost-user messages + FDs · mapped guest memoryvCPUs run underticket · turn · hold through used entryprepare(kind) · pause · resumeValidated bundleenterCommand mailbox + eventfd · Completed mailbox + eventfdBuilder gatherstaging / physical checkShare · Leasepublish under gateprepare · publish · fence · sync · read planSQE / CQEready · takeView::lookup · page hit / fillget · fill · leader / waiterplan(hash)Ready turns · Event/Replyload · prepare · write · create segmentinsert chunks · mark · sweepprepare · publish · punch pagespause · resumeinsert entrybackground turn per syscallblocking pread / pwrite / syncStaging · Governorscan · recoverone per imageinspectinspect · rebuild indexinspect · recoverinspect · fresh · liveopen · lockLocal variantStaging variantRaw variantBLAKE3Prepared::buildgoverned repairallocatenew segmentshared Lrueventfd signaleventfd wakefstatvfsroot fdtemp · sync · renamebefore_background_ioopen · write · sync · punch · reflinkFICLONEmetadata chargesmetadata chargespayload budgetBudgetAllocatorsegment fileschunk batchespages + COMMIT · reflinked manifestcatalog.v2reports · telemetryholdsrun-cas-*-vm launcher · spawns guestspawns cas-daemon · cas-host init · serve · SIGKILLLog::open on schedulescasctl censusssh · fio benchstaging-checkdedicated hosts (pending)host.json · telemetry.jsonlGuest VMLinux · unmodified+2QEMUstock · vhost-user-blk-pci+3Host LinuxNixOS · KVM · XFS · io_uring+4Store on disk<store-root>/ on one XFS filesystem+6Build, run, measureNix · casctl · cas-harness · evidence+14Image frontendBackend · one per image+10Local adapterLocal · runs on the frontend thread+2Image reactorthread local_async · owns the WAL + io_uring+2Compactor ownerthread cas-compactor · one per host+5Shared host stateSharedHost · one per process+3Host servicesupervisor · telemetry · reportsStore recoveryinspect everything · then repairStorage dispatchRaw · Staging v1 · Local · OpeningWAL (append log)Log · v2 packed batches+8Chunk storeStore · Reader · shared content+7ManifestCOW B+tree · block → hash+7Clean chunk cache256 MiB LRU · read fills only+2Budgetscharge before allocate · release on drop+1IO schedulerbulk opportunities · 3 demand : 1 backgroundDisk accountingGovernor · Staging · reserve RCatalogmembership · atomic publishSegment ticketsone namespace · highest header winsIO primitivesdirect · aligned · directory · encodingv1 staging loglegacy · 8 KiB slots · serialCensusoffline · 4 KiB and 16 KiBIO metricsthread-local scope counters
Diagram markup for this view · Mermaid
flowchart TB
  guest["Guest VM — Linux · unmodified"]
  qemu["QEMU — stock · vhost-user-blk-pci"]
  kernel["Host Linux — NixOS · KVM · XFS · io_uring"]
  disk["Store on disk — <store-root>/ on one XFS filesystem"]
  subgraph host["cas-host"]
    frontend["Image frontend — Backend · one per image"]
    local["Local adapter — Local · runs on the frontend thread"]
    reactor["Image reactor — thread local_async · owns the WAL + io_uring"]
    owner["Compactor owner — thread cas-compactor · one per host"]
    shared["Shared host state — SharedHost · one per process"]
    hostsvc["Host service — supervisor · telemetry · reports"]
    host-recovery["Store recovery — inspect everything · then repair"]
    storage-enum["Storage dispatch — Raw · Staging v1 · Local · Opening"]
  end
  subgraph core["cas-core"]
    append["WAL (append log) — Log · v2 packed batches"]
    store["Chunk store — Store · Reader · shared content"]
    manifest["Manifest — COW B+tree · block → hash"]
    cache["Clean chunk cache — 256 MiB LRU · read fills only"]
    budget["Budgets — charge before allocate · release on drop"]
    scheduler["IO scheduler — bulk opportunities · 3 demand : 1 background"]
    space["Disk accounting — Governor · Staging · reserve R"]
    catalog["Catalog — membership · atomic publish"]
    tickets["Segment tickets — one namespace · highest header wins"]
    direct["IO primitives — direct · aligned · directory · encoding"]
    staging-v1["v1 staging log — legacy · 8 KiB slots · serial"]
    census["Census — offline · 4 KiB and 16 KiB"]
    io-metrics["IO metrics — thread-local scope counters"]
  end
  tooling["Build, run, measure — Nix · casctl · cas-harness · evidence"]
  guest -->|"descriptor chains · kick eventfd"| qemu
  qemu -->|"vhost-user messages + FDs · mapped guest memory"| frontend
  guest -.->|"vCPUs run under"| kernel
  frontend -->|"ticket · turn · hold through used entry"| shared
  frontend -->|"prepare(kind) · pause · resume"| local
  frontend -->|"Validated bundle"| host-recovery
  local -->|"enter"| shared
  local -->|"Command mailbox + eventfd · Completed mailbox + eventfd"| reactor
  local -->|"Builder gather"| append
  local -->|"staging / physical check"| owner
  local -->|"Share · Lease"| budget
  reactor -->|"publish under gate"| shared
  reactor -->|"prepare · publish · fence · sync · read plan"| append
  reactor -->|"SQE / CQE"| kernel
  reactor -->|"ready · take"| scheduler
  reactor -->|"View::lookup · page hit / fill"| manifest
  reactor -->|"get · fill · leader / waiter"| cache
  reactor -->|"plan(hash)"| store
  reactor -->|"Ready turns · Event/Reply"| owner
  owner -->|"load · prepare · write · create segment"| append
  owner -->|"insert chunks · mark · sweep"| store
  owner -->|"prepare · publish · punch pages"| manifest
  owner -->|"pause · resume"| shared
  owner -->|"insert entry"| catalog
  owner -->|"background turn per syscall"| scheduler
  owner -->|"blocking pread / pwrite / sync"| direct
  owner -->|"Staging · Governor"| space
  hostsvc -->|"scan · recover"| host-recovery
  hostsvc -->|"one per image"| frontend
  host-recovery -->|"inspect"| catalog
  host-recovery -->|"inspect · rebuild index"| store
  host-recovery -->|"inspect · recover"| manifest
  host-recovery -->|"inspect · fresh · live"| append
  host-recovery -->|"open · lock"| tickets
  storage-enum -->|"Local variant"| local
  storage-enum -->|"Staging variant"| staging-v1
  storage-enum -->|"Raw variant"| kernel
  append -->|"BLAKE3"| store
  append -->|"Prepared::build"| manifest
  append -->|"governed repair"| space
  append -->|"allocate"| tickets
  store -->|"new segment"| tickets
  manifest -->|"shared Lru"| cache
  cache -->|"eventfd signal"| kernel
  scheduler -->|"eventfd wake"| kernel
  space -->|"fstatvfs"| kernel
  space -->|"root fd"| tickets
  catalog -->|"temp · sync · rename"| direct
  direct -->|"before_background_io"| scheduler
  direct -->|"open · write · sync · punch · reflink"| kernel
  append --> direct
  store --> direct
  manifest -->|"FICLONE"| direct
  append -->|"metadata charges"| budget
  manifest -->|"metadata charges"| budget
  cache -->|"payload budget"| budget
  store -->|"BudgetAllocator"| budget
  append -->|"segment files"| disk
  store -->|"chunk batches"| disk
  manifest -->|"pages + COMMIT · reflinked manifest"| disk
  catalog -->|"catalog.v2"| disk
  hostsvc -->|"reports · telemetry"| disk
  kernel -.->|"holds"| disk
  tooling -->|"run-cas-*-vm launcher · spawns guest"| qemu
  tooling -->|"spawns cas-daemon · cas-host init · serve · SIGKILL"| host
  tooling -->|"Log::open on schedules"| append
  tooling -->|"casctl census"| census
  tooling -->|"ssh · fio bench"| guest
  tooling -->|"staging-check"| staging-v1
  tooling -.->|"dedicated hosts (pending)"| kernel
  hostsvc -->|"host.json · telemetry.jsonl"| tooling

Top down, in seven layers

Read the diagram from the guest down. Each layer below links to its box; the parts inside are listed in the reference.

LayerWhat it isParts
Guest VM
Linux · unmodified
An unmodified Linux VM writes files; ext4 turns them into 4 KiB block requests and FLUSHes on one virtio-blk disk.2
QEMU
stock · vhost-user-blk-pci
Stock QEMU shares guest RAM with the daemon, sets the queues up over a Unix socket, retains the inflight memfd, and leaves the data path.3
cas-host
one process · crate cas-daemon
One process serves every image: a frontend and a reactor thread per image, one compactor thread, shared caches, budgets, scheduler and gates.30
cas-core
library · formats, indexes, accounting
The library holds the formats and their recovery parsers, the indexes and the copy-on-write manifest, the caches, the budgets and the disk accounting.38
Host Linux
NixOS · KVM · XFS · io_uring
XFS with O_DIRECT, fdatasync, reflinks and hole punching; io_uring per reactor; eventfds, memfds, flock and timerfds for control.4
Store on disk
<store-root>/ on one XFS filesystem
One store root: a catalog, shared chunk segments, a manifest and a private WAL per image, reflinked snapshots, archived rejects.6
Build, run, measure
Nix · casctl · cas-harness · evidence
Nix builds and boots everything; cas-harness and casctl launch, verify and bind results to a revision; docs/ keeps the records the site imports.14

The boundaries that matter. The guest never learns about hashes; QEMU never touches a block after setup. The frontend admits and completes; the reactor executes in one write order; the compactor is the only writer of shared content. cas-core has no threads and no policy: it is formats, indexes and accounting with invariants carried by types. The kernel supplies the only durability primitive, fdatasync on O_DIRECT files, and the design’s crash model is exactly what that primitive promises.

Three prefixes explain most of the arrows. A WRITE is published at P, the last mutation whose batch completed in order. A FLUSH waits for E, the last mutation an fdatasync covered. Compaction moves D, the last mutation the committed manifest represents. During serving D ≤ E ≤ P; the staging index covers (D, P], the manifest covers everything at or below D, and reclamation frees WAL payload below D once no reader or replay identity pins it.

One request, one path

The operation buttons above the diagram expand the parts a request passes through and number them. Each ends with what the operation guarantees.

OperationStepsWhat it guarantees
WRITE8The block is readable once its batch is published in order. Nothing is hashed yet, and nothing is durable until a FLUSH.
FLUSH6Every mutation admitted before the FLUSH is durable in the WAL. A failed sync never advances E and never acknowledges the FLUSH.
READ9Recent staged writes and ZERO ranges override the committed manifest. Only uncovered blocks reach the shared cache and chunk store, and one fetch per hash serves every waiting reader.
COMPACT8Chunks are durable before the manifest names them; the manifest is published before its WAL space is reclaimed. A newer guest write stays in staging and keeps precedence.
COLLECT7Every image pauses while the host reclaims chunk segments and manifest pages. If live data still fills usable space, write admission stays closed and the pressure is reported.
RECONNECT6The same guests keep running through a daemon replacement. It needs the original QEMU processes, guest RAM and inflight memfds; without them the host recovers cold from the catalog.
EXPERIMENT6A result exists only as a typed report bound to a source revision and executable hashes, kept under results/ and summarized in docs/. Every number on the site is imported from one of those files.

Every part, in words

The same text the panel shows, laid out layer by layer so it can be read straight through, searched, or printed. Paths link to 7a555f3.

Guest VM · Linux · unmodified — 2 parts

An ordinary Linux virtual machine. It sees one virtio-blk disk with 4 KiB blocks and never learns that the bytes behind it are content-addressed. Its own page cache, filesystem journal and fsync semantics are untouched.

  • Applications and fio write files; ext4 turns them into 4 KiB block requests and jbd2 journal writes; fsync becomes a virtio FLUSH.
  • The guest keeps a private page cache. A CAS cache hit is still copied into guest RAM; no memory is shared between guests.
  • In the lab the guests run under TCG emulation inside a KVM outer VM, so their timings are development measurements, not native ones.

Filesystem + page cacheext4 · jbd2 · fsync

The guest filesystem mounted on the CAS disk at /mnt/cas. It decides block layout, journaling and when a FLUSH is issued. Nearly everything it writes is 4 KiB aligned, which is why fixed 4 KiB chunks capture most duplicates.

  • A file write lands in the guest page cache and is written back later as block IO.
  • fsync or a journal commit issues a virtio-blk FLUSH; the guest cannot proceed until CAS acknowledges it durable.
  • A jbd2 journal write on a queue can sit ahead of unrelated reads on that queue; the read bypass below exists because of this.

nix/shared/workload.sh · crates/harnesses/fio/live.fio

virtio-blk driversplit virtqueues · 512-byte sectors

The guest kernel driver. It places 16-byte request headers, data buffers and a status byte into descriptor chains in guest memory, kicks a queue, and later consumes used entries.

  • Requests count in 512-byte sectors: sector 8 means byte 4096. CAS only accepts whole aligned 4 KiB blocks up to 1 MiB.
  • With the shared host the device advertises four queues of 256 entries; the reference backends advertise one queue of 128.
  • The driver must negotiate VERSION_1, BLK_SIZE and FLUSH or the backend refuses before any IO.
queues · depth
4 × 256 (shared host); 1 × 128 (reference)
max request
1 MiB, 16 segments of 64 KiB

docs/vhost-user-notes.md

QEMU · stock · vhost-user-blk-pci — 3 parts

The unmodified virtual machine monitor. It emulates the PCI device, shares guest RAM with the daemon and configures the virtqueues over a Unix socket. After setup it leaves the data path entirely.

  • Guest RAM is backed by a memfd or file with share=on so another process can map the same pages.
  • On connect QEMU sends SET_MEM_TABLE with the RAM file descriptors, negotiates features and protocol features, then SET_VRING_* for each queue.
  • Between GET_INFLIGHT_FD and SET_INFLIGHT_FD it retains the daemon's inflight memfd, which is what makes a daemon replacement possible without a reboot.
  • QEMU 10.2.4 distinguishes stop/start (SET with the retained fd) from a full device reset (GET then SET).

vhost-user socketUnix socket + SCM_RIGHTS

The control channel between QEMU and the daemon. Every message is a small struct; file descriptors for guest memory, kick and call eventfds and the inflight carrier travel as ancillary data.

  • QEMU is the frontend (client); the daemon listens and accepts exactly one connection per image.
  • Kick eventfds wake the daemon when the guest posts work; call eventfds interrupt the guest when a used entry is published.
  • The socket must live on a different filesystem from the storage root so its inode never counts against the governed disk.

docs/vhost-user-notes.md · docs/host-service.md

Shared guest RAMdescriptors · rings · payload

The guest's physical memory, mapped a second time inside the daemon. Descriptor tables, available and used rings and every data buffer live here, so a request is parsed and completed without copying it through the socket.

  • The daemon keeps its own accepted snapshot of the mapping and uses it for payload reads, status bytes and used-ring writes even while the framework replaces its atomic map.
  • A write's payload is gathered straight from these pages into its final WAL buffer: one copy on the host.
  • A read's response is copied back into these pages before the status byte is written.

docs/shared-frontend.md · docs/review/c2-copy-path.md

Retained inflight memfdCASIFL03 · survives daemon death

A sealed memory file the daemon creates and QEMU keeps open. It records which requests were discovered and admitted, per queue head, plus the published prefix P. A replacement daemon reads it back to resume the same guest.

  • The standard vhost-user inflight regions come first; a CAS trailer of one 4 KiB header and one 128-byte slot per queue head follows.
  • Slots move EMPTY → DISCOVERED → PREPARED → ACTIVE → EMPTY; a header FAILED flag freezes every transition except error completion.
  • It holds identities and sequence numbers only, never payload: payload is either in guest RAM or in the WAL.
magic · version
CASIFL03 · 3
max mapping
155,648 bytes (4 queues × 256)

docs/inflight-format.md · docs/storage-design.md

Host Linux · NixOS · KVM · XFS · io_uring — 4 parts

The host kernel does the actual storage work: it runs the guest under KVM, executes io_uring submissions against O_DIRECT files on XFS, and provides the durability primitive (fdatasync) that every promise in the design rests on.

  • CAS never uses the host page cache for payload: every segment is opened O_DIRECT with 4 KiB alignment checked through statx.
  • Reflinks, hole punching and FIEMAP make snapshots and reclamation cheap without moving live pages.
  • The declared crash model: writes covered by a successful fdatasync survive; anything after may be lost, torn or reordered.

XFS + block layerO_DIRECT · fdatasync · FICLONE · punch

A dedicated XFS filesystem on the test disk holds the whole store. It is chosen for reflink support, direct IO alignment reporting and hole punching, all of which the design depends on.

  • open(O_DIRECT | O_NOFOLLOW | O_NONBLOCK) plus flock on every segment, manifest and catalog file; STATX_DIOALIGN must report 4 KiB alignment or the open fails rather than falling back to buffered IO.
  • fallocate(KEEP_SIZE) preallocates segments so recovery can distinguish written length from reserved space; PUNCH_HOLE frees compacted WAL payload and dead manifest pages in place.
  • ioctl(FICLONE) reflinks a manifest to make a snapshot; FIEMAP verifies nothing remains beyond EOF after trimming.
  • fstatvfs on the root is the only source of truth for physical space; deleting a name is not counted until observed.

crates/cas/core/src/direct.rs · crates/cas/core/src/direct/fiemap.rs · nix/modules/disks.nix · docs/xfs-fixture.md · docs/physical-space.md

io_uringone ring per image reactor

The asynchronous IO interface each reactor thread uses for WAL appends, fence writes, fdatasync, chunk and page reads, and eventfd polls. Submissions are ordered by the reactor, not the kernel.

  • A ring of 256 entries with a registered wake eventfd; the reactor reaps completions by token and rejects stale tokens from retired generations.
  • An fdatasync is submitted only after the fence write it covers has completed, so no drain flag is needed on the host path; the raw reference backend uses IO_DRAIN instead.
  • A PollAdd on another reader's eventfd is how a waiter joins an in-flight chunk fetch without a thread.

crates/cas/daemon/src/local/reactor.rs · crates/cas/daemon/src/storage.rs · docs/async-io.md

KVM / TCGruns the guest CPU

Hardware virtualization for the guest. On Spark the lab runs one KVM outer VM that itself hosts the storage and the TCG-emulated inner guests, which is why every latency number so far is labelled a development measurement.

  • The host kernel has module loading disabled and no XFS driver, so the XFS test filesystem lives inside the KVM VM.
  • G1 requires a dedicated host and raw media before any p99 number counts.

docs/testbed.md · docs/ci-evaluation.md

IPC primitiveseventfd · memfd · flock · timerfd

The small kernel objects that carry control between threads and processes: eventfds for wakeups, memfds for shared memory, flock for exclusive ownership, timerfds for retry and recovery deadlines.

  • Every reactor, frontend and scheduler wake is an eventfd; a write of 1 that returns EAGAIN counts as already pending.
  • flock on the actual file description, not PID death, is what proves an old daemon's kernel IO can no longer touch a segment.
  • A monotonic timerfd wakes the frontend for the 100 ms admission retry and the 60 s recovery deadline even when the guest sends nothing.

crates/cas/core/src/eventfd.rs · crates/cas/daemon/src/deadline.rs · crates/cas/daemon/src/inflight/mapping.rs

Store on disk · <store-root>/ on one XFS filesystem — 6 parts

Everything durable lives under one root on a dedicated filesystem. Segment numbers come from one monotonic ticket namespace across chunk and staging segments, and the highest durable header is the allocation record: there is no counter file.

  • catalog/catalog.v2 names the images and snapshots that exist; loose files are not membership.
  • chunks/ holds shared immutable chunk segments; images/<id>/ holds a private manifest and staging WAL per image; snapshots/<id>/ holds reflinked manifests.
  • All integers are little-endian, every record is CRC32-checked, every IO offset is a 4 KiB multiple.
  • Rejected or torn suffixes are archived under rejected/ before any destructive repair.

catalog/catalog.v2CASCAT02 · membership

One small file listing every image and snapshot with its committed root, durable D and geometry. It is published atomically: write pending-<gen>-<attempt>.v2, fsync, rename, fsync the directory.

  • 64-byte header with a generation that increases on every change; 128-byte entries sorted by 16-byte id; whole-file CRC32.
  • Image entries carry image bytes; snapshot entries carry the source image plus the exact manifest commit and end they reflink.
  • Recovery reads only catalog.v2; pending files are never adopted or removed.

crates/cas/core/src/catalog.rs · crates/cas/core/src/catalog/format.rs · docs/catalog.md

chunks/segment-<ticket>.v2CASCHS02 · CASCHB02

Shared, immutable chunk segments. Each starts with a 4 KiB header, then batches of one 4 KiB header plus up to 63 chunks of exactly 4 KiB, each described by its BLAKE3 hash and CRC32. A chunk's address is its segment ticket and block index.

  • A full batch is 256 KiB; segments are preallocated to 64 MiB and may reach 256 MiB.
  • Ordinals are dense within a segment so a batch header's position is checkable against its number.
  • Collection copies live chunks into a fresh ticket and unlinks the old segment; the highest ticket is truncated to its header instead so the allocation record survives.

crates/cas/core/src/store/format.rs · crates/cas/core/src/store/file.rs · docs/storage-format.md · docs/chunk-store-io.md

images/<id>/manifest.v2CASMAN02 · COW B+tree

One append-only file per image mapping logical blocks to chunk hashes. A 4 KiB FILE header, then pages: LEAF (63 extents), BRANCH (252 children) and COMMIT. Children always sit at lower offsets than parents, so any COMMIT names a complete tree.

  • Leaf extents are (start, end, hash, kind): kind 1 is one hashed block, kind 2 is a ZERO range. Gaps read as zeros.
  • A COMMIT binds store, image, generation, root offset and height, and D, the mutation sequence the tree covers.
  • Old pages become garbage once no pinned root reaches them; reclamation hole-punches them without moving anything.

crates/cas/core/src/manifest/format.rs · crates/cas/core/src/manifest/file.rs · docs/storage-format.md · docs/manifest-editor.md

images/<id>/staging/segment-<ticket>.v2CASSEG02 · CASBAT02 · private WAL

The image's write-ahead log: preallocated 64 MiB segments of packed batches. Each DATA batch is a 4 KiB header of up to 63 descriptors plus up to 1 MiB of whole-block payload; a FENCE is a header alone carrying the mutation boundary a FLUSH made durable.

  • The segment header records store and image identity, writer epoch, number, capacity and the mutation sequence preceding it, so recovery chains segments without a directory index.
  • Every descriptor carries the request's serial, mutation sequence, attachment, queue and head, which is what live replay matches against the carrier.
  • Compacted payload is punched out below D; the batch headers stay until the whole segment can be unlinked.
isolated 4 KiB WRITE + FLUSH
12 KiB on disk
32 packed WRITEs + FLUSH
136 KiB for 128 KiB payload

crates/cas/core/src/append/format.rs · crates/cas/core/src/append/segment.rs · docs/storage-format.md · docs/wal-allocation.md

snapshots/<id>/manifest.v2FICLONE of an exact root

A snapshot is a reflink of the source manifest cut at the exact end of one COMMIT. It shares extents with the source until either side changes, and its catalog entry makes it a garbage-collection root.

  • Created only under a host quiescence, after the image was compacted through the cut.
  • A writable clone reflinks a snapshot and appends one COMMIT with a new image identity and D = 0; the old COMMITs stay as history but no longer match.

crates/cas/core/src/manifest/file/snapshot.rs · crates/cas/daemon/src/local/host/snapshots.rs · docs/snapshot-files.md · docs/host-snapshots.md

Outside the storesockets · reports · telemetry

Things that must not sit on the governed filesystem: vhost-user sockets, host.json and per-image reports, telemetry.jsonl, compaction-pause markers and the harness result directories.

  • cas-host refuses an endpoint whose socket parent or reports directory shares st_dev with the storage root.
  • Reports are written with create_new so a rerun never overwrites evidence.

crates/cas/daemon/src/host_service.rs · crates/cas/daemon/src/host_service/telemetry.rs

cas-host · one process · crate cas-daemon — 30 parts

The daemon that serves every image of one store. Per image it runs a vhost-user frontend thread and a reactor thread; per host it runs one compactor thread, a supervisor and a set of shared resources. It owns the guest protocol, ordering, admission and recovery; the formats and indexes come from cas-core.

  • cas-host init creates a store; cas-host --image <id>=<socket>... serves all catalog images, cold or retained.
  • The older cas-daemon binary serves one image with a chosen backend: raw io_uring, the v1 staging log, or the same local runtime; the suite keeps those as controls.
  • Failure is one shared domain: a socket worker error cancels every image. Image gates can fail one image, but the executable escalates.
  • Coordination uses threads, budgeted bounded channels, mutexes and eventfds. No cgroups, no async runtime.

Image frontendBackend · one per image

The vhost-user backend for one image: it maps guest memory, parses requests out of the rings, decides admission, keeps an owned record of every admitted request, and publishes completions. It runs on the framework's epoll worker thread under one mutex.

  • Every event: drain finished IO first, then visit each vring starting from a rotating queue, discovering and admitting requests.
  • Admission happens before any payload is copied; a waiting write stays in guest RAM with no mutation number.
  • Completion writes read payload, then the status byte, then add_used and the call eventfd, all through the accepted memory snapshot.
  • The struct is 30 fields spread across admission, frontier, lifecycle and recovery files that share its private state, which the September review called overgrown.

Backend VhostUserBackendMut PendingRequest Report

crates/cas/daemon/src/backend.rs · docs/shared-frontend.md · docs/frontend-allocation.md

vhost-user-backendrust-vmm 0.23.0 · vendored + patched

The pinned rust-vmm framework that speaks the socket protocol: it maps SET_MEM_TABLE regions, creates the vrings, registers kick eventfds in an epoll loop and calls the backend trait. Two narrow patches were added and pinned.

  • Patch 1: begin/end state-change hooks around every frontend message so the backend can quiesce before memory or queue changes.
  • Patch 2: GET/SET_INFLIGHT_FD forwarded to the backend instead of rejected, which is the whole live-recovery channel.
  • virtio-queue 0.18.0 is patched only to make DescriptorChain::new public so replay can walk a saved head whose ring slot has wrapped.
  • One socket thread decodes messages; one epoll worker owns all vrings of the image.

crates/vendor/README.md · crates/vendor/vhost-user-backend-0.23.0/src/handler.rs · crates/vendor/virtio-queue-0.18.0/CAS-PATCH.md · docs/vhost-user-notes.md

Request parserbytes, not descriptors

Decodes one descriptor chain into a typed Request: Read, Write, Zero, Flush, GetId, Unsupported or Invalid. It works on the byte stream, so a header split across descriptors or sharing one with data parses the same way.

  • Takes the first 16 readable bytes as the header and the last writable byte as the status, whatever the descriptor boundaries; tested at all 17 split positions.
  • Converts sectors to bytes and rejects anything not whole aligned 4 KiB blocks within capacity or above 1 MiB.
  • DISCARD and WRITE_ZEROES decode one 16-byte range; both become a ZERO mutation of at most 1 MiB. Unknown types answer UNSUPP, not an error.

request::parse Request Segment Completion

crates/cas/daemon/src/request.rs · docs/zero-discard.md

Discovery + read bypassFrontier · descriptor snapshots

Separates seeing a request from admitting it. Every descriptor on a queue is snapshotted and recorded DISCOVERED in the carrier as soon as it appears, so an independent read behind a write that is waiting for capacity can be admitted first.

  • A later request is eligible only if it is a read and overlaps no older discovered write, zero or discard; FLUSH conflicts with everything.
  • One ordinary head and one read candidate per queue are tried each pass, each through its own admission slot, so fairness tickets stay bounded.
  • Snapshot memory for the worst case, four queues of 256 chains, is reserved from the metadata budget up front: 6,297,600 bytes per frontend.
  • Discovered-but-unadmitted requests return to admission after a daemon replacement; they are never replayed as writes.

Frontier Discovered eligible conflicts

crates/cas/daemon/src/backend/frontier.rs · docs/read-progress.md

Queue admissionone waiting head per queue · 100 ms retry

Holds at most one waiting request per queue with the reason it waits: fairness turn, pending table full, or a storage pressure reason. A refused head keeps its descriptor in the guest ring and is retried on release wakeups and a 100 ms timer.

  • Asks the shared fair scheduler for a ticket and a turn, then asks storage to reserve credits; success commits the turn into the permit so the final credit drop wakes the next image.
  • Waits never expire into IOERR: the docs' five-second admission deadline was removed and the code says so explicitly.
  • Statistics per queue (started, resumed, canceled, longest wait) are what Update 03's admission-wait numbers come from.

QueueAdmission Waiting Reason

crates/cas/daemon/src/backend/admission.rs · docs/congestion-wait.md · docs/admission-release-wakeup.md

Pending tableowned request records

A fixed-capacity table of every admitted request, inserted before its payload is gathered or handed to storage and removed only when the completion returns ownership. Nothing admitted can disappear, and nothing can complete twice.

  • Capacity is the image request limit, 144 on the shared host, charged once at construction.
  • On failure every record whose queue is still valid gets one IOERR; the records stay for the terminal drain.

Pending<PendingRequest>

crates/cas/daemon/src/backend/pending.rs · docs/lifecycle.md · docs/control-tables.md

State-change bracketsquiesce · drain · rebase · resume

Wraps every QEMU configuration message in a pause. Before memory tables, queue geometry, enablement or reset change, the frontend stops admission, drains owned IO against the old state, and asks the reactor for a barrier; afterwards it validates cursors and resumes.

  • A drain has the 30 s IO deadline; timing out fails the image but never releases kernel-owned buffers or locks.
  • Queues touched by a change are marked blocked and rebase; a queue is served again only once its guest used index matches ours.
  • Device reset clears features and blocks every queue; a fresh GET_INFLIGHT_FD after serving starts a new writer epoch and carrier.

begin_change end_change rebase_queue validate_used_cursors

crates/cas/daemon/src/backend/lifecycle.rs · docs/lifecycle.md · docs/queue-setup-diagnostics.md

Inflight carrierthe memfd's state machine

The daemon side of the retained memfd: creates and seals it, records discovery and admission per head with acquire/release atomics, publishes the prefix P after ordered publication, and clears a slot only after the guest's used entry is written.

  • A PREPARED slot reserves its sequence even if the process dies before the counters update; reconciliation adopts it.
  • complete requires the request's publication boundary to be at or below P unless the image has failed.
  • Replacement: attach validates the header and every slot, reconcile compares each queue's used index (equal or one ahead), collects DISCOVERED slots in order and returns the replay set.

Carrier Entry Replay reconcile

crates/cas/daemon/src/inflight.rs · crates/cas/daemon/src/inflight/layout.rs · crates/cas/daemon/src/inflight/recovery.rs · crates/cas/daemon/src/inflight/mapping.rs · docs/inflight-format.md

Attachment sessionAwaitingFd → Replay → Waiting → Active

Negotiates the carrier with QEMU and drives a replacement daemon back to serving. Fresh attachments get a new epoch and carrier; a retained fd is reconciled, its missing mutations replayed from the original guest buffers, and its completions restored before admission opens.

  • Retained heads are re-decoded from the descriptor table with the patched DescriptorChain::new, then checked against the carrier's identity.
  • Standalone: replay runs on a cas-recovery thread under the 60 s deadline. Shared host: the frontend submits a Validated bundle and waits for the host coordinator to activate every image together.
  • On activation WRITE, ZERO and FLUSH completions are published at once (they are durable and fenced); READs re-enter the reactor; rejected entries get IOERR.

Session Phase activate_attachment finish_attachment Validated

crates/cas/daemon/src/backend/recovery.rs · docs/shared-live-recovery.md · docs/storage-design.md

Connection + reportService · Control

Owns one socket listener through connect, serve, disconnect and drain, and writes the final JSON report. The Control handle lets the supervisor snapshot the backend for telemetry or cancel it.

  • Registers the completion eventfd and timerfd on the epoll loop before taking the backend lock, because registration calls back into the backend.
  • A disconnect drains accepted IO without touching guest memory and records pending counts before and after.

Service Control FinalReport

crates/cas/daemon/src/service.rs · docs/disconnect-drain-evidence.md · docs/main-wrapper-integration.md

Faults + read tracingSIGSTOP points · CAS_TRACE_READS

Test-only hooks. Fault points stop the process at named write boundaries so the crash harness can cut at PREPARED, ACTIVE, before submit, after the append CQE, before and after sync, and around replay. Read tracing attributes each read's latency to admission, dispatch, IO and completion.

  • A pause publishes a JSON marker then raises SIGSTOP; the harness kills and replaces the daemon at that cut.
  • Tracing keeps 64-bucket histograms and the slowest examples; it is what located the 7.9 s admission wait behind an unrelated write.

crates/cas/daemon/src/fault.rs · crates/cas/daemon/src/read_trace.rs · crates/cas/daemon/src/local/host/fault.rs · docs/read-tracing.md · docs/shared-crash-controls.md

Local adapterLocal · runs on the frontend thread

The per-image storage adapter between the frontend and the reactor. It reserves credits, packs writes straight into their final WAL batch buffer, assigns mutation numbers, and hands commands to the reactor through a bounded mailbox.

  • Admit: enter the host quiescence gate, take a request credit, reserve WAL window space (which checks staging and physical capacity), then an append credit and a batch if needed. Any refusal is a typed pressure reason.
  • Gather: copy guest payload into the open Builder at its final position and record its CRC; the mutation sequence advances here.
  • Seal: a batch closes at 63 descriptors, 1 MiB, a non-write, a refusal, or the end of a processing pass. No timer holds a lone write.
  • A Permit travels with each request and releases, in order, its WAL slot, its admission entry and its fairness turn.

Local Shared Permit Packing Command

crates/cas/daemon/src/local.rs · crates/cas/daemon/src/local/pressure.rs · docs/wal-admission.md · docs/index-admission.md · docs/control-mailbox.md

WAL windowreserve before mutation

A per-image reservation window over the current WAL segment. It guarantees that every admitted write, plus its framing and a fence, fits in the segment and the staging interval index before a mutation number is assigned.

  • used = end + reserved + (issued − fenced + unsubmitted + initial fence) × 4 KiB; a candidate that would overflow marks rotation wanted and wakes the reactor.
  • Index pressure (unsubmitted mappings at capacity) refuses with WalIndex and asks for a compaction fence.
  • A dropped unused slot refunds itself and wakes the reactor.

Window Slot

crates/cas/daemon/src/local/window.rs · docs/wal-admission.md · docs/wal-allocation.md

Credit poolshost budgets · image shares

Request and byte budgets that bound what a guest can have in flight. Each image holds a Share of the host pools; a reservation takes the host lease first and rolls it back if the image cap refuses.

write requests
128 per image · 1,024 per host
read requests
8 per image · 64 per host
append bytes
8 MiB per image · 64 MiB per host
read bytes
8 MiB per image · 64 MiB per host
control reserve
8 / 64 KiB per image · 32 / 256 KiB per host
read reservation
bytes + 1 MiB scratch

HostPools Pools Share Credits

crates/cas/daemon/src/local/pools.rs · docs/daemon-owner-allocation.md

Image reactorthread local_async · owns the WAL + io_uring

One thread per image that executes storage IO in one write order. It appends batches, publishes completions contiguously, runs fence cohorts and fdatasync, resolves reads through staging, manifest and chunk store, and sequences background transactions with the compactor.

  • Each iteration: receive commands, reap completions, enforce deadlines, poll the background port, publish appends in order, refresh the window, advance ready work, dispatch new commands, submit bulk IO under the scheduler, then wait on two eventfds.
  • A later append never publishes before an earlier one completes; a slow write holds back later completions, in exchange for one visibility order.
  • A FLUSH captures a boundary and starts a fence cohort; later writes wait outside it so the sync stays finite.
  • While a rotation is pending, mutation dispatch stops but reads and covered flushes continue.

Reactor Work Pending Slots dispatch finish_ready

crates/cas/daemon/src/local/reactor.rs · crates/cas/daemon/src/local/reactor/slots.rs · docs/async-io.md · docs/control-tables.md · docs/physical-runtime.md

Read state machinestaging → manifest → cache → chunk

A read waits for its captured mutation boundary to publish, freezes the staged ranges and manifest root it will use, then walks the uncovered blocks one at a time: manifest page lookups through the page cache, then the shared chunk cache, then a coalesced fetch from the chunk store.

  • Staged WAL ranges win, including explicit ZERO; a partial read verifies the whole original payload CRC before copying a sub-range.
  • A cache miss claims a fetch leader for that hash; other readers register a waiter and poll the leader's eventfd instead of issuing their own IO.
  • The fetched block is verified by CRC and BLAKE3, offered to the cache as a separate charged copy, then copied into the response.
  • One 1 MiB cold read does not submit its 256 fetches at once; that sequential walk is an open cost.

Read Stage next_block chunk

crates/cas/daemon/src/local/reactor/read.rs · crates/cas/core/src/append/read.rs · docs/read-ownership.md · docs/coalesced-fills.md · docs/metadata-cache.md

Background portPort::poll · turns for the compactor

The reactor side of every background transaction. It decides when the image is ready for a compaction or rotation turn, hands the compactor a selection under the sequencer, installs receipts, and applies reclamation accounting.

  • A compaction turn is queued when durable data is uncompacted and either 100 ms have passed without a write, the oldest dirty data is 1 s old, or staging, index or host pressure forces it.
  • Events from the owner: Select, Published, Reclaimed, Allocate, Rotated, Deferred, Quiesce, Resume, CompactQuiescent; each answered by a typed Reply.
  • A background transaction that gets no reply for 30 s fails the image; the owner keeps whatever files and buffers it holds.

Port Event Reply Turn Ready

crates/cas/daemon/src/local/host.rs · docs/host-runtime.md · docs/compaction.md

Compactor ownerthread cas-compactor · one per host

The single background thread. It holds the chunk store writer and every image's mutable manifest, so all compaction, WAL rotation, collection and snapshot IO is serialized on it, with blocking file IO rather than io_uring.

  • Loop: if the disk is pressured and a second passed, collect; else take one Ready item within 50 ms: an image turn (compact or rotate), a collect request or a snapshot request.
  • Every direct read or write on this thread first asks the IO scheduler for a background opportunity.
  • Errors fail the host gate when the store or an account failed, otherwise only that image's gate; a panic fails the host.

Owner Endpoint Ready

crates/cas/daemon/src/local/host/worker.rs · docs/host-runtime.md · docs/host-scheduling.md

Compaction turnselect → load → prepare → write → publish → reclaim

Turns a bounded, durable WAL prefix into shared chunks and a new manifest root. Data is durable before the mapping that names it; the mapping is published before the WAL space it covers is reclaimed.

  • Select (reactor, under the sequencer): whole batches above D up to E, at most 1 MiB payload and 318 edits, resumed from a validated cursor.
  • Load and prepare (owner): read and CRC-verify the payload, drop versions covered by later edits, hash surviving nonzero blocks with BLAKE3, build the copy-on-write pages in memory.
  • Write: reserve a physical promise, insert missing chunks in batches of 63 and sync, append manifest pages and COMMIT and sync.
  • Publish (reactor): adopt the new View, advance D, drop staging mappings at or below D; then the owner punches or unlinks WAL payload not pinned by readers or replay identities.
settle · forced age
100 ms · 1 s
per transaction
≤ 1 MiB input, ≤ 318 edits, ≤ 128 MiB new pages

crates/cas/daemon/src/local/host/worker.rs · crates/cas/core/src/append/compaction.rs · crates/cas/core/src/append/compaction/output.rs · docs/compaction.md · docs/compaction-crash-cuts.md

WAL rotationnew segment · fresh attachment

Allocates the next staging segment for an image when its window would overflow or a fresh attachment needs a new writer epoch. The reactor prepares under its sequencer; the owner creates, preallocates and syncs the file; the reactor installs it.

  • The staging quota is reserved before the file exists; the promise is finished against the measured allocation.
  • A rollover fence syncs the old segment before the new one accepts data.

crates/cas/daemon/src/local/host/worker.rs · crates/cas/core/src/append/rotation.rs · docs/wal-allocation.md · docs/segment-allocation.md

Host collectionpause all · mark · copy · unlink · punch

Quiescent mark-and-sweep across every image. It pauses all admission including reads, drains accepted work, fences every image, marks chunks reachable from every manifest root, snapshot and pinned view, copies live chunks out of mixed segments, unlinks dead ones, and hole-punches dead manifest pages.

  • Triggered by an administrative request or automatically once a second while the physical governor is pressured.
  • Under pressure it alternates one quiescent compaction with another sweep until nothing advances, then reports capacity exhausted with write admission still closed.
  • The 30 s check sits between steps, not inside blocking syscalls or tree walks, so a long overwrite history can pause guests longer.

crates/cas/daemon/src/local/host/collection.rs · crates/cas/core/src/store/file/collection.rs · crates/cas/core/src/store/file/collection/sweep.rs · docs/host-collection.md · docs/chunk-collection.md · docs/host-quiescence.md

Snapshotcompact to a cut · reflink · catalog

Publishes an exact snapshot of one image under the same host pause: compact through the fenced cut, reflink the manifest at that COMMIT, insert the catalog entry, sync. It needs a catalog owner and a physical governor, so only recovered hosts can snapshot.

  • The cut is the image gate's durable sequence; compaction repeats until the manifest's D reaches it.
  • A clone gets its own image identity, WAL and sequence namespace but starts from the snapshot's mapping.

crates/cas/daemon/src/local/host/snapshots.rs · crates/cas/core/src/manifest/file/snapshot.rs · docs/host-snapshots.md · docs/snapshot-files.md

Capacity controlstaging 75 % / cap / 60 % · reserve R

The rules that connect disk accounting to admission. Staging pressure starts compaction at 75 % of an image or host quota, stops write admission at the quota and resumes below 60 %. Physical pressure preserves the background reserve and forces collection.

  • At startup every current segment, and their sum, must sit strictly below 60 % of the staging quota; at defaults that permits at most nine images.
  • A compaction reserves one segment plus chunk headers plus manifest bytes plus a 16 MiB margin before it writes.
reserve R
3S + M + 16 MiB = 336 MiB at defaults
staging quota
256 MiB per image · 1 GiB per host

crates/cas/daemon/src/local/host/capacity.rs · docs/live-capacity-control.md · docs/staging-capacity.md · docs/physical-collection-readiness.md

Shared host stateSharedHost · one per process

What every image shares: the chunk store reader, the two caches, the fetch registry, the admission scheduler, the quiescence gate, the failure gates and the accounts. Built once from one metadata budget before any image attaches.

SharedHost Resources Roots

crates/cas/daemon/src/local/host.rs · docs/host-runtime.md · docs/daemon-owner-allocation.md

Fair admissionFIFO per image · byte deficit round robin

Decides which guest request head may be admitted next across images. Eight heads per image (an ordinary head and a read candidate per queue), FIFO among an image's eligible heads, and a 1 MiB byte quantum rotated across images.

  • A ticket registers a head; a turn is granted only if that exact head is chosen, otherwise the chosen image's frontend is woken.
  • A refused turn now rotates its head behind the image's other heads and advances the cursor. Before the 14 September fix a refused write was re-chosen every retry and starved the read behind it for 8.1 s.
  • A release that races a refused turn still marks heads ready, closing the missed-wakeup bug found on 13 September.

Fair Ticket Turn Release

crates/cas/daemon/src/local/host/fair.rs · docs/host-scheduling.md · docs/admission-release-wakeup.md

Quiescence gateAdmission · counts live owners

Counts guest owners in flight and can stop new ones atomically. Collection and snapshots pause it, wait until the count is zero, run, and resume; dropping a started pause without finishing fails the gate closed.

Admission Entry Quiescence

crates/cas/daemon/src/local/host/admission.rs · docs/host-quiescence.md

Completion gatesHostGate → ImageGate

The failure words. Locking an image gate locks the host gate first and copies any shared failure in, so a completion can never be published past a poisoned store. The frontend holds the guard from its storage decision through the guest's used entry.

  • A store output failure poisons the host gate before the worker is told.
  • A WAL or manifest failure stays scoped to its image while the shared store is healthy.

HostGate Gate ImageState

crates/cas/daemon/src/local/state.rs · docs/host-runtime.md · docs/recovery-lock-readiness.md

Host servicesupervisor · telemetry · reports

The cas-host runtime around the images: validates endpoints, scans and recovers the store, attaches one frontend per image, supervises the socket threads every 10 ms, samples telemetry every 500 ms, and writes host.json plus one report per image.

  • Strict telemetry fails the run at 4,096 samples or 64 MiB; the interactive lab uses bounded rotation instead.
  • There is no control socket: control is the command line plus files.

crates/cas/daemon/src/host_service.rs · crates/cas/daemon/src/host_service/telemetry.rs · docs/host-service.md · docs/pressure-telemetry.md

Store recoveryinspect everything · then repair

Opens a store for serving. It locks tickets, inspects the catalog, chunk store, every manifest and WAL and every snapshot read-only, validates that every root's chunks exist and every required prefix is covered, and only then repairs: archive rejected suffixes, truncate, sync, write recovery fences.

  • Cold: no guest survives; each WAL rotates to a new writer epoch and fences before serving.
  • Retained: every frontend must first hand over its validated carrier state; the coordinator replays every image's missing mutations from surviving guest RAM, fences, and activates all images behind one barrier.
  • One missing or incompatible image blocks the whole host: the dependency graph is shared.

Inspection Checked Recovered RetainedHost

crates/cas/daemon/src/local/host/recovery.rs · crates/cas/daemon/src/local/host/recovery/frontend.rs · crates/cas/daemon/src/local/host/recovery/live.rs · crates/cas/daemon/src/local/host/initialize.rs · docs/shared-recovery.md · docs/shared-live-recovery.md · docs/host-initialization.md

Storage dispatchRaw · Staging v1 · Local · Opening

The enum the frontend talks to. Raw is a plain io_uring file backend and Staging the v1 serial log; both are kept as controls for the suite. Local is the runtime above; Opening is a locked, unrepaired image awaiting inflight negotiation.

  • The frontend still inspects the variant at several call sites to decide gather versus bounce copy and whether an error is fatal; the review asked for one submission entry point.
  • Restartable staging keeps the one-request-in-flight workaround from Update 01 and flushes after every write.

Storage Opening Permit Completed

crates/cas/daemon/src/storage.rs · crates/cas/daemon/src/storage/opening.rs · docs/rust-design.md

cas-core · library · formats, indexes, accounting — 38 parts

The storage library. It owns every byte format and its recovery parser, the in-memory indexes, the copy-on-write manifest tree, the caches, the budget system that charges before allocating, disk accounting and the bulk-IO arbiter. It has no threads of its own; the daemon drives it.

  • Two crate-wide constants: BLOCK_SIZE = 4096 and MAX_REQUEST_BYTES = 1 MiB. Virtio sectors are a separate 512-byte unit.
  • The invariants are carried by types: aligned buffers, leases that release on drop, permits that must be finished, receipts only successful IO can produce, formats validated on decode and re-decoded after encode.
  • The review found no data-loss, ordering or lock-order bug in the storage, replay, carrier or cache paths; every unsafe block has a justification.

WAL (append log)Log · v2 packed batches

The per-image write-ahead log. It encodes packed batches, tracks the three sequence counters issued ≥ published ≥ durable, publishes mutations into an interval index in order, writes fences for FLUSH, and offers selection, publication and reclamation interfaces to compaction.

  • A WRITE completes at published (P); a FLUSH waits for durable (E); compaction later moves the manifest's D forward. During serving D ≤ E ≤ P.
  • Nothing is ever overwritten in place; a segment is preallocated, appended, punched and finally unlinked.

Log Config Limits Status

crates/cas/core/src/append.rs · docs/storage-design.md · docs/wal-allocation.md

Batch codecBuilder · Header · SegmentHeader

The v2 byte layout and its validator. A batch is one 4 KiB header (64-byte envelope plus up to 63 descriptors of 64 bytes) followed by packed whole-block payload up to 1 MiB. A FENCE is a header alone. Decoding checks magic, version, CRC, counts, dense sequences, payload coverage and range arithmetic before any field is trusted.

  • Builder gathers guest bytes into their final position, records each payload CRC, then seals the header without moving payload.
  • follows() is the chain rule recovery, compaction and reclamation share: the next batch must fit before the segment's fence slot and continue the sequence.
descriptors per batch
63
max batch
4 KiB + 1 MiB
magic
CASSEG02 · CASBAT02

crates/cas/core/src/append/format.rs · docs/storage-format.md

Staging indexinterval map · preallocated nodes

A disjoint interval map from logical byte ranges to the newest published payload location or ZERO. It is what makes recent writes take precedence over the manifest on reads.

  • A pinned B-tree port over a preallocated node pool: 65,536 intervals reserve 13,174 slots of 1 KiB, about 12.9 MiB per image, charged to metadata up front so publication can never fail on memory.
  • Replacing a range can split one straddling predecessor and remove covered entries, so each descriptor costs at most two intervals.
  • Compaction drops every mapping at or below D.

crates/cas/core/src/append/index.rs · crates/cas/core/src/append/index/nodes.rs · docs/staging-metadata.md

Submission + cohortsissued · published · durable

The counters and the rules around them. prepare_append seals a batch at the next offset and advances issued; publish_append refuses unless the batch is the next in order; prepare_fence captures the boundary of a finite cohort; complete_sync advances durable only after every covered append published.

  • A batch that completes out of order waits as Pending until its predecessor publishes.
  • A 4 KiB fence slot is reserved at the end of every segment so a FLUSH can always be recorded without rolling over.
  • A failed sync never advances E; a poisoned log refuses everything after.

Submission Position Cohort

crates/cas/core/src/append/submission.rs · docs/persistence-model.md

Segments + pinspreallocated files · read pins · rotation

Segment files and their ownership. Creation encodes the header, preallocates the capacity, syncs the file and its directory before any data. Pins are one atomic per 4 KiB block so a read can hold the batch it depends on and a compaction scan can hold a whole segment.

  • Rotation is three-phase so file creation can run on the owner thread: prepare under the sequencer, create off it, install and recheck the boundary.
  • A fresh attachment's writer epoch is its own segment ticket.

crates/cas/core/src/append/segment.rs · crates/cas/core/src/append/segment/pins.rs · crates/cas/core/src/append/rotation.rs · docs/segment-allocation.md

Read plansReadPlan · covered mask

Builds the immutable plan a read executes: the staged ranges it covers (with pins acquired), a 256-bit mask of covered blocks, and a clone of the committed manifest View for the rest. Dropping the plan releases the pins.

  • A plan is Pending until the log has published the boundary the read captured.
  • Executing a range reads directly into the response when the whole payload is wanted, or into scratch with a full-payload CRC check otherwise.

crates/cas/core/src/append/read.rs · docs/read-ownership.md

WAL recoveryinspect · fresh · live replay

Reopens a log. Inspection replays every batch header in ticket order without trusting anything after the first failure, skips punched payload below D through intact headers, verifies payload above D, and records where the valid prefix ends. Repair archives the rejected suffix, truncates, syncs, and writes a recovery fence.

  • Cold: rotate to a new epoch and fence before serving.
  • Live: accept up to 1,024 retained identities, require the missing tail to be exactly the retained mutations above the prefix, re-verify every retained mutation at or below the prefix byte for byte against its descriptor, then replay one batch per mutation in original order.
  • A prefix below D is an error, never a shorter replay.

Recovery LivePlan LiveRecovery SharedRecovery

crates/cas/core/src/append/recovery.rs · crates/cas/core/src/append/shared.rs · docs/shared-recovery.md · docs/staging-base.md

Selection + outputSelection · Input · Prepared

The core of one compaction transaction. Selection captures the base View, a cursor and scan-pinned spans. Loading scans headers from the cursor, verifies payload above D and collects at most 318 edits and 1 MiB. Preparing drops covered edits, hashes survivors and builds the COW pages. Writing inserts chunks, then publishes the manifest, returning a receipt naming the exact new root.

crates/cas/core/src/append/compaction.rs · crates/cas/core/src/append/compaction/output.rs · docs/compaction.md · docs/compaction-crash-cuts.md

Reclamationpunch payload · unlink segments

Frees WAL space below D. A DATA batch entirely at or below D whose header block no reader pins and whose segment no scan holds has its payload hole-punched; a segment is unlinked only when fully covered, not current, below the ticket high-water mark and below the oldest live replay identity, with no other owners.

  • Punched batches keep their headers so recovery can chain past them.
  • Physical release is measured with st_blocks, never assumed.

crates/cas/core/src/append/reclaim.rs · docs/manifest-reclamation.md · docs/wal-allocation.md

Chunk storeStore · Reader · shared content

The shared, immutable content store under chunks/. One writer (the compactor) inserts batches of verified 4 KiB chunks; any number of readers resolve a hash to its segment and block without holding the writer through IO. An index entry is usable only after its batch is synced.

Store Reader Shared Segment

crates/cas/core/src/store.rs · crates/cas/core/src/store/file.rs · docs/chunk-store-io.md · docs/store-readers.md

Chunking + hashingfixed 4 KiB · BLAKE3 · zero = hole

Content identity. A chunk is one 4 KiB block named by its BLAKE3-256 hash; an all-zero block is not stored at all and becomes a manifest hole. Boundaries are fixed logical blocks; content-defined chunking remains an unimplemented census arm.

  • A one-byte edit makes a new chunk; identical blocks share even across unrelated images.
  • The index compares full hashes and does not byte-compare a duplicate against disk.

crates/cas/core/src/chunk.rs · docs/census.md

Hash indexhash → 48-bit ticket · 16-bit block

The in-memory map from 32-byte hash to a 64-bit store address, in a hashbrown table on the budgeted allocator with a GC mark bit per entry. Growth is reserved from the metadata budget before the chunks it would describe are written; denial leaves existing lookups intact.

  • Rebuilt on recovery from the inline hashes in every valid chunk record, so CRC checks suffice and BLAKE3 is not recomputed.
  • Every stored hash needs RAM, including dead entries awaiting collection: unique capacity is bounded by memory as well as disk.
address
segment ≤ 2⁴⁸−1, block ≤ 65,535 (256 MiB segment)

crates/cas/core/src/chunk_index.rs · docs/cache-table-capacity.md · docs/index-admission.md

Chunk batch codecCASCHS02 · CASCHB02 · 63 chunks

The chunk segment header and batch envelope mirror the WAL layout with their own magic. A 64-byte descriptor per chunk carries the hash, payload offset, length (exactly 4096) and CRC32; ordinals are dense per segment. No fence is needed: the data sync gates index publication.

crates/cas/core/src/store/format.rs · docs/storage-format.md

Insertdedupe · reserve · write · sync · publish

The writer path. Dedupe against the index and within the batch, reserve index growth, pick or create a segment, snapshot the offset under the lock, then without the lock preallocate, write and fdatasync, and only then publish each hash. Any exit with the output pending poisons the store.

crates/cas/core/src/store/file/insert.rs · docs/chunk-store-io.md

Readerplan → header → payload · verified

A cloneable handle that resolves a hash under a short mutex to a file, batch and address, then advances through header and payload IO on the caller's ring or thread. The header descriptor must repeat the hash; the payload must match its CRC.

  • Lookups return WouldBlock during collection so a reader never sees a segment mid-move.

crates/cas/core/src/store/file/read.rs · docs/store-readers.md

Collection + sweepmark bits · copying cleaner

Garbage collection of chunk segments. Marking sets the bit on every reachable index entry; victims are classified dead, live or mixed. A mixed victim's live chunks are re-read, verified, rehashed and appended into a fresh ticket, then the old segment is trimmed and unlinked. Unmarked entries are removed at the end.

  • Requires no outstanding read plan on any segment file, which is why the host quiesces first.
  • A crash before unlink leaves duplicate chunks; a crash after keeps the durable copy. Either way one copy of every reachable chunk survives.

crates/cas/core/src/store/file/collection.rs · crates/cas/core/src/store/file/collection/sweep.rs · docs/chunk-collection.md

Store inspectionrebuild index · archive tails

Opens chunks/ read-only, requires canonical names, walks every batch header for continuity and CRC, stops at the first invalid one, and rebuilds the hash index from the inline hashes. Recovery archives each tail and rejected creation, truncates, syncs and unlinks.

crates/cas/core/src/store/file/recovery.rs · docs/chunk-store-io.md

ManifestCOW B+tree · block → hash

The per-image map from logical block to content hash, as a persistent copy-on-write B+tree of checksummed 4 KiB pages in an append-only file. A COMMIT page ends every transaction; readers pin one committed root and never see a half-written tree.

  • It is a page-addressed tree, not a Merkle tree: pages are found by file offset, hashes live only in leaves.
  • Height is at most 8; a change touches at most 8H + 8 new pages, which bounds a transaction to under 90 MiB of new pages.

crates/cas/core/src/manifest.rs · docs/manifest-editor.md · docs/manifest-views.md

Page codecCASMAN02 · LEAF · BRANCH · COMMIT

The 64-byte page header (magic, kind, level, count, own offset, CRC) and the three bodies: 63 leaf extents of 64 bytes, 252 branch children of 16 bytes, or a COMMIT naming store, image, generation, root, height, D and image bytes. Every encoder re-decodes its output as a self-check.

crates/cas/core/src/manifest/format.rs · docs/storage-format.md

Tree + COW editorLookup · Prepared · 8H+8 bound

Traversal and mutation. Lookup is an IO-neutral state machine that asks for one page at a time and can be driven by a synchronous read or an io_uring completion. The editor applies up to 318 edits in memory, detaches subtrees a ZERO covers without reading them, splits only over-full nodes, and emits only pages reachable from the final root.

  • Every decoded node is checked against its parent's level, minimum key and bounds, so a duplicate child under the wrong key is rejected beyond the CRC.
  • A draft edit updates one path at a time; grouping edits by path is still open work.

crates/cas/core/src/manifest/tree.rs · crates/cas/core/src/manifest/tree/lookup.rs · crates/cas/core/src/manifest/tree/editor.rs · crates/cas/core/src/manifest/tree/node.rs · docs/manifest-editor.md · docs/manifest-views.md

Manifest fileManifest · View · Read

The locked owner of manifest.v2 and the immutable views over it. publish requires the prepared transaction to start at the current commit and end, reserves the successor pin, preallocates, writes, fdatasyncs, then swaps the commit; any failure leaves the owner poisoned with the old root still published.

  • A View clones the file handle, the commit and a root pin; a Read in flight keeps both alive after the View and even the Manifest are dropped.
  • inspect scans pages backwards from EOF and selects the newest COMMIT whose whole tree walks; a torn transaction leaves the earlier root.

crates/cas/core/src/manifest/file.rs · docs/manifest-views.md · docs/manifest-root-pins.md

Root registrypins · incarnation

A per-file registry of retained roots keyed by commit end. Every View, in-flight Read and snapshot holds a pin; page reclamation captures all of them to decide what is reachable. Each registry has a process-unique incarnation that is part of the page-cache key, so a reopened file cannot inherit stale pages.

crates/cas/core/src/manifest/file/pins.rs · docs/manifest-root-pins.md

Page cache16 MiB · verified pages only

A host-wide LRU of verified 4 KiB tree pages keyed by (incarnation, commit end, offset). Only a page that passed the checked descent may enter; a page a reader still holds cannot be evicted, and a refusal is counted rather than failed.

  • Because the key includes the commit end, every new publication misses on unchanged interior pages: an open design item.
  • It lives inside the 128 MiB foreground metadata budget, not beside it.

crates/cas/core/src/manifest/file/cache.rs · docs/metadata-cache.md

Page reclamationmark reachable · punch gaps

Frees dead manifest pages in place. It requires the exact owned EOF, re-verifies every retained COMMIT, walks each pinned root once collecting live offsets under the metadata budget, sorts and dedupes them, punches every gap in at most 128 MiB slices, and trims preallocation beyond EOF.

  • The file's logical end never shrinks: history is punched, not compacted away.

crates/cas/core/src/manifest/file/reclaim.rs · docs/manifest-reclamation.md

Snapshots + clonesFICLONE at an exact COMMIT

A snapshot is a reflink of a pinned View truncated to its commit end, synced with its directory. A clone reflinks a snapshot and appends one COMMIT for the new image with generation 1 and D = 0; the shared pages diverge by copy-on-write from then on.

crates/cas/core/src/manifest/file/snapshot.rs · docs/snapshot-files.md

Clean chunk cache256 MiB LRU · read fills only

A host-wide cache of verified chunk payload keyed by full hash. Only verified read fills populate it; writes and compaction never prewarm it. Eviction drops membership while a reader keeps its own clone and charge.

  • Shared by every image with no partition and no scan resistance: one guest's scan can displace another's hot set.
  • A fill recomputes BLAKE3 before insertion and makes a separate charged copy so the foreground read credit is not held forever.

crates/cas/core/src/cache.rs · docs/clean-cache.md

Intrusive LRUhashbrown table · fixed capacity

A hash table whose entries link to their older and newer neighbours, so promotion and eviction are lookups rather than scans. It reserves twice the resident capacity up front; the pinned hashbrown version rehashes tombstones in place, so churn never grows the table.

crates/cas/core/src/cache/lru.rs · docs/cache-table-capacity.md

Coalesced fills1 leader · bounded waiters · eventfd

One fetch per missing hash across the host. The first reader becomes the leader and fetches; concurrent readers become waiters holding a nonblocking eventfd the leader writes on completion or failure. A leader that drops without finishing publishes Interrupted.

bounds
1,024 leaders · 1,024 waiters

crates/cas/core/src/cache/fills.rs · crates/cas/core/src/eventfd.rs · docs/coalesced-fills.md

Budgetscharge before allocate · release on drop

Two-dimensional counters (bytes and requests) with a hard limit, and four ways to hold a charge: a raw Lease, a Share nesting an image cap inside a host cap, an Allocator that charges every layout before the global allocator sees it, and a BudgetArc whose control block is charged too.

  • The daemon creates one 128 MiB foreground metadata budget and one 128 MiB compaction budget and threads the same handles into every core object; there is no unbudgeted constructor on the Linux paths.
  • Growth of a table charges old and new storage for a moment, which is why index reservation happens before store IO.

crates/cas/core/src/budget.rs · crates/cas/core/src/budget/allocator.rs · crates/cas/core/src/budget/shared.rs · docs/shared-allocation.md · docs/daemon-owner-allocation.md

Bounded mailboxesQueue · channel · never grows

A fixed ring of slots reserved at construction and a mutex-plus-condvar channel over it. try_send never blocks; close rejects new publication while draining what is queued; a poisoned lock fails closed and wakes everyone. The daemon uses these for every reactor, owner and reply channel.

crates/cas/core/src/budget/queue.rs · crates/cas/core/src/budget/channel.rs · docs/control-mailbox.md

IO schedulerbulk opportunities · 3 demand : 1 background

Arbitrates who may hand the next bulk IO to the kernel: reactors with queued appends and reads, or the compactor thread's direct reads and writes. Every fourth opportunity is reserved for background; unused turns are borrowed by the other side.

  • A reactor marks itself ready and takes a turn only when it is the current selection; the scheduler wakes it through its bound eventfd.
  • The compactor waits at most 30 s per syscall; expiry fails that transaction, which the review flagged as a parameter to revisit.
  • FLUSH fences and eventfd polls bypass the arbiter entirely.

crates/cas/core/src/scheduler.rs · docs/host-scheduling.md · docs/congestion-wait.md

Disk accountingGovernor · Staging · reserve R

Physical and logical capacity. The Governor observes real allocation with fstatvfs and admits foreground promises below capacity − R and one exclusive background borrower within R. Staging is the per-image and host WAL quota with the 75 % / cap / 60 % hysteresis. Every promise is finished against a measured allocation or fails the account.

  • Recovery paths borrow a governed handle so repairs and archives are accounted like any other output.
  • The arithmetic-only Space is used by production only through the Governor.
reserve
R = 3S + M + 16 MiB
margin
16 MiB

crates/cas/core/src/space.rs · crates/cas/core/src/space/filesystem.rs · crates/cas/core/src/space/staging.rs · crates/cas/core/src/space/recovery.rs · docs/physical-space.md · docs/staging-capacity.md

Catalogmembership · atomic publish

The durable list of images and snapshots. Changes are prepared in memory with the next generation, written to a pending file, synced, renamed over catalog.v2 and the directory synced. Inspection reads only the live file and never adopts a pending one.

crates/cas/core/src/catalog.rs · crates/cas/core/src/catalog/format.rs · docs/catalog.md

Segment ticketsone namespace · highest header wins

The store-wide allocator for segment numbers. Opening scans chunks/, every image's staging/ and their rejected/ archives for the highest name; allocation is highest + 1 and only succeeds once the new header is durable. A failed creation poisons every allocator user.

crates/cas/core/src/segments.rs · docs/segment-allocation.md

IO primitivesdirect · aligned · directory · encoding

The shared low level. AlignedBuffer is a 4 KiB-aligned boxed slice that cannot be misaligned by construction. direct opens with O_DIRECT and flock, checks alignment via statx, rejects short IO, and wraps fallocate, punch, reflink, FIEMAP and sync. Directory locks, syncs, renames and archives rejected suffixes. encoding is the little-endian and CRC32 helper set.

  • Test-only fault injection can fail or pause any of these syscalls once, which is how the persistence oracle explores crash schedules.

crates/cas/core/src/aligned.rs · crates/cas/core/src/direct.rs · crates/cas/core/src/directory.rs · crates/cas/core/src/encoding.rs · docs/async-io.md

v1 staging loglegacy · 8 KiB slots · serial

The original single-file log from Update 01: fixed 8 KiB slots of one metadata block and one payload block, CRC per block, a FENCE that can only be a metadata block. Still used by casctl staging-check and the daemon's --backend staging control. No conversion path to v2 exists.

crates/cas/core/src/staging.rs · crates/cas/core/src/staging/format.rs · docs/history/spec-v1.md

Censusoffline · 4 KiB and 16 KiB

The offline measurement behind Update 01's sharing numbers. It hashes raw images at both chunk sizes with the same nonzero-BLAKE3 rule as the compactor and partitions each image's bytes into base, prior-image, within-image and unique classes. It uses ordinary collections and never runs in the daemon.

crates/cas/core/src/census.rs · docs/census.md

IO metricsthread-local scope counters

Fifteen operation counters (reads, writes, syncs, allocations, punches, hashing, lock waits, scheduler waits) accumulated inside a thread-local scope. The compactor wraps each turn in one; the counts feed the per-phase compaction telemetry.

crates/cas/core/src/io_metrics.rs · docs/pressure-telemetry.md

Build, run, measure · Nix · casctl · cas-harness · evidence — 14 parts

The outer system. Nix pins the toolchain, builds the binaries and assembles NixOS guests; cas-harness launches QEMU and the daemon, drives workloads and refuses evidence that is incomplete or unbound from its source; casctl is the interactive front door; every result lands in docs/ as a record the site imports.

  • A run is one Nix wrapper that execs the packaged harness with a JSON build record naming the guest launcher, the daemon and the source revision.
  • Nothing in this layer produces research-gate evidence: every report hard-codes paper_gate: null. Checkpoints C1–C5 are development milestones.

Rust workspacefour crates · pinned 1.98.1

cas-core, cas-daemon, cas-harness and cas-cli, edition 2024, with two vendored rust-vmm crates patched in through Cargo. The harness and CLI do not link the daemon crate: they run cas-daemon and cas-host by store path from build records.

  • Workspace lints deny unsafe_op_in_unsafe_fn and undocumented unsafe blocks.
  • just check runs rustfmt, clippy with warnings denied, the tests and a whitespace check; the same six checks open every checkpoint suite.
native tests on main
504 passed, 25 fixture-dependent ignored

Cargo.toml · rust-toolchain.toml · crates/harnesses/Cargo.toml · crates/cas/cli/Cargo.toml

casctlnamed labs · SSH · bench · census

The interactive CLI. casctl new boots a lab: a 4 GiB KVM outer VM owning an XFS disk and cas-host, with one to four TCG inner guests on private ext4 disks over one shared store. A user systemd unit owns the lab after the command returns.

  • Per lab: config.json, a Nix GC root for the pinned build, a sparse 4 GiB disk.raw, an ed25519 key and runs/<timestamp>/ with logs, memory samples and SSH config.
  • Inside the outer VM, lab-host runs cas-host init once, then cas-host with rotating telemetry, and launches each inner guest with its own socket; SSH reaches vmN through a loopback port and a ProxyJump.
  • bench runs four fio cases over SSH at QD1 and keeps every command, output and the memory and storage samples around it.
  • staging-check exercises the v1 log; census hashes raw images at 4 KiB and 16 KiB.
preflight
/dev/kvm, ≥ 25 GiB free, ≥ 6 GiB available
unit limits
MemoryMax 6G, no swap

crates/cas/cli/src/main.rs · crates/harnesses/src/lab.rs · crates/harnesses/src/lab/client.rs · crates/harnesses/src/lab/runtime.rs · crates/harnesses/src/lab/bench.rs · nix/lab/host.nix · nix/lab/guest.nix · docs/casctl.md

cas-harnesslaunch · verify · bind to source

The Rust experiment driver. Every subcommand owns its child process group, applies deadlines, decodes guest and daemon reports into typed records with hard acceptance conditions, and records the exact source and executables that produced them.

  • SIGINT and SIGTERM set a flag every poll loop checks; a child group is killed on drop, exit or timeout.
  • A run refuses to start if the running binary is not the one in the build record, and refuses to pass if the checkout changed during the run.

crates/harnesses/src/main.rs · crates/harnesses/src/process.rs · crates/harnesses/src/evidence.rs · crates/harnesses/src/source.rs · docs/testbed.md

vm runnersmoke · recovery · live recovery · reset

One daemon plus one KVM guest against a fresh 128 MiB scratch image. Smoke runs fio with verification; recovery kills the daemon after a confirmed FLUSH and reads back in a fresh guest; live recovery pauses the daemon at a named fault point, kills it, optionally cycles replacement daemons, and requires the same guest to finish; device reset unbinds and rebinds the driver twice.

guest
1024 MiB, 2 or 4 vCPUs, KVM
IO
64 MiB written and read; 128 MiB verified

crates/harnesses/src/vm.rs · crates/harnesses/src/vm/live.rs · crates/harnesses/src/vm/reset.rs · crates/harnesses/src/vm/interactive.rs · crates/harnesses/src/qemu.rs · docs/testbed.md · docs/shared-crash-controls.md

checkpoint suiteC1 14 · C2 16 · C3 32 · C4 40 · C5 41

The cumulative scenario inventory. Six source checks, then the reference and crash-point VM runs, the persistence model from C3, the XFS and shared fixtures with six compaction crash cuts from C4, and the competing-guest pressure fixture at C5. verify-suite re-scans everything from disk.

  • Each scenario runs from a Nix wrapper bound to the same source path and flake.lock; a mismatch fails the scenario.
  • The C5 suite last passed in full on b8399ad; four scheduler and compactor changes have merged since with focused checks only.

crates/harnesses/src/suite.rs · crates/harnesses/src/suite/scenarios.rs · nix/checkpoints.nix · nix/run-checkpoints.sh · docs/checkpoint-c4.md · docs/hosted-checkpoints.md

XFS + shared fixturesreflink proof · native tests in a guest · two guests

The XFS fixture boots a KVM guest with a real XFS disk, proves reflink and FIEMAP behaviour, and runs the native cas-core and cas-daemon test inventories there. The shared fixture nests two TCG guests with ext4 and SQLite over one cas-host, optionally SIGKILLs the host at a compaction cut and replaces it retained.

crates/harnesses/src/fixture.rs · crates/harnesses/src/fixture/checks.rs · crates/harnesses/src/shared.rs · crates/harnesses/src/shared/restart.rs · crates/harnesses/src/filesystem.rs · nix/fixture/workload.sh · nix/shared/outer.nix · docs/xfs-fixture.md · docs/shared-guest-fixture.md · docs/compaction-crash-cuts.md

pressure harness11 stages · file protocol · telemetry

Two guests and a controller exchange ready, request and completed files per stage: shared and disjoint reads, hot-set scans, displacement, four fio write shapes, calibration of the compactor's drain rate and a burst above it. It requires a sample with staging stopped, fresh-boot digests that match, and pool limits that equal the code defaults.

crates/harnesses/src/shared/pressure.rs · crates/harnesses/src/shared/pressure/checks.rs · crates/harnesses/src/pressure/guest.rs · docs/guest-pressure.md · docs/pressure-diagnosis.md

persistence oracle841 crash schedules · no VM

A deterministic model of the WAL crash contract. It builds a small log, records every complete-batch prefix as an oracle image, then for each of 841 schedules persists a subset of the unsynced tail sectors (truncations, reversals, torn sectors, header-only, payload-only, shuffles), cold-opens the copy and requires the recovered prefix to be at least the required one and to read exactly like its oracle.

crates/harnesses/src/persistence.rs · crates/harnesses/src/persistence/schedule.rs · crates/harnesses/src/persistence/oracle.rs · crates/harnesses/src/persistence/verify.rs · docs/persistence-model.md

census fleet + preflightcloud images · T0/T1/T2 · host inventory

The measurement behind Update 01: boots dated Ubuntu cloud images and clones through two upgrade epochs, normalizes their roots and runs the census. preflight records 14 host probes and check-disks proves the OS and data disks are distinct whole devices.

crates/harnesses/src/fleet.rs · crates/harnesses/src/host.rs · experiments/update-guest.sh · experiments/normalize-root.sh · nix/census.nix · docs/census.md

Nixflake · guests · fixtures · host modules

The flake builds the workspace with the pinned toolchain, one NixOS guest per backend, the fixture and lab VMs, the checkpoint bundle and the census tools, and exports host modules for dedicated test machines. Every wrapper carries a build record so the harness can bind evidence to a source revision.

flake.nix · nix/package.nix · nix/smoke.nix · nix/checkpoints.nix · nix/census.nix · nix/lab/default.nix · nix/fixture/default.nix · docs/testbed.md

NixOS guestssmoke · fixture · lab · shared

Prerendered guest systems: an ephemeral tmpfs root, the pinned fio jobs under /etc/cas, a oneshot that runs the workload and a finish hook that writes completion.json and powers off. The QEMU command line comes from the pinned qemu-vm module plus the per-backend disk device.

  • Raw: -drive cache=none,aio=io_uring with virtio-blk-pci and 4 KiB logical blocks.
  • Daemon and host: -chardev socket plus vhost-user-blk-pci with num-queues=1,queue-size=128 or 4×256, and shared memory enabled so guest RAM is mappable.
  • Results travel through a 9p share; SSH keys and host keys are exchanged over the same share.

nix/guest/default.nix · nix/guest/smoke.sh · nix/guest/finish.sh · nix/shared/guest.nix · nix/fixture/guest.nix · nix/lab/guest.sh · crates/harnesses/fio/smoke.fio

Test-host modulesdisko · XFS at /srv/cas-testbed

NixOS modules for the dedicated hosts G1 still needs: key-only SSH, performance governor, an ESP plus ext4 root on the OS disk and one XFS partition on the data disk, installed with nixos-anywhere from the template. The two hosts do not yet exist.

nix/modules/test-host.nix · nix/modules/disks.nix · nix/modules/bare-metal.nix · templates/test-host/flake.nix · nix/checks/host-config.nix · docs/testbed.md

CIGitHub mirror only

Three workflows on the GitHub mirror; Forgejo Actions stay off. implementation runs fmt, clippy and tests; nix formats, evaluates both architectures separately, builds every check and runs the smoke, recovery and live-recovery VMs plus the dev-vm SSH script; a manual dispatch runs the full C5 suite. pages would publish the site but is disabled.

.github/workflows/implementation.yml · .github/workflows/nix.yml · .github/workflows/pages.yml · docs/ci-evaluation.md · docs/hosted-checkpoints.md

Evidence chainresults/ → docs/validation → JSON → site

How a number reaches a page. Raw artifacts stay under ignored results/ on Spark or in an archive; each session appends to docs/validation.md and larger ones get their own record; measurements keep a README, the analysis script and its JSON; the site imports that JSON directly, so no number on a page is typed by hand.

  • TODO.md is the canonical tracker; an item is checked only when its acceptance record exists.
  • Retention: keep small evidence and failed attempts, delete VM disks after independent verification, never launch below 25 GiB free.

TODO.md · AGENTS.md · docs/validation.md · docs/artifact-retention.md · docs/measurements/integration-2026-09-15/display.py · playbook/src/lib/articles.ts · docs/validation.md · docs/artifact-retention.md

What the diagram does not show

  • Structure is not behaviour. A box and an arrow say what exists and who calls whom at 7a555f3. Whether a path works under load, crash or exhaustion is established only by the records in the validation history; the parts cite the notes that describe their acceptance, not proof of it.
  • Some notes are behind the code. The five-second admission deadline described in several design notes no longer exists; admission waits retry every 100 ms and never become IOERR. The manifest reclamation note describes bitmap windows the code replaced with one sorted live-offset list. The vhost-user note predates four queues, ZERO/DISCARD and retained recovery. The parts say so where it matters.
  • Nothing distributed exists yet. Remote reads, replication, ownership transfer and migration, the whole right-hand side of the study, have no box because they have no code. The census and the persistence oracle are the only measurement tools that touch the research questions directly.
  • Two engines, one drawing. The v1 staging log and the raw io_uring backend remain as controls for the checkpoint suite and are drawn faded; the diagram’s paths describe the local runtime that the shared host uses.

Revision and checks

How it was made. Five independent readers went through every first-party source file at 7a555f3 and reported each module’s types, invariants, constants and edges. Those reports became the 104 parts and 132 connections in the diagram. Where a design note and the code disagreed, the code won and the note is named in the part’s text. No experiment ran for this page.

Every path on this page links to 7a555f3, the current main after Update 03. The graph names 181 source files and 76 design notes, each checked to exist at that commit. The page ran pnpm check, pnpm build, a static check that every fragment link and pinned path resolves, and a browser check of the diagram’s expansion, selection and operation traces on Spark. Session record · Graph definition · Progress tracker.

The layout engine is ELK’s layered algorithm, the same one Mermaid uses for subgraph diagrams, loaded in the browser only when a view changes; the starting view is laid out at build time so the page is complete without JavaScript.