← index Update 02 · 14 September 2026

Private disks, shared bytes

CAS gives each VM a private writable disk while sharing identical stored blocks across images. This update explains the working single-host backend: how writes become durable, how reads find their bytes, and how background work recovers space.

The study asks whether sharing blocks by content can reduce storage and data movement across hosts enough to pay for indexing, copying and coordination. This backend lets us first measure those costs on one host. Remote reads, replication and migration remain to be built.

Latest finding: the merged scheduler lets independent reads pass blocked writes, and live recovery checks pass. Same-queue p99 is lower in this repeat, but multi-second outliers remain. Admission fairness is the next issue to resolve. Measurements and design tradeoffs ↓

Each image owns its writes. One host shares the chunks.

QEMU owns the virtual machine. cas-host owns its disk backend. Inside each guest, Linux turns file operations into block requests. CAS sees disk offsets and bytes; ext4 directories and SQLite transactions stay inside the guest.

CAS-HOST PROCESS · CAS-DAEMON CRATE guest RAM + eventfdguest RAM + eventfdchannel + eventfdchannel + eventfdin-process access QEMU + guest Aprivate filesystem / virtqueuesQEMU + guest Bprivate filesystem / virtqueuesImage A frontendsocket / parser / completionImage B frontendsocket / parser / completionImage A reactorprivate WAL + manifest viewImage B reactorprivate WAL + manifest viewCompactor threadchunks / manifests / catalogShared host resourcesindex / caches / budgets / gate
  1. QEMU + guest A private filesystem / virtqueues

    → Image A frontend · guest RAM + eventfd

  2. QEMU + guest B private filesystem / virtqueues

    → Image B frontend · guest RAM + eventfd

  3. Image A frontend socket / parser / completion

    → Image A reactor · channel + eventfd

  4. Image B frontend socket / parser / completion

    → Image B reactor · channel + eventfd

  5. Image A reactor private WAL + manifest view

    → Compactor thread

    → Shared host resources

  6. Image B reactor private WAL + manifest view

    → Compactor thread

    → Shared host resources

  7. Compactor thread chunks / manifests / catalog

    → Shared host resources · in-process access

  8. Shared host resources index / caches / budgets / gate

QEMU sets up each disk over a Unix socket. Requests use shared guest RAM and eventfds. Inside cas-host, each image keeps its own write order; the shared compactor publishes chunks and manifests.

Diagram markup · Mermaid
flowchart TB
  subgraph process["cas-host process · cas-daemon crate"]
    a["Image A frontend — socket / parser / completion"]
    b["Image B frontend — socket / parser / completion"]
    ra["Image A reactor — private WAL + manifest view"]
    rb["Image B reactor — private WAL + manifest view"]
    shared["Shared host resources — index / caches / budgets / gate"]
    worker["Compactor thread — chunks / manifests / catalog"]
  end
  ga["QEMU + guest A — private filesystem / virtqueues"]
  gb["QEMU + guest B — private filesystem / virtqueues"]
  ga -->|"guest RAM + eventfd"| a
  gb -->|"guest RAM + eventfd"| b
  a -->|"channel + eventfd"| ra
  b -->|"channel + eventfd"| rb
  ra --> worker
  rb --> worker
  ra --> shared
  rb --> shared
  worker -->|"in-process access"| shared

An image is one guest’s disk. Its write-ahead log (WAL) holds recent writes; its manifest maps settled blocks to content hashes. The shared chunk store holds the bytes named by those hashes.

The frontend accepts guest requests and returns completions. The image’s reactor drives IO in one write order. One shared compactor turns durable WAL data into chunks and updates manifests. These run in cas-host; the storage formats and indexes live in cas-core.

Processes, threads, crates and entry points
One host kernel
├─ QEMU A → guest Linux A → private filesystem
├─ QEMU B → guest Linux B → private filesystem
└─ cas-host                         one process, cas-daemon crate
   ├─ socket / queue worker × image  parses virtio; owns guest completion
   ├─ image reactor × image         WAL / read IO; one write order
   └─ cas-compactor                  shared chunk writer; manifests; GC
      shared: hash index, read caches, budgets, catalog

The guests have separate kernels and page caches.
The reactors and compactor are threads in the same Rust process.
Crate / directoryResponsibility
cas-daemonBoth cas-host and the older cas-daemon executable. Guest protocol, ordering, reactors, shared worker and recovery coordination.
cas-coreWAL, chunks, indexes, manifests, catalog, file IO, resource accounting and scheduling primitives.
cas-cliThe casctl CLI: VM commands, SSH, measurements and census. It calls the harness; it is outside the block IO path.
cas-harnessLaunches experiments, checks workloads, records source/artifact identity and verifies suites.
crates/vendorPinned rust-vmm extensions for inflight FDs, lifecycle hooks and saved-head descriptor walking. QEMU remains stock.
nix / experimentsBuild and launch reproducible host/guest fixtures. These are not extra storage servers.

Runtime coordination uses ordinary threads, bounded channels, mutexes and eventfds. The harness controls child process groups. The repository does not install per-image cgroup CPU, memory or IO limits.

The transport changes at each boundary

QEMU sets up the disk through a Unix vhost-user socket and passes file descriptors. Requests and payload live in shared guest RAM; kick/call eventfds announce work and completion. Inside CAS, channels move owned Rust values. Foreground disk IO uses io_uring; the compactor uses blocking file IO on its own thread.

Trace an operation across every boundary
  1. Guest Linux → Backendshared guest RAM + kick eventfd

    Retain validated descriptor spans. Reserve capacity before assigning a mutation sequence; waiting write payload stays in guest RAM.

  2. Backend → image reactorbounded Rust channel + eventfd

    Gather guest bytes into the final aligned WAL buffer. Move its ownership to the reactor.

  3. Image reactor → Linux / XFSio_uring WRITE on an O_DIRECT file

    Append the packed batch. Keep its buffer and file alive until the kernel completes.

  4. Image reactor → guestcompletion channel → used ring + call eventfd

    Publish the contiguous write prefix and staging mapping, then return the guest completion.

The block becomes readable after ordered publication. A FLUSH is still needed for durability.

Guest RAM, inflight memory and io_uring are different mappings
Guest driverposts head; reads statusShared guest RAMdescriptors / avail / used / dataCAS frontendmaps same backing pagesQEMUretains the carrier FDInflight memfdidentities + P; no payloadReplacement CASreconcile + replayCAS reactorhost FDs + aligned buffersHost io_uringSQEs / CQEs; no guest ringsHost kernelexecutes storage IO
  1. Guest driver posts head; reads status

    → Shared guest RAM

  2. Shared guest RAM descriptors / avail / used / data

    → CAS frontend

  3. CAS frontend maps same backing pages
  4. QEMU retains the carrier FD

    → Inflight memfd

  5. Inflight memfd identities + P; no payload

    → Replacement CAS

  6. Replacement CAS reconcile + replay
  7. CAS reactor host FDs + aligned buffers

    → Host io_uring

  8. Host io_uring SQEs / CQEs; no guest rings

    → Host kernel

  9. Host kernel executes storage IO

The guest-RAM mapping, retained recovery carrier and host io_uring mappings have different contents and lifetimes. None forwards a guest syscall.

Diagram markup · Mermaid
flowchart TB
  guest["Guest driver — posts head; reads status"]
  ram["Shared guest RAM — descriptors / avail / used / data"]
  backend["CAS frontend — maps same backing pages"]
  qemu["QEMU — retains the carrier FD"]
  carrier["Inflight memfd — identities + P; no payload"]
  replacement["Replacement CAS — reconcile + replay"]
  reactor["CAS reactor — host FDs + aligned buffers"]
  uring["Host io_uring — SQEs / CQEs; no guest rings"]
  kernel["Host kernel — executes storage IO"]
  guest --> ram
  ram --> backend
  qemu --> carrier
  carrier --> replacement
  reactor --> uring
  uring --> kernel

QEMU shares the guest-memory backing with CAS. GuestMemoryMmap gives the adapter checked access; another mapping does not copy all guest RAM. A KVM memory slot describes guest-physical address backing. It carries no block operation.

The separate inflight memfd holds request identities and the published prefix. QEMU retains that FD across a backend crash. It contains no write payload. Host io_uring submission/completion rings are a third mapping, shared with the host kernel.

Guest page caches remain private. A CAS cache hit still copies bytes into the requesting guest. The local runtime has no remote chunk transport or replication protocol.

Protocol reference: QEMU vhost-user.

WRITE makes bytes visible. FLUSH makes them durable.

The frontend reserves capacity before assigning a mutation number. It gathers the guest’s bytes into the final aligned append buffer, then sends that buffer to the reactor. The reactor appends to the image’s write-ahead log (WAL). Completed appends publish in order; a slow earlier write holds back later publication.

A FLUSH captures a fixed write boundary, writes a fence and syncs it. Later write submissions wait outside that cohort, so a continuous writer cannot make one FLUSH chase an ever-growing stream. Reads and other images can continue during this ordinary sync.

P published to the live viewE synced in the WALD committed in the manifest

During normal serving, D ≤ E ≤ P. A guest WRITE can complete at P. Its FLUSH waits for E. Compaction later moves D forward.

Follow one block through overwrite and compaction

P · published0
E · synced WAL0
D · manifest0

Start from an empty image. The counters below are mutation sequences, not byte offsets. This is an illustrative execution, not a recorded timing trace.

WAL layout, alignment and the actual copying cost

A v2 batch contains one 4 KiB header and up to 1 MiB of payload. The header holds up to 63 operation descriptors; a FENCE occupies its own 4 KiB header. Thus one isolated 4 KiB WRITE plus FLUSH encodes 12 KiB; 32 packed 4 KiB WRITEs plus FLUSH encode 136 KiB. This is format arithmetic, excluding segment creation, filesystem metadata and compaction.

Batching uses available work: it seals at capacity, FLUSH or the end of the queue-drain pass. There is no batching timer holding a lone write. CRC32 protects framing and payload. Recovery follows validated boundaries rather than scanning arbitrary payload for fence-like bytes.

The guest sector field counts 512-byte sectors: sector 8 means byte 4096. Accepted payload IO must cover whole aligned 4 KiB blocks, up to 1 MiB per request. Unsupported geometry is rejected. Multi-block requests have no whole-request atomicity guarantee.

O_DIRECT bypasses the ordinary host payload page cache. CAS still uses XFS, the host block layer and device driver. It checks STATX_DIOALIGN; it does not silently fall back to buffered IO. fdatasync supplies the durability boundary; namespace publication also syncs directories.

“One gather” describes guest-to-WAL-buffer copying. Compaction later reads and copies surviving bytes into chunk output. Short IO fails its owner; it is not retried as an arbitrary suffix. The 50 ms idle-sync setting is a policy trigger, not a hard durability deadline.

Linux contracts: open(2), fsync(2). Full byte format.

Recent writes take precedence over shared chunks

The image’s staging index first supplies its newest published writes or ZERO ranges. Only uncovered blocks consult the manifest and shared chunk index below. A read pins its chosen files and manifest root, so compaction cannot invalidate its sources halfway through.

Image manifestlogical block → content hash
Shared chunk indexcontent hash → segment + offset

Manifest content uses a shared clean chunk cache. On a miss, one reader fetches and verifies the hash; concurrent requests for that hash join its result. A separate manifest-page cache avoids repeated tree-page IO. Optional cache admission can fail while the verified read still succeeds.

Read IO, cache ownership and cold-read limitations
covereduncoveredhashhitmiss READ logical rangewait for captured boundaryStaging interval indexnewest published mutationsPinned WAL or ZEROCRC-checked data / zerosCaptured manifest ViewB+tree pages → hash / holeManifest page cacheor verified direct page IOShared chunk LRUfull 32-byte hash keyOwned responsecopy back to guest RAMCoalesced hash fetchone leader; bounded waitersHash index → segment + blockverify header, CRC and BLAKE3
  1. READ logical range wait for captured boundary

    → Staging interval index

  2. Staging interval index newest published mutations

    → Pinned WAL or ZERO · covered

    → Captured manifest View · uncovered

  3. Pinned WAL or ZERO CRC-checked data / zeros

    → Owned response

  4. Captured manifest View B+tree pages → hash / hole

    → Shared chunk LRU · hash

  5. Manifest page cache or verified direct page IO

    → Captured manifest View

  6. Shared chunk LRU full 32-byte hash key

    → Owned response · hit

    → Coalesced hash fetch · miss

  7. Owned response copy back to guest RAM
  8. Coalesced hash fetch one leader; bounded waiters

    → Hash index → segment + block

  9. Hash index → segment + block verify header, CRC and BLAKE3

    → Owned response

Staging has precedence, including an explicit ZERO. Only uncovered manifest hashes reach the shared cache and chunk store.

Diagram markup · Mermaid
flowchart TB
  read["READ logical range — wait for captured boundary"]
  staging["Staging interval index — newest published mutations"]
  payload["Pinned WAL or ZERO — CRC-checked data / zeros"]
  view["Captured manifest View — B+tree pages → hash / hole"]
  pages["Manifest page cache — or verified direct page IO"]
  cache["Shared chunk LRU — full 32-byte hash key"]
  response["Owned response — copy back to guest RAM"]
  fetch["Coalesced hash fetch — one leader; bounded waiters"]
  disk["Hash index → segment + block — verify header, CRC and BLAKE3"]
  read --> staging
  staging -->|"covered"| payload
  staging -->|"uncovered"| view
  pages --> view
  view -->|"hash"| cache
  cache -->|"hit"| response
  payload --> response
  cache -->|"miss"| fetch
  fetch --> disk
  disk --> response

WAL reads verify the original write CRC. A partial read may need scratch for that entire original payload. On the manifest path, the current reactor walks one block at a time: uncached page lookup, chunk header, then chunk payload. Multiple requests can overlap, but a single cold 1 MiB read does not submit all 256 chunk fetches at once.

The clean cache defaults to 256 MiB of payload and uses ordinary LRU. It has no per-image cache partition or scan-resistant policy. Only verified read fills populate it; writes and compaction do not prewarm it. The 16 MiB manifest-page payload cache consumes part of the 128 MiB foreground metadata budget.

Coalesced misses use bounded leader/waiter cells and eventfds; a reactor waits with io_uring PollAdd. The fetched buffer retains its original read credits through the final reader. Cache insertion makes a separate charged copy. Eviction drops cache membership, while an existing reader can retain the allocation.

Page keys bind a registry incarnation, committed end and offset. A new file/root cannot inherit stale cache meaning. Guest RAM, shared mappings, retained fetches and cache payload are different owners; summing every report repeats some of the same memory.

The on-disk files and their durable owners
<store-root>/                      one dedicated filesystem
├─ catalog/catalog.v2               active images and snapshot roots
├─ chunks/segment-<ticket>.v2       shared immutable chunk batches
├─ images/<image-id>/
│  ├─ manifest.v2                   logical blocks → content hashes
│  └─ staging/segment-<ticket>.v2   private write-ahead log (WAL)
└─ snapshots/<snapshot-id>/
   └─ manifest.v2                   immutable reflink of an exact root

Outside this filesystem: vhost-user sockets and reports.
Loose pending files are not catalog membership.

The manifest is an append-only copy-on-write B+tree of checksummed 4 KiB pages. Leaves contain extents; a COMMIT binds identity, root, geometry and D. Manifest owns mutation; View pins one committed root. It is not a content-hashed Merkle tree.

Store owns chunk output. Its shared Reader resolves immutable reads. The in-memory hash index is rebuilt from verified chunk records on recovery. Each address packs a 48-bit segment ticket and 16-bit block index. Moving a chunk changes that index, so GC need not rewrite every manifest.

Catalog defines complete image/snapshot membership. Publication syncs dependencies, writes and syncs a temporary catalog, renames it, then syncs its directory. Tickets allocates identities across chunk files, WALs and rejected names; reclamation preserves the highest durable identity to prevent reuse.

The compactor converts settled writes into shared content

The compactor takes a bounded, durable WAL prefix above D. It drops versions fully covered by later edits in that selection, verifies the input and hashes each surviving nonzero 4 KiB block with BLAKE3. Existing hashes reuse stored chunks; missing hashes get written and synced.

data is durable first Durable WAL prefixD < sequence ≤ ESurviving 4 KiB blocksverify CRC → BLAKE3Missing chunks → syncreuse existing hashesCOW pages + COMMITsync manifest filePublish new View + Dunder completion gateReclaim stagingwait for read / identity pins
  1. Durable WAL prefix D < sequence ≤ E

    → Surviving 4 KiB blocks

  2. Surviving 4 KiB blocks verify CRC → BLAKE3

    → Missing chunks → sync

  3. Missing chunks → sync reuse existing hashes

    → COW pages + COMMIT · data is durable first

  4. COW pages + COMMIT sync manifest file

    → Publish new View + D

  5. Publish new View + D under completion gate

    → Reclaim staging

  6. Reclaim staging wait for read / identity pins

A bounded, durable prefix is copied into the shared content store. Reclamation follows the durable manifest and the retirement of readers and replay identities.

Diagram markup · Mermaid
flowchart TB
  wal["Durable WAL prefix — D < sequence ≤ E"]
  hash["Surviving 4 KiB blocks — verify CRC → BLAKE3"]
  store["Missing chunks → sync — reuse existing hashes"]
  manifest["COW pages + COMMIT — sync manifest file"]
  publish["Publish new View + D — under completion gate"]
  reclaim["Reclaim staging — wait for read / identity pins"]
  wal --> hash
  hash --> store
  store -->|"data is durable first"| manifest
  manifest --> publish
  publish --> reclaim

Next it commits the image’s new manifest. Only then may the reactor install that committed view, advance D and reclaim covered WAL space once readers and recovery no longer need it. This order prevents a durable mapping from naming unsynced content. A newer guest overwrite remains in staging and keeps precedence.

Identical blocks share even when their images have no common ancestor. A one-byte edit creates a new 4 KiB chunk. Duplicate writes still consume private WAL space first; deduplication does not remove foreground admission costs.

Selection limits, scheduling and why this is fixed chunking

A validated segment/offset cursor resumes selection above the published D instead of rereading old framing. Recovery reconstructs that hint; a missing retained segment falls back to scanning. Manifest preparation grows as pages are needed, then emits only nodes reachable from the final root. Draft edits still update paths sequentially.

One selection covers whole batches with at most 1 MiB of payload and 318 mapping edits. It prepares the complete copy-on-write edit before chunk publication. Partial ZERO overlaps remain ordered; splitting them could exceed the reserved edit bound. Chunk output batches contain at most 63 blocks.

The worker uses typed selection, publication and reclamation replies to each reactor. The host gate precedes the image gate; those locks protect state publication and failure, not disk IO. A timed-out caller cannot refund the running worker’s buffers, files or physical-space promises.

Work prefers 100 ms of settling; a 1 s age trigger and staging/index pressure force progress attempts. A continuously writing guest can trigger a finite fence to supply durable input. These timers do not bound completion when storage is slow.

BLAKE3 chooses content identity. Boundaries are fixed logical blocks; content-defined chunking (CDC) remains unimplemented. All-zero content becomes a hole. The index compares full hashes, without byte-comparing incoming duplicates against disk. Compression, encryption and indexing payload in its original WAL location are absent from this path.

Freeing space takes more than deleting a mapping

There are three kinds of space to recover. WAL reclamation frees old private write payload after compaction. Chunk GC traces content still reachable through images, snapshots and readers. Manifest sweeping frees obsolete mapping pages.

StorageHow space is released
Private WALPunch eligible payload; unlink retired segments after readers and replay identities release them.
Shared chunksMark live content, copy live chunks out of selected segments, sync the copies, then delete the old segments.
Manifest pagesMark pages reachable from retained roots, then punch the gaps.

Ordinary compaction runs alongside guest IO. Host GC pauses new requests across all images. The shared worker drains admitted requests and fences the images before collection. Its file reads, writes, syncs, punches and unlinks run on that worker thread.

Snapshots use the same host pause. The target is compacted through its fenced cut, its exact manifest is reflinked, and catalog membership is made durable before success. A core clone starts with that mapping but gets its own image identity, WAL and sequence namespace.

This is a storage snapshot. It does not quiesce guest applications or make several images one atomic application transaction. Snapshot/clone primitives exist in Rust; the executable currently exposes initialization and serving, without a general live create/clone/delete/snapshot control API.

Reclamation limits and ZERO / DISCARD

Compaction cannot free every covered WAL allocation immediately. Readers pin original payload, retained replay needs operation identity, and the current/highest segment may need to keep its header. The governor accounts allocation still held after punching or unlinking.

Manifest sweeping walks each retained root once, collects page offsets under the metadata budget, then sorts and deduplicates them before punching gaps. Marking finishes before any page is punched. Memory scales with retained tree visits; allocation denial can still stop collection. The operation remains synchronous, and the current COMMIT keeps the file’s logical end high.

Snapshots retain reachable chunks and may share manifest extents through XFS reflinks. Neither logical deletion nor a requested punch count equals measured physical savings.

Negotiated WRITE_ZEROES and DISCARD both produce ZERO mutations, limited to one aligned range of at most 1 MiB. They read as zeros immediately after publication; physical release waits for later reclamation. Empty ranges consume no mutation number.

A full WAL makes new writes wait

Before admission, a waiting write stays in guest memory. CAS retains its descriptor without assigning a mutation sequence or copying its payload. After admission, a bounded host buffer carries it to the on-disk WAL; compaction drains that WAL in the background. The disk holds accepted writes, not an unlimited overflow queue.

A writer can fill its WAL faster than the compactor drains it. Healthy capacity waits stay pending until space returns. Reclaimed capacity wakes admission; a retry timer covers releases without a notification. Independent reads can now pass a blocked write on the same virtqueue. Overlapping ranges and FLUSH barriers still constrain admission. The latest experiment shows that distinction.

Which limit stops which work
PressureResponseWhat still progresses
Request / byte creditsWait before admission; final release wakes waiting images.Other eligible images and the separate control reserve.
WAL / staging indexFence a finite prefix, compact/reclaim, rotate through the worker.Eligible reads and covered FLUSHes; queue ordering can still block later requests.
Physical spacePreserve the progress reserve; collect the host.Already admitted owners drain. New guest IO pauses during collection.
Live data cannot shrinkReport exhausted capacity; keep physical write admission closed.Reads after a successful collection releases its host pause.
Chunk-index memoryNew unique chunks can fail compaction and its image.No index spill or automatic memory-pressure GC fallback exists.
IO timeout / integrity errorClose the affected failure gate; preserve uncertain owners.An activated image socket failure is isolated. Shared corruption and incomplete retained recovery still stop all endpoints.

Staging starts pressure compaction at 75% of its allocation-plus-promise cap, stops admission at the cap, and resumes below 60%. Physical space has a separate stop boundary that preserves the background reserve. GC must observe real filesystem release; deleting a name is insufficient.

Admission keeps FIFO order among eligible heads within an image and uses byte deficit round robin across images, with a 1 MiB quantum. A second scheduler gives compaction one in four ready bulk submission opportunities; either side can borrow unused turns. This controls submission opportunities; it guarantees neither device bandwidth nor p99 latency.

Budgets, ownership and deadlines in concrete terms
Default limitPer imageShared host
Write requests1281,024
Read requests864
Control reserve8 / 64 KiB32 / 256 KiB
Append bytes8 MiB64 MiB
Read/fetch bytes8 MiB64 MiB
MetadataShares host budgets128 MiB foreground + 128 MiB compaction
Staging allocation256 MiB1 GiB

These are code defaults, not a measured RSS total. Budget accounts capacity; Share imposes an image sublimit; leases follow owners. Budgeted collections account table capacity and old/new growth overlap. A read reserves response bytes plus 1 MiB of scratch allowance: the default byte budget admits seven simultaneous 4 KiB reads. Each frontend also reserves 6,297,600 bytes of metadata quota for descriptor snapshots; actual allocation depends on the discovered descriptors.

The default staging geometry permits at most nine catalog images: their 64 MiB current segments must total strictly less than 60% of the 1 GiB host staging cap. Increasing that cap changes this startup constraint; it does not establish tested concurrency at the larger image count.

The physical reserve is R = 3S + M + 16 MiB: 336 MiB with 64 MiB segments and a 128 MiB manifest transaction cap. One background owner may use it. The governed filesystem excludes unrelated writers; reports and sockets must live elsewhere.

Queued bulk IO keeps its bounded buffer until a submission turn is available. Device IO retains its 30 s deadline; a shared-fetch eventfd poll waits for its leader instead. Explicit terminal drain has a separate 30 s deadline and cancels dependency polls without dropping uncertain owners. Recovery bounds waiting at 60 s. These values are not hard bounds on every cleanup or background syscall.

Host collection pauses all admission, including controls from guests. Internally reserved fence/control work drains what was already accepted. Independent reads can pass a WAL-blocked write on the same or another virtqueue. Reads still wait for earlier overlapping work, FLUSH barriers and publication dependencies after admission. A control reserve does not bypass ordering dependencies.

Recovery depends on which machine state survived

EventRecovery contract
CAS dies; QEMU livesRetain guest RAM, rings and inflight FDs. Validate saved P and every image, replay missing owned mutations with original identities, then fence/sync before serving.
QEMU is gone / fresh bootCold recovery uses validated catalog, manifests, chunks and WAL. Complete unsynced writes may survive; only synced durability is promised.
Device reset / memory replacementDrain old queue, memory and IO owners before accepting replacement state.
Physical host power lossOnly stable storage survives. The sync contract is the model; this has not been established by a physical power-cut test.

Retained recovery is coordinated across the catalog. One missing or incompatible image can prevent the whole shared host from resuming. This is a consequence of repairing a shared dependency graph under one owner.

Replay identity, rejected tails and pinned library changes
shared barrier QEMU A + QEMU B surviveguest RAM + carrier FDs remainOld cas-host killedold IO owners must releaseNew cas-host: retainedlock + inspect all imagesValidate every carrierprefix P + original identityReplay missing mutationsoriginal per-image sequencesRecovery FENCE + syncfor every catalog imageRestore completionsthen enable ordinary admission
  1. QEMU A + QEMU B survive guest RAM + carrier FDs remain

    → Validate every carrier

  2. Old cas-host killed old IO owners must release

    → New cas-host: retained

  3. New cas-host: retained lock + inspect all images

    → Validate every carrier

  4. Validate every carrier prefix P + original identity

    → Replay missing mutations

  5. Replay missing mutations original per-image sequences

    → Recovery FENCE + sync

  6. Recovery FENCE + sync for every catalog image

    → Restore completions · shared barrier

  7. Restore completions then enable ordinary admission

Every image contributes its original memory, queue state and inflight carrier before the shared recovery barrier can open.

Diagram markup · Mermaid
flowchart TB
  qemu["QEMU A + QEMU B survive — guest RAM + carrier FDs remain"]
  old["Old cas-host killed — old IO owners must release"]
  new["New cas-host: retained — lock + inspect all images"]
  validate["Validate every carrier — prefix P + original identity"]
  replay["Replay missing mutations — original per-image sequences"]
  fence["Recovery FENCE + sync — for every catalog image"]
  resume["Restore completions — then enable ordinary admission"]
  old --> new
  qemu -.-> validate
  new --> validate
  validate --> replay
  replay --> fence
  fence -->|"shared barrier"| resume

The replacement acquires real storage locks; PID death does not establish that old kernel IO released its files. It inspects every required dependency before repair. Published-but-unsynced mutations must already be covered by the valid WAL/manifest prefix; completed guest buffers may have been reused.

The version-3 inflight carrier distinguishes discovered descriptors from admitted mutations. Discovered requests return to admission; they are not replayed as writes. Upgrading a version-2 retained attachment requires stopping its guest and making a fresh attachment. Still-owned missing mutations can be gathered from surviving guest RAM. Replay uses saved descriptor heads, original serials and mutation sequences. An old available-ring slot may already have wrapped, so replay cannot infer identity from that slot alone.

All recovery fences finish before the common serving barrier opens. Cold recovery archives rejected suffixes before repair and establishes a new durable writer epoch. Corruption of required data is an error, not permission to silently roll back to an older convenient prefix.

The vendored backend exposes inflight FD messages and brackets queue/memory lifecycle changes. The queue patch exposes descriptor walking from a saved head. These extensions are required by recovery; removing them as wrapper code would remove behavior tested by C3–C5.

Read bypass helps, but the long tail remains

The merged read scheduler has now run through live QEMU recovery/reset checks and the same mixed workload. The results separate working bypass from the remaining scheduling delay. The study’s benefit over ZFS or across hosts remains unmeasured.

Reads can bypass writes, but admission fairness still stalls them. The longest untraced read still took 6.83 s.

WorkloadEarlier read p99Merged read p99Earlier slowestMerged slowest
Read only2.57 ms–2.61 ms2.24 ms–2.28 ms56.91 ms10.34 ms
Reader + writer on CPU 092.80 ms–2.67 s24.51 ms–156.24 ms9.98 s6.83 s
Reader CPU 1, writer CPU 04.82 ms–6.39 ms5.60 ms–6.39 ms64.77 ms3.06 s

Untraced fio total latency. Each p99 range covers two guests; mixed cases have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows. Matched workloads ran in separate Spark sessions using nested TCG guests; these are development measurements.

The admission rule

Published firstWrite AWaiting for WAL space
Published laterRead BEligible for read admission

Read B can use its own read credits while Write A waits for WAL capacity. The write payload remains in guest RAM.

A and B stand for different disk ranges. CAS owns bounded descriptor snapshots before admission; it does not need a new guest queue, a lock-free queue or a new executor. After admission, the existing image-wide write-publication dependency still applies.

What the instrumentation says now

One matched 6.295-second read spent 6.293 seconds before admission and only 0.57 ms afterward. Its chunk was cached. The fairness scheduler kept refusing its turn; a separate Rust reproduction confirms that two blocked writers can repeatedly displace eligible reads without any disk IO.

Each image retries its oldest write before its eligible read. That retry marks the write ready again, which can displace the read the shared scheduler just selected. Two images can keep waking each other and repeating this until write capacity returns.

Next: preserve eligible-read selection across blocked-write retries, then repeat this workload. Overlap, FLUSH and finite-capacity rules still matter; this result does not justify a new cache or a lock-free queue.

Validation, resource costs and code

The packaged merge passed 499 native tests. All 13 live recovery/reset scenarios and 8 full 64 MiB seed checks passed. The recovery cases kept QEMU alive across backend replacement; they did not reproduce an overloaded deferred-read crash. Native frontend tests cover that ownership case.

The largest sampled lab memory peak was 3.93 GiB under a 6 GiB cap, with zero observed cgroup OOM or limit events. Disposable disks and keys were removed. Actual OOM/ENOSPC, sustained fairness, full guest-ring exhaustion, long-history GC and physical power loss remain separate gates.

Each frontend reserves about 6 MiB of descriptor metadata quota, with eight read-request slots separate from writes. The existing byte allowance fits seven simultaneous 4 KiB reads. Version-3 retained carriers require a fresh guest attachment when upgrading from version 2.

All four owners are in crates/cas/daemon. Raw-result analysis, per-job numbers and limitations.

Earlier: tracing the original queue-head stall

The original long read wait was at admission. One matched read took 7.938 seconds in fio: 7.933 seconds passed before it reached the head of its CAS queue, then 3.53 ms after admission. The reader and the blocked write touched different disk ranges.

1Guest
fio → Linux bio
2Virtqueue
shared descriptors
3CAS admission
wait behind a write
4Reactor → storage
cache, WAL or chunks
5Guest completion
used ring → fio

In these measurements, the frontend released the queue mutex while waiting, but stopped consuming the queue at its blocked write. Later reads could not reach their own read credits. A lock-free queue would preserve that obstruction unless its admission rule also changed.

fio total7937.964 ms
Before CAS queue head7932.921 ms
After CAS admission3.527 ms

Amber: before the CAS queue head. Blue: the rest of fio latency. Queue 0, disk offset 19038208. Selected slow IO, not a percentile sample; guest and CAS clocks are joined by request identity and only durations are compared.

Where the remaining time went
Guest bio submission → virtqueue publication
0.033 ms
fio time outside bio submission → block completion
1.054 ms
Virtqueue/completion time outside the CAS interval
0.428 ms
At the CAS queue head
0.001 ms
Frontend dispatch + command channel
0.018 ms
Reactor dispatch / publication dependency
2.322 ms
Read execution
1.165 ms
Response + guest notification
0.022 ms

The outside intervals combine work before and after an inner interval; they do not split one-way transport time. IO completion includes scheduling and reaping, not only device service. Full phase distributions, coverage and method.

Another CPU helps; the journal still uses its queue

Both separate-CPU probes found an 8 KiB journal write from jbd2 blocking unrelated reads on queue 1 for 1.4–2.4 seconds. The reactor already separates reads from commands. The obstacle is earlier, where the frontend discovers requests.

Guest workloadRead p99Slowest read
Read only2.57–2.61 ms56.57–56.91 ms
Reader + writer on CPU 092.8 ms–2.67 s2.99–9.98 s
Reader CPU 1, writer CPU 04.82–6.39 ms21.84–64.77 ms

fio total latency with CAS tracing and guest probes off. Ranges cover both guests; mixed workloads have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten seconds per job. These are nested-VM development measurements, not native-drive latency or a latency guarantee.

Cold read-only p99 was 3.69–4.18 ms; warm p99 was 1.97–2.02 ms. Warm mixed reads still reached p99 of 2.1–2.3 seconds with zero cache misses. Faster payload reads alone do not remove this queueing delay.

The design this led to

PR #53 now retains bounded descriptor snapshots before storage admission. Independent reads can pass blocked writes; overlapping reads and FLUSH barriers keep their order. Waiting write payloads stay in guest RAM. The existing reactor and compactor still execute the work.

These timing results predate the scheduler change. The live comparison above tests the merged implementation. Implementation, crate paths and limitations.

The cost is a fixed descriptor metadata quota and explicit dependency/recovery state. The new retained-carrier format also requires a fresh attachment when upgrading an existing guest. It also cannot help a read that the guest has not published. Cache sharding and thread pinning remain separate options for measured downstream costs. Alternatives, pros and cons, and required acceptance tests.

What was measured, and where to read the code

Instrumentation covers the guest bio/request path, queue admission reasons, publication dependencies, IO stages, cache lock waits and telemetry freshness. Histograms cover all traced completions; a bounded recent tail supplies examples. Guest probes measurably increase latency, so the table uses unprobed runs. The report also compares process-cold and warm caches.

No scheduling policy changed in this experiment. Sustained native-host performance, full guest-ring exhaustion, remote reads, GC pressure and actual OOM/ENOSPC remain separate gates. The disposable labs were stopped and their disks and keys removed.

Earlier: locating the admission stall inside CAS

A cached read waited 7.922 seconds behind writes. Once admitted, it finished in 1.34 ms. The long wait came before the read entered storage. 14 September instrumented probe.

Guest ringWrite waiting
WAL capacity
ReadRead

The frontend stops consuming this virtqueue at its blocked write. The queue mutex is released; the requests remain ordered behind that head. Measured queue-lock acquisition peaked at 0.187 ms. Making the queue lock-free would not change this admission rule.

Behind a blocked write7.922 s99.95% of observed time
After admission1.340 msThrough guest notification

Another 2.625 ms passed before admission. Total observed: 7926.004 ms. Timing starts when CAS sees the request in the ring; earlier guest waiting is excluded.

Inside the 1.340 ms after admission

Queue 0, request 88351, disk offset 20860928.

Frontend dispatch + command channel
0.017 ms
Reactor dispatch / write dependency
0.536 ms
Prepare + advance the read
0.036 ms
IO submission scheduler
0.001 ms
Manifest-page IO
0.709 ms
Other reactor time
0.006 ms
Response channel + guest completion
0.035 ms

All six selected tail reads hit the chunk cache: no WAL or chunk-payload IO, and no shared-fetch wait. IO timing includes kernel scheduling and completion reaping. These examples do not establish cold-read latency.

Another CPU helps; it does not reserve a read queue

WorkloadRead maxCompletion p99
Read only65–67 ms2.54–2.61 ms
Reader + writer, same CPU2.71–7.93 s325–2,869 ms
Reader CPU 1, writer CPU 022 ms–2.21 s4.62–6.46 ms

Each guest mapped CPU 0 to queue 0 and CPU 1 to queue 1. A write still reached queue 1 in each separate-CPU pass, holding reads there for about 2.2 seconds. The trace confirms the write and queue; its guest thread or filesystem source remains unmeasured.

Ranges cover both guests and two mixed passes. 4 KiB/QD8 reads, 1 MiB/QD32 writes, ten seconds per job; caches enabled, WALs drained between passes. Two-vCPU TCG guests on Spark. All 24 fio jobs and restart seed checks passed; no GC or OOM occurred. Test disks and keys were removed.

Trace coverage, source paths and remaining measurements

This probe captured 385,700 of 385,860 reads. The follow-up fixes the cursor/peek observation race and passes 477 native tests; it was not rerun in a VM. Only the 32 slowest completed reads per image are retained. Status snapshots can lag stage boundaries, and tracing overhead is unmeasured.

Next: define bounded read progress past a waiting write, preserving overlap, FLUSH and recovery rules. Cold reads, read-credit waits, telemetry freshness and long-history GC still need measurement.

Earlier: less compaction work and automatic write resumption

The preceding one-vCPU repeat measured the compaction changes and reproduced 10.5-second reads. It motivated the request-level instrumentation above.

Writes resumed when WAL space returned. All 20 fio jobs completed without IO errors. Both logs drained, and both 64 MiB seeds verified after load and a full restart. Measured on 14 September at eeff4b5.

Mixed-read latency still reached 10.5 seconds. No GC occurred. Both guests sent every request through queue 0, where writes waited for WAL space. Later reads can wait behind them; other-queue scheduling had no work to serve here. The evidence points to queue blocking, but does not time every phase of each read.

Initialized bytesMiB / input MiB
Before
73.7
After
5.1
Background readscalls / input MiB
Before
740.3
After
481.7
Final manifest pagespages / input MiB
Before
752.5
After
103.5
Syncscalls / input MiB
Before
5.9
After
6.2

Work per MiB of compacted input: 878 MiB before, 1,173 MiB after. Counts include rotation and reclamation; neither run had GC. One repetition with different backlog and timing. These are operation counts, not physical-device traffic or a statistical speedup.

Source revisions, drain observations and remaining costs

Before: bdb26fc. After: eeff4b5, including the compaction changes in PR #49. The full report retains commands, counters and comparison limits.

The sampled drain rate rose from 13.1 to 21.6 MiB/s. Initialization fell to 0.34 s, while reads took 40.20 s and syncs 31.71 s across the new progression. Reclamation still revisits retained WAL history. Measure its reads and syncs by caller, then reduce that repeated work. Within-queue read bypass needs an explicit ordering contract first.

Earlier runs: how congestion became an IO error

Removing queue expiry let an earlier repeat complete through 12.7 / 13.6-second waits with zero errors, but mixed reads reached 11.3 seconds and one GC paused the host for 3.17 seconds. The trace below predates that fix.

Earlier failure, shown below. The 256 MiB image quota stayed closed for about 8.6–13.6 seconds while the WAL drained; queued heads expired after five seconds. The fix keeps capacity waits pending and wakes admission when reclamation restores space. A separate native test also held an accepted write before submission for 31 seconds and completed it after release.

Per-image staging allocation during overloadAllocation approaches 256 MiB. The image stops admitting writes, then drains toward the 60 percent reopen threshold. Orange shading marks sampled stopped admission. The host-wide quota remains open throughout.256 MiB cap153.6 MiB reopen threshold0 s23.2 s

214.1 MiB · write quota closed · 1 blocked queue head · 1 guest IO error

Orange = write quota closed. Reserving the next 8 MiB segment can stop writes before allocation reaches 256 MiB. Samples are about 500 ms apart; a threshold crossing and a new allocation can occur between them.

Historical compactor phase totals
Read the WAL23.8%
Hash + prepare tree24.1%
Write / sync chunks13.1%
Publish manifest9.0%
Publish D + reclaim WAL30.0%

Elapsed phase time across this lab, including seeding and drain. Disk reservation was under 0.1%. Preparation combines hashing and tree edits; these are not CPU samples or a hash-only profile.

The wake-up gap now has a regression test

The original trace showed Guest 2’s quota reopen while its head stayed blocked. A deterministic test then reproduced the missing notification without unrelated IO. It passes after the fix; this does not establish the exact thread interleaving of the historical guest run.

What disk and memory controls established

With a restricted physical-space budget, GC could not reopen writes, but all 256 MiB of seeded data remained readable. Native memory-budget tests refused index growth and compaction preparation before output; a read-page allocation denial failed its image. Actual kernel OOM and unexpected filesystem ENOSPC remain untested.

Unmodified build: EIO at 1 MiB / QD32 after the earlier stages. Instrumented repeat: EIO already at 4 KiB / QD64. The stage varies with backlog; neither is a universal queue-depth limit. No physical-space pressure, GC or cgroup OOM explained these failures. Seed checks passed, and writes resumed after drain.

Experiment, controls and evidence · measured at c11e106. New instrumentation in cas-daemon: local/pressure.rs and local/host/statistics.rs.

The code now resumes WAL selection from a validated cursor, grows preparation buffers as needed and writes only pages reachable from the final manifest root. The comparison measures the resulting reduction in work per compacted MiB. Buffer initialization counts are cumulative work, not peak RAM or deduplication savings.

Lab geometry, memory and measurement limits

Spark ran one 4 GiB KVM VM with a 4 GiB XFS lab disk and two inner 512 MiB, one-vCPU guests using TCG emulation. The service had a 6 GiB memory limit and no swap allowance. Its whole-service peak was 2.81 GiB; sampled storage PSS peaked at 441.90 MiB. These overlap and must not be added.

There were no cgroup OOM events or limit hits. This does not test actual OOM or unexpected filesystem ENOSPC. No GC occurred in that run, so its read stalls cannot be attributed to a GC pause. One repetition with different backlog and timing does not establish a statistical speedup or sustained device rate.

The lab disk and temporary keys were deleted after verification. Small reports and receipts remain. Commands, revisions, verification and cleanup.

Earlier baseline · 13 September · CAS versus two raw backends

These measurements predate the congestion and compaction changes. At e15d733, each arm had two guests, one active and one idle. Controls used QEMU raw storage or the reference raw io_uring daemon.

Median of each run’s p99 · lower is better

QEMU raw 1.52 ms
Run range 1.38–1.6 ms
Raw daemon 2.09 ms
Run range 1.7–3 ms
CAS 1.58 ms
Run range 0.96–2.8 ms

5 × 5s per case · 32 MiB file · direct IO, QD1 · caches retained between jobs

Memory observed across each lab session
BackendLab peakStorage PSS
QEMU raw2,301 MiBIn QEMU
Raw daemon2,329 MiB25 MiB
CAS2,770 MiB248 MiB

Lab peak: cgroup memory, including the outer VM. Storage PSS: sampled inner daemon peak. The columns overlap; do not add them. No cgroup OOM kills occurred.

CAS’s median fdatasync p99 was 9.76 ms, versus 3.62 ms raw and 3.78 ms with the reference daemon. Sequential reads reached 437 MiB/s, versus 842 MiB/s raw. These observations motivated profiling; a new matched comparison is needed to describe the current backend.

All 60 timed jobs passed. CAS completed 487 compaction batches, with no GC or cache eviction. Its lab peaked at 2,770 MiB of cgroup memory. No lab recorded an OOM kill.

CAS advertised four queues and the controls one; each guest had one vCPU and jobs ran at QD1. Caches and telemetry remained enabled. These nested-lab results do not establish native NVMe performance or the study’s latency gate.

Each plotted value is the median of five per-run results. The fdatasync p99 measures sync separately from the preceding write; adding their p99s would not yield a transaction p99. Process PSS/RSS was sampled every 500 ms; the service’s memory peak also includes the outer VM.

Method, per-run values and evidence receipts.

What remains before the study can draw conclusions

  • Memory bounds are incomplete. Report construction, framework allocations, thread stacks and final resident-memory attribution remain outside the closed audit. Passing budget counters cannot close that gap.
  • GC can stall every guest. Its sweep is synchronous, and deadline checks do not interrupt each blocking syscall or tree walk. Background watchdogs and guest-side timeouts still apply. Pause duration after a long overwrite history needs measurement.
  • The index limits unique capacity. Every stored hash needs RAM, including entries awaiting GC. Table growth overlaps old and new allocations. Disk headroom alone does not guarantee another unique write can be compacted.
  • Shared failures still stop the host. An image socket failure after activation now leaves its peers running; a protocol test checks that boundary. Shared corruption and incomplete retained recovery remain host-wide failures. There is no automatic image recovery.
  • One worker couples the images. Chunk writes, rotation, snapshots and collection share that thread. A stalled background operation delays the others. The read cache is also shared LRU, so a scan can displace another guest’s hot blocks.
  • Read isolation remains incomplete. Bounded bypass, overlap rules and retained recovery are implemented. Live mixed runs still show seconds-long fairness waits. The scheduler must preserve an eligible read’s opportunity when blocked writers retry.

The current system is single-host storage with one writable attachment per image. Remote reads, replicated durability, migration/handoff, RDMA, CDC, compression, guest-memory sharing and general online lifecycle tooling remain outside the implemented path.

Next: resolve the shared admission fairness interaction and repeat the same mixed workload. Then attribute remaining reclamation reads/syncs and long-history GC pauses; finish the allocation audit and actual capacity-failure tests. Dedicated-media, ZFS and two-host comparisons are still required to answer the research questions. Compaction measurements and code paths · Full storage audit.

Development checks establish behavior at specific revisions

The records cover write ordering, durable FLUSH, compaction crash cuts, shared recovery and fresh-boot integrity. They support using this backend for further experiments. Dedicated-media performance, physical power loss and complete resident-memory accounting remain unaccepted.

Tested revisions, checkpoint counts and open gates

Architecture links use 4b50554, the merged read scheduler tested here. Each older experiment keeps its own source revision; its timings do not describe this binary. Source links require Forgejo repository access.

RecordWhat it establishes
14 September merged read scheduler13 live recovery/reset scenarios, 50 workload jobs and eight full seed checks on 4b50554. Bypass occurs; a new two-image Rust reproduction confirms remaining admission starvation. Latency acceptance stays open.
14 September read attribution24 fio jobs on fc738c2; traced same-queue waits and separate-CPU controls. A subsequent observation-race correction passed 477 native tests; 25 fixture-dependent ignores. Scheduler behavior is unchanged.
14 September pressure repeat20 fio jobs, quota resumption, drain and restart verification on eeff4b5. Same-queue read stalls remain.
14 September storage follow-up471 native tests passed; 25 fixture-dependent tests were ignored in that run. Separately, 185 selected XFS tests and two-image live recovery passed. The record gives each tested source.
13 September integrated run41/41 scenarios on b8399ad, including 178 XFS fixture tests, six compaction crash cuts, eleven competing-guest stages and fresh-boot checks. This predates the pressure fixes.

C1–C3 cover the reference, copy/format and ordering/recovery checkpoints. C4’s 40 functional scenarios and C5’s integrated inventory passed, but aggregate closure still requires the final allocation audit and dedicated-media repetitions. The study’s G1–G6 research gates are separate.

The suite verifier is a separate repository command. Scenario counts include build/style checks and nested groups; 41 scenarios are not 41 independent fault models. Earlier failures and retries remain in their session records.

Bulk VM disks were removed after verification. Logs and receipts remain at the locations recorded by each session; they are not a portable archive of the deleted data. This editorial update runs no new storage experiment.

Open allocation audit · Validation history · Progress tracker

What changed since Update 01

Update 01 describes the earlier fixed 8 KiB log and serial write-through recovery. The current host uses a packed WAL, concurrent retained recovery, a shared compactor and verified read caches. The current limits and source revisions are recorded above.

Start a VM. Write to its disk. Reopen it.

casctl new demo --count 2
casctl ls
casctl shell demo
casctl ssh demo/2 -- lsblk
casctl status demo
casctl bench demo --case flush
casctl stop demo
casctl start demo
casctl stop demo
casctl rm demo

demo/1 and demo/2 have private ext4 disks at /mnt/cas, backed by one shared CAS store. stop retains their data. rm deletes the stopped lab’s disk and archives its small results. The guest root is temporary.

Setup, transport and ownership

On Spark, build with nix build .#casctl --out-link result-cli. Use ./result-cli/bin/casctl in place of casctl above, starting with doctor to check prerequisites. A user service owns each lab after the command exits. The CLI generates an SSH key locally and pins guest host keys; the private key stays on Spark.

The lab puts a dedicated XFS filesystem and cas-host inside a 4 GiB KVM VM. Its inner guests use TCG emulation, 512 MiB each. SSH reaches them through a loopback port and the outer host. Block IO uses the same shared-memory virtqueues, eventfds and io_uring paths described above.

Interactive telemetry retains the current and previous files through bounded rotation. Strict checkpoint recording instead fails at 4,096 samples or 64 MiB.

The backing disk is capped at 4 GiB. The user service has a 6 GiB memory limit, and the runner checks for 25 GiB of remaining host disk space. Guest count is chosen at creation; separate new commands create separate stores.

Commands and limits · crates/cas/cli/src/main.rs · crates/harnesses/src/lab.rs · nix/lab/host.nix