How CAS changed
Update 05 · 24 September 2026
CAS shares identical disk blocks across VM images. Writes enter a per-image log. Background compaction moves durable data into shared chunks.
Under write pressure, reads stalled for seconds. The changes below separate the capacity problem, the compaction cost, and the scheduler bug.
Writes enter the log before shared storage
Each guest has its own disk image. QEMU exposes the guest's requests to CAS through shared-memory queues. Admission reserves capacity before CAS accepts a write.
Guest writes
Background compaction
The write-ahead log, or WAL, holds recent writes. A WRITE completion means the write is published. An ordered FLUSH makes the log durable on the local host.
The compactor hashes durable data in 4 KiB chunks. It stores each missing chunk once. A manifest maps each image's disk blocks to chunk hashes. CAS commits that map before reclaiming the covered log data.
Reads use the newest mapping
A recently written block
A block outside the WAL overlay
A read can combine recent WAL data with shared chunks. Holes return zeros. CAS keeps source pins until the read completes, so reclamation cannot remove data still in use.
Durability and source
Published writes, durable WAL, and compacted mappings are separate boundaries. The manifest cannot cover writes beyond the durable WAL. Reclamation also waits for reader and replay pins. Host garbage collection is separate from ordinary background compaction.
A full log should delay a write
The first failure was a capacity timeout. Writes filled their WAL quota faster than compaction could reclaim it. After five seconds, an admission wait became a guest IO error.
Before · a full log becomes an error
After · wait for capacity
CAS now keeps the request pending. Its unadmitted payload stays in guest RAM. Reclamation wakes admission when space returns. A periodic retry covers releases without a wake notification.
The reproduced timeout errors disappeared. Writes resumed after waits of 12.66 and 13.60 seconds. Reads still took up to 11.34 seconds.
Measurement scope and source
14 September, two bounded Spark repeats. The two mixed-read maxima were 11.34 and 11.15 seconds. This fixes the reproduced capacity error, not all IO errors or maximum latency. Device and lifecycle deadlines still exist.
The compactor repeated work it could skip
Selection rescanned old log prefixes. Buffer preparation initialized the maximum capacity. Manifest output included pages superseded by later edits in the same batch.
Before · repeat old work
After · keep the useful work
Selection now resumes from a validated cursor. Buffers grow as needed. The final output keeps only reachable new manifest pages.
| Work | Before | After |
|---|---|---|
| Initialized bytesMiB per input MiB | 73.67 | 5.08 |
| Background readscalls per input MiB | 740.35 | 481.7 |
| Final manifest pagespages per input MiB | 752.51 | 103.5 |
| Syncscalls per input MiB | 5.95 | 6.23 |
Repeated work fell. Sync calls did not. Mixed-read maxima were still 10.50 and 9.57 seconds.
These runs had different backlogs. The table counts work per MiB. It does not establish a device speedup or a read-latency improvement.
What still scans
The cursor is a validated hint with a safe scan fallback. Reclamation still walks retained history. Manifest edits can still create temporary paths in memory; final emission omits superseded pages.
A blocked write hid reads behind it
The frontend stopped at a write that could not get WAL capacity. A later read of unrelated blocks stayed undiscovered, even though it needed no write space.
Before · discovery stops at the write
After · discover, then select
The frontend now retains a bounded set of request descriptors, then selects eligible reads. It does not copy an unlimited queue of write payloads.
A read still waits behind an earlier overlapping write, zero, or discard. It also respects FLUSH barriers. This preserves the image's ordering rules.
Bypass occurred in the measured runs. Read p99 fell in the matched same-CPU comparison, but one read still took 6.834 seconds.
The remaining trace exposed another problem. A read could be discovered and eligible, yet still lose every scheduler turn.
Measurement scope and source
14 September. The comparison used matched settings in separate sessions. It is distinct from the same-session control below. After admission, reads also retain the existing image-wide publication dependency.
The scheduler kept retrying the refused write
With two images, each turn could retry the same capacity-blocked write. The scheduler moved to the other image, then repeated. Eligible reads in both images waited behind those retries.
Before · the refused write stays first
After · rotate the refused ticket
A refused, uncommitted ticket now moves to the back of its image's list. On the next visit, an eligible read can run. The scheduler still shares service between images using byte deficits.
| Measurement | Before rotation | After rotation |
|---|---|---|
| Mixed-read p99 | 219 to 287 ms | 9 to 20 ms |
| Independent bypass admission, max | 8.110 s | ≤ 0.137 s |
Read p99 improved. Writers continued to progress. One fixed run still had a 25.529-second read.
The 137 ms bound above applies only to independent bypass admission. It is not a bound on all read latency. The remaining maximum stalls need tracing.
What this result establishes
The same-session control supports the scheduler change. It does not establish unchanged writer throughput or solve every read stall. The untraced ordinary-head maxima are consistent with FLUSH or WAL-capacity blocking, but their individual causes remain unproven.
CloudLab removes the nested guest setup
The Spark results above came from a development fixture. Workload guests used TCG emulation inside an outer KVM VM. Other work also ran on Spark.
Spark · historical measurements
CloudLab · protocol checks
The CloudLab host has an AMD EPYC 9354P, 32 cores, 192 GB RAM, and two 800 GB NVMe drives. CAS runs on the physical host. Guests run directly under KVM.
All three backends passed direct KVM data checks on file-backed XFS. Dedicated NVMe measurements have not started.
Raw XFS is the storage baseline. Passthrough adds the daemon path without CAS storage. CAS adds the log, sharing, and compaction. The runner uses the same guest, workload, and resource limits for all three.
Native runner settings
- Each guest has two vCPUs, 2 GiB RAM, and one queue of 128 entries.
- Each run has an 8 GiB cgroup limit and no swap.
- The runner pins CPU affinity and memory policy to the storage NUMA node.
- fio uses direct IO and a 512 MiB working set. CRC checks verify the data before and after.
- CAS has a 16 MiB clean cache. Jobs retain cache state within a repeat. Each repeat starts with new storage.
- Read-only jobs wait for compaction to catch up. Write completion and fdatasync latency are separate measurements.
- The pressure case uses separate reader and writer guests. It does not test same-image FLUSH stalls.
Runner and method · Workload · a1bbdf3
The next result needs dedicated storage
- Run raw XFS, passthrough, and CAS on the dedicated NVMe. Measure latency, physical storage, device writes, and total memory.
- Rerun recovery on that exact revision. Complete the allocation audit.
- Trace the remaining maximum read stalls.
- Add a ZFS comparison before claiming a storage advantage.
Cross-host sharing, remote reads, replicated durability, and migration remain unimplemented.
Measurement records
The explanations above stand on their own. These dated panels retain the detailed Spark measurements. Compare values within each experiment.
13 September · raw, passthrough, and CAS
Median of each run’s p99 · lower is better
5 × 5s per case · 32 MiB file · direct IO, QD1 · caches retained between jobs
Memory observed across each lab session
| Backend | Lab peak | Storage PSS |
|---|---|---|
| QEMU raw | 2,301 MiB | In QEMU |
| Raw daemon | 2,329 MiB | 25 MiB |
| CAS | 2,770 MiB | 248 MiB |
Lab peak: cgroup memory, including the outer VM. Storage PSS: sampled inner daemon peak. The columns overlap; do not add them. No cgroup OOM kills occurred.
14 September · capacity and compaction
Writes resumed when WAL space returned. All 20 fio jobs completed without IO errors. Both logs drained, and both 64 MiB seeds verified after load and a full restart. Measured on 14 September at eeff4b5.
Mixed-read latency still reached 10.5 seconds. No GC occurred. Both guests sent every request through queue 0, where writes waited for WAL space. Later reads can wait behind them; other-queue scheduling had no work to serve here. The evidence points to queue blocking, but does not time every phase of each read.
Work per MiB of compacted input: 878 MiB before, 1,173 MiB after. Counts include rotation and reclamation; neither run had GC. One repetition with different backlog and timing. These are operation counts, not physical-device traffic or a statistical speedup.
Source revisions, drain observations and remaining costs
Before: bdb26fc. After: eeff4b5, including the compaction changes in PR #49. The full report retains commands, counters and comparison limits.
The sampled drain rate rose from 13.1 to 21.6 MiB/s. Initialization fell to 0.34 s, while reads took 40.20 s and syncs 31.71 s across the new progression. Reclamation still revisits retained WAL history. Measure its reads and syncs by caller, then reduce that repeated work. Within-queue read bypass needs an explicit ordering contract first.
Earlier runs: how congestion became an IO error
Removing queue expiry let an earlier repeat complete through 12.7 / 13.6-second waits with zero errors, but mixed reads reached 11.3 seconds and one GC paused the host for 3.17 seconds. The trace below predates that fix.
Earlier failure, shown below. The 256 MiB image quota stayed closed for about 8.6–13.6 seconds while the WAL drained; queued heads expired after five seconds. The fix keeps capacity waits pending and wakes admission when reclamation restores space. A separate native test also held an accepted write before submission for 31 seconds and completed it after release.
214.1 MiB · write quota closed · 1 blocked queue head · 1 guest IO error
Orange = write quota closed. Reserving the next 8 MiB segment can stop writes before allocation reaches 256 MiB. Samples are about 500 ms apart; a threshold crossing and a new allocation can occur between them.
Historical compactor phase totals
Elapsed phase time across this lab, including seeding and drain. Disk reservation was under 0.1%. Preparation combines hashing and tree edits; these are not CPU samples or a hash-only profile.
The wake-up gap now has a regression test
The original trace showed Guest 2’s quota reopen while its head stayed blocked. A deterministic test then reproduced the missing notification without unrelated IO. It passes after the fix; this does not establish the exact thread interleaving of the historical guest run.
What disk and memory controls established
With a restricted physical-space budget, GC could not reopen writes, but all 256 MiB of seeded data remained readable. Native memory-budget tests refused index growth and compaction preparation before output; a read-page allocation denial failed its image. Actual kernel OOM and unexpected filesystem ENOSPC remain untested.
Unmodified build: EIO at 1 MiB / QD32 after the earlier stages. Instrumented repeat: EIO already at 4 KiB / QD64. The stage varies with backlog; neither is a universal queue-depth limit. No physical-space pressure, GC or cgroup OOM explained these failures. Seed checks passed, and writes resumed after drain.
Experiment, controls and evidence · measured at c11e106. New instrumentation in cas-daemon: local/pressure.rs and local/host/statistics.rs.
14 September · read traces and bypass
A cached read waited 7.922 seconds behind writes. Once admitted, it finished in 1.34 ms. The long wait came before the read entered storage. 14 September instrumented probe.
WAL capacityReadRead
The frontend stops consuming this virtqueue at its blocked write. The queue mutex is released; the requests remain ordered behind that head. Measured queue-lock acquisition peaked at 0.187 ms. Making the queue lock-free would not change this admission rule.
Another 2.625 ms passed before admission. Total observed: 7926.004 ms. Timing starts when CAS sees the request in the ring; earlier guest waiting is excluded.
Inside the 1.340 ms after admission
Queue 0, request 88351, disk offset 20860928.
- Frontend dispatch + command channel
- 0.017 ms
- Reactor dispatch / write dependency
- 0.536 ms
- Prepare + advance the read
- 0.036 ms
- IO submission scheduler
- 0.001 ms
- Manifest-page IO
- 0.709 ms
- Other reactor time
- 0.006 ms
- Response channel + guest completion
- 0.035 ms
All six selected tail reads hit the chunk cache: no WAL or chunk-payload IO, and no shared-fetch wait. IO timing includes kernel scheduling and completion reaping. These examples do not establish cold-read latency.
Another CPU helps; it does not reserve a read queue
| Workload | Read max | Completion p99 |
|---|---|---|
| Read only | 65–67 ms | 2.54–2.61 ms |
| Reader + writer, same CPU | 2.71–7.93 s | 325–2,869 ms |
| Reader CPU 1, writer CPU 0 | 22 ms–2.21 s | 4.62–6.46 ms |
Each guest mapped CPU 0 to queue 0 and CPU 1 to queue 1. A write still reached queue 1 in each separate-CPU pass, holding reads there for about 2.2 seconds. The trace confirms the write and queue; its guest thread or filesystem source remains unmeasured.
Ranges cover both guests and two mixed passes. 4 KiB/QD8 reads, 1 MiB/QD32 writes, ten seconds per job; caches enabled, WALs drained between passes. Two-vCPU TCG guests on Spark. All 24 fio jobs and restart seed checks passed; no GC or OOM occurred. Test disks and keys were removed.
Trace coverage, source paths and remaining measurements
This probe captured 385,700 of 385,860 reads. The follow-up fixes the cursor/peek observation race and passes 477 native tests; it was not rerun in a VM. Only the 32 slowest completed reads per image are retained. Status snapshots can lag stage boundaries, and tracing overhead is unmeasured.
Next: define bounded read progress past a waiting write, preserving overlap, FLUSH and recovery rules. Cold reads, read-credit waits, telemetry freshness and long-history GC still need measurement.
crates/cas/daemon/src/backend.rs· queue admission and guest completioncrates/cas/daemon/src/read_trace.rs· measured observercrates/cas/daemon/src/local/reactor.rs· dependency, scheduler and IO phases- Trace field guide and corrected observer · commands, failures and cleanup
The original long read wait was at admission. One matched read took 7.938 seconds in fio: 7.933 seconds passed before it reached the head of its CAS queue, then 3.53 ms after admission. The reader and the blocked write touched different disk ranges.
fio → Linux bio 2Virtqueue
shared descriptors 3CAS admission
wait behind a write 4Reactor → storage
cache, WAL or chunks 5Guest completion
used ring → fio
In these measurements, the frontend released the queue mutex while waiting, but stopped consuming the queue at its blocked write. Later reads could not reach their own read credits. A lock-free queue would preserve that obstruction unless its admission rule also changed.
Amber: before the CAS queue head. Blue: the rest of fio latency. Queue 0, disk offset 19038208. Selected slow IO, not a percentile sample; guest and CAS clocks are joined by request identity and only durations are compared.
Where the remaining time went
- Guest bio submission → virtqueue publication
- 0.033 ms
- fio time outside bio submission → block completion
- 1.054 ms
- Virtqueue/completion time outside the CAS interval
- 0.428 ms
- At the CAS queue head
- 0.001 ms
- Frontend dispatch + command channel
- 0.018 ms
- Reactor dispatch / publication dependency
- 2.322 ms
- Read execution
- 1.165 ms
- Response + guest notification
- 0.022 ms
The outside intervals combine work before and after an inner interval; they do not split one-way transport time. IO completion includes scheduling and reaping, not only device service. Full phase distributions, coverage and method.
Another CPU helps; the journal still uses its queue
Both separate-CPU probes found an 8 KiB journal write from jbd2 blocking unrelated reads on queue 1 for 1.4–2.4 seconds. The reactor already separates reads from commands. The obstacle is earlier, where the frontend discovers requests.
| Guest workload | Read p99 | Slowest read |
|---|---|---|
| Read only | 2.57–2.61 ms | 56.57–56.91 ms |
| Reader + writer on CPU 0 | 92.8 ms–2.67 s | 2.99–9.98 s |
| Reader CPU 1, writer CPU 0 | 4.82–6.39 ms | 21.84–64.77 ms |
fio total latency with CAS tracing and guest probes off. Ranges cover both guests; mixed workloads have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten seconds per job. These are nested-VM development measurements, not native-drive latency or a latency guarantee.
Cold read-only p99 was 3.69–4.18 ms; warm p99 was 1.97–2.02 ms. Warm mixed reads still reached p99 of 2.1–2.3 seconds with zero cache misses. Faster payload reads alone do not remove this queueing delay.
The design this led to
PR #53 now retains bounded descriptor snapshots before storage admission. Independent reads can pass blocked writes; overlapping reads and FLUSH barriers keep their order. Waiting write payloads stay in guest RAM. The existing reactor and compactor still execute the work.
These timing results predate the scheduler change. The live comparison above tests the merged implementation. Implementation, crate paths and limitations.
The cost is a fixed descriptor metadata quota and explicit dependency/recovery state. The new retained-carrier format also requires a fresh attachment when upgrading an existing guest. It also cannot help a read that the guest has not published. Cache sharding and thread pinning remain separate options for measured downstream costs. Alternatives, pros and cons, and required acceptance tests.
What was measured, and where to read the code
Instrumentation covers the guest bio/request path, queue admission reasons, publication dependencies, IO stages, cache lock waits and telemetry freshness. Histograms cover all traced completions; a bounded recent tail supplies examples. Guest probes measurably increase latency, so the table uses unprobed runs. The report also compares process-cold and warm caches.
No scheduling policy changed in this experiment. Sustained native-host performance, full guest-ring exhaustion, remote reads, GC pressure and actual OOM/ENOSPC remain separate gates. The disposable labs were stopped and their disks and keys removed.
crates/cas/daemon/src/backend.rs— virtqueue discovery and admissioncrates/cas/daemon/src/read_trace.rs— phases, gate reasons and retained examplescrates/cas/daemon/src/local/reactor/read.rs— read executioncrates/cas/core/src/cache.rs— shared chunk-cache ownershipcrates/harnesses/probes/read-path.bt— guest request identities
Reads can bypass writes, but admission fairness still stalls them. The longest untraced read still took 6.83 s.
| Workload | Earlier read p99 | Merged read p99 | Earlier slowest | Merged slowest |
|---|---|---|---|---|
| Read only | 2.57 ms–2.61 ms | 2.24 ms–2.28 ms | 56.91 ms | 10.34 ms |
| Reader + writer on CPU 0 | 92.80 ms–2.67 s | 24.51 ms–156.24 ms | 9.98 s | 6.83 s |
| Reader CPU 1, writer CPU 0 | 4.82 ms–6.39 ms | 5.60 ms–6.39 ms | 64.77 ms | 3.06 s |
Untraced fio total latency. Each p99 range covers two guests; mixed cases have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows. Matched workloads ran in separate Spark sessions using nested TCG guests; these are development measurements.
The admission rule
Read B can use its own read credits while Write A waits for WAL capacity. The write payload remains in guest RAM.
A and B stand for different disk ranges. CAS owns bounded descriptor snapshots before admission; it does not need a new guest queue, a lock-free queue or a new executor. After admission, the existing image-wide write-publication dependency still applies.
What the instrumentation says now
One matched 6.295-second read spent 6.293 seconds before admission and only 0.57 ms afterward. Its chunk was cached. The fairness scheduler kept refusing its turn; a separate Rust reproduction confirms that two blocked writers can repeatedly displace eligible reads without any disk IO.
Each image retries its oldest write before its eligible read. That retry marks the write ready again, which can displace the read the shared scheduler just selected. Two images can keep waking each other and repeating this until write capacity returns.
Next: preserve eligible-read selection across blocked-write retries, then repeat this workload. Overlap, FLUSH and finite-capacity rules still matter; this result does not justify a new cache or a lock-free queue.
Validation, resource costs and code
The packaged merge passed 499 native tests. All 13 live recovery/reset scenarios and 8 full 64 MiB seed checks passed. The recovery cases kept QEMU alive across backend replacement; they did not reproduce an overloaded deferred-read crash. Native frontend tests cover that ownership case.
The largest sampled lab memory peak was 3.93 GiB under a 6 GiB cap, with zero observed cgroup OOM or limit events. Disposable disks and keys were removed. Actual OOM/ENOSPC, sustained fairness, full guest-ring exhaustion, long-history GC and physical power loss remain separate gates.
Each frontend reserves about 6 MiB of descriptor metadata quota, with eight read-request slots separate from writes. The existing byte allowance fits seven simultaneous 4 KiB reads. Version-3 retained carriers require a fresh guest attachment when upgrading from version 2.
backend/frontier.rs— discover, check dependencies, admit independent readslocal/pools.rs— separate read creditsinflight.rsandbackend/recovery.rs— retain and recover ownershipread_trace.rs— phase timings and admission observations
All four owners are in crates/cas/daemon. Raw-result analysis, per-job numbers and limitations.
15 September · scheduler control
Independent reads waited up to 8.1 s in admission before the fix and at most 0.137 s after it. Three back-to-back runs in one host session, with the same workload.
| Workload | Arm | Read p99 | Slowest read | Reads per 10 s | Write MiB/s |
|---|---|---|---|---|---|
| Read only | Before (main before the review) | 5.5 ms | 282 ms–284 ms | 68k–68k | — |
| After, pass 1 | 4.9 ms–5.1 ms | 227 ms–228 ms | 64k–67k | — | |
| After, pass 2 | 4.0 ms–4.1 ms | 117 ms–120 ms | 75k–75k | — | |
| Reader and writer on CPU 0 | Before (main before the review) | 219 ms–287 ms | 8.1 s–24.1 s | 2.3k–5.2k | 8.8–21 |
| After, pass 1 | 11 ms–20 ms | 307 ms–25.5 s | 30k–33k | 8.5–24 | |
| After, pass 2 | 8.6 ms–12 ms | 325 ms–697 ms | 42k–46k | 11–23 | |
| Reader CPU 1, writer CPU 0 | Before (main before the review) | 20 ms–20 ms | 188 ms–9.1 s | 17k–31k | 19–21 |
| After, pass 1 | 17 ms–24 ms | 106 ms–11.7 s | 20k–37k | 17–20 | |
| After, pass 2 | 11 ms–21 ms | 113 ms–5.6 s | 19k–46k | 27–28 |
fio total latency, 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows, two guests per arm; ranges cover both guests and both passes of each mixed workload. Every job and every final 64 MiB CRC check passed; no storage errors. Nested TCG guests on a busy shared host: A Documents archive (tar | zstd) and a guest-removal job ran on Spark throughout; load average 6–9 on 20 cores. Only arms within this session are comparable.
Scheduler counters at the end of each arm
- Independent-read admission, longest wait
- 8.1 s / 8.1 s
- Ordinary queue head, longest wait
- 23.1 s / 22.0 s
- Reads that passed a blocked head
- 3,705 / 3,616
- Refused scheduler turns
- 1,345,363 / 1,321,083
- Admitted requests
- 147,575 / 132,730
Per image (guest 1 / guest 2), from the daemon’s own telemetry at the end of the arm. Reads at the head of their own queue count with the ordinary heads.