← index Update 03 · 15 September 2026

Reads no longer wait on writers

The single-host backend is functionally complete and now reviewed end to end. This update reports that review, the two runtime defects it found, and a live measurement of the fix against the previous scheduler under identical conditions. It then says plainly how far the research study itself has progressed: not yet past its first gate.

Update 02 explains how the backend works. Nothing in the architecture changed here; this page is about its state.

Latest finding: with the previous scheduler, independent reads waited up to 8.1 s for admission while two images retried blocked writes. After the fix the same counter peaked at 137 ms, same-CPU read p99 fell from hundreds of milliseconds to 9–20 ms, and writers continued to progress. Multi-second tails from FLUSH barriers behind WAL-blocked writes remain. Measurements ↓

Where the study stands

The implementation checkpoints C0–C5 describe the local backend; the research gates G1–G6 describe the study. The first column is done in development form. The second has not started, because it needs dedicated hardware and a ZFS comparator that do not yet exist.

CheckpointStateOpen
C0–C3acceptedDesign, build-bound validator, packed append, four queues, retained live recovery.
C4functional scenarios passed: 40 scenariosFinal allocation audit.
C5scenarios passed on b8399ad: 41 scenariosMemory accounting, dedicated-media repetitions, rerun on current main.
GateWhat it requiresState
G1Passthrough within 10% of raw p99not started: no dedicated host, no R0 baseline
G2Recovery and ordering suitedevelopment suites pass; not rerun on current main
G3Single-host table (R0, R1, R3)not started: no ZFS R1 profile
G4Two-host tablenot started
G5Transport matrixnot started
G6One-command reproductionnot started

The progress tracker has 130 items checked and 29 open: Next 9, Local storage 6, Architecture review 7, Two hosts 4, Publication 3. Native checks on current main pass 504 tests with 25 fixture-dependent tests ignored. The consolidated C5 suite last ran on b8399ad; the scheduler and compactor have changed since, and it has not been rerun.

One pass over every file

Six independent reviewers each read a slice of the roughly 59k lines of first-party Rust, every file in full. They agreed on the shape. The invariants are carried by types: aligned buffers, budgeted allocation that charges before it allocates, permits and tickets that release on drop, receipts only successful IO can produce, formats validated on decode and round-tripped on encode. Every unsafe block has a justification that holds. No data-loss, ordering or lock-order bug was found in the storage, replay, carrier or cache paths.

The seams are overgrown. The vhost backend is a 34-field struct spread over four files sharing private state; a storage enum’s variant checks leak into it; three state machines run 200 lines or more mixing events with policy; the same small helpers were copied three to nine times; reports were untyped JSON beside a typed pattern; and four different things are called “admission”.

What changed

PRKindChange
#55runtime fixRefused admission heads rotate behind their image’s eligible reads
#56runtime fixA denied image attachment stays attachable without a stale admission wake
#57documentationEleven stale status statements corrected
#58refactorcas-core: shared helpers, direct IO routing, narrower visibility, dead module removed
#59refactorcas-daemon: typed reports, one gate helper, no production panics on a poisoned gate
#60refactorHarness: typed crash points and checkpoints, one helper per duplicated idiom

All six merged through #61 as 62 commits: +250 / −388 lines in cas-core, +792 / −284 in cas-daemon, +810 / −489 in the harness, CLI and Nix. Each refactor commit is behaviour-preserving; report JSON was checked byte for byte against the old output, and CLI help text is unchanged.

The two runtime defects

Read starvation with two images. A refused write’s retry marked itself ready again and, as the oldest head, was chosen ahead of the eligible read behind it. Each refusal handed the turn to the other image, which repeated the pattern, so neither read ran until write capacity returned. The refused head now rejoins behind its image’s other heads. The recorded reproduction that failed after 64 visits passes within four; retry self-healing is unchanged, so no write can strand.

A denied attachment lost the image. Attaching an image took it out of its slot before cloning descriptors, charging the metadata budget and binding its wakes. A denial in any of those left the image permanently unattachable and its admission wake marked attached, which would fail the next host quiescence. Fallible steps now run first; the regression test fails on the old ordering at the stale-wake assertion.

Design decisions left open on purpose
  • Storage dispatch. Give the storage enum one submission entry point so the backend stops inspecting its variant.
  • Background deadline. The scheduler’s fixed 30 s background wait can mark the chunk store failed when a demand reactor stalls. Make it a parameter or a retryable refusal before the first byte.
  • Page-cache key. It includes the manifest end offset, so every new publication misses on unchanged interior pages.
  • Permit ordering. Forgetting one call on a staging permit fails the whole account; a typestate would make it a compile error.
  • Two staging logs. The v1 log is still used by the older daemon path and the CLI beside the v2 write-ahead log.
  • Unused accounting. Three disk-reservation methods have no production caller and are now test-only; delete or wire them in.

Scheduler fix record · Attach fix record · Integration record

The fix, measured against a control

The integrated source passed its native checks and all 13 live QEMU recovery and reset cases. Then the same two-guest mixed workload from Update 02 ran three times in one session: once on the previous runtime as a control and twice on the new one.

Independent reads waited up to 8.1 s in admission before the fix and at most 0.137 s after it. Three back-to-back runs in one host session, with the same workload.

WorkloadArmRead p99Slowest readReads per 10 sWrite MiB/s
Read onlyBefore (main before the review)5.5 ms282 ms–284 ms68k–68k—
After, pass 14.9 ms–5.1 ms227 ms–228 ms64k–67k—
After, pass 24.0 ms–4.1 ms117 ms–120 ms75k–75k—
Reader and writer on CPU 0Before (main before the review)219 ms–287 ms8.1 s–24.1 s2.3k–5.2k8.8–21
After, pass 111 ms–20 ms307 ms–25.5 s30k–33k8.5–24
After, pass 28.6 ms–12 ms325 ms–697 ms42k–46k11–23
Reader CPU 1, writer CPU 0Before (main before the review)20 ms–20 ms188 ms–9.1 s17k–31k19–21
After, pass 117 ms–24 ms106 ms–11.7 s20k–37k17–20
After, pass 211 ms–21 ms113 ms–5.6 s19k–46k27–28

fio total latency, 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows, two guests per arm; ranges cover both guests and both passes of each mixed workload. Every job and every final 64 MiB CRC check passed; no storage errors. Nested TCG guests on a busy shared host: A Documents archive (tar | zstd) and a guest-removal job ran on Spark throughout; load average 6–9 on 20 cores. Only arms within this session are comparable.

Scheduler counters at the end of each arm

Independent-read admission, longest wait
8.1 s / 8.1 s
Ordinary queue head, longest wait
23.1 s / 22.0 s
Reads that passed a blocked head
3,705 / 3,616
Refused scheduler turns
1,345,363 / 1,321,083
Admitted requests
147,575 / 132,730

Per image (guest 1 / guest 2), from the daemon’s own telemetry at the end of the arm. Reads at the head of their own queue count with the ordinary heads.

What did not change. Slowest reads of 5–25 s appear in every arm, including the control. The same-CPU ones coincide with ordinary-head waits of the same length, consistent with writes and FLUSH barriers queued behind WAL capacity while compaction drained at about half its earlier rate on the loaded host; the separate-CPU ones occur with almost no bypasses and are a different mechanism. The second pass had no same-CPU read above 0.7 s. These tails are the next scheduling item.

What the numbers do not say

  • Nothing here is a research result. Every number in this repository comes from TCG guests nested in a KVM VM on one shared workstation. G1 needs a dedicated host and a raw baseline on real media; neither exists.
  • Performance is far from the hypothesis thresholds. The only CAS-versus-raw comparison, on 13 September, put fdatasync p99 at 9.76 ms against 3.62 ms and sequential reads at 437 against 842 MiB/s. The study allows 20% on p99. Whether that gap is emulation or design is unknown until native media is measured.
  • Read isolation is better, not solved. The starvation loop is gone; the FLUSH-barrier and separate-CPU tails are not. Any latency claim is meaningless with multi-second outliers.
  • The consolidated suite is stale. C5 passed on b8399ad. Four scheduler and compactor changes have merged since, each with its own focused checks but no full rerun.
  • Process weight. Each result is written four times, the design notes are unindexed, the validation log is out of order, and about 60 merged worktrees still hold cited evidence. The records are honest; their volume is consuming velocity the gates need.

Next

  1. Bound the remaining tails: FLUSH barriers behind WAL-blocked writes and separate-CPU maxima, with a traced repeat of this workload.
  2. Rerun the consolidated C5 suite on current main and close the C4/C5 allocation audit.
  3. Secure two dedicated hosts and experiment disks; measure raw XFS and passthrough latency for G1; boot the pinned ZFS profile for G3.
  4. Decide the open design items above; archive merged worktrees after moving their cited evidence; index the design notes.

Evidence and revisions

RecordWhat it establishes
15 September integration504 native tests, 13 live recovery/reset cases and three mixed-IO arms with a control on d64f5f4, merged as 4e890f2.
Comparison tablesPer-stage fio and telemetry for control and both integration passes; the script that produced them.
14 September read schedulerBypass working on 4b50554; the two-image starvation reproduced natively.
13 September C541 of 41 scenarios on b8399ad, independently verified.

Links point at 4e890f2, current main. Raw lab artifacts and receipts stay on Spark under the integration worktree; bulk VM data was removed after verification per the retention policy. Progress tracker · Validation history.