Index

How CAS changed

Update 05 · 24 September 2026

CAS shares identical disk blocks across VM images. Writes enter a per-image log. Background compaction moves durable data into shared chunks.

Under write pressure, reads stalled for seconds. The changes below separate the capacity problem, the compaction cost, and the scheduler bug.

Writes enter the log before shared storage

Each guest has its own disk image. QEMU exposes the guest's requests to CAS through shared-memory queues. Admission reserves capacity before CAS accepts a write.

Guest writes

Guest writesQEMU shares request descriptors with CAS. Admission reserves capacity before the write enters the per-image write-ahead log. FLUSH makes the ordered log durable.QEMU guestvirtual block diskCAS frontendshared descriptorsAdmissionreserve capacityImage WALappend writesrequestWRITE publishes; FLUSH persists.

Background compaction

Background compactionThe compactor hashes durable log data, stores missing chunks, and commits a manifest. Covered log data can be reclaimed only after reader and replay pins release.Durable WALfixed 4 KiB chunksShared storewrite missing hashesManifestcommit block → hashReclaim WALwait for pinssync chunksFreed space wakes admission.
The write path and the background path. FLUSH does not wait for compaction.

The write-ahead log, or WAL, holds recent writes. A WRITE completion means the write is published. An ordered FLUSH makes the log durable on the local host.

The compactor hashes durable data in 4 KiB chunks. It stores each missing chunk once. A manifest maps each image's disk blocks to chunk hashes. CAS commits that map before reclaiming the covered log data.

Reads use the newest mapping

A recently written block

A recently written blockA captured read plan first checks the write-ahead log. Recent mappings take precedence over the manifest.Read blockcaptured read planWAL overlaynewest mappingRecent data comes from the log.

A block outside the WAL overlay

A block outside the WAL overlayThe manifest maps the block to a content hash. The shared cache serves a hit; a miss loads and verifies the chunk from the store. Holes return zeros.Manifestblock → hashCacheshared chunksStoreload + verifymissOne request can read both sources.
The WAL overlay takes precedence. For older data, the manifest identifies the shared chunk.

A read can combine recent WAL data with shared chunks. Holes return zeros. CAS keeps source pins until the read completes, so reclamation cannot remove data still in use.

Durability and source

Published writes, durable WAL, and compacted mappings are separate boundaries. The manifest cannot cover writes beyond the durable WAL. Reclamation also waits for reader and replay pins. Host garbage collection is separate from ordinary background compaction.

Storage design · Read routing · Manifest publication

A full log should delay a write

The first failure was a capacity timeout. Writes filled their WAL quota faster than compaction could reclaim it. After five seconds, an admission wait became a guest IO error.

Before · a full log becomes an error

Before · a full log becomes an errorA write reaches a full WAL quota. The admission deadline expires after five seconds and returns an IO error to the guest.WriteWAL quota fullWait5 s deadlineGuest IOERRspace is still full

After · wait for capacity

After · wait for capacityThe request stays pending and its payload stays in guest RAM. WAL reclamation restores capacity and wakes admission. A periodic retry covers releases without notifications.Pending writepayload in guest RAMReclaim WALspace restoredRetry admissioncontinue writewakeno timeout error
The change removes the capacity deadline. It does not make space available sooner.

CAS now keeps the request pending. Its unadmitted payload stays in guest RAM. Reclamation wakes admission when space returns. A periodic retry covers releases without a wake notification.

The reproduced timeout errors disappeared. Writes resumed after waits of 12.66 and 13.60 seconds. Reads still took up to 11.34 seconds.

Measurement scope and source

14 September, two bounded Spark repeats. The two mixed-read maxima were 11.34 and 11.15 seconds. This fixes the reproduced capacity error, not all IO errors or maximum latency. Device and lifecycle deadlines still exist.

Capacity-wait records · Admission and retry

The compactor repeated work it could skip

Selection rescanned old log prefixes. Buffer preparation initialized the maximum capacity. Manifest output included pages superseded by later edits in the same batch.

Before · repeat old work

Before · repeat old workEach selection walks retained WAL prefixes. Preparation initializes maximum buffers and final output includes intermediate manifest pages.Scan retained WALold prefix + new recordsPrepare and emitfull buffers + intermediate pagesRepeat for the next batch.

After · keep the useful work

After · keep the useful workSelection resumes from a validated cursor. Buffers grow as needed. Final manifest output retains only reachable new pages. Reclamation still walks history.Resume selectionvalidated WAL cursorPrepare and emitgrow buffers + final pagesSame durability order.
A cursor skips processed framing. Smaller preparation and final-page emission reduce work per batch.

Selection now resumes from a validated cursor. Buffers grow as needed. The final output keeps only reachable new manifest pages.

Work per MiB compacted · Spark, 14 September
WorkBeforeAfter
Initialized bytesMiB per input MiB73.675.08
Background readscalls per input MiB740.35481.7
Final manifest pagespages per input MiB752.51103.5
Syncscalls per input MiB5.956.23

Repeated work fell. Sync calls did not. Mixed-read maxima were still 10.50 and 9.57 seconds.

These runs had different backlogs. The table counts work per MiB. It does not establish a device speedup or a read-latency improvement.

What still scans

The cursor is a validated hint with a safe scan fallback. Reclamation still walks retained history. Manifest edits can still create temporary paths in memory; final emission omits superseded pages.

Compaction records · Selection cursor

A blocked write hid reads behind it

The frontend stopped at a write that could not get WAL capacity. A later read of unrelated blocks stayed undiscovered, even though it needed no write space.

Before · discovery stops at the write

Before · discovery stops at the writeA capacity-blocked write at the guest queue head prevents discovery of an independent read behind it.Writecapacity blockedReadnot discoveredguest queueThe read needs no WAL space.It still cannot reach admission.

After · discover, then select

After · discover, then selectCAS retains a bounded set of request descriptors, then selects independent reads. It preserves overlapping-write dependencies and FLUSH barriers.Writestill pendingReadindependentRead admissionretained frontend requestspayload stays put
Discovery and admission are separate steps. Only independent reads can pass the blocked write.

The frontend now retains a bounded set of request descriptors, then selects eligible reads. It does not copy an unlimited queue of write payloads.

A read still waits behind an earlier overlapping write, zero, or discard. It also respects FLUSH barriers. This preserves the image's ordering rules.

Bypass occurred in the measured runs. Read p99 fell in the matched same-CPU comparison, but one read still took 6.834 seconds.

The remaining trace exposed another problem. A read could be discovered and eligible, yet still lose every scheduler turn.

Measurement scope and source

14 September. The comparison used matched settings in separate sessions. It is distinct from the same-session control below. After admission, reads also retain the existing image-wide publication dependency.

Read-bypass records · Discovery and ordering checks

The scheduler kept retrying the refused write

With two images, each turn could retry the same capacity-blocked write. The scheduler moved to the other image, then repeated. Eligible reads in both images waited behind those retries.

Before · the refused write stays first

Before · the refused write stays firstImage A retries a refused write, then image B does the same. On the next visit to either image, the write remains ahead of its eligible read.Writeretry firstReadeligibleWriteretry firstReadeligibleABEach visit retries the same head.

After · rotate the refused ticket

After · rotate the refused ticketA refused, uncommitted scheduler ticket moves behind other tickets in its image. The image selector retains byte-deficit fairness. Mutation publication and FLUSH order do not change.Readnext visitWriteretry laterReadnext visitWriteretry laterABScheduler order, not disk order.
A and B are images. These are scheduler tickets, not guest queue positions or write publication order.

A refused, uncommitted ticket now moves to the back of its image's list. On the next visit, an eligible read can run. The scheduler still shares service between images using byte deficits.

Same-session, same-CPU control · Spark, 15 September
MeasurementBefore rotationAfter rotation
Mixed-read p99219 to 287 ms9 to 20 ms
Independent bypass admission, max8.110 s≤ 0.137 s

Read p99 improved. Writers continued to progress. One fixed run still had a 25.529-second read.

The 137 ms bound above applies only to independent bypass admission. It is not a bound on all read latency. The remaining maximum stalls need tracing.

What this result establishes

The same-session control supports the scheduler change. It does not establish unchanged writer throughput or solve every read stall. The untraced ordinary-head maxima are consistent with FLUSH or WAL-capacity blocking, but their individual causes remain unproven.

Same-session records · Refused-ticket rotation

CloudLab removes the nested guest setup

The Spark results above came from a development fixture. Workload guests used TCG emulation inside an outer KVM VM. Other work also ran on Spark.

Spark · historical measurements

Spark · historical measurementsWorkload guests used TCG emulation inside an outer KVM VM. CAS and XFS ran in that VM on a virtual disk backed by a host file. Spark also ran other work.Inner guestTCG emulationCAS + file-backed XFSinside the outer KVM VMSpark physical hostNested development fixture

CloudLab · protocol checks

CloudLab · protocol checksThe guest runs directly under KVM. CAS runs on the physical EPYC host. Completed checks used file-backed XFS on the OS storage path. Dedicated NVMe benchmarks have not run.Guestdirect KVMCAS + file-backed XFSon the physical hostCloudLab EPYC hostDedicated NVMe is the next test.
The completed CloudLab checks change the execution path. They do not yet measure dedicated-NVMe performance.

The CloudLab host has an AMD EPYC 9354P, 32 cores, 192 GB RAM, and two 800 GB NVMe drives. CAS runs on the physical host. Guests run directly under KVM.

All three backends passed direct KVM data checks on file-backed XFS. Dedicated NVMe measurements have not started.

Raw XFS is the storage baseline. Passthrough adds the daemon path without CAS storage. CAS adds the log, sharing, and compaction. The runner uses the same guest, workload, and resource limits for all three.

Native runner settings
  • Each guest has two vCPUs, 2 GiB RAM, and one queue of 128 entries.
  • Each run has an 8 GiB cgroup limit and no swap.
  • The runner pins CPU affinity and memory policy to the storage NUMA node.
  • fio uses direct IO and a 512 MiB working set. CRC checks verify the data before and after.
  • CAS has a 16 MiB clean cache. Jobs retain cache state within a repeat. Each repeat starts with new storage.
  • Read-only jobs wait for compaction to catch up. Write completion and fdatasync latency are separate measurements.
  • The pressure case uses separate reader and writer guests. It does not test same-image FLUSH stalls.

Runner and method · Workload · a1bbdf3

The next result needs dedicated storage

  • Run raw XFS, passthrough, and CAS on the dedicated NVMe. Measure latency, physical storage, device writes, and total memory.
  • Rerun recovery on that exact revision. Complete the allocation audit.
  • Trace the remaining maximum read stalls.
  • Add a ZFS comparison before claiming a storage advantage.

Cross-host sharing, remote reads, replicated durability, and migration remain unimplemented.

Measurement records

The explanations above stand on their own. These dated panels retain the detailed Spark measurements. Compare values within each experiment.

13 September · raw, passthrough, and CAS

Median of each run’s p99 · lower is better

QEMU raw 1.52 ms
Run range 1.38–1.6 ms
Raw daemon 2.09 ms
Run range 1.7–3 ms
CAS 1.58 ms
Run range 0.96–2.8 ms

5 × 5s per case · 32 MiB file · direct IO, QD1 · caches retained between jobs

Memory observed across each lab session
BackendLab peakStorage PSS
QEMU raw2,301 MiBIn QEMU
Raw daemon2,329 MiB25 MiB
CAS2,770 MiB248 MiB

Lab peak: cgroup memory, including the outer VM. Storage PSS: sampled inner daemon peak. The columns overlap; do not add them. No cgroup OOM kills occurred.

14 September · capacity and compaction

Writes resumed when WAL space returned. All 20 fio jobs completed without IO errors. Both logs drained, and both 64 MiB seeds verified after load and a full restart. Measured on 14 September at eeff4b5.

Mixed-read latency still reached 10.5 seconds. No GC occurred. Both guests sent every request through queue 0, where writes waited for WAL space. Later reads can wait behind them; other-queue scheduling had no work to serve here. The evidence points to queue blocking, but does not time every phase of each read.

Initialized bytesMiB / input MiB
Before
73.7
After
5.1
Background readscalls / input MiB
Before
740.3
After
481.7
Final manifest pagespages / input MiB
Before
752.5
After
103.5
Syncscalls / input MiB
Before
5.9
After
6.2

Work per MiB of compacted input: 878 MiB before, 1,173 MiB after. Counts include rotation and reclamation; neither run had GC. One repetition with different backlog and timing. These are operation counts, not physical-device traffic or a statistical speedup.

Source revisions, drain observations and remaining costs

Before: bdb26fc. After: eeff4b5, including the compaction changes in PR #49. The full report retains commands, counters and comparison limits.

The sampled drain rate rose from 13.1 to 21.6 MiB/s. Initialization fell to 0.34 s, while reads took 40.20 s and syncs 31.71 s across the new progression. Reclamation still revisits retained WAL history. Measure its reads and syncs by caller, then reduce that repeated work. Within-queue read bypass needs an explicit ordering contract first.

Earlier runs: how congestion became an IO error

Removing queue expiry let an earlier repeat complete through 12.7 / 13.6-second waits with zero errors, but mixed reads reached 11.3 seconds and one GC paused the host for 3.17 seconds. The trace below predates that fix.

Earlier failure, shown below. The 256 MiB image quota stayed closed for about 8.6–13.6 seconds while the WAL drained; queued heads expired after five seconds. The fix keeps capacity waits pending and wakes admission when reclamation restores space. A separate native test also held an accepted write before submission for 31 seconds and completed it after release.

Per-image staging allocation during overloadAllocation approaches 256 MiB. The image stops admitting writes, then drains toward the 60 percent reopen threshold. Orange shading marks sampled stopped admission. The host-wide quota remains open throughout.256 MiB cap153.6 MiB reopen threshold0 s23.2 s

214.1 MiB · write quota closed · 1 blocked queue head · 1 guest IO error

Orange = write quota closed. Reserving the next 8 MiB segment can stop writes before allocation reaches 256 MiB. Samples are about 500 ms apart; a threshold crossing and a new allocation can occur between them.

Historical compactor phase totals
Read the WAL23.8%
Hash + prepare tree24.1%
Write / sync chunks13.1%
Publish manifest9.0%
Publish D + reclaim WAL30.0%

Elapsed phase time across this lab, including seeding and drain. Disk reservation was under 0.1%. Preparation combines hashing and tree edits; these are not CPU samples or a hash-only profile.

The wake-up gap now has a regression test

The original trace showed Guest 2’s quota reopen while its head stayed blocked. A deterministic test then reproduced the missing notification without unrelated IO. It passes after the fix; this does not establish the exact thread interleaving of the historical guest run.

What disk and memory controls established

With a restricted physical-space budget, GC could not reopen writes, but all 256 MiB of seeded data remained readable. Native memory-budget tests refused index growth and compaction preparation before output; a read-page allocation denial failed its image. Actual kernel OOM and unexpected filesystem ENOSPC remain untested.

Unmodified build: EIO at 1 MiB / QD32 after the earlier stages. Instrumented repeat: EIO already at 4 KiB / QD64. The stage varies with backlog; neither is a universal queue-depth limit. No physical-space pressure, GC or cgroup OOM explained these failures. Seed checks passed, and writes resumed after drain.

Experiment, controls and evidence · measured at c11e106. New instrumentation in cas-daemon: local/pressure.rs and local/host/statistics.rs.

14 September · read traces and bypass

A cached read waited 7.922 seconds behind writes. Once admitted, it finished in 1.34 ms. The long wait came before the read entered storage. 14 September instrumented probe.

Guest ringWrite waiting
WAL capacity
ReadRead

The frontend stops consuming this virtqueue at its blocked write. The queue mutex is released; the requests remain ordered behind that head. Measured queue-lock acquisition peaked at 0.187 ms. Making the queue lock-free would not change this admission rule.

Behind a blocked write7.922 s99.95% of observed time
After admission1.340 msThrough guest notification

Another 2.625 ms passed before admission. Total observed: 7926.004 ms. Timing starts when CAS sees the request in the ring; earlier guest waiting is excluded.

Inside the 1.340 ms after admission

Queue 0, request 88351, disk offset 20860928.

Frontend dispatch + command channel
0.017 ms
Reactor dispatch / write dependency
0.536 ms
Prepare + advance the read
0.036 ms
IO submission scheduler
0.001 ms
Manifest-page IO
0.709 ms
Other reactor time
0.006 ms
Response channel + guest completion
0.035 ms

All six selected tail reads hit the chunk cache: no WAL or chunk-payload IO, and no shared-fetch wait. IO timing includes kernel scheduling and completion reaping. These examples do not establish cold-read latency.

Another CPU helps; it does not reserve a read queue

WorkloadRead maxCompletion p99
Read only65–67 ms2.54–2.61 ms
Reader + writer, same CPU2.71–7.93 s325–2,869 ms
Reader CPU 1, writer CPU 022 ms–2.21 s4.62–6.46 ms

Each guest mapped CPU 0 to queue 0 and CPU 1 to queue 1. A write still reached queue 1 in each separate-CPU pass, holding reads there for about 2.2 seconds. The trace confirms the write and queue; its guest thread or filesystem source remains unmeasured.

Ranges cover both guests and two mixed passes. 4 KiB/QD8 reads, 1 MiB/QD32 writes, ten seconds per job; caches enabled, WALs drained between passes. Two-vCPU TCG guests on Spark. All 24 fio jobs and restart seed checks passed; no GC or OOM occurred. Test disks and keys were removed.

Trace coverage, source paths and remaining measurements

This probe captured 385,700 of 385,860 reads. The follow-up fixes the cursor/peek observation race and passes 477 native tests; it was not rerun in a VM. Only the 32 slowest completed reads per image are retained. Status snapshots can lag stage boundaries, and tracing overhead is unmeasured.

Next: define bounded read progress past a waiting write, preserving overlap, FLUSH and recovery rules. Cold reads, read-credit waits, telemetry freshness and long-history GC still need measurement.

The original long read wait was at admission. One matched read took 7.938 seconds in fio: 7.933 seconds passed before it reached the head of its CAS queue, then 3.53 ms after admission. The reader and the blocked write touched different disk ranges.

1Guest
fio → Linux bio
2Virtqueue
shared descriptors
3CAS admission
wait behind a write
4Reactor → storage
cache, WAL or chunks
5Guest completion
used ring → fio

In these measurements, the frontend released the queue mutex while waiting, but stopped consuming the queue at its blocked write. Later reads could not reach their own read credits. A lock-free queue would preserve that obstruction unless its admission rule also changed.

fio total7937.964 ms
Before CAS queue head7932.921 ms
After CAS admission3.527 ms

Amber: before the CAS queue head. Blue: the rest of fio latency. Queue 0, disk offset 19038208. Selected slow IO, not a percentile sample; guest and CAS clocks are joined by request identity and only durations are compared.

Where the remaining time went
Guest bio submission → virtqueue publication
0.033 ms
fio time outside bio submission → block completion
1.054 ms
Virtqueue/completion time outside the CAS interval
0.428 ms
At the CAS queue head
0.001 ms
Frontend dispatch + command channel
0.018 ms
Reactor dispatch / publication dependency
2.322 ms
Read execution
1.165 ms
Response + guest notification
0.022 ms

The outside intervals combine work before and after an inner interval; they do not split one-way transport time. IO completion includes scheduling and reaping, not only device service. Full phase distributions, coverage and method.

Another CPU helps; the journal still uses its queue

Both separate-CPU probes found an 8 KiB journal write from jbd2 blocking unrelated reads on queue 1 for 1.4–2.4 seconds. The reactor already separates reads from commands. The obstacle is earlier, where the frontend discovers requests.

Guest workloadRead p99Slowest read
Read only2.57–2.61 ms56.57–56.91 ms
Reader + writer on CPU 092.8 ms–2.67 s2.99–9.98 s
Reader CPU 1, writer CPU 04.82–6.39 ms21.84–64.77 ms

fio total latency with CAS tracing and guest probes off. Ranges cover both guests; mixed workloads have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten seconds per job. These are nested-VM development measurements, not native-drive latency or a latency guarantee.

Cold read-only p99 was 3.69–4.18 ms; warm p99 was 1.97–2.02 ms. Warm mixed reads still reached p99 of 2.1–2.3 seconds with zero cache misses. Faster payload reads alone do not remove this queueing delay.

The design this led to

PR #53 now retains bounded descriptor snapshots before storage admission. Independent reads can pass blocked writes; overlapping reads and FLUSH barriers keep their order. Waiting write payloads stay in guest RAM. The existing reactor and compactor still execute the work.

These timing results predate the scheduler change. The live comparison above tests the merged implementation. Implementation, crate paths and limitations.

The cost is a fixed descriptor metadata quota and explicit dependency/recovery state. The new retained-carrier format also requires a fresh attachment when upgrading an existing guest. It also cannot help a read that the guest has not published. Cache sharding and thread pinning remain separate options for measured downstream costs. Alternatives, pros and cons, and required acceptance tests.

What was measured, and where to read the code

Instrumentation covers the guest bio/request path, queue admission reasons, publication dependencies, IO stages, cache lock waits and telemetry freshness. Histograms cover all traced completions; a bounded recent tail supplies examples. Guest probes measurably increase latency, so the table uses unprobed runs. The report also compares process-cold and warm caches.

No scheduling policy changed in this experiment. Sustained native-host performance, full guest-ring exhaustion, remote reads, GC pressure and actual OOM/ENOSPC remain separate gates. The disposable labs were stopped and their disks and keys removed.

Reads can bypass writes, but admission fairness still stalls them. The longest untraced read still took 6.83 s.

WorkloadEarlier read p99Merged read p99Earlier slowestMerged slowest
Read only2.57 ms–2.61 ms2.24 ms–2.28 ms56.91 ms10.34 ms
Reader + writer on CPU 092.80 ms–2.67 s24.51 ms–156.24 ms9.98 s6.83 s
Reader CPU 1, writer CPU 04.82 ms–6.39 ms5.60 ms–6.39 ms64.77 ms3.06 s

Untraced fio total latency. Each p99 range covers two guests; mixed cases have two passes. 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows. Matched workloads ran in separate Spark sessions using nested TCG guests; these are development measurements.

The admission rule

Published firstWrite AWaiting for WAL space
Published laterRead BEligible for read admission

Read B can use its own read credits while Write A waits for WAL capacity. The write payload remains in guest RAM.

A and B stand for different disk ranges. CAS owns bounded descriptor snapshots before admission; it does not need a new guest queue, a lock-free queue or a new executor. After admission, the existing image-wide write-publication dependency still applies.

What the instrumentation says now

One matched 6.295-second read spent 6.293 seconds before admission and only 0.57 ms afterward. Its chunk was cached. The fairness scheduler kept refusing its turn; a separate Rust reproduction confirms that two blocked writers can repeatedly displace eligible reads without any disk IO.

Each image retries its oldest write before its eligible read. That retry marks the write ready again, which can displace the read the shared scheduler just selected. Two images can keep waking each other and repeating this until write capacity returns.

Next: preserve eligible-read selection across blocked-write retries, then repeat this workload. Overlap, FLUSH and finite-capacity rules still matter; this result does not justify a new cache or a lock-free queue.

Validation, resource costs and code

The packaged merge passed 499 native tests. All 13 live recovery/reset scenarios and 8 full 64 MiB seed checks passed. The recovery cases kept QEMU alive across backend replacement; they did not reproduce an overloaded deferred-read crash. Native frontend tests cover that ownership case.

The largest sampled lab memory peak was 3.93 GiB under a 6 GiB cap, with zero observed cgroup OOM or limit events. Disposable disks and keys were removed. Actual OOM/ENOSPC, sustained fairness, full guest-ring exhaustion, long-history GC and physical power loss remain separate gates.

Each frontend reserves about 6 MiB of descriptor metadata quota, with eight read-request slots separate from writes. The existing byte allowance fits seven simultaneous 4 KiB reads. Version-3 retained carriers require a fresh guest attachment when upgrading from version 2.

All four owners are in crates/cas/daemon. Raw-result analysis, per-job numbers and limitations.

15 September · scheduler control

Independent reads waited up to 8.1 s in admission before the fix and at most 0.137 s after it. Three back-to-back runs in one host session, with the same workload.

WorkloadArmRead p99Slowest readReads per 10 sWrite MiB/s
Read onlyBefore (main before the review)5.5 ms282 ms–284 ms68k–68k—
After, pass 14.9 ms–5.1 ms227 ms–228 ms64k–67k—
After, pass 24.0 ms–4.1 ms117 ms–120 ms75k–75k—
Reader and writer on CPU 0Before (main before the review)219 ms–287 ms8.1 s–24.1 s2.3k–5.2k8.8–21
After, pass 111 ms–20 ms307 ms–25.5 s30k–33k8.5–24
After, pass 28.6 ms–12 ms325 ms–697 ms42k–46k11–23
Reader CPU 1, writer CPU 0Before (main before the review)20 ms–20 ms188 ms–9.1 s17k–31k19–21
After, pass 117 ms–24 ms106 ms–11.7 s20k–37k17–20
After, pass 211 ms–21 ms113 ms–5.6 s19k–46k27–28

fio total latency, 4 KiB/QD8 reads and 1 MiB/QD32 writes, ten-second windows, two guests per arm; ranges cover both guests and both passes of each mixed workload. Every job and every final 64 MiB CRC check passed; no storage errors. Nested TCG guests on a busy shared host: A Documents archive (tar | zstd) and a guest-removal job ran on Spark throughout; load average 6–9 on 20 cores. Only arms within this session are comparable.

Scheduler counters at the end of each arm

Independent-read admission, longest wait
8.1 s / 8.1 s
Ordinary queue head, longest wait
23.1 s / 22.0 s
Reads that passed a blocked head
3,705 / 3,616
Refused scheduler turns
1,345,363 / 1,321,083
Admitted requests
147,575 / 132,730

Per image (guest 1 / guest 2), from the daemon’s own telemetry at the end of the arm. Reads at the head of their own queue count with the ordinary heads.