Plan
Fourteen weeks at about 320 hours, which is 23 a week against the 8 the course credit corresponds to.
The plan is sized to the work, and the descoping order defines what is removed if it slips.
Development checkpoints and research gates
The schedule below is the research plan. The architecture page records the dated C0–C5 implementation status; TODO.md is the canonical checklist. C0–C3 are accepted, the C4 functional inventory passes, and C5 has validated shared guests, caches and integrated scheduling. Pressure workloads and final memory accounting remain open.
Development checks run on Spark with disposable KVM/XFS and nested filesystem guests. G1–G6 and dedicated-media repetitions remain pending under their own acceptance conditions. Each session retains concise results and failures; large storage artifacts follow the archive and disk-headroom policy.
Hardware
The planned measurement testbed is two CloudLab c6525-100g nodes (Utah), to be reserved as a pair.
Each node has an AMD EPYC 7402P with 24 cores at 2.80 GHz, 128 GB ECC DDR4-3200, two 1.6 TB PCIe 4.0 NVMe SSDs, and a ConnectX-5 Ex 100 GbE with one port on the experiment network.
One NVMe holds the system and results, and the other is the device under test.
The pair is one hop through a single switch.
RoCE between two of these nodes works on the lossy fabric, since BPF-oF ran it there (page 04).
The testbed uses the repository's pinned NixOS host configuration. R1 adds a compatible kernel and the pinned OpenZFS userspace/module, built and boot-verified together before measurement. The host closure, locks and runtime versions are archived per cohort. The dedicated pair and its installation path remain pending.
An experiment expires after a few hours unless it is extended, so every run is scripted to complete inside one sitting.
CloudLab is free for research.
A project is opened by a faculty member and reviewed by CloudLab staff, so the sponsor opens it before Sep 9.
The fallback is two OVHcloud Advance-4 2026 servers (EPYC 4585PX, 16 cores, 64 GB DDR5 ECC, 2 × 960 GB NVMe) on a 25 Gbps private link, which loses the RDMA arm and replaces the 100 GbE fabric with 25 GbE.
Schedule
| Weeks | Build | Measure |
|---|---|---|
| 1–2 | vhost-user-blk daemon in passthrough: staging log, FLUSH, replay. Kernel and ZFS image. | R0, with the drive's read and fdatasync times; passthrough within 10% of R0 p99 (G1). Thresholds frozen. zdb -S phase 0 on the synthetic fleet. |
| 3–5 | Compactor with settle window, store, index, manifests, watermark, governor, recovery. Three chunk-size arms. | kill -9 recovery and the three ordering tests pass (G2). First capture numbers. |
| 6–7 | R1 configured, both volblocksize arms. R2 if time permits. | Page 02 table complete (G3), sweep before every capacity number. |
| 8–9 | Protocol with separate GET and PUT connections, rendezvous placement, k, segment PUT with durable ack, HAS, pins, surplus copies, sweep. Provisioning; migration with the fenced handoff. | Replicated mode on two nodes. |
| 10 | Partitioned mode. Fleet class over TCP. | Page 03 table complete (G4). |
| 11–12 | nvmet exports, RoCE configuration, busy-polling and blocking daemon, depth prefetch, profile prefetch. | Transport matrix and prefetch sweeps (G5). Partitioned boot storm. |
| 13–14 | Report; reproducibility pack (G6). |
Gates
G1. Passthrough daemon under stock QEMU within 10% of R0 p99 by the end of week 2. If this slips, everything after it slips, and the sponsor is informed that week.
G2. The recovery and ordering tests on page 01 pass before any daemon number is reported.
G3. Page 02 table complete: R0, R1 at two block sizes, R3 at three chunk sizes; latency, capture, index, amplification; variance beside every number.
G4. Page 03 table complete: both modes, every flow, each read against the bound its row names.
G5. Transport matrix complete for every non-stretch probe, null and file, memory and NVMe, with the RoCE counters printed beside every RDMA number.
G6. One command rebuilds the fleet from dated archives, and one command reruns every table on a fresh pair.
Descoping order
When the schedule slips, items come off from the top.
- ibverbs daemon arm, and with it fleet class over RDMA.
- Super-chunk placement.
- R2 dm-vdo.
- Profile prefetch (depth prefetch stays), and with it the boot-storm clause of hypothesis 3.
- Fleet class over TCP. Hypothesis 4 is then reported as untested, with the literature's numbers as the estimate.
- Partitioned mode. Replicated mode alone still gives hypothesis 2's transfer result, and the remote read of hypothesis 3 is then measured with the local copy disabled so that the read is forced to the peer.
Not removed under any slip: page 02, the nvmet TCP probe, and the daemon over TCP. The RDMA probes go if RoCE configuration exceeds its budget or the fallback hardware is used.
Risks
- Daemon overrun. The largest risk and the reason G1 is at week 2. Protocol plumbing comes from maintained crates, so the hours go to the components listed as new code on page 01.
- RoCE configuration. GID selection, MTU, adaptive retransmission on a lossy fabric. Budgeted at 8 hours. If it exceeds 20, the RDMA rows are dropped and the TCP rows stand.
- Node availability. 36 nodes of this type exist, so the pair is reserved in week 1 for every measurement week.
- Correctness debt. The defects that stall or corrupt a guest are known from a prior implementation, and each has a test on page 01 and hours in weeks 3 to 5, before any number is taken.
- O_DIRECT alignment. Final append buffers satisfy the backing filesystem's alignment requirements. Verify addresses, offsets and lengths independently of the guest block size; report payload copies and any buffered fallback.
- Known configuration pitfalls. The 100G interface stays down unless the profile declares a link on it; the ZFS pitfalls are in the R1 table on page 02.
- Census realism. Scripted drift is not real drift. The fleet is built from real dated archives, the scripts are published, and the numbers it supplies are bounds the daemon is read against, not claims about fleets in the wild.
Logistics
CS 4993, 1 credit.
Expectations in writing before Sep 9.
Thirty minutes of sponsor time every two weeks, with G1 as a scheduled meeting.
Future work
Availability.
Fleet class is the seed of replication before acknowledgment. With it and k ≥ 2 on N ≥ 3 the system has a failure model, which needs membership, failure detection, and rebalancing, none of which this study touches.
Placement and reclamation.
Super-chunk placement for locality. A cache policy that weighs a chunk's owner distance. An on-disk copy-on-read tier for chunks that are cold at their owner. Reference counts kept as derived state, with the sweep as the auditor, so an overwrite frees space at once as it does in ZFS.
The same split elsewhere.
Prefix caching in LLM serving (vLLM, SGLang, Mooncake) names cached KV blocks by a hash chain over the whole token history, so two requests share only along a common prefix. That is lineage.
The same document after two different preambles is computed twice. That is the cross-host case here, and its size on a real trace is unmeasured.