← index
page 05 / 06

Plan

Fourteen weeks at about 320 hours, which is 23 a week against the 8 the course credit corresponds to.
The plan is sized to the work, and the descoping order defines what is removed if it slips.

Development checkpoints and research gates

The schedule below is the research plan. The architecture page records the dated C0–C5 implementation status; TODO.md is the canonical checklist. C0–C3 are accepted, the C4 functional inventory passes, and C5 has validated shared guests, caches and integrated scheduling. Pressure workloads and final memory accounting remain open.

Development checks run on Spark with disposable KVM/XFS and nested filesystem guests. G1–G6 and dedicated-media repetitions remain pending under their own acceptance conditions. Each session retains concise results and failures; large storage artifacts follow the archive and disk-headroom policy.

Hardware

The planned measurement testbed is two CloudLab c6525-100g nodes (Utah), to be reserved as a pair.
Each node has an AMD EPYC 7402P with 24 cores at 2.80 GHz, 128 GB ECC DDR4-3200, two 1.6 TB PCIe 4.0 NVMe SSDs, and a ConnectX-5 Ex 100 GbE with one port on the experiment network.
One NVMe holds the system and results, and the other is the device under test.
The pair is one hop through a single switch.

RoCE between two of these nodes works on the lossy fabric, since BPF-oF ran it there (page 04).
The testbed uses the repository's pinned NixOS host configuration. R1 adds a compatible kernel and the pinned OpenZFS userspace/module, built and boot-verified together before measurement. The host closure, locks and runtime versions are archived per cohort. The dedicated pair and its installation path remain pending.
An experiment expires after a few hours unless it is extended, so every run is scripted to complete inside one sitting.

CloudLab is free for research.
A project is opened by a faculty member and reviewed by CloudLab staff, so the sponsor opens it before Sep 9.
The fallback is two OVHcloud Advance-4 2026 servers (EPYC 4585PX, 16 cores, 64 GB DDR5 ECC, 2 × 960 GB NVMe) on a 25 Gbps private link, which loses the RDMA arm and replaces the 100 GbE fabric with 25 GbE.

Schedule

WeeksBuildMeasure
1–2vhost-user-blk daemon in passthrough: staging log, FLUSH, replay. Kernel and ZFS image.R0, with the drive's read and fdatasync times; passthrough within 10% of R0 p99 (G1). Thresholds frozen. zdb -S phase 0 on the synthetic fleet.
3–5Compactor with settle window, store, index, manifests, watermark, governor, recovery. Three chunk-size arms.kill -9 recovery and the three ordering tests pass (G2). First capture numbers.
6–7R1 configured, both volblocksize arms. R2 if time permits.Page 02 table complete (G3), sweep before every capacity number.
8–9Protocol with separate GET and PUT connections, rendezvous placement, k, segment PUT with durable ack, HAS, pins, surplus copies, sweep. Provisioning; migration with the fenced handoff.Replicated mode on two nodes.
10Partitioned mode. Fleet class over TCP.Page 03 table complete (G4).
11–12nvmet exports, RoCE configuration, busy-polling and blocking daemon, depth prefetch, profile prefetch.Transport matrix and prefetch sweeps (G5). Partitioned boot storm.
13–14Report; reproducibility pack (G6).

Gates

G1. Passthrough daemon under stock QEMU within 10% of R0 p99 by the end of week 2. If this slips, everything after it slips, and the sponsor is informed that week.

G2. The recovery and ordering tests on page 01 pass before any daemon number is reported.

G3. Page 02 table complete: R0, R1 at two block sizes, R3 at three chunk sizes; latency, capture, index, amplification; variance beside every number.

G4. Page 03 table complete: both modes, every flow, each read against the bound its row names.

G5. Transport matrix complete for every non-stretch probe, null and file, memory and NVMe, with the RoCE counters printed beside every RDMA number.

G6. One command rebuilds the fleet from dated archives, and one command reruns every table on a fresh pair.

Descoping order

When the schedule slips, items come off from the top.

  1. ibverbs daemon arm, and with it fleet class over RDMA.
  2. Super-chunk placement.
  3. R2 dm-vdo.
  4. Profile prefetch (depth prefetch stays), and with it the boot-storm clause of hypothesis 3.
  5. Fleet class over TCP. Hypothesis 4 is then reported as untested, with the literature's numbers as the estimate.
  6. Partitioned mode. Replicated mode alone still gives hypothesis 2's transfer result, and the remote read of hypothesis 3 is then measured with the local copy disabled so that the read is forced to the peer.

Not removed under any slip: page 02, the nvmet TCP probe, and the daemon over TCP. The RDMA probes go if RoCE configuration exceeds its budget or the fallback hardware is used.

Risks

Logistics

CS 4993, 1 credit.
Expectations in writing before Sep 9.
Thirty minutes of sponsor time every two weeks, with G1 as a scheduled meeting.

Future work

Availability.
Fleet class is the seed of replication before acknowledgment. With it and k ≥ 2 on N ≥ 3 the system has a failure model, which needs membership, failure detection, and rebalancing, none of which this study touches.

Placement and reclamation.
Super-chunk placement for locality. A cache policy that weighs a chunk's owner distance. An on-disk copy-on-read tier for chunks that are cold at their owner. Reference counts kept as derived state, with the sweep as the auditor, so an overwrite frees space at once as it does in ZFS.

The same split elsewhere.
Prefix caching in LLM serving (vLLM, SGLang, Mooncake) names cached KV blocks by a hash chain over the whole token history, so two requests share only along a common prefix. That is lineage.
The same document after two different preambles is computed twice. That is the cross-host case here, and its size on a real trace is unmeasured.