Remote read
Page 04 measures the one place the network enters guest latency in local class, a cold read whose chunk lives on another host, then how much of it prefetch removes, then the round trip fleet class adds to FLUSH.
Where the time goes in a remote read
A 4 KiB random read at QD1 from an enterprise NVMe SSD completes in about 80 µs: 81.6 µs on a PM1725 in Systor '17 and about 80 µs on a PM1735 in blk-switch. R0 measures the testbed's drive in week 1.
On 100 GbE the transport sits on top.
Against a null device, kernel nvme-rdma added 12.1 µs and kernel nvme-tcp 21.4 µs for a 4 KiB read on ConnectX-5 (SPDK 24.05). A raw RDMA round trip is 2 to 5 µs, from 32 bytes to 4 KiB (eRPC; SPDK). On two CloudLab c6525-100g nodes, the testbed's node type, BPF-oF measured average round trips of 18 µs over nvme-rdma and 30 µs over nvme-tcp on kernel 5.12.
A userspace daemon over kernel TCP has no published measurement as a remote read target. From the kernel TCP round-trip floor of 13 to 23 µs (Homa; Zuo et al.) plus a file read, we estimate 20 to 30 µs when it polls and more when it sleeps.
The testbed replaces every one of these figures.
On these figures the difference between RDMA and TCP is about 9 µs on a read of about 100 µs.
The larger factor, about 4x, is whether the chunk is in the owner's memory or on its disk.
If the figures hold, a chunk from a peer's memory over TCP arrives before one from local NVMe, and hypothesis 3 tests this.
Every host sends its reads of a chunk to the same k owners, so we predict a chunk read by many guests is hot at its owner, and a remote read in that case is the memory row. The owner's hit rate is reported under the boot storm below.
A caution from a prior implementation by the author: a peer round trip over QUIC with TLS on a bonded 25 GbE link measured 108 µs at p50 and 257 µs at p99 (unpublished). The daemon here uses kernel TCP with TCP_NODELAY on 100 GbE.
Transport probes
The architecture's transport is the daemon over kernel TCP.
The other rows exist to show what the kernel stack and the userspace hop each cost.
| Probe | What it isolates | Code |
|---|---|---|
ib_read_lat -s 4096 | the hardware floor | none |
| nvme-rdma export | kernel block path over RDMA; owner's store exported by nvmet as a file-backed namespace, buffered_io on for memory, off for media | configuration |
| nvme-tcp export | same over kernel TCP | configuration |
| daemon, TCP, busy-polling | the architecture, without the wakeup | the daemon |
| daemon, TCP, blocking | the architecture as deployed; the scheduler wakeup is the cost | the daemon |
| daemon, ibverbs two-sidedstretch | the userspace hop without the kernel stack | ~40 h |
The nvmet export is a probe and not the architecture. It exposes the raw store, needs the reader to know offsets, and has no place for authentication.
It is in the table because the difference between it and the daemon over the same TCP is the cost of the userspace hop, with the delta between SPDK's userspace target and the kernel's, 1 µs on TCP and 2.7 µs on RDMA in the 24.05 reports, as the reference point.
Method
- Same two hosts, NIC, drive, and kernel for every row. Kernel, firmware, MTU, IRQ affinity, interrupt moderation, C-states, busy-poll, and PFC state recorded.
- The link is measured before any remote number:
ib_read_latfor the RDMA floor and a TCP ping-pong for the kernel floor, both recorded beside the rows. - Two targets per row: a null device for fabric plus stack alone, and the real file for end to end. Each from the owner's memory and from its NVMe.
- Two load states for the file rows: quiet, and with
PUTtraffic running on its own connection at the ship rate from page 03, because a cold read in deployment competes with compaction. The difference is what the read-priority rule on page 01 buys. - 4 KiB, 16 KiB, 64 KiB. p50, p99, p99.9. Five runs of 30 s, caches dropped between, medians with spread.
- QD sweep 1, 4, 16, 64 for throughput and CPU per IOPS on both ends. Kernel TCP costs 2.5 to 3x the CPU of RDMA at equal IOPS in the SPDK 24.05 reports and in i10, and the ratio measured here is reported.
- RoCE hardware counters (
out_of_sequence,packet_seq_err,local_ack_timeout_err) printed beside every RDMA number, showing zero retransmits on a fabric with no PFC.
Prefetch
The depth sweep reads sequentially through the manifest with P chunks in flight, P doubling from 1 to 512 at 4 KiB and from 1 to 32 at 64 KiB.
The bandwidth-delay point is about 250 KB for the fabric (100 Gb/s × 20 µs) and about 1.2 MB with media under it (100 Gb/s × 100 µs), so we predict that about 20 chunks of 64 KiB or 300 of 4 KiB in flight hide the remote read.
Success is remote sequential throughput within 10% of local, the throughput clause of hypothesis 3.
Profile prefetch records the chunk sequence of one boot and replays it on later boots.
DADI, REAP, FaaSnap, VMTorrent, and Nydus each prefetch a recorded access profile (page 06).
It is budgeted at one day.
Under a guest workload
Partitioned boot storm at n = 16, with and without profile prefetch, against the same storm in replicated mode.
The run reports guest p99, host device reads per guest byte, the fraction of reads served by the peer, and the fraction of those the peer answered from memory, so the per-read cost and the miss rate can be multiplied and the hotness prediction on page 00 is checked.
The gap between partitioned with prefetch and replicated is the residual cost of one copy per chunk.
The FLUSH round trip
Fleet class puts one round trip and one remote fdatasync in front of every FLUSH acknowledgment. On the figures above the transport is about a tenth of a cold read over RDMA and a fifth over TCP, because 80 µs of media sits beneath the read. A FLUSH has no media to hide behind, so we predict the round trip and the peer's fdatasync are most of its cost, and hypothesis 4 bounds it.
It is measured here with the same discipline as the read rows: write p99 at QD1 for local class, for fleet class over the daemon on TCP, and for fleet class over ibverbs if that arm lands, with the peer's fdatasync time reported separately so the transport's share is visible.
RDMA on this testbed
The CloudLab fabric is lossy. No PFC or ECN is documented on the shared switches, and the one published RoCE measurement on this node type does not say whether either was on.
Adaptive retransmission is enabled on the NIC and the counters above show whether the runs were clean.
ConnectX-5 cannot do io_uring zero-copy receive, so that option is unavailable.