← index
page 04 / 06

Remote read

Page 04 measures the one place the network enters guest latency in local class, a cold read whose chunk lives on another host, then how much of it prefetch removes, then the round trip fleet class adds to FLUSH.

Where the time goes in a remote read

A 4 KiB random read at QD1 from an enterprise NVMe SSD completes in about 80 µs: 81.6 µs on a PM1725 in Systor '17 and about 80 µs on a PM1735 in blk-switch. R0 measures the testbed's drive in week 1.
On 100 GbE the transport sits on top.
Against a null device, kernel nvme-rdma added 12.1 µs and kernel nvme-tcp 21.4 µs for a 4 KiB read on ConnectX-5 (SPDK 24.05). A raw RDMA round trip is 2 to 5 µs, from 32 bytes to 4 KiB (eRPC; SPDK). On two CloudLab c6525-100g nodes, the testbed's node type, BPF-oF measured average round trips of 18 µs over nvme-rdma and 30 µs over nvme-tcp on kernel 5.12.
A userspace daemon over kernel TCP has no published measurement as a remote read target. From the kernel TCP round-trip floor of 13 to 23 µs (Homa; Zuo et al.) plus a file read, we estimate 20 to 30 µs when it polls and more when it sleeps.
The testbed replaces every one of these figures.

On these figures the difference between RDMA and TCP is about 9 µs on a read of about 100 µs.
The larger factor, about 4x, is whether the chunk is in the owner's memory or on its disk.
If the figures hold, a chunk from a peer's memory over TCP arrives before one from local NVMe, and hypothesis 3 tests this.
Every host sends its reads of a chunk to the same k owners, so we predict a chunk read by many guests is hot at its owner, and a remote read in that case is the memory row. The owner's hit rate is reported under the boot storm below.

A caution from a prior implementation by the author: a peer round trip over QUIC with TLS on a bonded 25 GbE link measured 108 µs at p50 and 257 µs at p99 (unpublished). The daemon here uses kernel TCP with TCP_NODELAY on 100 GbE.

local NVMe read ≈ 80 µs null target over nvme-rdma, stack alone ≈ 12 µs null target over nvme-tcp, stack alone ≈ 21 µs peer memory, daemon on TCP, estimate 20 to 30 µs shorter than the local NVMe barthe case hash placement makes common peer NVMe over RDMA ≈ 92 µs peer NVMe over nvme-tcp ≈ 101 µs peer NVMe, daemon on TCP ≈ 110 µs 15 to 40% over localthe cold case; prefetch is measured against it
Literature values for one 4 KiB read at QD1, in microseconds, before the testbed measures them. The daemon rows are estimates.

Transport probes

The architecture's transport is the daemon over kernel TCP.
The other rows exist to show what the kernel stack and the userspace hop each cost.

ProbeWhat it isolatesCode
ib_read_lat -s 4096the hardware floornone
nvme-rdma exportkernel block path over RDMA; owner's store exported by nvmet as a file-backed namespace, buffered_io on for memory, off for mediaconfiguration
nvme-tcp exportsame over kernel TCPconfiguration
daemon, TCP, busy-pollingthe architecture, without the wakeupthe daemon
daemon, TCP, blockingthe architecture as deployed; the scheduler wakeup is the costthe daemon
daemon, ibverbs two-sidedstretchthe userspace hop without the kernel stack~40 h

The nvmet export is a probe and not the architecture. It exposes the raw store, needs the reader to know offsets, and has no place for authentication.
It is in the table because the difference between it and the daemon over the same TCP is the cost of the userspace hop, with the delta between SPDK's userspace target and the kernel's, 1 µs on TCP and 2.7 µs on RDMA in the 24.05 reports, as the reference point.

Method

Prefetch

The depth sweep reads sequentially through the manifest with P chunks in flight, P doubling from 1 to 512 at 4 KiB and from 1 to 32 at 64 KiB.
The bandwidth-delay point is about 250 KB for the fabric (100 Gb/s × 20 µs) and about 1.2 MB with media under it (100 Gb/s × 100 µs), so we predict that about 20 chunks of 64 KiB or 300 of 4 KiB in flight hide the remote read.
Success is remote sequential throughput within 10% of local, the throughput clause of hypothesis 3.

Profile prefetch records the chunk sequence of one boot and replays it on later boots.
DADI, REAP, FaaSnap, VMTorrent, and Nydus each prefetch a recorded access profile (page 06).
It is budgeted at one day.

Under a guest workload

Partitioned boot storm at n = 16, with and without profile prefetch, against the same storm in replicated mode.
The run reports guest p99, host device reads per guest byte, the fraction of reads served by the peer, and the fraction of those the peer answered from memory, so the per-read cost and the miss rate can be multiplied and the hotness prediction on page 00 is checked.
The gap between partitioned with prefetch and replicated is the residual cost of one copy per chunk.

The FLUSH round trip

Fleet class puts one round trip and one remote fdatasync in front of every FLUSH acknowledgment. On the figures above the transport is about a tenth of a cold read over RDMA and a fifth over TCP, because 80 µs of media sits beneath the read. A FLUSH has no media to hide behind, so we predict the round trip and the peer's fdatasync are most of its cost, and hypothesis 4 bounds it.
It is measured here with the same discipline as the read rows: write p99 at QD1 for local class, for fleet class over the daemon on TCP, and for fleet class over ibverbs if that arm lands, with the peer's fdatasync time reported separately so the transport's share is visible.

RDMA on this testbed

The CloudLab fabric is lossy. No PFC or ECN is documented on the shared switches, and the one published RoCE measurement on this node type does not say whether either was on.
Adaptive retransmission is enabled on the NIC and the counters above show whether the runs were clean.
ConnectX-5 cannot do io_uring zero-copy receive, so that option is unavailable.