What a read actually costs: measuring Swarmfile on a real 8 GB file
Our homepage makes one claim everything else rests on: opening part of a huge file should cost what you read, not what the file weighs. Claims like that are cheap, so we measured it: on our staging service, with real machines, and with the method written down so you can check our working. Along the way we found, and fixed, a second of pure waste on every cold read, and a local network that wasn't being used as one. This post has the numbers, the fix, and the limits.
The setup#
- Service: Swarmfile's staging environment, the same code as production, on Cloudflare Workers with block storage in R2.
- Writer: a Windows 11 x64 desktop. It creates the test file on its Swarmfile drive and uploads it.
- Reader: a Mac (Apple silicon, macOS 26) on the same office network, which has never seen the file.
- The file: pseudo-random bytes, so nothing compresses or deduplicates. We ran it at 1 GiB and 8 GiB.
- Timing: each read is a single
ddcall on the mounted drive, and the time isdd's own figure: the reader's local clock, not ours. - Bytes moved: read from the reader engine's own fetch counters, split by where the bytes came from: the cloud, or another machine.
The reads, in order:
- List the file. Browsing should be metadata only.
- Peer phase (writer still online): a cold 64 KiB read, a cold 1 MiB read from the middle, and 15 × 1 MiB reads spread across the file, roughly what scrubbing a timeline does.
- Cloud phase: stop the writer's engine, then repeat the same three reads at different, never-read offsets. Every byte now has to come from the cloud.
- Warm re-read of something already read.
The results#
Bytes moved don't follow the file's size. Cloud-only reads, before any fix:
| Read | 1 GiB file | 8 GiB file |
|---|---|---|
| List the file | 0 | 0 |
| Cold 64 KiB | 1.2 MiB | 2.5 MiB |
| Cold 1 MiB from the middle | 3.2 MiB | 2.3 MiB |
| Scrub, 15 × 1 MiB | 35 MiB | 33 MiB |
| Warm re-read | 0 | 0 |
Time, 8 GiB file, every byte from the cloud, before and after the fix described below:
| Read | Before | After |
|---|---|---|
| Cold 64 KiB | 2.2 s | 0.96 s |
| Cold 1 MiB from the middle | 0.75 s | 0.49 s |
| Scrub, 15 × 1 MiB | 26.9 s | 10.8 s |
| Warm re-read | < 1 ms | < 1 ms |
When the writer was still online and served the bytes itself over the office network, the same reads took 0.1-0.2 s each at both file sizes (a 15-read scrub: 2.0-2.5 s), and moved about the same 2-4 MiB per read.
Two things hold at every size:
- A read moves roughly what it reads. Every cold read moved 1-4 MiB, whether the file was 1 GiB or 8 GiB. It's a little more than the read itself because Swarmfile fetches whole chunks (about 1 MiB on average, content-defined), so a 1 MiB read that straddles two chunks pulls both.
- Browsing is free and re-reading is instant. Listing moved nothing; a warm re-read came from local cache in under a millisecond.
The second we were wasting#
The first runs had an odd shape: cloud reads got slower at 8 GiB (1.8 s per scrub step, against 0.8 s at 1 GiB) even though they moved the same bytes. File size had nothing to do with it. When we stopped the writer to force cloud-only reads, the reader still had it on its list of peers, and on every block it tried that peer first, waited a full second for an answer that never came, and only then asked the cloud. A peer that had answered recently only ever earned a two-second timeout, so the reader kept trying it.
That's not a lab quirk. It's what happens when a teammate closes their laptop. So we fixed it:
- a peer that stops answering is now skipped after a failure or two, backing off further each time;
- the wait before asking the cloud adapts to how fast your peers normally answer, instead of a flat second;
- the chunks a single read needs are fetched at the same time, not one after another.
The "After" column above is the same 8 GiB test with those fixes (and the ones in the next section): the first cold read from the cloud went from 2.2 s to under a second, and scrubbing from 1.8 s per seek to 0.72 s.
The local network wasn't local#
Our first runs counted every byte the reader got from the writer as coming from a remote peer, even though the two machines sat on the same office network. Chasing that down found a bigger problem: local discovery (the mDNS announcements machines use to find each other on a network) wasn't announcing anything. Each machine registered itself without its network addresses, so there was nothing to announce, and machines only ever found each other through our cloud service. Transfers still went machine-to-machine, but not as local traffic, and not without the cloud's help.
We fixed the announcement, and made discovery use only real local network adapters, not VPNs, virtual machines or container bridges. The Windows installer now adds a firewall rule for Swarmfile's local traffic (home and work networks only). Once local discovery worked, one more problem showed up: a teammate's machine that had just gone offline was now first in line for every read, and each read waited on it before asking the cloud. A machine that stops answering is now dropped after its first failed attempt, and the cloud is asked within 0.3 s.
In the final run, every byte the reader fetched from the writer came over the local network, and cloud reads with the writer gone were as fast as they'd ever been.
The limits, plainly#
- Tested to 8 GiB. Each file's list of chunks (its manifest) grows by roughly 0.15-0.2 MiB per GiB, and the first read of a file fetches it. Somewhere around 20-30 GB a file switches to a segmented manifest, so a read fetches only the small piece of the list it needs. That code path is covered by tests but we haven't benchmarked it on real machines yet. Until we have, we won't quote "opens in seconds" for single files above about 20 GB.
- One run per size. These are single runs, not averages. Cloud latency in particular varies run to run.
- Cloud seeking isn't local seeking. About 0.7 seconds per seek from the cloud is fine for opening files and jumping around a project; it isn't smooth timeline scrubbing. Scrubbing is snappy when a teammate's machine already has the data (0.1-0.2 s), which is the case the office-first design is built for.
- Local discovery needs a network that allows it. Machines find each other with multicast. Some corporate networks and guest Wi-Fi block it, and a machine's own firewall can too. Swarmfile's doctor now warns when it sees teammates in the same office through the cloud but never on the local network, and a setting (
lan_from_office) lets an office skip multicast entirely. - Our lab, not yours. The harness that ran this needs our lab machines and staging credentials, so you can't rerun it as-is. The core measurement doesn't need our lab, though: public projects on Explore can be streamed anonymously with ordinary HTTP range requests, so reading a range of a big public file and counting the bytes is something anyone can do with
curl.
What's next#
Reading straight through a file (playing a video, copying it) now fetches ahead in the background, so it doesn't pay a round trip per chunk. Seeking to a new spot still costs about 0.5-0.7 s from the cloud, and most of that is two round trips per block: one to our service, which checks access and hands back a signed link, and one to storage. Getting signed links for many blocks in a single request is next. When it lands, we'll rerun this and post the numbers, including the ones that don't flatter us.