Blog / Losing a machine shouldn't cost you a shard

Losing a machine shouldn't cost you a shard

Swarmfile erasure-codes file content when it's worth it (files of at least 64 KiB, when there are peers to hold the shards) and keeps the shards alongside the full chunk, not instead of it. Each coded chunk is split into 10 data shards plus 4 parity shards using Reed-Solomon 10+4, and any 10 of those 14 reconstruct the original data. That much is a fairly standard resilience technique. The more interesting question is where those 14 shards actually live, and that has two different answers depending on whether you've turned on deliberate placement.

"Adaptive" means it turns itself off when it isn't needed#

Erasure coding isn't unconditionally on. The engine only encodes a file when it decides the redundancy is actually worth the extra shard traffic, and the trigger is a specific, literal threshold: 2 or more LAN peers present. Below that, the file is small (under 64 KiB) doesn't bother, and the exact behavior branches on why you'd want EC at all:

  • Fewer than 2 LAN peers. This is the pure WAN case. Erasure coding here is about masking WAN latency, not machine loss, so it kicks in once there are 2+ WAN or seed peers to spread shards across. A more precise variant of this same check uses measured round-trip time directly (a p50 RTT threshold of 80ms) when the engine has enough RTT data, falling back to the simpler peer-count rule when it doesn't.
  • 2+ LAN peers, deliberate placement off. The engine backs off entirely. The reasoning: LAN peers already cache blocks opportunistically (see the LAN-first P2P post), so redundancy is already a natural side effect of normal reads, so spending extra bandwidth to erasure-code on top of that buys little.
  • 2+ LAN peers, deliberate placement on. This is the one case where a LAN-heavy office actually wants EC the most, not the least, because now the office isn't relying on "whoever happened to fetch it," it's placing shards on purpose. Which is the other half of this post.

Deliberate placement: real rendezvous hashing, not a label#

Opportunistic caching means a shard's location is an accident of who read it first. Deliberate placement (ec_lan_placement) makes it a decision: given an office roster, every shard is assigned to specific machines by rendezvous hashing (also called highest-random-weight hashing). It's a real implementation, not a marketing gloss over a generic hash. For a given shard and a candidate peer, the engine computes a score from the first 8 bytes of BLAKE3(shard_cid || peer_id); each shard goes to the k peers with the highest score against it, ties broken on peer-id byte order for full determinism. k is the replication factor (3 by default), so each of the 14 shards in a chunk gets 3 specific, named machines as its home, not 3 copies of the whole file.

Rendezvous hashing has one property that makes it worth using over something simpler like a modulo assignment: adding or removing a peer reshuffles only the shards that peer is actually involved in. Add one machine to a 21-peer office running replication-3, and only about 14% of shards move, matching the expected K/(N+1) ratio almost exactly, instead of a naive scheme that could reshuffle everything on every roster change.

Keeping placement honest without babysitting it#

Two things drive reconciliation, running side by side: a reactive pass that fires when the office's peer roster actually changes (debounced: the same roster has to be observed twice in a row before the engine acts on it, so one flaky mDNS blip doesn't trigger a placement scramble), and an unconditional safety-net pass every 5 minutes by default, which catches newly-added files that never triggered a roster change at all, or a roster-change notification that got missed. Both converge on the same routine: diff what a machine should hold against what it does hold, then fetch or release shards to close the gap, without evicting anything else already cached to make room.

You can see exactly what this decided for any file with swarmfile shards <path>. It lists every shard's role (data or parity), which peers it's assigned to, and whether this machine is holding it right now. ec-placement enable/disable/status toggles the feature itself; there's no separate manual rebalance command, because the reconciler is designed to never need one.

What this buys you that adaptive alone doesn't#

Adaptive erasure coding is a statistical bet: on average, enough peers will have enough shards. Deliberate placement is a specific guarantee: these three machines hold this shard, chosen deterministically, so losing any single one of them doesn't touch your ability to reconstruct the file. It's still a weaker guarantee than a self-hosted seed node, which pins the entire file tree unconditionally rather than 3 of 14 shards. Deliberate placement is for spreading resilience across the ordinary machines already in your office, not a replacement for a seed node if you need a stronger floor than that.