Browse docs
Docs / Reference / Storage Format

Storage Format

This page specifies how Swarmfile stores a file's bytes in object storage: how a file is split into blocks, how blocks are named, how the list of blocks (the manifest) is encoded, how blocks are encrypted, and where they land in a bucket. It is written so that someone holding the stored objects (and, for an encrypted project, the project key) can reconstruct a file with no Swarmfile software involved. It describes the format the current code writes and reads. Where something isn't pinned down, this page says so rather than guessing.

It covers file content only. File names, folder structure, versions, branches, commits, ACLs, and comments live in the project's metadata on the hub, not inside blocks, and their formats aren't specified here. In practice you need that metadata to know which manifest belongs to which path: each file entry records its manifest's identifier as rootCid.

If you only want your files back, you don't need this page: copy them off the mount, or use Branch Mirror to export real files to your own bucket. See Data Portability & Offboarding.

Overview#

  • A file is split into chunks, usually with content-defined chunking (FastCDC) averaging 1 MiB.
  • Each chunk is stored as one block, named by its CID: the BLAKE3 hash of the chunk's plaintext bytes.
  • A manifest, a small JSON document listing the chunks in file order, is itself stored as a block with its own CID. That CID is the file entry's rootCid.
  • A very large file's manifest is split into segment blocks under a small segmented root (manifest format v2).
  • On encrypted projects, every block, manifests included, is stored as AES-256-GCM ciphertext under a key derived from the project key and the block's CID. The CID is still the plaintext hash.
  • Some chunks also get Reed-Solomon 10+4 erasure-coded shards, stored as extra blocks in addition to the chunk itself.
entry.rootCid ──► manifest block (flat v1)  ──► chunk blocks (in file order)
             └──► segmented root (v2) ──► segment blocks ──► chunk blocks

Chunking#

Default parameters#

ConstantValue
FastCDC minimum chunk262,144 bytes (256 KiB)
FastCDC average (target) chunk1,048,576 bytes (1 MiB)
FastCDC maximum chunk4,194,304 bytes (4 MiB)
Fixed-size chunk (when fixed chunking is used)1,048,576 bytes (1 MiB)

The CDC implementation is FastCDC "v2020" from the Rust fastcdc crate (3.2.x), with its default normalization (level 1) and gear table. The hub has a byte-identical TypeScript port for its git-LFS ingest path, pinned by golden tests against the Rust chunker. Identical chunk boundaries across both means identical CIDs, which is what lets them dedupe against each other.

You never need to re-run the chunker to read a file. The manifest records every chunk's offset and size, so a reader just follows the manifest. The chunking parameters only matter if you want to reproduce Swarmfile's CIDs for new data.

File-type profiles#

On the mount's write path, the chunking strategy depends on the file extension (case-insensitive):

ExtensionsStrategy
everything not listed below (incl. .rvt, .dwg, .max, .ifc, .drp, .settings, .prproj)CDC 256 KiB / 1 MiB / 4 MiB
.pdf, .dwfFixed 1 MiB
.slog, .rws, .dat, .xml, .avb, .avp, .msn, .pmsm, .nksmFixed 1 MiB, plus the small-file fast path

Small-file fast path: a file with one of the fast-path extensions that is smaller than 50 KiB (51,200 bytes) is stored as a single chunk, with the strategy recorded as fixed and chunk_size equal to the file's size (minimum 1).

The profiles also set read-ahead behavior. That doesn't affect the stored format.

Other writers can pick other strategies. For example, swarmfile-seed uses fixed-size chunks unless you pass --cdc, and small browser uploads are written as a single chunk. Whichever strategy was used is recorded in the manifest's chunking field. The per-extension table above may change. It is not part of the stable format.

Block identifiers (CIDs)#

A CID is the BLAKE3-256 hash of the block's plaintext bytes, written as 64 lowercase hex characters. There is no multihash or multibase prefix, no version byte, and no domain separation.

Block kindCID is BLAKE3 of…
Chunkthe chunk's plaintext bytes
Flat manifestthe manifest's exact serialized JSON bytes
Manifest segment / segmented rootthat block's exact serialized JSON bytes
Erasure shardthe shard's bytes as stored (see Erasure coding)

Because a CID hashes plaintext, encrypting a block never changes its CID, and identical content within a project always has one CID. Encryption uses a fresh random nonce every time, so the same plaintext encrypted twice gives different stored bytes. Dedup happens at the CID level, not the ciphertext level.

Verification rule: after you fetch a block (and decrypt it, if it's encrypted), blake3(plaintext) must equal the CID you asked for. The only exception is an erasure shard, which is verified against its stored bytes.

Manifests#

Manifests are JSON as written by Rust's serde_json::to_vec: compact, with no whitespace and fields in the order shown below. Because a manifest's CID is the hash of its exact bytes, re-serializing a manifest with different formatting or key order gives a different CID. To read a manifest, any JSON parser works. To reproduce a CID, you have to match the serialization byte for byte.

Flat manifest (format v1)#

Used whenever the whole manifest fits in one block, which covers almost every file. It has no version field. Version 1 is implicit.

{
  "total_size": 3355443,
  "chunk_size": 1048576,
  "chunks": [
    {"cid": "9f1c…e2a0", "offset": 0,       "size": 1210334},
    {"cid": "4b7d…0c19", "offset": 1210334, "size": 862117},
    {"cid": "e03a…77f4", "offset": 2072451, "size": 1282992}
  ],
  "chunking": {"type": "cdc", "min": 262144, "avg": 1048576, "max": 4194304}
}

(Formatted and with CIDs shortened for readability. Real CIDs are 64 hex characters, and stored bytes have no whitespace.)

FieldTypeMeaning
total_sizeu64File length in bytes. Equals the sum of every chunk's size.
chunk_sizeu32For fixed: the chunk size. For cdc: the target average, not a real chunk size. Kept for backward compatibility. Use chunking and each chunk's size instead.
chunksarray of ChunkRefChunks in file order. An empty file has [].
chunkingobjectHow the file was chunked (see below). If absent, readers treat it as {"type":"fixed","chunk_size":1048576}.

chunking is an internally tagged object:

  • {"type": "fixed", "chunk_size": N}
  • {"type": "cdc", "min": N, "avg": N, "max": N}

ChunkRef:

FieldTypeMeaning
cidstringCID of the chunk's plaintext (64 lowercase hex).
offsetu64Byte offset of this chunk in the file.
sizeu32Plaintext length of this chunk.
erasureobject, optionalErasure-coding group for this chunk (see Erasure coding). Left out when there are no shards.

Chunks are contiguous: each chunk's offset equals the previous chunk's offset + size. The same CID can appear more than once when a file repeats content. Its block is stored once and referenced from each position.

Legacy browser-upload shape#

Small files uploaded through the web app before August 2026 may have a manifest in an older envelope, which readers still accept:

{"version": 1, "type": "file", "size": 35, "blocks": [{"offset": 0, "size": 35, "cid": "…"}]}

blocks maps one-to-one onto chunks, and size onto total_size (if size is missing, total the block sizes). Treat it as fixed chunking. New writes never produce this shape.

Segmented manifest (format v2)#

When a flat manifest would be too big for one block (see Size thresholds), the chunk list is split into segment blocks, and the entry's rootCid points at a small segmented root instead. A flat manifest's bytes never change when this happens. A file is only segmented if it can't be stored flat.

Segmented root:

{
  "manifest_version": 2,
  "total_size": 5497558138880,
  "chunking": {"type": "cdc", "min": 262144, "avg": 1048576, "max": 4194304},
  "segments": [
    {"cid": "a41e…9d03", "first_offset": 0,           "chunk_count": 31207},
    {"cid": "07bc…41fe", "first_offset": 32730906112, "chunk_count": 31188}
  ]
}
FieldTypeMeaning
manifest_versionu32Always 2.
total_sizeu64File length in bytes.
chunkingobjectSame as in the flat manifest, but always present here.
segmentsarray of SegmentRefSegments in file order.

SegmentRef: cid (the segment block's CID), first_offset (file offset of the segment's first chunk, so a range read can binary-search to the segment it needs), and chunk_count (how many ChunkRefs the segment holds).

Segment block:

{"chunks": [ {"cid": "…", "offset": 0, "size": 1048576}, … ]}

A segment holds one contiguous slice of the ChunkRef list, in the same shape as the flat manifest's chunks, including any erasure objects. To rebuild the full chunk list, concatenate the segments' chunks in segments order. If a segment's length doesn't match its chunk_count, treat that as corruption.

There is one level of segmentation. The root is never nested. A file whose segmented root would itself exceed the block limit is refused at write time. By the code's own estimate, that ceiling is hundreds of TiB for a single file.

Telling the shapes apart#

A root block is parsed like this, in order:

  1. Parse as a flat manifest, which needs chunks and also accepts the legacy blocks shape. If that works, it's flat.
  2. Otherwise parse as a segmented root, which needs manifest_version, total_size, chunking, and segments. If that works and manifest_version == 2, it's segmented.
  3. Otherwise the block is an error. That includes a segmented root with any manifest_version other than 2. Readers fail loudly and never guess.

The two canonical shapes are mutually exclusive: a flat manifest never has segments or manifest_version, and a segmented root never has chunks.

Size thresholds#

LimitValue
Maximum stored size of any manifest, segment, or root block4,194,368 bytes (4 MiB + 64)
Target plaintext size per segment¾ of that limit, 3,145,776 bytes

The flat-or-segmented decision is made on the stored size: after encryption on an encrypted project, where encryption adds 29 bytes per block. Segments are packed greedily in file order. Each ChunkRef costs its compact-JSON length plus 1 byte, a new segment starts when the next ref would push the open one past the target, and a ref is never split. Readers don't need any of this. It only matters if you want to reproduce the exact segment and root CIDs for a file.

To give a sense of scale: a bare ChunkRef is about 100 bytes, so a flat manifest holds roughly 40,000 chunks (about 40 GB at the 1 MiB CDC average). A ChunkRef carrying erasure metadata is about 10× larger, so erasure-coded files segment much sooner.

Encryption#

Which projects are encrypted#

Each project has a fixed encryption tier, chosen when it's created:

TierBlocks at restWho holds the project key
managed (default on paid plans)AES-256-GCM ciphertextThe hub, which stores the key wrapped by a hub-held key-encryption key (KEK) and hands it over TLS to authenticated clients that are allowed to have it
e2e (opt-in, Pro and above)AES-256-GCM ciphertextOnly enrolled user devices and, if configured, the org's recovery key. The hub stores only wrapped copies it can't open.
plaintextUnencryptedNobody. There's no key.

The plaintext tier is the only tier available on the Free (no-card) plan, and every public project must use it. Free-plan and public-project blocks are stored as plain bytes. A paid-plan org that asks for plaintext on a private project gets managed instead.

What is not encrypted, even on the managed tier#

  • Directly uploaded git-LFS objects. An LFS object uploaded by the stock git-lfs client through a presigned single PUT, or an object of 64 MiB or more uploaded by Swarmfile's built-in LFS transfer agent (as a multipart upload), is stored as one plaintext object under the lfs/ keyspace (see Where blocks are stored). Smaller objects pushed through the agent, and objects that go through the hub's native ingest, become ordinary encrypted blocks with a normal manifest. LFS is refused outright on e2e projects. See Git-LFS and Security Architecture.

  • Some hub-generated derivatives. The hub renders previews of managed content in memory. Thumbnails, video poster frames and point-cloud preview images are then stored encrypted (see Hub-generated preview images). These stay unencrypted:

    • scrubbable video proxies (a 720p H.264 transcode) at _videoproxy/{rootCid}, kept up to 30 days;
    • whole-file copies the hub assembles to serve browser downloads and previews at _materialized/{rootCid}, kept up to 24 hours;
    • short-lived render inputs and outputs the preview containers read and write (_videosrc/, _videoposter/, _pcpreview/), deleted when that render finishes;
    • legacy preview images written before previews were encrypted. These are plaintext JPEGs at the bare key {sha256}. They aren't rewritten or deleted, and stop being used once the file's content changes.

    None of these exist for e2e projects, because the hub can't read their content. On the plaintext tier, previews are plaintext like everything else.

  • Legacy managed blocks. Some managed-tier content written by older web uploads was stored as plaintext. Readers handle this by checking whether blake3(stored bytes) == CID before trying to decrypt (see below).

Block ciphertext layout#

offset  length  field
0       1       version      0x02 (current), 0x01 (legacy, read-only), or 0x82-0xC0 (a rotated key generation)
1       12      nonce        random, fresh for every encryption
13      N+16    ciphertext   AES-256-GCM(plaintext), with the 16-byte GCM tag appended

Encryption adds exactly 29 bytes to each block (1 + 12 + 16). There is no associated data (AAD).

The version byte also says which generation of the project key sealed the block (see Key rotation): 0x01 and 0x02 mean generation 1, and 0x80 + g means generation g, for g from 2 to 64 (bytes 0x82 to 0xC0). Every other value is reserved, and readers reject it.

The per-block key is derived with HKDF-SHA256 from the 32-byte project key of that generation:

Version byteHKDF input keyHKDF saltHKDF infoOutput
0x02generation-1 project key (32 bytes)the CID as its 64-character lowercase hex ASCII string (64 bytes, not the 32 raw hash bytes)swarmfile-block-encryption/v1 (ASCII)32 bytes, the AES-256 key
0x82-0xC0the project key of generation version − 0x80same as 0x02same as 0x0232 bytes
0x01 (legacy)generation-1 project keynone (RFC 5869 default: 32 zero bytes)the CID's hex ASCII string32 bytes

The CID used in the derivation is the block's own CID: the chunk's CID for a chunk, and the manifest's, segment's, or root's CID for those blocks. Every block is sealed under a different key, and a block can't be decrypted under the wrong CID.

A byte-identical implementation of this layout exists in the desktop engine (Rust), the hub (TypeScript/Web Crypto), the hub's preview containers, and the web app. The web app decrypts only E2E share content, which never uses a rotated key generation.

How the project key is held (who can decrypt)#

This section only explains who can decrypt. It's not needed to parse the format.

  • Managed: the hub generates a random 32-byte project key and stores it AES-256-GCM-wrapped under its KEK. An authorized, authenticated client gets it over TLS as 64 hex characters (GET {hub}/orgs/{orgId}/keys/{projectId} → {"projectId": "…", "key": "<hex>"}). Once the project's key has been rotated, the response also carries every generation: {"projectId": "…", "key": "<generation-1 hex>", "currentGeneration": 2, "keys": [{"generation": 1, "key": "<hex>"}, {"generation": 2, "key": "<hex>"}]}. key is always generation 1. The desktop engine may cache the key locally, protected by DPAPI on Windows or file permissions on macOS/Linux. With the keys and the stored blocks, you can decrypt the project yourself.
  • E2E: the client creating the project generates the key and wraps it separately to each enrolled device's X25519 public key, and to the org's recovery public key if one is configured. Each wrapped copy is an envelope: [0x01][nonce:12][ephemeral X25519 public key:32][AES-256-GCM ciphertext + tag], where the AES key is HKDF-SHA256(X25519(device private key, ephemeral public key), salt = nonce, info = "swarmfile-e2e-envelope/v1"). The org recovery private key is split into Shamir shares that are handed out when it's created and never stored by the hub. A quorum of shares recovers the project key. See End-to-End Encryption.

Key rotation#

A managed project's key can be rotated by the project's owners and admins (POST /keys/:projectId/rotate). Rotation creates a new key generation and makes it current. It never replaces or deletes an earlier generation:

  • Every block records the generation that sealed it in its version byte. Everything written before a project's first rotation is generation 1 (0x01 or 0x02), so a rotation doesn't change or re-encrypt any stored block.
  • New blocks are sealed under the current generation (0x80 + g). Readers look up the generation the byte names and use that key. There's no trial decryption.
  • Every retained generation is stored wrapped under the hub's KEK, like the generation-1 key. A KEK rotation re-wraps all of them.
  • Retired generations are kept, so blocks sealed under them stay readable. Nothing re-encrypts existing blocks under the new generation yet. Until that exists, an old generation can't be destroyed without losing the blocks it sealed.

A reader that predates rotation rejects a 0x82+ block as an unsupported version. It never decrypts one into wrong bytes. Such a reader still opens every generation-1 block, because key in the key response stays generation 1. For the same reason, an older engine keeps writing 0x02 blocks under generation 1, which every reader can open.

Why a marker in the version byte, and not trial decryption. Trying each generation's key until the GCM tag verifies would also have worked without a format change. But AES-GCM checks its tag only after processing the whole block, so every read of an older block would cost one full decryption per newer generation, and a missing key would look like a corrupt block. Putting the generation in the byte that already selects the key derivation keeps the layout and the 29-byte overhead unchanged, costs nothing on reads, and makes a missing key a specific, recoverable error. The trade-off is a limit of 64 generations per project.

Hub-generated preview images#

On a managed project, the thumbnails, video poster frames and point-cloud preview images the hub generates are JPEGs stored with the block ciphertext layout and the project key. Where a content block uses its CID, a preview uses the SHA-256 of the JPEG bytes, as 64 lowercase hex characters. That id is both the HKDF salt and the object name, at p/{projectId}/{sha256} (behind the same org prefix or bucket as blocks). To read one, decrypt it like a content block (a 0x02 block, or a rotated generation's 0x82+ block) with the SHA-256 hex standing in for the CID, then check that sha256(plaintext) equals that hex.

A preview written before this scheme is a plain JPEG at the bare key {sha256}. Readers try the scoped key first, then the bare key. On the plaintext tier, previews are always plain JPEGs at the bare key.

Erasure coding#

Some chunks also get Reed-Solomon erasure coding over GF(2⁸), with 10 data shards + 4 parity shards. Any 10 of the 14 shards rebuild the chunk.

When it's applied. Erasure coding is not applied to every block. The engine decides for each save (SWARMFILE_EC_ENABLED = on, off, or auto, default auto). In auto mode, shards are produced only for files of at least 64 KiB, and only when the machine has a peer-to-peer fabric with other peers, favoring WAN and seed peers. It's skipped when two or more LAN peers are present, unless deliberate LAN shard placement is turned on, and when WAN peers are measured as fast (p50 RTT ≤ 80 ms). Uploads with no P2P fabric (cloud-only mode, and bulk migration by default) don't produce shards. These rules may change. A reader should just follow each ChunkRef's erasure field.

Shards are extra. The whole chunk block is always stored too. Shards add redundancy and are never a substitute for the chunk.

Encoding. The input is the chunk's stored bytes: the ciphertext envelope on an encrypted project, or the plaintext otherwise.

  1. shard_size = ceil(len(stored) / 10).
  2. Zero-pad stored to 10 × shard_size and split it into data shards 0-9.
  3. Compute parity shards 10-13 with Reed-Solomon (10, 4) over GF(2⁸). The engine uses the Rust reed-solomon-erasure crate (galois_8). Its exact encoding matrix is defined by that crate and isn't restated here.
  4. Each shard is stored as a block whose CID is blake3(shard bytes). Shards are not encrypted again.

erasure object (on a ChunkRef):

{
  "original_cid": "9f1c…e2a0",
  "original_size": 1210363,
  "shard_size": 121037,
  "shards": [
    {"cid": "…", "shard_index": 0,  "shard_size": 121037},
    …
    {"cid": "…", "shard_index": 13, "shard_size": 121037}
  ],
  "encrypted": true
}
FieldMeaning
original_cidThe chunk's CID, the same as the ChunkRef's cid.
original_sizeLength of the stored bytes that were encoded, so on an encrypted project it includes the 29-byte envelope. It is not the plaintext size.
shard_sizeShard length, the same for all 14.
shardsExactly 14 entries. shard_index 0-9 are data shards and 10-13 are parity.
encryptedWhether the encoded bytes were ciphertext. If absent, it defaults to false.

Reconstructing from shards: collect at least 10 of the 14 shards, checking each against its CID. Run Reed-Solomon reconstruction, join data shards 0-9 in order, and truncate to original_size. The result is the chunk's stored bytes. Decrypt and verify them exactly as you would a directly fetched chunk block.

Where blocks are stored#

Blocks sit in an S3-compatible bucket, and the object key is built from the CID. The layout depends on the org's storage tier:

Storage tierObject key for a project's block
Shared bucket (Free plan and default){orgId}/p/{projectId}/{cid}
Dedicated bucket (org-exclusive R2 bucket)p/{projectId}/{cid}
Bring your own storage{configured prefix}/p/{projectId}/{cid}, or p/{projectId}/{cid} with no prefix configured

This applies the same way to chunks, manifests, segments, roots, and erasure shards.

Legacy bare keys. Blocks written before project scoping, or by clients that didn't name a project, sit at the bare key {cid} (behind the same org prefix or bucket as above). Readers try the project-scoped key first and fall back to the bare key. Take the scoped key first: on encrypted tiers, a bare object with the same CID can hold another project's ciphertext, which this project's key can't decrypt.

Other keyspaces, under the same org prefix or bucket:

  • lfs/{projectId}/{oid}: directly uploaded git-LFS objects, addressed by git-lfs's oid (the SHA-256 of the file, 64 hex characters). Each is one plaintext object holding the whole file. An older managed-tier shape may also exist: the object at that key has custom metadata enc: managed-v1 and holds a small JSON {"enc":"managed-v1","chunks":N,"size":S}, with 4 MiB ciphertext pieces at lfs/{projectId}/{oid}/c/{n}. Each piece uses the block ciphertext layout above, with the ASCII string {oid}-{n} (n counting from 0) taking the place of the CID in the key derivation. New writes never create that shape.
  • t/{orgId}/{projectId}/{cid} and c/{orgId}/{projectId}/{hash}: tree and commit objects for version history. Their format isn't covered here.
  • p/{projectId}/{sha256}: encrypted preview images for a managed project. They share the project-scoped namespace with blocks but are named by SHA-256, not BLAKE3.
  • {sha256} (bare), _videoproxy/, _materialized/: unencrypted derived renderings (see above).

The physical bucket name, the account behind it, and anything else about the shared bucket's arrangement are operational details. They aren't specified here and may change.

Reconstructing a file by hand#

You need:

  • the file's rootCid, from the project's metadata;
  • read access to the stored objects (your dedicated or BYOS bucket, or blocks fetched through the hub's block API);
  • for a managed or e2e project, the 32-byte project key, or every generation of it if a managed project's key has been rotated.

Then, in this order:

  1. Define fetch(cid): read the object at the project-scoped key, falling back to the bare key. Then:
    • no project key (plaintext tier): the stored bytes are the plaintext;
    • with a managed key: if blake3(stored) == cid, the block is legacy plaintext (a tolerance for managed projects only - an E2E or static-key reader must treat a plaintext block as an error), so use it as-is. Otherwise its first byte must be 0x01, 0x02, or 0x82-0xC0. Pick the key generation that byte names (generation 1 for 0x01/0x02, byte − 0x80 otherwise), derive the per-block key for that version, and AES-256-GCM-decrypt stored[13..] with nonce stored[1..13] (the tag is the last 16 bytes).
    • Verify that blake3(plaintext) == cid. Stop on any mismatch or GCM authentication failure.
  2. Fetch the root: root = JSON.parse(fetch(rootCid)).
  3. Classify it using the rules in Telling the shapes apart.
    • Flat: chunks = root.chunks (or the legacy root.blocks).
    • Segmented (manifest_version == 2): for each entry in root.segments, in order, run seg = JSON.parse(fetch(entry.cid)), check that len(seg.chunks) == entry.chunk_count, and append seg.chunks.
  4. Fetch each chunk: for each ChunkRef in order, data = fetch(ref.cid) and check that len(data) == ref.size. If a chunk block is missing and the ref has an erasure group, rebuild the chunk's stored bytes from any 10 shards (see above), then decrypt and verify them as in step 1.
  5. Assemble: write each chunk's bytes at ref.offset. Since chunks are contiguous and in order, you can simply concatenate them.
  6. Check the total: the output length must equal total_size. The manifest has no separate whole-file hash. Integrity comes from the per-block CID checks above, and the manifest itself is covered by rootCid.

A 1-byte range read only needs the chunks that overlap it. With a segmented root, binary-search first_offset to find which segments to fetch.

Versioning & stability#

Stable: a stored object is never rewritten to upgrade its format, so existing data stays readable without migration. Changing anything below would change CIDs for existing content, so the code treats each of these as a breaking change:

  • BLAKE3-256 hex CIDs over plaintext;
  • the flat manifest's field names and serialization;
  • the segmented root and segment shapes, with manifest_version: 2;
  • the [version][nonce][ciphertext+tag] envelope and the 0x02 key derivation;
  • the key-generation version bytes 0x82-0xC0, which use the 0x02 derivation under that generation's key;
  • readers accepting the legacy manifest shape and 0x01 ciphertext.

How format versions are signalled:

WhatSignal
Flat vs segmented manifestWhich fields are present, plus manifest_version (absent = v1, 2 = segmented). Readers reject unknown versions.
Ciphertext envelopeThe first byte (0x01 legacy, 0x02 current, 0x80 + g for key generation g ≥ 2). Readers reject unknown versions.
ChunkingThe manifest's chunking.type (fixed / cdc).
E2E key envelopeIts first byte (0x01).

May change without notice (none of this affects reading data that's already stored):

  • the default chunking parameters and the per-extension profiles;
  • when erasure coding is applied;
  • the block size limit and the ¾ segment target (which change which files get segmented and what their manifest CIDs are);
  • bucket and prefix arrangement beyond the key shapes documented above;
  • how git-LFS objects are stored.

Not specified here: file and folder metadata, tree and commit objects, and anything about Reed-Solomon beyond "10+4 over GF(2⁸) via reed-solomon-erasure". If you depend on one of these, contact us.