Storage Format
This page specifies how Swarmfile stores a file's bytes in object storage: how a file is split into blocks, how blocks are named, how the list of blocks (the manifest) is encoded, how blocks are encrypted, and where they land in a bucket. It is written so that someone holding the stored objects (and, for an encrypted project, the project key) can reconstruct a file with no Swarmfile software involved. It describes the format the current code writes and reads. Where something isn't pinned down, this page says so rather than guessing.
It covers file content only. File names, folder structure, versions, branches, commits, ACLs, and comments live in the project's metadata on the hub, not inside blocks, and their formats aren't specified here. In practice you need that metadata to know which manifest belongs to which path: each file entry records its manifest's identifier as rootCid.
If you only want your files back, you don't need this page: copy them off the mount, or use Branch Mirror to export real files to your own bucket. See Data Portability & Offboarding.
Overview#
- A file is split into chunks, usually with content-defined chunking (FastCDC) averaging 1 MiB.
- Each chunk is stored as one block, named by its CID: the BLAKE3 hash of the chunk's plaintext bytes.
- A manifest, a small JSON document listing the chunks in file order, is itself stored as a block with its own CID. That CID is the file entry's
rootCid. - A very large file's manifest is split into segment blocks under a small segmented root (manifest format v2).
- On encrypted projects, every block, manifests included, is stored as AES-256-GCM ciphertext under a key derived from the project key and the block's CID. The CID is still the plaintext hash.
- Some chunks also get Reed-Solomon 10+4 erasure-coded shards, stored as extra blocks in addition to the chunk itself.
entry.rootCid ──► manifest block (flat v1) ──► chunk blocks (in file order)
└──► segmented root (v2) ──► segment blocks ──► chunk blocks
Chunking#
Default parameters#
| Constant | Value |
|---|---|
| FastCDC minimum chunk | 262,144 bytes (256 KiB) |
| FastCDC average (target) chunk | 1,048,576 bytes (1 MiB) |
| FastCDC maximum chunk | 4,194,304 bytes (4 MiB) |
| Fixed-size chunk (when fixed chunking is used) | 1,048,576 bytes (1 MiB) |
The CDC implementation is FastCDC "v2020" from the Rust fastcdc crate (3.2.x), with its default normalization (level 1) and gear table. The hub has a byte-identical TypeScript port for its git-LFS ingest path, pinned by golden tests against the Rust chunker. Identical chunk boundaries across both means identical CIDs, which is what lets them dedupe against each other.
You never need to re-run the chunker to read a file. The manifest records every chunk's offset and size, so a reader just follows the manifest. The chunking parameters only matter if you want to reproduce Swarmfile's CIDs for new data.
File-type profiles#
On the mount's write path, the chunking strategy depends on the file extension (case-insensitive):
| Extensions | Strategy |
|---|---|
everything not listed below (incl. .rvt, .dwg, .max, .ifc, .drp, .settings, .prproj) | CDC 256 KiB / 1 MiB / 4 MiB |
.pdf, .dwf | Fixed 1 MiB |
.slog, .rws, .dat, .xml, .avb, .avp, .msn, .pmsm, .nksm | Fixed 1 MiB, plus the small-file fast path |
Small-file fast path: a file with one of the fast-path extensions that is smaller than 50 KiB (51,200 bytes) is stored as a single chunk, with the strategy recorded as fixed and chunk_size equal to the file's size (minimum 1).
The profiles also set read-ahead behavior. That doesn't affect the stored format.
Other writers can pick other strategies. For example, swarmfile-seed uses fixed-size chunks unless you pass --cdc, and small browser uploads are written as a single chunk. Whichever strategy was used is recorded in the manifest's chunking field. The per-extension table above may change. It is not part of the stable format.
Block identifiers (CIDs)#
A CID is the BLAKE3-256 hash of the block's plaintext bytes, written as 64 lowercase hex characters. There is no multihash or multibase prefix, no version byte, and no domain separation.
| Block kind | CID is BLAKE3 of… |
|---|---|
| Chunk | the chunk's plaintext bytes |
| Flat manifest | the manifest's exact serialized JSON bytes |
| Manifest segment / segmented root | that block's exact serialized JSON bytes |
| Erasure shard | the shard's bytes as stored (see Erasure coding) |
Because a CID hashes plaintext, encrypting a block never changes its CID, and identical content within a project always has one CID. Encryption uses a fresh random nonce every time, so the same plaintext encrypted twice gives different stored bytes. Dedup happens at the CID level, not the ciphertext level.
Verification rule: after you fetch a block (and decrypt it, if it's encrypted), blake3(plaintext) must equal the CID you asked for. The only exception is an erasure shard, which is verified against its stored bytes.
Manifests#
Manifests are JSON as written by Rust's serde_json::to_vec: compact, with no whitespace and fields in the order shown below. Because a manifest's CID is the hash of its exact bytes, re-serializing a manifest with different formatting or key order gives a different CID. To read a manifest, any JSON parser works. To reproduce a CID, you have to match the serialization byte for byte.
Flat manifest (format v1)#
Used whenever the whole manifest fits in one block, which covers almost every file. It has no version field. Version 1 is implicit.
{
"total_size": 3355443,
"chunk_size": 1048576,
"chunks": [
{"cid": "9f1c…e2a0", "offset": 0, "size": 1210334},
{"cid": "4b7d…0c19", "offset": 1210334, "size": 862117},
{"cid": "e03a…77f4", "offset": 2072451, "size": 1282992}
],
"chunking": {"type": "cdc", "min": 262144, "avg": 1048576, "max": 4194304}
}
(Formatted and with CIDs shortened for readability. Real CIDs are 64 hex characters, and stored bytes have no whitespace.)
| Field | Type | Meaning |
|---|---|---|
total_size | u64 | File length in bytes. Equals the sum of every chunk's size. |
chunk_size | u32 | For fixed: the chunk size. For cdc: the target average, not a real chunk size. Kept for backward compatibility. Use chunking and each chunk's size instead. |
chunks | array of ChunkRef | Chunks in file order. An empty file has []. |
chunking | object | How the file was chunked (see below). If absent, readers treat it as {"type":"fixed","chunk_size":1048576}. |
chunking is an internally tagged object:
{"type": "fixed", "chunk_size": N}{"type": "cdc", "min": N, "avg": N, "max": N}
ChunkRef:
| Field | Type | Meaning |
|---|---|---|
cid | string | CID of the chunk's plaintext (64 lowercase hex). |
offset | u64 | Byte offset of this chunk in the file. |
size | u32 | Plaintext length of this chunk. |
erasure | object, optional | Erasure-coding group for this chunk (see Erasure coding). Left out when there are no shards. |
Chunks are contiguous: each chunk's offset equals the previous chunk's offset + size. The same CID can appear more than once when a file repeats content. Its block is stored once and referenced from each position.
Legacy browser-upload shape#
Small files uploaded through the web app before August 2026 may have a manifest in an older envelope, which readers still accept:
{"version": 1, "type": "file", "size": 35, "blocks": [{"offset": 0, "size": 35, "cid": "…"}]}
blocks maps one-to-one onto chunks, and size onto total_size (if size is missing, total the block sizes). Treat it as fixed chunking. New writes never produce this shape.
Segmented manifest (format v2)#
When a flat manifest would be too big for one block (see Size thresholds), the chunk list is split into segment blocks, and the entry's rootCid points at a small segmented root instead. A flat manifest's bytes never change when this happens. A file is only segmented if it can't be stored flat.
Segmented root:
{
"manifest_version": 2,
"total_size": 5497558138880,
"chunking": {"type": "cdc", "min": 262144, "avg": 1048576, "max": 4194304},
"segments": [
{"cid": "a41e…9d03", "first_offset": 0, "chunk_count": 31207},
{"cid": "07bc…41fe", "first_offset": 32730906112, "chunk_count": 31188}
]
}
| Field | Type | Meaning |
|---|---|---|
manifest_version | u32 | Always 2. |
total_size | u64 | File length in bytes. |
chunking | object | Same as in the flat manifest, but always present here. |
segments | array of SegmentRef | Segments in file order. |
SegmentRef: cid (the segment block's CID), first_offset (file offset of the segment's first chunk, so a range read can binary-search to the segment it needs), and chunk_count (how many ChunkRefs the segment holds).
Segment block:
{"chunks": [ {"cid": "…", "offset": 0, "size": 1048576}, … ]}
A segment holds one contiguous slice of the ChunkRef list, in the same shape as the flat manifest's chunks, including any erasure objects. To rebuild the full chunk list, concatenate the segments' chunks in segments order. If a segment's length doesn't match its chunk_count, treat that as corruption.
There is one level of segmentation. The root is never nested. A file whose segmented root would itself exceed the block limit is refused at write time. By the code's own estimate, that ceiling is hundreds of TiB for a single file.
Telling the shapes apart#
A root block is parsed like this, in order:
- Parse as a flat manifest, which needs
chunksand also accepts the legacyblocksshape. If that works, it's flat. - Otherwise parse as a segmented root, which needs
manifest_version,total_size,chunking, andsegments. If that works andmanifest_version == 2, it's segmented. - Otherwise the block is an error. That includes a segmented root with any
manifest_versionother than2. Readers fail loudly and never guess.
The two canonical shapes are mutually exclusive: a flat manifest never has segments or manifest_version, and a segmented root never has chunks.
Size thresholds#
| Limit | Value |
|---|---|
| Maximum stored size of any manifest, segment, or root block | 4,194,368 bytes (4 MiB + 64) |
| Target plaintext size per segment | ¾ of that limit, 3,145,776 bytes |
The flat-or-segmented decision is made on the stored size: after encryption on an encrypted project, where encryption adds 29 bytes per block. Segments are packed greedily in file order. Each ChunkRef costs its compact-JSON length plus 1 byte, a new segment starts when the next ref would push the open one past the target, and a ref is never split. Readers don't need any of this. It only matters if you want to reproduce the exact segment and root CIDs for a file.
To give a sense of scale: a bare ChunkRef is about 100 bytes, so a flat manifest holds roughly 40,000 chunks (about 40 GB at the 1 MiB CDC average). A ChunkRef carrying erasure metadata is about 10× larger, so erasure-coded files segment much sooner.
Encryption#
Which projects are encrypted#
Each project has a fixed encryption tier, chosen when it's created:
| Tier | Blocks at rest | Who holds the project key |
|---|---|---|
managed (default on paid plans) | AES-256-GCM ciphertext | The hub, which stores the key wrapped by a hub-held key-encryption key (KEK) and hands it over TLS to authenticated clients that are allowed to have it |
e2e (opt-in, Pro and above) | AES-256-GCM ciphertext | Only enrolled user devices and, if configured, the org's recovery key. The hub stores only wrapped copies it can't open. |
plaintext | Unencrypted | Nobody. There's no key. |
The plaintext tier is the only tier available on the Free (no-card) plan, and every public project must use it. Free-plan and public-project blocks are stored as plain bytes. A paid-plan org that asks for plaintext on a private project gets managed instead.
What is not encrypted, even on the managed tier#
-
Directly uploaded git-LFS objects. An LFS object uploaded by the stock
git-lfsclient through a presigned single PUT, or an object of 64 MiB or more uploaded by Swarmfile's built-in LFS transfer agent (as a multipart upload), is stored as one plaintext object under thelfs/keyspace (see Where blocks are stored). Smaller objects pushed through the agent, and objects that go through the hub's native ingest, become ordinary encrypted blocks with a normal manifest. LFS is refused outright one2eprojects. See Git-LFS and Security Architecture. -
Some hub-generated derivatives. The hub renders previews of managed content in memory. Thumbnails, video poster frames and point-cloud preview images are then stored encrypted (see Hub-generated preview images). These stay unencrypted:
- scrubbable video proxies (a 720p H.264 transcode) at
_videoproxy/{rootCid}, kept up to 30 days; - whole-file copies the hub assembles to serve browser downloads and previews at
_materialized/{rootCid}, kept up to 24 hours; - short-lived render inputs and outputs the preview containers read and write (
_videosrc/,_videoposter/,_pcpreview/), deleted when that render finishes; - legacy preview images written before previews were encrypted. These are plaintext JPEGs at the bare key
{sha256}. They aren't rewritten or deleted, and stop being used once the file's content changes.
None of these exist for
e2eprojects, because the hub can't read their content. On theplaintexttier, previews are plaintext like everything else. - scrubbable video proxies (a 720p H.264 transcode) at
-
Legacy managed blocks. Some managed-tier content written by older web uploads was stored as plaintext. Readers handle this by checking whether
blake3(stored bytes) == CIDbefore trying to decrypt (see below).
Block ciphertext layout#
offset length field
0 1 version 0x02 (current), 0x01 (legacy, read-only), or 0x82-0xC0 (a rotated key generation)
1 12 nonce random, fresh for every encryption
13 N+16 ciphertext AES-256-GCM(plaintext), with the 16-byte GCM tag appended
Encryption adds exactly 29 bytes to each block (1 + 12 + 16). There is no associated data (AAD).
The version byte also says which generation of the project key sealed the block (see Key rotation): 0x01 and 0x02 mean generation 1, and 0x80 + g means generation g, for g from 2 to 64 (bytes 0x82 to 0xC0). Every other value is reserved, and readers reject it.
The per-block key is derived with HKDF-SHA256 from the 32-byte project key of that generation:
| Version byte | HKDF input key | HKDF salt | HKDF info | Output |
|---|---|---|---|---|
0x02 | generation-1 project key (32 bytes) | the CID as its 64-character lowercase hex ASCII string (64 bytes, not the 32 raw hash bytes) | swarmfile-block-encryption/v1 (ASCII) | 32 bytes, the AES-256 key |
0x82-0xC0 | the project key of generation version − 0x80 | same as 0x02 | same as 0x02 | 32 bytes |
0x01 (legacy) | generation-1 project key | none (RFC 5869 default: 32 zero bytes) | the CID's hex ASCII string | 32 bytes |
The CID used in the derivation is the block's own CID: the chunk's CID for a chunk, and the manifest's, segment's, or root's CID for those blocks. Every block is sealed under a different key, and a block can't be decrypted under the wrong CID.
A byte-identical implementation of this layout exists in the desktop engine (Rust), the hub (TypeScript/Web Crypto), the hub's preview containers, and the web app. The web app decrypts only E2E share content, which never uses a rotated key generation.
How the project key is held (who can decrypt)#
This section only explains who can decrypt. It's not needed to parse the format.
- Managed: the hub generates a random 32-byte project key and stores it AES-256-GCM-wrapped under its KEK. An authorized, authenticated client gets it over TLS as 64 hex characters (
GET {hub}/orgs/{orgId}/keys/{projectId}→{"projectId": "…", "key": "<hex>"}). Once the project's key has been rotated, the response also carries every generation:{"projectId": "…", "key": "<generation-1 hex>", "currentGeneration": 2, "keys": [{"generation": 1, "key": "<hex>"}, {"generation": 2, "key": "<hex>"}]}.keyis always generation 1. The desktop engine may cache the key locally, protected by DPAPI on Windows or file permissions on macOS/Linux. With the keys and the stored blocks, you can decrypt the project yourself. - E2E: the client creating the project generates the key and wraps it separately to each enrolled device's X25519 public key, and to the org's recovery public key if one is configured. Each wrapped copy is an envelope:
[0x01][nonce:12][ephemeral X25519 public key:32][AES-256-GCM ciphertext + tag], where the AES key isHKDF-SHA256(X25519(device private key, ephemeral public key), salt = nonce, info = "swarmfile-e2e-envelope/v1"). The org recovery private key is split into Shamir shares that are handed out when it's created and never stored by the hub. A quorum of shares recovers the project key. See End-to-End Encryption.
Key rotation#
A managed project's key can be rotated by the project's owners and admins (POST /keys/:projectId/rotate). Rotation creates a new key generation and makes it current. It never replaces or deletes an earlier generation:
- Every block records the generation that sealed it in its version byte. Everything written before a project's first rotation is generation 1 (
0x01or0x02), so a rotation doesn't change or re-encrypt any stored block. - New blocks are sealed under the current generation (
0x80 + g). Readers look up the generation the byte names and use that key. There's no trial decryption. - Every retained generation is stored wrapped under the hub's KEK, like the generation-1 key. A KEK rotation re-wraps all of them.
- Retired generations are kept, so blocks sealed under them stay readable. Nothing re-encrypts existing blocks under the new generation yet. Until that exists, an old generation can't be destroyed without losing the blocks it sealed.
A reader that predates rotation rejects a 0x82+ block as an unsupported version. It never decrypts one into wrong bytes. Such a reader still opens every generation-1 block, because key in the key response stays generation 1. For the same reason, an older engine keeps writing 0x02 blocks under generation 1, which every reader can open.
Why a marker in the version byte, and not trial decryption. Trying each generation's key until the GCM tag verifies would also have worked without a format change. But AES-GCM checks its tag only after processing the whole block, so every read of an older block would cost one full decryption per newer generation, and a missing key would look like a corrupt block. Putting the generation in the byte that already selects the key derivation keeps the layout and the 29-byte overhead unchanged, costs nothing on reads, and makes a missing key a specific, recoverable error. The trade-off is a limit of 64 generations per project.
Hub-generated preview images#
On a managed project, the thumbnails, video poster frames and point-cloud preview images the hub generates are JPEGs stored with the block ciphertext layout and the project key. Where a content block uses its CID, a preview uses the SHA-256 of the JPEG bytes, as 64 lowercase hex characters. That id is both the HKDF salt and the object name, at p/{projectId}/{sha256} (behind the same org prefix or bucket as blocks). To read one, decrypt it like a content block (a 0x02 block, or a rotated generation's 0x82+ block) with the SHA-256 hex standing in for the CID, then check that sha256(plaintext) equals that hex.
A preview written before this scheme is a plain JPEG at the bare key {sha256}. Readers try the scoped key first, then the bare key. On the plaintext tier, previews are always plain JPEGs at the bare key.
Erasure coding#
Some chunks also get Reed-Solomon erasure coding over GF(2⁸), with 10 data shards + 4 parity shards. Any 10 of the 14 shards rebuild the chunk.
When it's applied. Erasure coding is not applied to every block. The engine decides for each save (SWARMFILE_EC_ENABLED = on, off, or auto, default auto). In auto mode, shards are produced only for files of at least 64 KiB, and only when the machine has a peer-to-peer fabric with other peers, favoring WAN and seed peers. It's skipped when two or more LAN peers are present, unless deliberate LAN shard placement is turned on, and when WAN peers are measured as fast (p50 RTT ≤ 80 ms). Uploads with no P2P fabric (cloud-only mode, and bulk migration by default) don't produce shards. These rules may change. A reader should just follow each ChunkRef's erasure field.
Shards are extra. The whole chunk block is always stored too. Shards add redundancy and are never a substitute for the chunk.
Encoding. The input is the chunk's stored bytes: the ciphertext envelope on an encrypted project, or the plaintext otherwise.
shard_size = ceil(len(stored) / 10).- Zero-pad
storedto10 × shard_sizeand split it into data shards 0-9. - Compute parity shards 10-13 with Reed-Solomon (10, 4) over GF(2⁸). The engine uses the Rust
reed-solomon-erasurecrate (galois_8). Its exact encoding matrix is defined by that crate and isn't restated here. - Each shard is stored as a block whose CID is
blake3(shard bytes). Shards are not encrypted again.
erasure object (on a ChunkRef):
{
"original_cid": "9f1c…e2a0",
"original_size": 1210363,
"shard_size": 121037,
"shards": [
{"cid": "…", "shard_index": 0, "shard_size": 121037},
…
{"cid": "…", "shard_index": 13, "shard_size": 121037}
],
"encrypted": true
}
| Field | Meaning |
|---|---|
original_cid | The chunk's CID, the same as the ChunkRef's cid. |
original_size | Length of the stored bytes that were encoded, so on an encrypted project it includes the 29-byte envelope. It is not the plaintext size. |
shard_size | Shard length, the same for all 14. |
shards | Exactly 14 entries. shard_index 0-9 are data shards and 10-13 are parity. |
encrypted | Whether the encoded bytes were ciphertext. If absent, it defaults to false. |
Reconstructing from shards: collect at least 10 of the 14 shards, checking each against its CID. Run Reed-Solomon reconstruction, join data shards 0-9 in order, and truncate to original_size. The result is the chunk's stored bytes. Decrypt and verify them exactly as you would a directly fetched chunk block.
Where blocks are stored#
Blocks sit in an S3-compatible bucket, and the object key is built from the CID. The layout depends on the org's storage tier:
| Storage tier | Object key for a project's block |
|---|---|
| Shared bucket (Free plan and default) | {orgId}/p/{projectId}/{cid} |
| Dedicated bucket (org-exclusive R2 bucket) | p/{projectId}/{cid} |
| Bring your own storage | {configured prefix}/p/{projectId}/{cid}, or p/{projectId}/{cid} with no prefix configured |
This applies the same way to chunks, manifests, segments, roots, and erasure shards.
Legacy bare keys. Blocks written before project scoping, or by clients that didn't name a project, sit at the bare key {cid} (behind the same org prefix or bucket as above). Readers try the project-scoped key first and fall back to the bare key. Take the scoped key first: on encrypted tiers, a bare object with the same CID can hold another project's ciphertext, which this project's key can't decrypt.
Other keyspaces, under the same org prefix or bucket:
lfs/{projectId}/{oid}: directly uploaded git-LFS objects, addressed by git-lfs'soid(the SHA-256 of the file, 64 hex characters). Each is one plaintext object holding the whole file. An older managed-tier shape may also exist: the object at that key has custom metadataenc: managed-v1and holds a small JSON{"enc":"managed-v1","chunks":N,"size":S}, with 4 MiB ciphertext pieces atlfs/{projectId}/{oid}/c/{n}. Each piece uses the block ciphertext layout above, with the ASCII string{oid}-{n}(ncounting from 0) taking the place of the CID in the key derivation. New writes never create that shape.t/{orgId}/{projectId}/{cid}andc/{orgId}/{projectId}/{hash}: tree and commit objects for version history. Their format isn't covered here.p/{projectId}/{sha256}: encrypted preview images for a managed project. They share the project-scoped namespace with blocks but are named by SHA-256, not BLAKE3.{sha256}(bare),_videoproxy/,_materialized/: unencrypted derived renderings (see above).
The physical bucket name, the account behind it, and anything else about the shared bucket's arrangement are operational details. They aren't specified here and may change.
Reconstructing a file by hand#
You need:
- the file's
rootCid, from the project's metadata; - read access to the stored objects (your dedicated or BYOS bucket, or blocks fetched through the hub's block API);
- for a
managedore2eproject, the 32-byte project key, or every generation of it if a managed project's key has been rotated.
Then, in this order:
- Define
fetch(cid): read the object at the project-scoped key, falling back to the bare key. Then:- no project key (plaintext tier): the stored bytes are the plaintext;
- with a managed key: if
blake3(stored) == cid, the block is legacy plaintext (a tolerance for managed projects only - an E2E or static-key reader must treat a plaintext block as an error), so use it as-is. Otherwise its first byte must be0x01,0x02, or0x82-0xC0. Pick the key generation that byte names (generation 1 for0x01/0x02,byte − 0x80otherwise), derive the per-block key for that version, and AES-256-GCM-decryptstored[13..]with noncestored[1..13](the tag is the last 16 bytes). - Verify that
blake3(plaintext) == cid. Stop on any mismatch or GCM authentication failure.
- Fetch the root:
root = JSON.parse(fetch(rootCid)). - Classify it using the rules in Telling the shapes apart.
- Flat:
chunks = root.chunks(or the legacyroot.blocks). - Segmented (
manifest_version == 2): for each entry inroot.segments, in order, runseg = JSON.parse(fetch(entry.cid)), check thatlen(seg.chunks) == entry.chunk_count, and appendseg.chunks.
- Flat:
- Fetch each chunk: for each ChunkRef in order,
data = fetch(ref.cid)and check thatlen(data) == ref.size. If a chunk block is missing and the ref has anerasuregroup, rebuild the chunk's stored bytes from any 10 shards (see above), then decrypt and verify them as in step 1. - Assemble: write each chunk's bytes at
ref.offset. Since chunks are contiguous and in order, you can simply concatenate them. - Check the total: the output length must equal
total_size. The manifest has no separate whole-file hash. Integrity comes from the per-block CID checks above, and the manifest itself is covered byrootCid.
A 1-byte range read only needs the chunks that overlap it. With a segmented root, binary-search first_offset to find which segments to fetch.
Versioning & stability#
Stable: a stored object is never rewritten to upgrade its format, so existing data stays readable without migration. Changing anything below would change CIDs for existing content, so the code treats each of these as a breaking change:
- BLAKE3-256 hex CIDs over plaintext;
- the flat manifest's field names and serialization;
- the segmented root and segment shapes, with
manifest_version: 2; - the
[version][nonce][ciphertext+tag]envelope and the0x02key derivation; - the key-generation version bytes
0x82-0xC0, which use the0x02derivation under that generation's key; - readers accepting the legacy manifest shape and
0x01ciphertext.
How format versions are signalled:
| What | Signal |
|---|---|
| Flat vs segmented manifest | Which fields are present, plus manifest_version (absent = v1, 2 = segmented). Readers reject unknown versions. |
| Ciphertext envelope | The first byte (0x01 legacy, 0x02 current, 0x80 + g for key generation g ≥ 2). Readers reject unknown versions. |
| Chunking | The manifest's chunking.type (fixed / cdc). |
| E2E key envelope | Its first byte (0x01). |
May change without notice (none of this affects reading data that's already stored):
- the default chunking parameters and the per-extension profiles;
- when erasure coding is applied;
- the block size limit and the ¾ segment target (which change which files get segmented and what their manifest CIDs are);
- bucket and prefix arrangement beyond the key shapes documented above;
- how git-LFS objects are stored.
Not specified here: file and folder metadata, tree and commit objects, and anything about Reed-Solomon beyond "10+4 over GF(2⁸) via reed-solomon-erasure". If you depend on one of these, contact us.