Security Architecture
This is the deep technical companion to Security - written to be handed directly to a security team running a vendor review. It states plainly what each layer protects against, who can decrypt what, and where the limits of that protection are. If your review needs something not covered here, contact us and we'll answer specifics.
The one-paragraph version#
On a paid plan (Starter and above), private project content is encrypted with AES-256-GCM before it is stored - on the machine that writes it for anything saved through the drive, and by the hub for browser uploads to a managed project (the hub holds a managed project's key; an end-to-end project's key never reaches it). Peer-to-peer nodes and our cloud object storage see only ciphertext addressed by a BLAKE3 content hash, with a few documented exceptions: the git-LFS direct-upload path (see below); server-side video proxies and the whole-file copies assembled to serve browser downloads (both retained short-term - see "Worth knowing about server-side previews" below); and any public project, which is stored unencrypted by design - see Unencrypted tier. Free-plan projects are all public and use that tier because they have no other, and a paying org can still choose it deliberately for a public project that needs the anonymous-download paths (which read unencrypted content only). Because a public project's content is public, its plaintext blocks may also be served to peers on the org's private swarm - peers reach only what the anonymous read routes already serve, peer serving is read-only, and every fetched block is BLAKE3-verified before it is used. A paying org's private projects are encrypted - managed by default, E2E opt-in - which is also what makes them eligible for peer-to-peer transfer. The only party that can turn ciphertext back into your files is whoever holds the project key - and you choose who that is. On the default managed tier that party is the Swarmfile hub, so we can give you server-side previews, search, and proxies - managed means encrypted at rest with hub-held keys, not zero-knowledge. On the opt-in end-to-end (E2E) tier the key never reaches our servers in a form we can open, so we structurally cannot read your content. Both tiers are fully supported; the difference is a deliberate trade you make per project, not a weakness in one of them.
The Free plan is the one plan with no at-rest encryption at all: its projects are stored unencrypted, with no project key - the trade-off for no-card, no-cost storage. See Unencrypted tier below. On Starter and above, a private project gets managed encryption by default (E2E opt-in); a public project is unencrypted-only on every plan, because public releases and archives still read unencrypted content only. The unencrypted tier is therefore Free-plan projects plus a paying org's deliberately-public projects.
Threat model - what each tier defends against#
| Adversary | Managed tier | E2E tier | Unencrypted tier (Free plan; a paid org's explicit public-project opt-out) |
|---|---|---|---|
| A malicious/curious peer on your swarm | ✅ sees only ciphertext + BLAKE3 CIDs | ✅ same | ❌ sees plaintext content - which is public by design (peers see only what anonymous readers see) |
| The storage provider or anyone who exfiltrates a storage bucket | ✅ ciphertext at rest, no keys co-located | ✅ same | ❌ plaintext at rest |
| Someone who subpoenas or compromises the Swarmfile hub | ⚠️ hub holds the wrapping key and can decrypt | ✅ hub holds only un-openable wrapped blobs | ❌ nothing to compromise - already plaintext |
| A Swarmfile employee with production access | ⚠️ same as above - possible, audited, not prevented | ✅ prevented by construction | ❌ same as above |
| A non-member of the project (incl. other orgs, other members) | ✅ blocked by the block ACL on every hub fetch | ✅ blocked by ACL and has no wrapped key | ✅ nothing confidential to reach (public by design); hub fetches stay ACL-gated, and swarm peers can only fetch blocks the anonymous read routes already serve |
| A lost password / lost device | ✅ recoverable (hub re-wraps) | ⚠️ recoverable only via the org recovery key | ✅ no key to lose |
Exception - git-LFS objects. The block properties above describe project content written through
the drive/engine. A project used as a git-LFS server is not uniformly
encrypted at rest: objects uploaded through an agent's chunked block path are encrypted on a managed
project (the desktop agent client-side; the standalone agent's plaintext blocks are encrypted by the
hub on write), but two direct-upload paths store the object unencrypted - a stock git-lfs push, and
the desktop app's built-in transfer agent for objects of 64 MiB or more. That is one of the documented
exceptions to "ciphertext at rest", and it is a deliberate trade: LFS is an add-on surface, and
encrypting and decrypting multi-GB artifacts on every transfer isn't worth it. A binary that must be
encrypted at rest belongs on the mounted drive (managed or E2E), not routed through LFS.
The two ⚠️ rows on the managed tier are the entire reason the E2E tier exists. If "the vendor can technically decrypt our footage" fails your review, the answer is the E2E tier - not a promise about the managed one. The unencrypted tier is a different trade entirely: it is not a weaker version of managed encryption, it is the deliberate absence of at-rest encryption in exchange for a free, no-card tier - access is still gated by the same block ACL every other tier uses, but there is no encryption layer to defend the storage boundary at all.
How block encryption works (both tiers)#
Encryption of the bytes is identical on both tiers; only custody of the project key differs.
- Each project has one 32-byte project master key.
- Every content block gets its own key, derived from the project key and bound to the block's content (the CID is the BLAKE3 hash of the plaintext chunk), so a key is never reused across blocks.
- The block is sealed with AES-256-GCM and keyed in object storage by its CID. (The exact block format is in the Storage Format spec.)
- On read, the reader re-derives the key, decrypts, and re-verifies the BLAKE3 CID before any byte reaches an application - so a tampered or corrupted block is rejected, not served.
Content-defined chunking means identical content across file revisions shares a CID and deduplicates, and the per-project key scoping means dedup never crosses a project boundary.
Managed tier (default)#
The project master key is generated once, wrapped by a hub-held key-encryption-key (KEK), and stored as ciphertext in the org's metadata store. The hub unwraps it only when it needs to act on your behalf - serving a download, generating a thumbnail, or transcoding a preview proxy.
What this protects: your data at rest against the storage layer, a leaked bucket, a lost disk, or a peer on your swarm. What it does not protect against: the hub itself, since the hub custodies the KEK. That is the standard posture of essentially every managed B2B file service (it is what lets us offer previews and search) - and it is why, between two members of the same org, the confidentiality boundary is the block ACL, not cryptography. Every block fetch is authorized per CID against the project's permission model; encryption is the at-rest guarantee, the ACL is the member-to-member guarantee.
Web uploads (small files sent from the browser over TLS) are encrypted by the hub the instant they arrive, before anything is written to storage - so managed data is always ciphertext at rest, matching what the desktop engine writes client-side.
Worth knowing about server-side previews. Because the hub can decrypt managed content, it renders previews on the server: thumbnails, video poster frames, point-cloud images, image/PDF previews, and scrubbable video proxies. The rendering happens in memory. What gets stored depends on the kind of preview:
- Thumbnails, poster frames and point-cloud preview images are stored encrypted under the project's key, with the same AES-256-GCM block encryption as your files, and decrypted only to serve a request that already passed the file's read check. The hub can still decrypt them, exactly as it can decrypt the source files. That is what managed encryption means.
- Still stored unencrypted: scrubbable video proxies (a 720p transcode, kept up to 30 days), the whole-file copies assembled to serve browser downloads and previews (kept up to 24 hours), and the short-lived inputs and outputs of a preview render, deleted when the render finishes.
So a compromise of the storage layer could still expose a video proxy or a recently downloaded file, even though the source blocks stay sealed. The E2E tier removes this residual entirely: on E2E the hub can't decrypt, so no server-side previews or derived copies exist. If preview confidentiality at rest matters for your content, choose E2E.
Key rotation#
Two things can rotate: the KEK and the project key. Rotating the KEK only re-wraps the stored project keys. The keys themselves don't change, so no block needs re-encrypting.
A managed project's key can also be rotated - by the project's owners, its admins, or its creator, from the dashboard's project settings. This adds a new key generation and makes it current. Earlier generations are kept, so every block written before the rotation stays readable, and each block records which generation sealed it (see Storage Format). New content is sealed under the new generation by every client that knows about key generations - the Desktop App, the hub itself, and the git-LFS agent. A rotation doesn't re-encrypt content that's already stored, though, and an old generation is never destroyed. So after a suspected key compromise, content written before the rotation is still protected only by the old key.
An E2E project rotates the same way - the generation model is identical - but the hub never
holds or serves the key, so a member's own device performs the rotation: it mints a new generation
and re-wraps it to every remaining device plus the org recovery key, and the hub only stores the
envelopes the device produced. The same caveat applies: existing content is not
re-encrypted, so it stays readable under the old generation. Revoking a single device is separate -
the hub revokes that device's wrap(s), and the wrapped-key route then answers it a typed
key_access_revoked.
Every key-encryption-key (KEK) is a ring, not a single value: new values are wrapped under the active entry and reads try every entry, so a rotation never makes stored data unreadable. The same discipline covers the managed-tier project-key KEK and every credential the hub stores for a customer.
A rotation is two-step, and the old entry is never dropped in the same step that adds its replacement: stored values are re-sealed under the new entry first, and only then is the old key retired.
Unencrypted tier#
The Free plan - no card required - is restricted to a third tier that is not a variant of managed encryption: it has no project key at all. Blocks are written and stored exactly as they arrive, no AES-256-GCM sealing, no KEK, nothing to wrap or unwrap. A public project on any plan also has to be on this tier, for the same underlying reason: anonymous public serving only works from unencrypted blocks. This is the tier's design, chosen deliberately so a free tier can exist without a per-project key - and, for public projects, because there is nothing to hide from anonymous visitors anyway.
What stays the same as every other tier: the block ACL. An unencrypted project's content is not encrypted, but it is not unauthenticated either - every hub block fetch is still authorized per CID against the project's permission model, the same gate managed and E2E projects use. Plaintext is peer-served only for public projects, and then only with content the anonymous read routes already serve, while a private project's cleartext is never served there. Plaintext removes the storage-at-rest guarantee; it does not remove access control on the hub.
Encryption tier is fixed at project creation and immutable afterward. A project created on The Free plan stays unencrypted even after the org upgrades to a paid plan - upgrading does not retroactively encrypt existing content. A new private project created after upgrading gets managed encryption normally, the same as any Starter/Pro project. If a specific project needs encryption at rest, create it on a paid plan; don't rely on a later upgrade to change an existing Free-plan project's guarantee.
The Free plan can only create public projects (private projects require Starter or above), so its
unencrypted-tier content is also, by definition, intentionally public-visible. On a paid plan a project
is private and managed-encrypted by default, and only enters this tier if it is created as public
- visibility is fixed at creation, never flipped, so this is a create-time choice: the CLI's
swarmfile project create <name> --public (which pins the unencrypted tier with it) or the create API's
"visibility": "public" + "encryptionTier": "plaintext". The Desktop App's New Project dialog and
the web dashboard both create private projects by default, with a Publish publicly card for a
project meant to be open (unencrypted projects remain the Free-plan shape there). See
Permissions for how project visibility works; it's a separate setting from,
but always paired with the unencrypted tier.
Keeping paths out of the public view#
An unencrypted public project's content is public by design, but a publisher chooses which content
actually goes public. A release reads a .swarmfile/publicignore.yml committed at the tag's tree and
leaves matched paths out of the public release entirely - no browse entry, raw stream, download,
archive, file search or Explore - while members keep seeing everything. The same policy can mark
platform-derived data (currently CI-run logs) member-only. The policy file itself is never published,
an invalid file fails the publish instead of publishing everything, and a change applies to future
releases only, so an already-published release needs a new tag to pick it up. The full syntax and the
enforcement points are in Publishing Releases.
Two boundaries worth stating plainly: the filter applies to the public release surface only - share links are not covered, so a link still serves what it points at - and an org policy can only tighten this exposure, never loosen it. It is a publisher control, not a retroactive takedown: copies already downloaded stay where they are.
End-to-end (E2E) tier (opt-in, per project)#
On an E2E project the project key is never wrapped with a hub-held KEK. Instead:
- Each member has an X25519 keypair, generated and enrolled per device; the private key never leaves the member's machine. (Deriving the same keypair from a password via Argon2id - so a member could recover access on a new device without a per-device re-enrollment step - is on the roadmap, not implemented today: every device holds its own independently-generated keypair.)
- The project key is wrapped to each member's public key. Granting a new member = an existing member's client wraps the project key to the newcomer's public key and uploads the (un-openable) wrapped blob. The hub distributes wrapped blobs but never sees the plaintext project key.
This is a content-only guarantee, and we say so up front: file and folder contents become unreadable to the server, but filenames, folder structure, and file sizes stay server-visible so that search, ACLs, browsing, and quota accounting keep working. Encrypting metadata too would break all of those and is a separate, much larger effort - nearly every B2B "E2E" product draws this same line.
To be precise about which metadata: what the server sees is the entry-level metadata it keeps in its database - the name, its place in the folder tree, and the file's total size. It does not see the file's internal structure. The per-file block manifest (the ordered list of content-chunk hashes and their sizes) is itself encrypted alongside the content - the engine runs it through the same block-encryption as the chunks - so on the E2E tier the server holds the total size but the chunk-level layout is opaque to it. (This is why an E2E share is fetched and reassembled entirely in the recipient's browser: the server can't read the manifest to hand out a chunk list, let alone assemble the file.)
One more thing the server does see: block names. Each stored block is named by the BLAKE3 hash of its plaintext chunk - that is what lets identical content deduplicate. So the server can tell when two stored blocks are the same, and someone who already holds a copy of a file could chunk it the way the engine does and check whether those blocks exist in a project. What the server does not keep for E2E content is any whole-file hash: the SHA-256 that git-LFS uses to address a file is never stored for an E2E project, so E2E files can't be matched against published lists of file hashes.
Share links on an E2E project#
A public share link has no account and no device keypair, so it can't receive the project key wrapped
the member way. Instead the link carries the project key wrapped to a per-share random "link key"
that lives only in the URL #fragment - browsers never send a fragment to a server, so the hub
distributes the wrapped blob but never sees the link key or the project key. The recipient's browser
unwraps it and decrypts the file locally.
The link carries the project master key, not a per-file key - the block cipher derives every block's key from the project key, so there is no narrower key to hand out. Two consequences follow:
- The share's block endpoint is scoped to the shared file. The hub serves a share's raw ciphertext blocks through one endpoint, restricted to the shared file's own block set: its manifest plus the chunk hashes the creator's browser enumerated at share time (the hub can't read the E2E manifest to derive them itself).
- That boundary is enforced at the serving layer. Because the recipient genuinely holds the project key, the hub's scoping of the link to the shared file's blocks is an access control, not cryptographic isolation. Per-file share keys - which would make that isolation cryptographic - are a planned change. Treat a single-file share as granting access to that file and nothing else: scope links with expiry and revocation, and share narrowly. Sharing is an explicit trust decision.
The trades E2E makes (know these before you turn it on)#
- Server-side previews/proxies are disabled for E2E content - the server can't decrypt to render them. Image/PDF preview can move to in-browser decryption; server-generated video posters and point-cloud rasters cannot. This is a deliberate, headline trade-off.
- SSO/SCIM grants authentication, not key access. A provisioned user can sign in but sees nothing until an existing member's client (or the optional key-granter, below) wraps the project key to them; when no grant was queued at all (a group-added member, a failed fan-out), the member's own engine asks the hub to open one.
- Recovery is via an org recovery key, or not at all. Losing every device holding a wrapped copy means losing the data. The enterprise-viable answer is an org recovery key: project keys are also wrapped to a recovery public key whose private half is escrowed via Shamir secret sharing (2-of-3 by default) across your admins. This reintroduces a party that can decrypt, so "E2E with recovery" is a point on a spectrum - customers who choose no-recovery accept unrecoverable loss, and we describe it that way at the point of choice. Rotating the recovery key does not re-wrap existing projects: a project stays sealed to the key in force when it was created, recovery accepts shares from any recovery key the org has created, and the previous shares must be kept or that project becomes unrecoverable.
Optional always-on key-granter#
For orgs that add many members while their existing members are offline, an opt-in key-granter service can hold an escrowed recovery share and fulfill pending grants on a tight cadence. It is off by default and, like the recovery key, is a decrypting party you are choosing to introduce for operational convenience - enrolled explicitly per org, never automatically.
Integrity, availability, and network isolation#
- Content integrity. Every block is BLAKE3-verified on arrival; corrupt data from a bad peer, transport, or disk never reaches the application layer.
- Erasure coding. Blocks are protected with adaptive Reed-Solomon 10+4 - any 10 of 14 shards reconstruct the data, so up to 4 slow or offline peers never block a read.
- Private P2P swarm. The peer transport requires a pre-shared key and a custom protocol identifier; a random internet host cannot dial your blocks even if it learns a peer's address. Peers that connect directly do learn each other's network addresses, which is inherent to peer-to-peer transfer; the cloud-only mode below removes that exposure.
- Cloud-only (hub-only) data-plane mode. An org policy (or engine-local
SWARMFILE_HUB_ONLY) disables P2P entirely - QUIC never binds, no peer discovery, no gossip, no peer-IP exposure - and all reads fall through to HTTPS against cloud storage. Re-verified fail-closed on every switch between projects or orgs, not just at startup, so a policy check that can't complete never leaves P2P silently enabled. Toggling it is itself an audited event. - Ransomware detection. The system watches for the signature of encryption malware - a large
number of distinct entries overwritten (by content hash) in a short window, scoped
per-(user, machine) and counting distinct entries rather than raw history rows. The response is
two-tier: crossing the alert bar (a configurable number of distinct entries) only
records an alert-only incident and notifies org owners - it does not block - while crossing the
block bar (a higher multiple of it) is what hard-quarantines the user, halting further writes
even mid-lock. Writes are attributed to the machine that made them: a staged commit carries its
submitting machine, its modifications count in the overwrite signal (the check runs as the commit
lands, so a single 500-entry commit from one machine blocks immediately), and its creates are
creates, never overwrites - while the server-side re-applies a restore, checkout, revert or merge
performs (content the project already held) are excluded, so a legitimate 5,000-file restore does
not trip those bars. A
git pushcarries the pushing machine like any other client batch, so push sweeps count as well. Two accepted exceptions remain: web-app conflict tools (including the batch auto-resolve) carry no machine and are not counted; the separate create-and-delete signature is evaluated inline as a direct delete lands, while a merge that both adds and removes hundreds of entries - applied server-side with no inline check - and an over-cap folder-delete subtree reach it only through the daily scan, so that coverage is best-effort. A delete-only sweep has no signature at all - deletes alone are how anyone empties a folder, and the conjunction is what distinguishes it from encryption-to-new-names. An in-place sweep of hundreds of distinct files from one client can still reach either bar; orgs that routinely rewrite that many files in place can tune the thresholds. Alert-only incidents auto-expire after 7 days, blocks require a human to clear. - Storage blast-radius containment. Every paid org gets dedicated storage by default: its own bucket, reached through a credential scoped to that bucket alone, so a credential minted for your org cannot reach another org's bucket. See Dedicated Storage Isolation.
- Data residency (jurisdiction pinning). An org's metadata and dedicated block storage can be pinned to a real, infrastructure-enforced jurisdiction - EU or US - chosen once at signup and permanent thereafter; free and self-serve on every paid plan, not an Enterprise upsell. An individual project can also pin its own region at creation, independent of the org's: it gets its own metadata namespace and its own region bucket, so its content is never served from the org's storage. FedRAMP / FedRAMP-High storage regions are available on Enterprise, provisioned through sales (Swarmfile is not FedRAMP-authorized). This is a platform-level restriction, not an application-level convention: jurisdiction-pinned storage lives in a namespace invisible to requests that don't carry the matching jurisdiction. See Data Residency.
For the strictest buyers: self-hosted control plane#
Where "the vendor's infrastructure touches our data at all" is disqualifying (some government, defense, and critical-infrastructure work), the E2E tier may still not be enough because metadata is server-visible. For those cases a self-hosted control plane is the Enterprise option: the same code running on infrastructure you control, on a self-hostable runtime compatible with the platform the hosted service is built on. It is available on Enterprise today; a fully air-gapped deployment is on the roadmap - contact us to scope what it would involve for your environment. It is the top of the spectrum; the E2E tier is the middle; managed is the default.
Certifications and controls on the roadmap#
- SOC 2 is on our roadmap (Type I first, Type II thereafter); we have not yet begun a formal audit. In the meantime we answer security questionnaires directly and walk teams through this architecture on request.
- SIEM forwarding - a managed Splunk/Datadog/S3 stream of audit, access, and quarantine events - is planned. Today, a Pro-plan owner can export the activity feed to CSV (up to 100k rows per export), and outbound webhooks can push events to an endpoint you control on a periodic schedule, not in real time.
- Legal hold & retention immutability - a per-entry legal-hold flag overriding retention/trash purge - is planned.
- DLP & egress controls - watermarking, share-link domain allowlists, and export-policy gating - are planned.
If any of these is a hard requirement for you today, contact us to talk through sequencing. For the permission model that is live today, see Permissions; for the E2E setup walkthrough, see End-to-End Encryption Setup.