# Security Architecture

This is the deep technical companion to [Security](https://swarmfile.com/docs/admin/security) - written to be handed
directly to a security team running a vendor review. It states plainly what each layer protects
against, who can decrypt what, and where the limits of that protection are. If your review needs
something not covered here, [contact us](https://swarmfile.com/contact) and we'll answer specifics.

## The one-paragraph version

On a **paid plan** (Starter and above), **private** project content is encrypted with AES-256-GCM
before it is stored - on the machine that writes it for anything saved through the drive, and by the
hub for browser uploads to a managed project (the hub holds a managed project's key; an end-to-end
project's key never reaches it). Peer-to-peer nodes and our cloud object storage see only ciphertext
addressed by a BLAKE3 content hash, with a few documented exceptions: the
[git-LFS](https://swarmfile.com/docs/guides/git-lfs) direct-upload path (see below); server-side **video proxies** and the
**whole-file copies** assembled to serve browser downloads (both retained short-term - see
"Worth knowing about server-side previews" below); and any
**public** project, which is stored unencrypted by design - see [Unencrypted tier](#unencrypted-tier).
Free-plan projects are all public and use that tier because they have no other, and a paying org
can still choose it deliberately for a public project that needs the anonymous-download paths
(which read unencrypted content only). Because a public project's content is public, its plaintext
blocks may also be served to peers on the org's private swarm - peers reach only what the anonymous
read routes already serve, peer serving is read-only, and every fetched block is BLAKE3-verified
before it is used. A paying org's **private** projects are encrypted - managed by default, E2E
opt-in - which is also what makes them eligible for peer-to-peer transfer. The **only** party that can turn ciphertext back into your files is whoever holds
the project key - and *you choose who that is*. On the default **managed** tier that party is the
Swarmfile hub, so we can give you server-side previews, search, and proxies - managed means
*encrypted at rest with hub-held keys*, not zero-knowledge. On the opt-in
**end-to-end (E2E)** tier the key never reaches our servers in a form we can open, so we structurally
cannot read your content. Both tiers are fully supported; the difference is a deliberate
trade you make per project, not a weakness in one of them.

The **Free** plan is the one plan with no at-rest encryption at all: its projects are stored
**unencrypted**, with no project key - the trade-off for no-card, no-cost storage. See
[Unencrypted tier](#unencrypted-tier) below. On Starter and above, a **private** project gets managed
encryption by default (E2E opt-in); a **public** project is unencrypted-only on every plan, because
public releases and archives still read unencrypted content only. The unencrypted tier is therefore
Free-plan projects plus a paying org's deliberately-public projects.

## Threat model - what each tier defends against

| Adversary | Managed tier | E2E tier | Unencrypted tier (Free plan; a paid org's explicit public-project opt-out) |
|---|---|---|---|
| A malicious/curious **peer** on your swarm | ✅ sees only ciphertext + BLAKE3 CIDs | ✅ same | ❌ sees plaintext content - which is public by design (peers see only what anonymous readers see) |
| The **storage provider** or anyone who exfiltrates a storage bucket | ✅ ciphertext at rest, no keys co-located | ✅ same | ❌ plaintext at rest |
| Someone who **subpoenas or compromises the Swarmfile hub** | ⚠️ hub holds the wrapping key and *can* decrypt | ✅ hub holds only un-openable wrapped blobs | ❌ nothing to compromise - already plaintext |
| A **Swarmfile employee** with production access | ⚠️ same as above - possible, audited, not prevented | ✅ prevented by construction | ❌ same as above |
| A **non-member of the project** (incl. other orgs, other members) | ✅ blocked by the block ACL on every hub fetch | ✅ blocked by ACL *and* has no wrapped key | ✅ nothing confidential to reach (public by design); hub fetches stay ACL-gated, and swarm peers can only fetch blocks the anonymous read routes already serve |
| A **lost password / lost device** | ✅ recoverable (hub re-wraps) | ⚠️ recoverable only via the org recovery key | ✅ no key to lose |

**Exception - git-LFS objects.** The block properties above describe project content written through
the drive/engine. A project used as a [git-LFS](https://swarmfile.com/docs/guides/git-lfs) server is not uniformly
encrypted at rest: objects uploaded through an agent's chunked block path are encrypted on a managed
project (the desktop agent client-side; the standalone agent's plaintext blocks are encrypted by the
hub on write), but two direct-upload paths store the object unencrypted - a stock `git-lfs push`, and
the desktop app's built-in transfer agent for objects of 64 MiB or more. That is one of the documented
exceptions to "ciphertext at rest", and it is a deliberate trade: LFS is an add-on surface, and
encrypting and decrypting multi-GB artifacts on every transfer isn't worth it. A binary that must be
encrypted at rest belongs on the mounted drive (managed or E2E), not routed through LFS.

The two ⚠️ rows on the managed tier are the entire reason the E2E tier exists. If "the vendor can
technically decrypt our footage" fails your review, the answer is the E2E tier - not a promise about
the managed one. The unencrypted tier is a different trade entirely: it is not a weaker version of
managed encryption, it is the deliberate absence of at-rest encryption in exchange for a free,
no-card tier - access is still gated by the same block ACL every other tier uses, but there is no
encryption layer to defend the *storage* boundary at all.

## How block encryption works (both tiers)

Encryption of the bytes is identical on both tiers; only *custody of the project key* differs.

- Each project has one 32-byte **project master key**.
- Every content block gets its own key, derived from the project key and bound to the block's
  content (the CID is the BLAKE3 hash of the *plaintext* chunk), so a key is never reused
  across blocks.
- The block is sealed with **AES-256-GCM** and keyed in object storage by its CID. (The exact
  block format is in the [Storage Format](https://swarmfile.com/docs/reference/storage-format) spec.)
- On read, the reader re-derives the key, decrypts, and re-verifies the BLAKE3 CID before any byte
  reaches an application - so a tampered or corrupted block is rejected, not served.

Content-defined chunking means identical content across file revisions shares a CID and deduplicates,
and the per-project key scoping means dedup never crosses a project boundary.

## Managed tier (default)

The project master key is generated once, **wrapped by a hub-held key-encryption-key (KEK)**, and
stored as ciphertext in the org's metadata store. The hub unwraps it only when it needs to act on your
behalf - serving a download, generating a thumbnail, or transcoding a preview proxy.

What this protects: your data at rest against the storage layer, a leaked bucket, a lost disk, or a
peer on your swarm. What it does **not** protect against: the hub itself, since the hub custodies the
KEK. That is the standard posture of essentially every managed B2B file service (it is what
lets us offer previews and search) - and it is why, between two members of the same org, the
confidentiality boundary is the **block ACL**, not cryptography. Every block fetch is authorized per
CID against the project's permission model; encryption is the at-rest guarantee, the ACL is the
member-to-member guarantee.

Web uploads (small files sent from the browser over TLS) are encrypted by the hub the instant they
arrive, before anything is written to storage - so managed data is always ciphertext at rest,
matching what the desktop engine writes client-side.

**Worth knowing about server-side previews.** Because the hub *can* decrypt managed content, it
renders previews on the server: thumbnails, video poster frames, point-cloud images, image/PDF
previews, and scrubbable video proxies. The rendering happens in memory. What gets stored depends on
the kind of preview:

- **Thumbnails, poster frames and point-cloud preview images are stored encrypted** under the
  project's key, with the same AES-256-GCM block encryption as your files, and decrypted only to serve
  a request that already passed the file's read check. The hub can still decrypt them, exactly as it
  can decrypt the source files. That is what managed encryption means.
- **Still stored unencrypted:** scrubbable **video proxies** (a 720p transcode, kept up to 30 days),
  the **whole-file copies** assembled to serve browser downloads and previews (kept up to 24 hours),
  and the short-lived inputs and outputs of a preview render, deleted when the render finishes.

So a compromise of the storage layer could still expose a video proxy or a recently downloaded file,
even though the source blocks stay sealed. The **E2E tier** removes this residual entirely: on E2E the
hub can't decrypt, so no server-side previews or derived copies exist. If preview confidentiality at
rest matters for your content, choose E2E.

### Key rotation

**Two things can rotate: the KEK and the project key.** Rotating the KEK only re-wraps the stored
project keys. The keys themselves don't change, so no block needs re-encrypting.

A managed project's key can also be rotated - by the project's owners, its admins, or its creator,
from the dashboard's project settings. This adds a new key *generation* and makes it current. Earlier generations are kept, so
every block written before the rotation stays readable, and each block records which generation
sealed it (see [Storage Format](https://swarmfile.com/docs/reference/storage-format#key-rotation)). New content is
sealed under the new generation by every client that knows about key generations - the Desktop
App, the hub itself, and the git-LFS agent. A rotation doesn't re-encrypt content that's already stored,
though, and an old generation is never destroyed. So after a suspected key compromise, content
written before the rotation is still protected only by the old key.

An **E2E** project rotates the same way - the generation model is identical - but the hub never
holds or serves the key, so a member's own device performs the rotation: it mints a new generation
and re-wraps it to every remaining device plus the org recovery key, and the hub only stores the
envelopes the device produced. The same caveat applies: existing content is not
re-encrypted, so it stays readable under the old generation. Revoking a single device is separate -
the hub revokes that device's wrap(s), and the wrapped-key route then answers it a typed
`key_access_revoked`.

Every key-encryption-key (KEK) is a **ring**, not a single value: new values are
wrapped under the active entry and reads try every entry, so a rotation never
makes stored data unreadable. The same discipline covers the managed-tier
project-key KEK and every credential the hub stores for a customer.

A rotation is two-step, and the old entry is never dropped in the same step
that adds its replacement: stored values are re-sealed under the new entry
first, and only then is the old key retired.

## Unencrypted tier

The Free plan - no card required - is restricted to a third tier that is not a variant of managed
encryption: it has **no project key at all**. Blocks are written and stored exactly as they arrive, no
AES-256-GCM sealing, no KEK, nothing to wrap or unwrap. A **public project on any plan** also has to be
on this tier, for the same underlying reason: anonymous public serving only works from unencrypted
blocks. This is the tier's design, chosen deliberately so a free tier can exist without a per-project
key - and, for public projects, because there is nothing to hide from anonymous visitors anyway.

What stays the same as every other tier: the **block ACL**. An unencrypted project's content is not
encrypted, but it is not *unauthenticated* either - every hub block fetch is still authorized per CID
against the project's permission model, the same gate managed and E2E projects use. Plaintext is
peer-served only for **public** projects, and then only with content the anonymous read routes already
serve, while a private project's cleartext is never served there. Plaintext removes the
storage-at-rest guarantee; it does not remove access control on the hub.

**Encryption tier is fixed at project creation and immutable afterward.** A project created on
The Free plan stays unencrypted even after the org upgrades to a paid plan - upgrading does not retroactively
encrypt existing content. A new **private** project created after upgrading gets managed encryption
normally, the same as any Starter/Pro project. If a specific project needs encryption at rest, create
it on a paid plan; don't rely on a later upgrade to change an existing Free-plan project's guarantee.

The Free plan can only create **public** projects (private projects require Starter or above), so its
unencrypted-tier content is also, by definition, intentionally public-visible. On a paid plan a project
is private and managed-encrypted by default, and only enters this tier if it is created as **public**
- visibility is fixed at creation, never flipped, so this is a create-time choice: the CLI's
`swarmfile project create <name> --public` (which pins the unencrypted tier with it) or the create API's
`"visibility": "public"` + `"encryptionTier": "plaintext"`. The Desktop App's New Project dialog and
the web dashboard both create private projects by default, with a **Publish publicly** card for a
project meant to be open (unencrypted projects remain the Free-plan shape there). See
[Permissions](https://swarmfile.com/docs/admin/permissions) for how project visibility works; it's a separate setting from,
but always paired with the unencrypted tier.

### Keeping paths out of the public view

An unencrypted public project's content is public by design, but a publisher chooses *which* content
actually goes public. A release reads a `.swarmfile/publicignore.yml` committed at the tag's tree and
leaves matched paths out of the public release entirely - no browse entry, raw stream, download,
archive, file search or Explore - while members keep seeing everything. The same policy can mark
platform-derived data (currently CI-run logs) member-only. The policy file itself is never published,
an invalid file fails the publish instead of publishing everything, and a change applies to future
releases only, so an already-published release needs a new tag to pick it up. The full syntax and the
enforcement points are in [Publishing Releases](https://swarmfile.com/docs/guides/publishing-releases#keeping-paths-out-of-a-release).

Two boundaries worth stating plainly: the filter applies to the public release surface only - **share
links are not covered**, so a link still serves what it points at - and an org policy can only
*tighten* this exposure, never loosen it. It is a publisher control, not a retroactive takedown:
copies already downloaded stay where they are.

## End-to-end (E2E) tier (opt-in, per project)

On an E2E project the project key is **never wrapped with a hub-held KEK**. Instead:

- Each member has an **X25519 keypair**, generated and enrolled per device; the private key never
  leaves the member's machine. (Deriving the same keypair from a password via Argon2id - so a member
  could recover access on a new device without a per-device re-enrollment step - is on the roadmap,
  not implemented today: every device holds its own independently-generated keypair.)
- The project key is **wrapped to each member's public key**. Granting a new member = an existing
  member's client wraps the project key to the newcomer's public key and uploads the (un-openable)
  wrapped blob. The hub distributes wrapped blobs but never sees the plaintext project key.

This is a **content-only** guarantee, and we say so up front: file and folder *contents* become
unreadable to the server, but **filenames, folder structure, and file sizes stay server-visible** so
that search, ACLs, browsing, and quota accounting keep working. Encrypting metadata too would break
all of those and is a separate, much larger effort - nearly every B2B "E2E" product draws this same
line.

To be precise about which metadata: what the server sees is the *entry-level* metadata it keeps in
its database - the name, its place in the folder tree, and the file's total size. It does **not** see
the file's internal structure. The per-file block **manifest** (the ordered list of content-chunk
hashes and their sizes) is itself encrypted alongside the content - the engine runs it through the
same block-encryption as the chunks - so on the E2E tier the server holds the total size but the
chunk-level layout is opaque to it. (This is why an E2E share is fetched and reassembled entirely in
the recipient's browser: the server can't read the manifest to hand out a chunk list, let alone
assemble the file.)

One more thing the server does see: **block names**. Each stored block is named by the BLAKE3 hash
of its *plaintext* chunk - that is what lets identical content deduplicate. So the server can tell
when two stored blocks are the same, and someone who already holds a copy of a file could chunk it
the way the engine does and check whether those blocks exist in a project. What the server does
**not** keep for E2E content is any whole-file hash: the SHA-256 that git-LFS uses to address a file
is never stored for an E2E project, so E2E files can't be matched against published lists of
file hashes.

### Share links on an E2E project

A public share link has no account and no device keypair, so it can't receive the project key wrapped
the member way. Instead the link carries the project key wrapped to a **per-share random "link key"**
that lives only in the URL `#fragment` - browsers never send a fragment to a server, so the hub
distributes the wrapped blob but never sees the link key or the project key. The recipient's browser
unwraps it and decrypts the file locally.

The link carries the **project master key**, not a per-file key - the block cipher derives every
block's key from the project key, so there is no narrower key to hand out. Two consequences follow:

- **The share's block endpoint is scoped to the shared file.** The hub serves a share's raw ciphertext
  blocks through one endpoint, restricted to the shared file's own block set: its manifest plus the
  chunk hashes the creator's browser enumerated at share time (the hub can't read the E2E manifest to
  derive them itself).
- **That boundary is enforced at the serving layer.** Because the recipient genuinely holds the project
  key, the hub's scoping of the link to the shared file's blocks is an access control, not
  cryptographic isolation. Per-file share keys - which would make that isolation cryptographic - are a
  planned change. Treat a single-file share as granting access to that file and nothing else: scope
  links with expiry and revocation, and share narrowly. Sharing is an explicit trust decision.

### The trades E2E makes (know these before you turn it on)

- **Server-side previews/proxies are disabled** for E2E content - the server can't decrypt to render
  them. Image/PDF preview can move to in-browser decryption; server-generated video posters and
  point-cloud rasters cannot. This is a deliberate, headline trade-off.
- **SSO/SCIM grants authentication, not key access.** A provisioned user can sign in but sees nothing
  until an existing member's client (or the optional key-granter, below) wraps the project key to
  them; when no grant was queued at all (a group-added member, a failed fan-out), the member's own
  engine asks the hub to open one.
- **Recovery is via an org recovery key, or not at all.** Losing every device holding a wrapped copy
  means losing the data. The enterprise-viable answer is an **org recovery key**: project keys are
  *also* wrapped to a recovery public key whose private half is escrowed via **Shamir secret sharing
  (2-of-3 by default)** across your admins. This reintroduces a party that *can* decrypt, so "E2E
  with recovery" is a point on a spectrum - customers who choose no-recovery accept unrecoverable
  loss, and we describe it that way at the point of choice. Rotating the recovery key does not
  re-wrap existing projects: a project stays sealed to the key in force when it was created,
  recovery accepts shares from any recovery key the org has created, and the previous shares must
  be kept or that project becomes unrecoverable.

### Optional always-on key-granter

For orgs that add many members while their existing members are offline, an opt-in **key-granter**
service can hold an escrowed recovery share and fulfill pending grants on a tight cadence. It is
off by default and, like the recovery key, is a decrypting party you are choosing to introduce for
operational convenience - enrolled explicitly per org, never automatically.

## Integrity, availability, and network isolation

- **Content integrity.** Every block is BLAKE3-verified on arrival; corrupt data from a bad peer,
  transport, or disk never reaches the application layer.
- **Erasure coding.** Blocks are protected with adaptive Reed-Solomon 10+4 - any 10 of 14 shards
  reconstruct the data, so up to 4 slow or offline peers never block a read.
- **Private P2P swarm.** The peer transport requires a pre-shared key and a custom protocol
  identifier; a random internet host cannot dial your blocks even if it learns a peer's address.
  Peers that connect directly do learn each other's network addresses, which is inherent to
  peer-to-peer transfer; the cloud-only mode below removes that exposure.
- **Cloud-only (hub-only) data-plane mode.** An org policy (or engine-local `SWARMFILE_HUB_ONLY`) disables P2P
  entirely - QUIC never binds, no peer discovery, no gossip, no peer-IP exposure - and all reads fall
  through to HTTPS against cloud storage. Re-verified fail-closed on every switch between projects or
  orgs, not just at startup, so a policy check that can't complete never leaves P2P silently enabled.
  Toggling it is itself an audited event.
- **Ransomware detection.** The system watches for the signature of encryption malware - a large
  number of *distinct entries* overwritten (by content hash) in a short window, scoped
  per-(user, machine) and counting distinct entries rather than raw history rows. The response is
  two-tier: crossing the **alert bar** (a configurable number of distinct entries) only
  records an alert-only incident and notifies org owners - it does *not* block - while crossing the
  **block bar** (a higher multiple of it) is what hard-quarantines the user, halting further writes
  even mid-lock. Writes are attributed to the machine that made them: a staged commit carries its
  submitting machine, its modifications count in the overwrite signal (the check runs as the commit
  lands, so a single 500-entry commit from one machine blocks immediately), and its creates are
  creates, never overwrites - while the server-side re-applies a restore, checkout, revert or merge
  performs (content the project already held) are excluded, so a legitimate 5,000-file restore does
  not trip those bars. A `git push` carries the pushing machine like any other client batch, so push sweeps
  count as well. Two accepted exceptions remain: web-app conflict tools (including the batch
  auto-resolve) carry no machine and are not counted; the separate
  create-and-delete signature is evaluated inline as a direct delete lands, while a merge that both
  adds and removes hundreds of entries - applied server-side with no inline check - and an over-cap
  folder-delete subtree reach it only through the daily scan, so that coverage is best-effort. A
  delete-only sweep has no signature at all - deletes alone are how anyone empties a folder, and
  the conjunction is what distinguishes it from encryption-to-new-names. An in-place sweep of hundreds of distinct files from one client can still reach either
  bar; orgs that routinely rewrite that many files in place can tune the thresholds. Alert-only
  incidents auto-expire after 7 days, blocks require a human to clear.
- **Storage blast-radius containment.** Every **paid** org gets **dedicated storage** by default: its own
  bucket, reached through a credential scoped to that bucket alone, so a credential minted for your org
  cannot reach another org's bucket. See [Dedicated Storage Isolation](https://swarmfile.com/docs/admin/dedicated-storage).
- **Data residency (jurisdiction pinning).** An org's metadata and dedicated block storage can be
  pinned to a real, infrastructure-enforced jurisdiction - EU or US - chosen once at signup and
  permanent thereafter; free and self-serve on every paid plan, not an Enterprise upsell.
  An individual project can also pin its **own** region at creation, independent of the org's: it
  gets its own metadata namespace and its own region bucket, so its content is never served from the
  org's storage.
  FedRAMP / FedRAMP-High storage regions are available on Enterprise, provisioned through sales
  (Swarmfile is not FedRAMP-authorized). This is a
  platform-level restriction, not an application-level convention: jurisdiction-pinned storage lives
  in a namespace invisible to requests that don't carry the matching jurisdiction.
  See [Data Residency](https://swarmfile.com/docs/admin/data-residency).

## For the strictest buyers: self-hosted control plane

Where "the vendor's infrastructure touches our data at all" is disqualifying (some government,
defense, and critical-infrastructure work), the E2E tier may still not be enough because metadata is
server-visible. For those cases a **self-hosted control plane** is the Enterprise option: the same
code running on infrastructure you control, on a self-hostable runtime compatible with the platform
the hosted service is built on. It is available on Enterprise today; a fully air-gapped deployment is
on the roadmap - [contact us](https://swarmfile.com/contact) to scope what it would involve for your environment. It is the top of the spectrum; the E2E tier is the
middle; managed is the default.

## Certifications and controls on the roadmap

- **SOC 2** is on our roadmap (Type I first, Type II thereafter); we have not yet begun a formal
  audit. In the meantime we answer security questionnaires directly and walk teams through this
  architecture on request.
- **SIEM forwarding** - a managed Splunk/Datadog/S3 stream of audit, access, and quarantine events - is
  planned. Today, a Pro-plan owner can export the [activity feed](https://swarmfile.com/docs/admin/operations) to CSV (up to
  100k rows per export), and [outbound webhooks](https://swarmfile.com/docs/admin/webhooks) can push events to an endpoint
  you control on a periodic schedule, not in real time.
- **Legal hold & retention immutability** - a per-entry legal-hold flag overriding retention/trash
  purge - is planned.
- **DLP & egress controls** - watermarking, share-link domain allowlists, and export-policy gating -
  are planned.

If any of these is a hard requirement for you today, [contact us](https://swarmfile.com/contact) to talk through
sequencing. For the permission model that *is* live today, see
[Permissions](https://swarmfile.com/docs/admin/permissions); for the E2E setup walkthrough, see
[End-to-End Encryption Setup](https://swarmfile.com/docs/admin/end-to-end-encryption).
