Blog / Swarmfile: A streaming version-control system

Swarmfile: A streaming version-control system

The TLDR version#

Swarmfile is a modern Version Control System (VCS) built to hold millions of files of multi-petabyte content at practically unlimited scale, scoped to bucket-storage with LAN+WAN peer-to-peer sharing and seeding for minimal latency when cloning repositories.

We use virtual filesystem mounts to expose repositories as local drives - so you can have a multi-petabyte cloud-backed (or EC-sharded) drive that appears quickly and is available to work with, while also taking up minimal local disk-space.

I've mentioned that we support large files, we support git-LFS too, but we also support small files such as code. We're coders ourselves and we wanted to build something that felt like Git. A lot of the concepts still travel over (branches, merging, tags, releases etc.) but there are entirely new concepts that we've tried to document heavily to give you the best start at an entirely new VCS and how it is relevant in the era of multi-agent AI.

Our documentation is here or you can put all of our documentation into an LLM and ask it questions.

The Longer Version#

Why did we build this?#

While operating another startup in late 2025 my cofounder and I (we're both called Matt) got into a conversation with a potential customer who was in desperate need for a solution to a long-term problem.

They worked with very large PDF files that they exported to SharePoint (from the US) and they relied on a manual pipeline of work driven by a large team in India. The Indian team would download, modify and then re-upload the PDF files back to SharePoint for our client to download, double-check, annotate and then re-upload for more changes.

Each round-trip took about nine hours, and colleagues at each office had to reduce their bandwidth consumption during these periods.

If we break that down then:

  1. The organization is paying for SharePoint storage.
  2. Both offices experience reduced productivity during the transfer window.
  3. Egress and bandwidth cost real money for whole file transfers.
  4. In many cases only a few MiB of the files require alteration.
  5. Some colleagues had to download the same files as other colleagues - resulting in duplicated bandwidth for the same file.
  6. Additionally they paid for VPN usage and egress/transfer usage.

I did some research and I figured that the iroh ecosystem could solve a lot of these problems. A Rust-based local engine could coordinate seeding and sharing at both LAN and WAN scale, sharing only the altered chunks when they're needed. To stretch this theoretical system further I figured that each user could optionally offload some of that work to a local-NAS with an LRU cache, or even shard it using Erasure Coding across members of the iroh swarm.

As our implementation progressed we realized that Enterprises need secure controls (e.g. OIDC SSO, SCIM etc.), which led to the build-out of an end-to-end encryption layer, VCS and more, which is essentially where we're at today.

The original use-case (the US-India one) now looks like:

  1. The organization is paying for an S3-API-compatible bucket.
  2. All colleagues see file-changes almost immediately, regardless of location.
  3. Actual file-downloads only download the parts of a file that the application needs, so it feels fast.
  4. No VPN needed, because of E2E encryption at rest, and encryption in transit meets their needs.
  5. The same ACL permission-set is still applied to their files.
  6. Nobody waits nine hours, and bandwidth contention is a thing of the past.

It's fast, very fast, and we've tested it at a scale of millions of files, with changesets and merges in the hundreds of thousands of files, at folder-depths of thousands.

Why would you use this?#

There are a lot of use-cases and capabilities which I've only really touched on very briefly in the past few paragraphs. But you can think of Swarmfile like this:

  • The convenience of Dropbox, with the version-control of Git. It can handle Petabytes of storage, and millions of files. It can take up no space on your machine. Changes from colleagues come through quickly. It is made for multi-agent working.

All of that covers a bunch of complex use-cases which are too numerous to list here - but here are some of the ones unique to Swarmfile:

  • While a user is uploading a file, you can skip to the part of the file that you want to see as the upload is in progress. - Imagine your colleague is uploading a large video, normally you'd have to wait until the file is entirely uploaded to watch some random part of the file. With Swarmfile you can influence which chunks of the file get uploaded so that you can see the meaningful parts sooner. This applies to files of any type and of any size.
  • Multiple AI Agents can work on the same code on a local machine. - Imagine you've started 7 instances of Claude Code and you're furiously working through multiple new features in parallel. Suddenly, Claude has to stop because it realizes that there's multiple overlapping agents editing the same files. With Swarmfile each branch can be mounted as its own drive so that agents can work without treading on each other's toes, and they never have to stop. Or, you can mount the same branch multiple-times and rely on file-locks.
  • Work disconnected, even for large files - Most current working practices dictate that a file must go back to the cloud (or to some central store) before colleagues can see it. With Swarmfile's LAN peer seeding, we can spread your files across your network peers using Erasure-Coding so that files and their changes are transmitted without ever needing to touch the cloud. You can think of the Swarmfile cloud-storage as the backstop for obtaining fresh chunks, not the thing every clone has to go through. In these scenarios it makes sense to have a NAS locally available as a peer for warm-reads.
  • Ransomware is a solved-problem - Right now ransomware works by encrypting company data and then issuing a demand to be able to decrypt it. Swarmfile has built-in ransomware detection to catch mass-file alteration and stop it. Even if we miss it then every file and every change has a Merkle-DAG backed revision history which can be used to rollback file changes.
  • Don't waste compute waiting for "cloning" - Currently if you're cloning, forking or downloading a project from a VCS then you're paying a few different taxes - first of all you have to wait to download the files and (if you're running a cloud-hosted agent) then you're paying for the compute while the download happens. With Swarmfile, neither tax is paid in full since all files are available almost immediately when the mount appears, and files are only downloaded when they're read/needed.

Hopefully we've outlined enough here to catch your curiosity AND your attention, but there are more features to cover and more that we plan to deliver over the coming months and years.

What about Open Source, and vendor lock-in?#

In short, Open Sourcing is underway - part of Swarmfile is already public - and vendor lock-in is something we actively want to avoid.

So on the topic of Open Source, we've already open-sourced the git-LFS transfer agent - the code for swarmfile-lfs is public today - and we will be releasing a lot more in the new year (2027). But we know that Open Source requires a lot of ground-work and capacity-planning that we don't have yet since we're a startup. We're worried that it would overwhelm us.

Regarding vendor lock-in:

  • We already support git-LFS, like many other providers.
  • We have a Git-native interface, there's no reason to stop using Git-tooling.
  • Your data is your own, you can extract it whenever you like.
  • You can bring your own bucket (S3-API-compatible).
  • You can keep a running-copy of all changes in an S3-API-compatible bucket, that we update continuously (it's scoped to a branch).
  • We are already Open Sourcing our stuff - swarmfile-lfs, our git-LFS transfer agent, is public today, published from the monorepo at every release.
  • CI can run on our compute or yours - hosted CI runs your jobs out of the box, and the self-hosted runner plus webhooks let you run work on your own machines and push events to endpoints you control.
  • There are no esoteric requirements to start today - we do not, and will never mandate any one particular ecosystem (e.g. Windows, Azure, AWS etc.).

There's also an interesting story to tell regarding Open Source today on Swarmfile - we donate a $0.50 levy from every paying seat, every month, to an Open Collective project that you choose, or one is chosen randomly for you.

Every paying seat can choose where that $0.50 levy goes, and they get to choose each month.

It is our aspiration to become a Fiscal Host with Open Collective so that projects that chose to Open Source with Swarmfile can more easily collect a proportionate pro-rata share (think Spotify) of that $0.50 levy based on usage (e.g. from clones, forks, release downloads).

But that is a long-road and requires Merchant of Record integration which we are not yet equipped to deal with. But in the meantime we can provide transparency through Open Collective donations.

Is it familiar?#

Yes. It's just a drive (or a folder) - it's also Git compatible, so there's no reason to stop using Git.

We also have approximate feature-parity across our API, CLI and Desktop apps for Windows, Linux and macOS.

The starting point is the "mount" or (shared-drive). If a user doesn't want to dig into anything technical with Swarmfile then they never need to know anything more than just their shared-drive. They can save and read files from S:\ and just do their day-to-day work like they normally do without any problem.

If you're coming from Git then familiarity will be important - most of our tooling has similar wording to Git and will try to point you in the right direction where the wording isn't exactly the same. Also the shared-drive might feel a little "spooky", since we instantly "mount a Project in a volume" (instead of "cloning a repo to a folder"). But you can still reach for Git if you want.

If you open our Desktop app today then you will find that it has the approximate look and feel of both Dropbox Desktop and GitHub Desktop - we hope that this interface is idiomatic and familiar for all users and it supports features for all levels of technical capability, including more advanced features that a user might expect to be CLI only.

For a more advanced user there is a command-line-interface (CLI) which offers slightly more features than the Desktop app and can be used to instrument CI runners, render-farms etc. and we have comprehensive documentation to support it.

Broadly, much of the wording and tooling is similar to Git with some deviations. Swarmfile exists across a lot of competing spaces, we could have borrowed terminologies from SVN, Perforce, Mercurial, AccuRev or others, but Git is familiar to the authors and hopefully to many of our users.

Are we Git?#

No, but we're Git-native and our collaboration is GitHub-shaped. What that means is that we have an opinionated Git-native surface as an easy on-ramp for adoption, although you won't get all the advantages that Swarmfile has to offer by using this.

As a streaming, peer-to-peer virtual-filesystem we're a big departure from Git and it gives us the freedom to do things which Git cannot do, while delivering the guarantees that Git offers.

  • Millions of files - tested and proven. Our limit isn't space, it's time - having more files means that some operations may take more time.
  • Thousands of folders deep - tested and proven. If you need two thousand folders of depth with ACL permissions applied, then we can do it.
  • Massive merges - proven to hundreds of thousands of files, we wouldn't recommend it, but as above the limitation is time.
  • Secure - granular permissions means that users never see files you don't want them to see.
  • Native binary object storage - we were initially built for large binary blobs.
  • Native S3-compatible bucket write - anything S3-API compatible.
  • Native ACL permissions integration - we work well with Windows.
  • And more...

You can use Git today with Swarmfile by installing our tools and then cloning a project with git clone swarmfile://<org>/<project>.

We've just started#

October 2026 is our go-live period, we've just started and we intend to keep going. We're a bootstrapped startup hoping that we've made enough of an impression for you to give us a try - and in return we'll keep you updated, we'll keep improving, we'll reward open-source and we'll listen to your needs.

Hope to see you in the swarm.

Matt.