Browse docs
Docs / Guides / Coding Agents

Coding agents (mounted workspaces)

A Swarmfile project mounts as a real filesystem path. That makes it a good workspace for a coding agent: the agent reads and writes ordinary files, and every change is versioned, lockable, and shared - no git clone, no git worktree per task, and no full re-install of dependencies for every parallel checkout.

The shape this guide builds:

  • one engine process per machine (or per agent when you want process isolation - see below),
  • one project API key so the engine runs headless (no browser sign-in),
  • one mount per agent, each with a label and (optionally) its own branch,
  • one shared scratch cache per project per machine, so a second mount reuses the first one's downloaded packages,
  • and the durable create queue, which is what makes a burst of file creations from an agent safe.

The same primitives appear in the Desktop App - this guide is the scripted, headless form, and every command in it was run against real mounts on macOS and Windows before it was published (see Measured behavior near the end).

What each agent gets, and what it must not be given#

Give an agent:

  • The mount path. It's a normal POSIX/Windows path (/private/tmp/sf-ex-b/…, X:\…). Everything the agent does happens there; there is no separate "sync" step.
  • The scratch dir (below) as the place for tool caches.
  • Exit codes (below) so its automation can tell "you need to do something" (2) from "something failed" (1).
  • One SWARMFILE_* environment set per engine - never two engines sharing a cache directory.

Do not hand an agent:

  • A second mount point that points at the same directory another agent uses - mount points are exclusive to one engine process.
  • A personal refresh token copied off a workstation. Use a project API key: it's revocable, it's walled off to one project, and it can't act as the human who minted it.

Platform matrix: how many mounts per engine#

PlatformSeveral mounts in one engine?Recommended shape for N agents
Windows (WinFsp)Yes - mounts open takes a drive letter (Y:, Z:, …), auto for the next free one, or an empty folder (a folder with anything in it is refused)One engine, N mounts, one API key
Linux (FUSE)Yes - mounts open takes directory paths (same engine path as Windows; macOS/Linux use a folder, not a drive letter)One engine, N mounts, one API key
macOS (FUSE-T)Yes - mounts open takes directory paths; the FUSE-T library is patched for multiple live mounts (N is bounded by the go-nfsv4 helper port range below)One engine, N mounts, one API key

An engine can serve every agent's mount: auth is one API key, and the shared scratch cache (below) is per machine, so the dependency cache is still paid for once. Running one engine per agent remains a valid isolation choice (a wedged mount helper then can't affect siblings), just no longer a macOS requirement. Each FUSE-T mount costs one go-nfsv4 helper and one port from a fixed 18-port loopback range (52100-52117, shared machine-wide), so a Mac tops out at 18 mounts per machine - if an open fails saying the range is exhausted, close a mount or clear leaked go-nfsv4 processes.

1. Headless engine with a project API key#

The commands in this guide use the swarmfile CLI, which ships with every platform installer (Linux .deb, the Windows MSI, the macOS .pkg). A headless box needs only that package - no GUI sign-in - and the CLI resolves the running engine's control socket on its own. See CLI: swarmfile for the per-platform install locations.

Mint a key as an org admin or owner, from any running engine:

swarmfile api-key create agent-node-1                 # this mount's project
swarmfile api-key create agent-node-2 --project proj_… # or an explicit project id

--project takes a project id, not a slug (a slug comes back as an opaque "hub returned unexpected status 502" - an error-surface wart worth fixing later). The secret is printed once (sf_key_…). List and revoke keys with swarmfile api-key list and swarmfile api-key revoke <id>. A key is walled off to exactly one project, so a compromised agent node can't reach anything else in the org.

Keys count against your plan's headless-key allowance - Starter 2 per seat (minimum 5), Pro 5 per seat (minimum 10) - and a key pack adds 10 for $20/month. One key per machine/fleet is the recommended pattern (revocation radius and its own rate-limit bucket); a 402 with code: keys_cap means the org is at its cap, so revoke an unused key or add a pack under Settings → Billing. Agents running through a workstation's existing mount need no key at all.

To give each agent its own engine (process isolation; otherwise one engine can host every agent's mount, as the matrix above recommends), point each at its own cache dir and mount point:

export SWARMFILE_API_KEY=sf_key_…
export SWARMFILE_ORG_ID=org_…
export SWARMFILE_PROJECT_ID=proj_…
export SWARMFILE_CACHE_DIR="$HOME/.swarmfile/agent-1"   # one per engine
export SWARMFILE_MOUNT_POINT="$HOME/agents/agent-1"     # one per engine
export SWARMFILE_MACHINE_ID=agent-node-1
swarmfile-engine

With SWARMFILE_API_KEY set, the engine skips sign-in entirely - no browser, no refresh token. (If both an API key and a refresh token are set, the key wins and the engine says so in its log.) See the api-key CLI reference for the full flags, and Deployment Topologies for putting these on dedicated boxes.

2. One mount per agent (and targeting it)#

The engine's first mount is the boot mount, id default. Additional mounts are full, symmetric mounts - each with its own changelists, metadata, and lock manager:

swarmfile mounts list
swarmfile mounts open Y: --label agent-2 --branch feature-x
swarmfile mounts open Z: --label agent-3 --project proj_other
swarmfile mounts status agent-2-id
swarmfile mounts current agent-2-id

The examples use Windows drive letters; on macOS and Linux the mount point is a directory path instead (e.g. swarmfile mounts open /Volumes/agent-2 --label agent-2). Everything else - labels, ids, --branch, --project, status, current - is the same on every platform.

  • --label is for humans; the id (a long hex string, or default for the boot mount) is what scripts use - take it from mounts list --json.
  • --branch <name> pins that mount to a branch. It opens (or reuses) a divergent (project, branch) scope, so the mount works that branch while every other mount stays on its own - and can be switched later with mounts branch. A unique directory root per agent (e.g. /agents/<name>/…) is still good hygiene, but no longer required: the engine now brings missing ancestor directories into existence on a directory create, so a fresh branch is safe to write into anywhere.
  • --project <id> opens a mount on a different project in the same org.
  • --exclusive opts that mount into exclusive-write mode: while it has a file open for writing, another mount opened with the same flag on the same project+branch has its write-open refused (the app sees EAGAIN, or "file in use" on Windows) - a rename, delete, or atomic-save replace of that file is refused too - instead of both editing the file and letting whichever closes last win. Off by default, fixed at open, and shown as a lock badge in the drive list. Open every agent mount with it when several agents share one tree: two mounts of one engine are a single identity to the hub, so ordinary locks do not separate them. An engine older than the flag ignores it silently - the CLI warns when a successful open didn't take it, so upgrade the engine if that warning appears.
  • Every command that isn't under mounts targets the current mount unless you pass the global flag:
swarmfile --mount agent-2-id status
swarmfile --mount agent-2-id commit -m "agent 2 results"

Close a mount with swarmfile mounts close <id>. It refuses while a changelist is open (--force abandons it) and refuses the boot mount outright (exit 2, is_default_mount) - closing that one means shutting down the engine.

3. Shared dependency cache#

Every mount of a project on a machine can share one scratch directory for tool caches (a pnpm store, a cargo registry, a pip wheel cache). Declare what run should redirect, in a committed file at the project root:

# .swarmfile/cache.yml
caches:
  - name: pnpm
    env: PNPM_HOME
    subdir: pnpm
  - name: cargo-registry
    env: CARGO_HOME
    subdir: cargo

swarmfile scratch init writes a commented starter file. Then run installs and builds through run, so the tool sees the shared cache:

swarmfile mounts scratch-dir default     # the shared dir + declared redirects
swarmfile run -- pnpm install            # PNPM_HOME -> shared scratch
swarmfile run -- cargo build --release   # CARGO_HOME -> shared scratch

run passes stdio and the exit code through, and it never refuses to run your command over a cache problem - with no config, a bad entry, or an unreachable engine it warns and runs the command unmodified.

Maintenance:

swarmfile scratch list
swarmfile scratch gc --older-than-days 14 --dry-run
swarmfile scratch clear <name> --force

One scratch dir per project per machine, shared regardless of branch. Only redirect read-mostly, content-addressed caches - not build output directories, which would contend between concurrent mounts. Nothing in a scratch cache is unrecoverable project data; clearing one only costs a slower next build. Full reference: scratch and run.

4. Create bursts and the pending-create queue#

File creation no longer waits for the hub. On the mounted drive, a create (or mkdir, or symlink) returns at once under a locally minted id; a background worker confirms it with the hub (8 in flight by default) and peers on the LAN see the name by gossip before the hub does. Creates that are due together go to the hub together, up to 50 in one request, and a new file's first save travels with its create instead of in requests of its own. This is what makes an agent's scaffolding burst safe and fast: 5,000 new files reach the hub in under half a minute without losing files (measured below), and a create made while the hub is unreachable simply queues.

What to check, and what an agent will see:

swarmfile status                 # "new files: N waiting to reach the hub (oldest …)" while queued
swarmfile --json status | jq .pendingCreates
pendingCreates fieldMeaning
queuedCreates this machine made that the hub hasn't confirmed. 0 = drained.
canceledDeleted locally before the hub heard of them; no hub call needed.
stuckQueued and failed 6+ times - hub refusing or unreachable for these.
oldestAgeSecsAge of the oldest queued create. Growing with queued unchanged means nothing is draining.
peerUnconfirmedEntries peers gossiped as created that this machine hasn't seen confirmed.
itemsThe oldest 20 queued creates (path, type, attempts, age).

Rules of the road while something is pending:

  • A name collision on the hub renames, it doesn't overwrite. If the hub already has a different file at that path when the queued create lands, your copy is kept as name (conflict-xxxxxxxx).ext and the error ring says so. A folder with the same name is merged into the hub's folder (like mkdir -p) - files you created inside it land there, and the hub's existing contents appear after a few seconds.
  • Operations on a just-created file wait, then answer 409. Locking, ACL reads, share-link creation and comment writes on a pending file wait up to 8 seconds for the create to confirm and then proceed. A change (a lock, share, ACL or comment write) sends the file's create at once, even while its first save is still uploading, so it doesn't wait behind that save. If the hub can't be reached in that time they answer 409 with code: "pending_create" - retry shortly. For an entry another machine created and the hub hasn't confirmed, the refusal is immediate.
  • A create the hub refuses for good is retracted from peers and reported once in the error ring (create failed for "…"). An empty file, a symlink, or a folder is removed. A file you had already written content into is kept on this computer only (never uploaded) so the work isn't lost - the message says so. Content in files under a refused folder is discarded, and the count is in the message.
  • Deleting or renaming a just-created file is passed to other machines immediately. A file a peer created that never reaches the hub is removed after 6 hours with a message.
  • A new file can reach the hub up to a minute after it appears locally. While a new file's first save is still uploading, its create waits for it (up to 60 seconds) so the two go to the hub as one request. Other machines on the LAN see the file at once by gossip, and the web dashboard's Files view shows "N files arriving from …" until the hub has them.

Tunables, rarely needed:

VariableDefaultEffect
SWARMFILE_CREATE_DISPATCH_CONCURRENCY8 (1-64)Creates sent to the hub at once when they go one by one. Lower it if the hub keeps answering 429; raise it on an unthrottled hub.
SWARMFILE_PENDING_CREATE_WAIT_SECS8 (0-120)How long lock/ACL/share/comment writes wait for a pending create before 409 pending_create. 0 refuses at once.
SWARMFILE_CREATE_BATCHonSend due creates in batches of up to 50 per request. 0 sends each on its own.
SWARMFILE_CREATE_COALESCEonSend a new file's first save with its create (one request). 0 sends the create, then locks and commits separately.
SWARMFILE_CREATE_HOLD_UPLOADING_SECS60 (0-600)How long a create waits for its first save while that save is uploading. The upload finishing sends it at once, so this only caps a stall.
SWARMFILE_CREATE_HOLD_SECS2 (0-10)How long a create waits for a file that is open for writing but not saved yet. 0 turns off all waiting.
SWARMFILE_ARRIVING_REPORTSonTell the project how many new files are on their way (at most once every 30 s), for the dashboard's "files arriving" note. 0/false/off turns it off.

These are client-side. The hub's own ceilings are per-org and per-user (plan defaults, owner/admin configurable), and a single API key can carry its own request limit (swarmfile api-key rate-limit) - the right lever when a fleet of agents shares one key and needs more headroom than one person's window. swarmfile admin rate-limits usage reports this minute's usage against them; see Organizations, Projects & Members.

How fast can one user create files?#

The per-user ceiling is a number of requests per minute, and it depends on the plan. Defaults:

PlanPer-user defaultOwner/admin can raise it toPer-org default
Free100 / minfixed300 / min
Starter300 / min500 / min1,000 / min
Pro500 / min1,000 / min1,000 / min
Enterprise500 / min5,000 / min10,000 / min

What counts against it is requests, not bytes, and block uploads are not counted at all:

Kind of new fileWhat it costs
A new file with content (saved when it is created)one create that carries its first save: 1 request, or a tenth of one when it goes out in a batch
An empty file or folder1 request, or a tenth of one in a batch
A file created empty and written latercreate, take the write lock, then one commit that also gives the lock back: 3 requests

A batch of up to 50 creates is charged one request per ten files, so a burst that queues up costs about a tenth of a request per file. Every machine signed in as the same user shares one allowance, and the app paces itself to the budget the hub reports (and to its share of it when several machines are busy) instead of running into the limit and waiting out a whole minute.

Measured on a Pro organization at default limits (500 a minute):

  • 5,000 empty files from one machine reached the hub in 25 seconds, in about 100 requests, with none lost and no 429s. Before batching, the same burst took about 29 minutes.
  • Two machines writing small files for 90 seconds (about 930 new files) finished syncing about 4 seconds after the writers stopped, with no locks taken, no 429s and no lost updates. Before, a similar run took 14 minutes.
  • On macOS, writing many small files through the mounted drive runs at a few files a second. At that pace, the local write is the bound, not the hub.

Nothing is lost when a limit is reached; the queue simply drains more slowly.

If a fleet needs more, give it its own API key with its own ceiling (swarmfile api-key rate-limit) instead of raising everyone's, or ask an owner or admin to raise the org and user limits (swarmfile admin rate-limits; a raise can be time-boxed with --for 6h).

5. Parking and switching#

An agent's scope can move without losing queued or staged work:

  • Branch switch (hot). swarmfile branch switch <name> moves the boot mount; swarmfile mounts branch <id> <branch> [--wait] moves one independent mount. Both refuse (exit 2) while a changelist is open. Creates queued before the switch land on the branch they were made under - the scope is captured when the create is enqueued, not when it drains.
  • Project/org switch. swarmfile workspace switch --org <id> [--project <id>] moves the whole engine. A same-org switch usually takes a hot path; the fallback restarts the engine process (the OS supervisor relaunches it). Queued creates and their corrective renames/deletes stay in the project captured at enqueue.
  • Park. swarmfile changelist park sets staged work aside so a branch or project can be left; swarmfile changelist parked lists it (age, files) and swarmfile changelist resume restores it. Parked work is a set per branch, and it's kept alive on the hub, so it survives the switch. changelist discard --yes abandons it.
  • Upstream refresh. A branch is a fork: it does not see later work on main. There is no main→branch merge exposed today - swarmfile merge and mr both go into the parent - so plan agent branches to be short-lived, or recreate the branch from current main when it needs the latest. A merge-down surface is a known gap.
  • On a TTY, workspace switch prompts before abandoning staged/queued work and offers to park first; --yes/--json skip the prompt for scripts.

The tray mirrors all of this: the mount switcher lists every mount with its branch and per-row switch - and the tray menu's Open Another Drive… opens the same Open a drive dialog (its Prevent another drive from editing a file this one has open checkbox, under Advanced, is the same --exclusive opt-in), so a second mount is one menu click away - while a Parked changes panel lists, resumes (all at once or one at a time), and discards parked work. The dialog puts the drive at the next free drive letter on Windows and in a folder next to your main drive on macOS and Linux, unless you choose otherwise under Advanced. If a mount can't open, the form says why straight away - for a drive letter that's taken it offers the next free one - and the failed mount never becomes the current one.

6. What to hand an agent (a copy-paste preamble)#

You are working in a Swarmfile mount, not a git checkout.
- Workspace path: <mount path>          (writes sync automatically; there is no commit-required step)
- Tool caches:    run builds via `swarmfile run -- <cmd>` from the mount
- Status:         `swarmfile status` / `swarmfile --json status` (jq .pendingCreates)
- Exit codes:     0 success, 1 failure, 2 a refusal you should act on
- If an operation answers 409 with code "pending_create", the file is still
  syncing: wait a moment and retry. Do not delete and recreate it.
- Don't run two engines against the same cache dir or mount point.

Measured behavior#

Numbers from a test run against a local hub and real mounts (macOS and Windows):

ObservationMeasurement
Creates with the hub unreachablereturn locally in about 1 s; queue depth rises instead of failing
Peer visibility by gossip (LAN)new names listed by the peer in under 1 s
Pending-create guarda lock/ACL/comment on a still-syncing file answers 409 pending_create after about 9 s
1,000-file burst, rate-limited hubdrains with 0 files lost
5,000-file burst, same hublocal writes return in 7.3 s, with 0 files lost (drain time with batched creates: see How fast can one user create files?)
Dispatch concurrency 1 / 8 / 32a 200-create drain took 17.5 s / 31.2 s / 25.7 s, with zero 429s and zero files lost at each setting
A 402 on every createretried, not terminal: with the hub refusing everything, creates stay readable locally and the queue drains once the fault clears, with nothing lost
Hot branch switch with 10 queued createsaccepted in about 0.6 s; all 10 landed on the branch captured when they were created
Hot project switch with 9 queued createstakes the restart fallback; the queue still lands all 9 in the project captured at enqueue, including a rename made before the switch
Staged save → park → parked → resumeround-trips with the staged entry intact
Shared scratch cachePNPM_HOME/CARGO_HOME point into one per-project dir; scratch gc removes stale dirs and reports the space

Upgrading a fleet#

The queue is always on - there is no feature flag - so deploy in this order:

  1. Hub first. A hub that predates the queue answers a queued create with its own id, which the engine reads as a name collision and conflict-renames the file.
  2. Engines second, and upgrade every machine on a LAN together: the gossip signature covers the pending flag, so an older peer's gossip fails verification (nothing is lost; peers fall back to the hub until both sides are upgraded).

If a queue looks stuck or you need to triage rows, contact support with the output of swarmfile status and swarmfile --json status - queue fields, stuck-row triage, and what not to delete are support-side detail.