Diagnostics & Error Reference
When something goes wrong, Swarmfile names the problem in one of a few places. This page maps the surfaces - doctor checks, status counters, health endpoints, and error codes - to plain meanings and next steps.
Start here#
| Surface | What it is | How to reach it |
|---|---|---|
| Desktop App → Settings → Diagnostics | Live checks, report file, Repair drive, Reinstall…, update check | The tray's settings |
| Status pill (header) | Mount location, sync state, the hub's reason for any pause | Click it |
swarmfile status | The same snapshot as JSON or prose: counters, peers, bandwidth | CLI |
swarmfile-doctor | Standalone probe suite that runs with no engine and no sign-in | CLI |
swarmfile doctor | Runs the same probes through a live engine | CLI |
| Hub error codes | Machine-readable reasons on API refusals (bring-your-own scripts) | JSON bodies |
swarmfile-doctor writes a JSON report on every run and returns a verdict:
swarmfile-doctor --json --report doctor-report.json # full detail
swarmfile-doctor --repair --yes # reinstall the package (keeps your queue)
Exit codes: 0 everything passed, 1 at least one failure, 2 warnings only. The report defaults to ~/.cache/swarmfile/last-doctor.json (macOS/Linux) or %USERPROFILE%\Swarmfile\last-doctor.json (Windows).
Engine logs live in the cache directory (~/.cache/swarmfile on macOS and Linux, %LOCALAPPDATA%\swarmfile on Windows) as engine.log; the Desktop App writes tray.log beside it, and macOS mount-helper output lands in ~/Library/Logs/fuse-t/fuse-t.log. Diagnostics has an Open log directory action. Increase detail with RUST_LOG=swarmfile_engine=debug.
Status counters#
swarmfile status (and the Desktop App's panels) report these; the names are the JSON fields from swarmfile status --json:
| Counter | Meaning | What to do |
|---|---|---|
pendingUploads | Blocks queued on this machine awaiting cloud upload, plus closed files whose save is still running | Wait; check for connection trouble |
peerPendingUploads | In-flight uploads reported by other machines on this branch | Usually informational |
pendingCreates | Files/folders created locally and queued for the hub (an object with the count, stuck rows at six or more failed attempts, and the oldest age) | They land automatically; see Coding Agents |
commitsPending | Bytes uploaded but the commit hasn't landed (retrying, rate-limited, quota-held, or locked) | Check quotaHeldCommits and connectivity |
quotaHeldCommits | Commits refused on billing grounds, retried on a 15-minute cadence; spendCapHeldCommits is the subset held by the org's spend cap, which the tray names ("Waiting on spend cap") | Raise the limit or free space; retries resume |
commitsRefused | Commits the hub keeps refusing for other reasons (a subset of commitsPending) | Investigate the refusal; usually policy or conflict |
abandonedGroups | Upload groups that exhausted retries or hit a plan cap | swarmfile sync stuck lists them; re-save after fixing the cause |
failedWrites | Guest-session writes permanently given up (guest replay only) | Re-open the guest session and retry |
activeLocks | Write locks this engine currently holds | Informational |
checkoutLocks | Long-lived per-file checkout locks (explicit, git-LFS) you hold | Release with unlock / offline return when done |
scopeReservations | Offline scope reservations this engine holds - one row can cover a whole subtree; Pack & Go uses these rather than per-file locks | swarmfile offline status |
checkoutRenewFailures | Renewal passes that re-issued nothing - reservations at risk of lapsing | Reconnect and re-prepare before the lease ends |
checkoutExpiringSoon | Reservations expiring within 24 hours | Extend or return; see Working Offline |
peers | LAN/office/seed peers currently reachable, with RTT | More peers = faster reads for shared files |
bandwidth.bytesIn / bandwidth.bytesOut | Transfer volume for this session | Useful when a link is saturated |
blockStore | Pinned blocks and local cache usage | swarmfile cache status / hydrate status |
auth | Session state ("signed out" when there's no usable session) | Sign in again |
recentErrors | A ring of the most recent engine errors | Often the fastest pointer to the root cause |
lastSequenceId | Latest committed sequence known here | Diagnostic for history lag |
mounted / mountState / mountError | Whether the mount is live and, if not, why (a stable code plus a user-facing sentence) | See the mount errors below |
Health endpoints#
There is no default health port - the server starts only when SWARMFILE_HEALTH_PORT is set (typical choices: 8080 in container deployments, 9100 for a seed node). It binds 0.0.0.0:<port> and serves exactly:
GET /livez- a cheap{"status":"ok"}for liveness probes (Kubernetes and friends). Allowed from anywhere.GET /healthz- the detailed snapshot (version, uptime, peer counts and RTT, pin/cache stats, encryption state). Loopback-only: a non-loopback caller gets403 {"error":"loopback only"}, because the detail exposes fleet topology.- Anything else -
404.
Doctor checks#
swarmfile-doctor reports one row per check. The whole report is capped at 10 seconds; if the cap fires, a synthetic report check fails, and a probe that couldn't complete counts as a failure rather than a silent pass. Each of the three network checks - hub https /health, oidc token, r2 path - retries one fast transient failure (a connection error or a 5xx) before it reports Fail, so a brief hub rollout blip doesn't read as an outage; a definitive refusal like an expired token's 401 is reported on the first reply. The ones that name a common failure:
| Check | Failing usually means |
|---|---|
engine-running | No engine for this user; sign in / start the app, or expect the standalone checks below |
install-layout / updater-helper-task | A partial or relocated install (on macOS, a pre-R41 layout the updater can no longer reach); reinstall (helper task is Windows) |
dns <host> | DNS resolution for the hub/identity/storage host failed |
hub https /health | The hub host is unreachable or blocked; proxy/firewall issue |
hub deep health | The hub is up but a dependency (storage, workers) is degraded |
oidc token | The sign-in session expired, was revoked, or the account isn't a member of the configured org; sign in again |
hub deep health | The hub is up but one of its dependencies (D1, R2, the IdP issuer, its own clock) is degraded |
r2 path | Cloud storage credentials, presigning, or the block path failed for this org |
iroh quic dial / mdns listen | Peer-to-peer path blocked (firewall, VLAN, VPN); falls back to the hub |
host firewall / macos local network | OS firewall or macOS Local Network consent blocks LAN discovery |
uploads in flight / in-flight streaming | Uploads or reads waiting on in-flight bytes exist - expected during activity, informative otherwise |
fuse-t mount helper / mount | The macOS mount helper died, or the mount point is wedged or absent; Repair drive |
duplicate mount stack | Two engines are competing for the same mount point |
exclusive writes | Never fails - informational: names the drives in the opt-in exclusive-write mode |
history retention | Never fails - informational: the active project's retention mode (every saved version kept and cold-archived, or named commits only). "Not checked" without a project scope or on an older hub |
upload quota | Blocks are parked because the org is at a billing limit - see the count/bytes/age in the row |
plan | Subscription inactive, a feature isn't in the plan (e.g. seed mode), or a metered dimension is ≥80% |
pack and go policy | A SWARMFILE_PACK_AND_GO_POLICY value that isn't allowed/disabled - failing closed |
previous run | The last engine run ended abnormally (panic or unexpected exit) - the row points at engine.log |
url scheme | swarmfile:// links (Open in app, sign-in, Explorer's Swarmfile menu) won't reach this install: no app is registered for the scheme, another app or install owns it, or (macOS) several copies of Swarmfile compete for it and a link may start an old one. Fix link handling re-registers this install; on Windows, a registration pointing at another install needs Repair from Settings › Apps |
install-layout | Warning only: the engine isn't where the installer puts it - moved out of the app bundle, or a pre-R41 macOS install whose LaunchAgent still runs a stale /usr/local/bin copy. On macOS it also warns when /usr/local/bin is missing the CLI-tool links (a .dmg-only install the tray couldn't repair), since swarmfile is then not on PATH. |
Engine and mount errors (what the app shows)#
| Code / message | Meaning | Fix |
|---|---|---|
not_signed_in | The engine has no usable session | Sign in from the app |
not_a_member | The session doesn't include the configured org/project | Switch organization or ask for an invite |
project_deleted | The mounted project was permanently deleted | Mount another project; contact support about recovery within retention |
config_parse_error | A config.json couldn't be read, so the engine came up without a project | The message names the file and the parse position - fix it and restart; see Engine Config File |
mount_failed / no_free_drive_letter | Windows could not attach the mount | Free a drive letter or choose one in the app; swarmfile mounts open auto |
drive_letter_in_use | The chosen letter (or Auto's candidates) are all taken | The message names what holds it; pick another |
mount_point_not_empty | Windows refuses to mount over a non-empty folder | Choose an empty folder |
pack_and_go_disabled | An administrator disabled bulk offline reservation | Use a per-file lock or ask your admin |
confirmation_required | A destructive CLI command was run without --yes in a non-interactive context | Re-run with -y after checking what it will delete |
usage_error | A CLI invocation didn't parse (unknown command, flag, or argument); emitted as JSON when --json is present, with clap's kind | Fix the invocation - see Scripting against the CLI |
changelists_open | A branch/project switch was refused because staged work is open | Submit, cancel, or --park first |
is_default_mount | You tried to close the boot mount on its own | Close another mount, or quit the engine instead |
shared_branch_context | The mount shares the engine's branch context | Re-open it with mounts open --branch <name> to make it independently switchable |
already_running / project_switch_running | Another switch or long job of that kind is in flight | Wait for it to finish, then retry |
pending_create | A lock/comment/share on a file that is still being created | Retry after it syncs (swarmfile status → pendingCreates) |
Hub error codes (for scripts and integrations)#
API refusals carry a machine-readable code. The ones most worth handling, in the order you'll meet them:
| Code | Meaning | Typical handling |
|---|---|---|
bad_request, invalid_request, invalid_state, invalid_name, bad_cid, invalid_cid, bad_hash | Malformed or out-of-order request | Fix the request; don't retry blindly |
credential_invalid, email_verification_required, email_not_verified, session_expired, session_revoked, unauthenticated, terms_acceptance_required | Auth/session problem | Sign in again; verify email for share actions; accept updated terms |
api_key_* (api_key_grant, api_key_org_mismatch, api_key_project_mismatch, api_key_scope_forbidden) | API-key scope doesn't cover this operation | Use a key scoped to the project/org, or mint a new one |
forbidden, not_member, admin_role_required, idp_binding_rejected | Not permitted for this account/role | Escalate to an admin |
org_mfa_required, org_email_verification_required | The org's authentication policy refuses this credential | Enroll an authenticator app / verify your email under Account settings, then sign in again - the refusal body's enrollUrl points there. Don't retry unchanged; a personal access token must be re-minted from a compliant session |
owner_mfa_required, mfa_status_unavailable | Turning on "require MFA" needs the acting owner's own confirmed factor (or the factor-status probe was unreachable) | Enroll an authenticator app first, then retry the policy change |
identity_conflict | An external-IdP sign-in conflicts with an existing Swarmfile account: its subject names a first-party user, or a principal provisioned by a different issuer | Contact the organization's administrator; sign in through the org's own SSO rather than a built-in Swarmfile account |
entry_not_found, entry_deleted, project_not_found, project_deleted, path_not_found, not_a_file, source_gone, target_gone | The target is gone | Re-list; treat as terminal |
exists, name_conflict, id_conflict, name_taken | Collision or illegal name | Rename / resolve the conflict |
entry_uploading | The entry has no committed version yet | Wait for upload, then retry |
chunk_missing, cid_mismatch, checksum_mismatch, size_mismatch, manifest_unreadable, manifest_too_large, tree_object_corrupt, tree_object_missing, commit_object_corrupt, commit_object_missing | Content integrity/structure problem | Re-upload the file; report if persistent |
conflict_detected, conflict_unresolved, merge_conflict, stale_merge | Concurrent edits, or the target moved since the merge/conflict was prepared | Resolve the conflict (swarmfile conflicts resolve), then retry |
(no code; the envelope's error text says target_changed / merge_candidate_changed) | The file changed again after the conflict was recorded, or the automatic merge result changed after you previewed it | Re-read the file and resolve against the newer version; run the merge attempt again before accepting. Match the text, not a code - the hub sends these two without one |
locked | Someone else holds the file, or a machine's Pack & Go scope reservation covers it | Wait, or request an unlock (swarmfile unlock-request); a scope reservation is released by its owner with swarmfile offline return |
commit_mode_locked, policy_enforced | An admin lock requires a different commit mode, or refuses the setting | Check the enforced mode with swarmfile config get-commit-mode, then use it |
commit_race, head_changed, merge_in_progress, fork_in_progress | History moved under you, or a merge/fork is running | Re-read, wait, then retry |
branch_not_found, branch_mismatch, branch_locked, branch_has_open_mr, branch_has_active_children, merge_main | Branch rules prevented the action | Use the allowed path (merge the MR, delete children first, never merge into main directly) |
not_approved, self_review, is_draft, not_draft, protected_branch | Merge-request gates | Get approvals, un-draft, don't self-review |
project_archived, storage_exceeded, quota_exceeded, spend_cap_reached, plan_restricted, keys_cap, subscription_inactive, too_many_entries, too_many_descendants, registry_too_large, batch_too_large, body_too_large, merge_too_large, too_many_cids | Plan or size limit | Free space, upgrade, or split the request/project; merge_too_large means the merge exceeded the single-project merge capacity - land the branch in smaller pieces - otherwise retry after the allowance changes. spend_cap_reached is the dollar ceiling: the envelope carries capCents/usedCents/remainingCents (and keyLimitCents/keyId when a per-key budget was the binding wall) - raise the cap or free space and held writes resume on their own |
project_storage_full | This project hit its metadata ceiling (files and versions combined) | Delete or purge files, delete stale branches, shorten history retention, or switch the project to Named commits only (the same guidance the owner notification carries) - see Project storage limit |
checkout_required | The org hasn't completed subscription checkout | Finish checkout in Billing before creating projects |
rate_limited | Request budget exceeded | Back off and retry; see Rate limits |
grace_period_inaccessible | The org is in the post-unpaid 28-day window | Resubscribe - see Data Portability |
quarantined | Ransomware detection quarantined the account; writes stop until an owner releases it (on Windows the drive stays mounted meanwhile, with saves held locally) | Review and release from the quarantine panel - see Ransomware quarantine |
project_key_unavailable, project_key_rotation_unsupported | Encryption key unavailable, or the project can't rotate (plaintext storage, or it has reached the maximum key generations) | Check the project's encryption tier; contact support if the cap is reached |
e2e_upload_unsupported, e2e_unsupported, e2e_not_supported, encryption_tier_not_public | The operation isn't allowed for this tier | Use the supported tier (e.g. LFS is rejected on E2E) |
jurisdiction_incompatible | The destination isn't in the org's pinned jurisdiction | Use storage/a destination in the pinned region |
share_too_large, block_not_in_share, block_session_expired, raw_byte_budget_exceeded | Share-link limits or a public-streaming budget | Reduce scope, create a new share, or wait out the budget |
too_many_webhooks, blocked_target, invalid_url | Webhook limits or an unsafe target URL | Delete unused webhooks; use a public HTTPS endpoint |
preview_budget_exhausted, preview_too_large, preview_type_unsupported | Preview generation budget, size, or type limit | Wait for the daily reset; serve the file directly or download it |
too_many_active_jobs, too_many_concurrent_reads | Temporary concurrency limit | Back off and retry |
invite_unavailable, org_not_empty, installation_exists, already_provisioning, storage_provisioning | Admin lifecycle state | Complete or wait for the pending operation |
announce_unverified | A seed node's announce failed node authentication | Re-enroll the node; see Self-Hosted Seed Nodes |
org_migration_frozen, project_unreachable, async_unavailable | A migration/maintenance state | Retry after it completes |
not_found, forbidden | Generic - the meaning depends on the operation | Use your caller's contextual fallback; don't guess |
The full set is generated from the hub's own routes and shipped in contracts/hub-schemas.json and the OpenAPI contract; an unrecognized code is preserved verbatim rather than dropped, so a script can always log it.
Exit codes#
| Binary | Codes |
|---|---|
swarmfile | 0 success; 1 failure (unreachable engine, transport error, no product equivalent); 2 actionable refusal or a usage error - under --json the two are distinguishable by code (usage_error vs the refusal's code) - see Scripting against the CLI |
swarmfile-doctor | 0 all passed; 1 ≥1 failure; 2 warnings only. A usage error under --json prints the shared {"code":"usage_error",…} body, as every standalone tool now does |
swarmfile-runner | The engine's code (the alias re-execs it); 127 if the sibling engine is missing |
Where to go next#
- Troubleshooting - symptom-first workflows.
- Network Requirements - the hosts and ports the checks probe.
- Environment Variable Index -
SWARMFILE_HEALTH_PORT,RUST_LOG, and the diagnostic switches.