# swarmfile-migrate

`swarmfile-migrate` bulk-imports data you already have into a Swarmfile project: a local directory, a flat list of file paths, or an S3-compatible bucket. It's a standalone binary - no running engine required - meant for the first-time import of an existing archive, not for day-to-day file operations. Every desktop installer carries it, and the Linux `.deb`, macOS `.pkg` and Windows installers put it on `PATH` (`/usr/bin` on Linux, `/usr/local/bin` on macOS, or the Windows install folder); a macOS `.dmg`-only install needs one `.pkg` (or **Diagnostics → Reinstall**) run for that entry - so if the desktop app is installed there's nothing extra to download.

This page is the flag reference. For a practical walkthrough with a worked example, see [Migrating Existing Data](https://swarmfile.com/docs/guides/migrating-existing-data).

> Run with `--dry-run` first. It walks the source and reports what would be migrated without uploading anything, which is the safe way to check your exclude patterns and source/destination settings before committing to a real run.

## Source selection

Pick exactly one source: a local directory, a flat file list, or an S3-compatible bucket.

| Flag | Description |
|---|---|
| `--source` | Local directory source |
| `--source-list` | Flat file-list source |
| `--source-root` | Root to resolve relative paths in `--source-list` against |
| `--source-s3-bucket` | S3-compatible source bucket |
| `--source-s3-prefix` | Prefix within the S3 bucket to migrate |
| `--source-s3-endpoint` | S3-compatible endpoint URL |
| `--source-s3-region` | Default `us-east-1` |

> S3 support ships in every released installer - `--source-s3-*` works out of the box, with no extra download. The AWS SDK it needs is large and lives in the `swarmfile-migrate` tool rather than the Swarmfile engine; a from-source build can opt out with `--no-default-features`, and a binary built that way fails immediately (naming the missing feature) instead of partway through a migration.

## Destination

| Flag | Description |
|---|---|
| `--root-path` | Destination path inside the project. Default `/` |
| `--hub-url` | Your hub URL, e.g. `https://hub.swarmfile.com` (or your self-hosted hub's URL). Defaults to `$SWARMFILE_HUB_URL` if set, otherwise the hosted hub (`https://hub.swarmfile.com`) - so a machine whose environment already names its hub does not need this flag |
| `--org-id` | Target organization |
| `--project-id` | Target project. Defaults to `$SWARMFILE_PROJECT_ID` if set; with neither, entries are created **project-less** - invisible to any mount that scopes its feed to a project (which every engine does by default), so set one for a normal import |
| `--oidc-issuer` | OIDC issuer for authenticated hub access |
| `--oidc-client-id` | OIDC client ID |
| `--oidc-refresh-token` | OIDC refresh token |

### Authentication

The simplest and primary way to authenticate is a **project-scoped API key** in the environment: set `SWARMFILE_API_KEY=sf_key_…` and `swarmfile-migrate` picks it up automatically - the same headless credential every other Swarmfile binary honors - with no auth flags at all. This is the recommended path for a CI runner or a render-farm import box.

The `--oidc-issuer` / `--oidc-client-id` / `--oidc-refresh-token` flags are the alternative: pass **all three together** to authenticate with an OIDC refresh token instead. If you pass no API key and not all three OIDC flags, hub calls go out unauthenticated and will fail against any hub that enforces auth.

## Exclusions

| Flag | Description |
|---|---|
| `--exclude` | Glob pattern to exclude. Repeatable. Matching a directory prunes its whole subtree. |
| `--exclude-from` | File containing exclude patterns, one per line |

## Reliability

| Flag | Description |
|---|---|
| `--concurrency` | File-level concurrency. Default `8` |
| `--block-concurrency` | Max concurrent block uploads, shared across all files. Default `4` |
| `--retry` | Max retry attempts per file - both in-run (a transient failure retries immediately, with backoff, before the run moves on to the next file) and across separate re-invocations of this tool against the same `--state-db`. Default `3`; `0` disables retries either way |
| `--verify` | Verify each migrated file against the source with a CID/hash comparison. On by default (`true`); pass `--verify=false` to skip - verifying means reading the source a second time, so skipping it roughly halves I/O for a very large migration where you trust the transport |
| `--state-db` | Local state DB path. Tracks what's already uploaded and enables resuming a re-run. Default `<source>/.swarmfile-migration/migrate.db`; for a `--source-list` source, it sits alongside the list file instead |
| `--report-json` | Write a JSON report of the run |
| `--start-at` | Resume marker to start from |

## Safety

| Flag | Description |
|---|---|
| `--dry-run` | Report what would be migrated without uploading anything |
| `--force` | Re-migrate files already registered/uploaded in a prior run, instead of skipping them |

## Encryption & erasure coding

There's no `--encrypt` flag - migrated data is encrypted the same way any other write to the project is, driven by environment variables rather than a migration-specific setting: `SWARMFILE_ENCRYPTION_KEY`/`SWARMFILE_ENCRYPTION_PROJECT_ID`, or simply `--project-id` combined with OIDC credentials, resolves the key. Encryption is on by default; if no key source is configured for the invocation, the tool warns and migrates the data as plaintext rather than refusing to run - unless encryption was explicitly requested, in which case it fails instead. Erasure coding is controlled by `SWARMFILE_EC_ENABLED` (`on`/`off`/`auto`): default `auto` (also what an unset or unrecognized value falls back to) **never erasure-codes during migration**, because `auto` is evaluated from live peer/network topology and a bulk import has none - set `SWARMFILE_EC_ENABLED=on` (or `true`) to force erasure coding on the migrated blocks.

## Examples

**Step 1 - preview with `--dry-run`.** Always start here: it walks the source and reports what would be migrated without uploading anything, so you can confirm your exclude patterns and source/destination settings before committing to a real run:

```bash
swarmfile-migrate --source /Volumes/nas/ProjectArchive \
  --org-id acme-films --project-id feature-01 \
  --exclude "*.cache" --exclude "**/tmp/**" \
  --dry-run
```

**Step 2 - the real run.** Once the dry-run looks right, drop `--dry-run` (post-upload verification is on by default, so no flag needed). Here `SWARMFILE_API_KEY` supplies the credential and `$SWARMFILE_HUB_URL` the hub, so neither needs a flag:

```bash
export SWARMFILE_API_KEY=sf_key_…
swarmfile-migrate --source /Volumes/nas/ProjectArchive \
  --org-id acme-films --project-id feature-01 \
  --exclude "*.cache" --exclude "**/tmp/**" \
  --state-db ./migrate-state.db
```

S3-compatible bucket (preview it with `--dry-run` the same way first):

```bash
swarmfile-migrate --source-s3-bucket project-archive \
  --source-s3-prefix renders/2025 \
  --source-s3-endpoint https://s3.us-west-000.backblazeb2.com \
  --hub-url https://hub.swarmfile.com \
  --org-id acme-films --project-id feature-01 \
  --dry-run
```
