Runner (Headless CI)
A runner is a headless Swarmfile engine that watches one branch and
executes a job when a commit or tag matches a rule - the same shape as
GitHub Actions or a self-hosted CI runner, but built on the engine's
existing content-addressed blocks instead of a full git clone. Each rule
match materializes the branch into a plain checkout directory (not a
live mount) before running the job. By default that is every file in the
tree - the checkout directory persists between runs, so a re-checkout only
re-fetches blocks that actually changed since last time (an unchanged file
is a fast local copy, not a network fetch), but the first run against a
large project, or one where most files change often, still pays for the full
tree. A top-level sparse: list narrows the
checkout to the paths a job needs, so a huge repo only materializes the part
it uses.
What it is not#
This self-hosted runner has no job queue, no matrix builds, and no artifact storage - it is the trigger→execute→report loop only. For those, use Swarmfile's separate Hosted CI, which runs jobs on Swarmfile-managed containers. Run history is inspectable via the CLI, the Desktop App, and the web dashboard (below).
A runner is also deliberately not a mounted engine: it serves no control
socket and never presents a drive, so the swarmfile CLI cannot talk to it.
It is configured entirely by environment variables plus the
.swarmfile/runner.yml in the branch it watches, and a job starts only when
a matching commit or tag lands - this runner has no external "run this now"
API. On a machine that does run a mounted engine, swarmfile tail streams this
project's activity in real time; the org's outbound webhooks are delivered
on a periodic schedule, not a real-time build trigger; treat them as a
reaction/integration channel.
The watched branch must be protected. .swarmfile/runner.yml decides what
runs on the runner host, so a runner executes it only on a branch with
effective branch protection
- otherwise anyone who can push to that branch could run commands on the
machine. SWARMFILE_RUNNER_ALLOW_UNPROTECTED=1 overrides that and is not a
production option - use it only on a trusted or throwaway runner.
For a job that isn't trigger-driven, a machine running a mounted engine can
use swarmfile materialize <ref> --out <dir> - the same checkout, written
to a plain directory with no FUSE, pinned to whichever ref you name.
Where rules live#
Rules are not configured on the hub, in the Desktop App, or via the CLI -
there is no swarmfile runner rules command, because there is nothing on
the hub to create. A rule is a plain YAML file, committed into the branch
it governs, at:
.swarmfile/runner.yml
This is deliberate: a rule change is reviewable in the same diff as the code it gates, exactly like a GitHub Actions workflow file. Whoever can commit to the branch controls what runs - the same trust boundary as any other file on that branch, not a new one.
# Optional, top-level: check out only these paths (see "Checkout scope").
sparse:
- "src/**"
- "*.md"
rules:
- name: build-and-test
on: commit # commit | tag - default: commit
branches: # glob patterns; omit to match every branch
- main
- "release/*"
paths: # optional - matches the full relative path
- "*.dwg"
run: "scripts/ci.sh" # shell command, run with the checkout as CWD
timeout: 1800 # seconds - default 1800 (30 min); on expiry the command's whole process tree is killed
A project can define several rules; each is evaluated independently against every commit/tag on the watched branch.
paths: matches each changed file's full relative path. *.dwg
matches plan.dwg wherever it lives in the tree; assets/** anchors to
that folder. A leading ! negates (gitignore-style, last match wins), so
paths: ["src/**", "!src/docs/**"] skips doc-only changes; the same applies
to branches:. Leave paths: off to run on every commit to a matching
branch.
on: tag rules ignore paths:. A tag names a single commit; there's
no "files changed since the last tag" the way there is for a fresh commit.
Checkout scope (sparse)#
By default a matched run checks out the whole branch. For a large repo, add
a top-level sparse: list to materialize only the paths a job needs:
sparse:
- "src/**"
- "assets/models/*.bin"
- "*.md"
rules:
- name: build
run: "make -C src"
A path is checked out when it matches a glob, or when a directory in its
ancestry matches. A glob with no / matches by name at any depth (*.md
finds docs/notes.md); a glob containing / is anchored to the repo root
(assets/models/*.bin does not match vendor/assets/models/x.bin). A
leading ! excludes a path an earlier glob included (gitignore-style, last
match wins), so ["src/**", "!src/**/*.test.ts"] checks out src/ without
its tests; the top-level exclude: key is sugar for those ! negations, and
with no sparse: it means "everything except". Paths that match nothing are
neither written nor fetched: a directory the filter cannot reach is skipped
without being walked (and a frozen commit tree is not even fetched), and a
directory holding no matching file is not created. Only anchored globs
enable that pruning - a bare name like *.md can match at any depth, so the
walk cannot skip anything; anchor with a / (docs/*.md) when you can.
sparse: is top-level, not per-rule: a runner has one checkout
directory and materializes it once per event for every rule that matched, so
a per-rule scope could not be honored. Omit it (or leave it empty) for a
full checkout. Changing the list between runs converges the directory: paths
that drop out of scope are removed, the newly included ones are fetched. At
most 64 globs of 1024 characters each, and at least one non-! glob; an
invalid list fails the affected runs with the reason (visible in
swarmfile runner runs) rather than silently checking out everything.
Running a runner#
A runner is a normal swarmfile-engine process started with
SWARMFILE_RUNNER_MODE=true and a project API key (see
CLI: swarmfile for swarmfile api-key create - the key
counts against your plan's headless-key allowance; see
Billing & Plans). Like
seed mode, it never mounts a drive and never prompts for interactive
sign-in; unlike seed mode, it materializes a real checkout directory (under
its cache dir) because a job needs actual files on disk to run against -
not a FUSE/WinFsp mount, which would require a kernel driver on a CI box
that may not have one.
export SWARMFILE_API_KEY=sf_key_...
export SWARMFILE_PROJECT_ID=proj_xyz789
export SWARMFILE_BRANCH=main
export SWARMFILE_RUNNER_MODE=true
swarmfile-engine
swarmfile-runner is the same thing as its own binary: it sets
SWARMFILE_RUNNER_MODE=true (when unset) and re-execs the sibling
swarmfile-engine, so the CI role reads as its own tool without an env-var
prefix. It ships with the installers - /usr/bin/swarmfile-runner via the
Linux .deb, inside the macOS app bundle (with a
/usr/local/bin symlink), and next to swarmfile-engine.exe from the Windows
MSI. On any engine binary the env-var form above is equivalent. See
swarmfile-runner for its flags and acquisition.
It takes the same env and flags:
SWARMFILE_API_KEY=sf_key_... SWARMFILE_PROJECT_ID=proj_xyz789 swarmfile-runner
SWARMFILE_BRANCH is the same setting a normal branch-pinned mount already
uses (see Branches & Merging) - "which
branch does this runner watch" needs no setting of its own.
On startup the runner replays any commits it missed while it was offline (via the branch's commit log), then listens live. A crash or restart never silently skips a commit.
Before checking anything out, a triggered run waits up to 120 seconds for a
file that is still uploading on the branch - a mid-change tree would be the
wrong thing to build. If the upload is still live at the deadline, the run is
not started: it is recorded as a failure whose log says it was skipped for a
live upload (visible in swarmfile runner runs), rather than building from a
partial tree. There is no separate skipped status. A shutdown
during the wait ends it immediately, treated the same way.
Checking out a ref in your own CI#
If you already have a CI system, you don't need runner mode to get a checkout:
swarmfile materialize <ref> --out <dir> writes a ref's tree to a plain
directory with no mount and no FUSE, so it works on a container with no
kernel driver. Pin by hash - a tag is a name that can be repointed by
delete + recreate, and a branch moves; the 64-hex hash is the only immutable
pin (a commit's hash is shown in the web commit log, MR header, and release
page, and on swarmfile log/show):
swarmfile materialize hash:$SWARMFILE_COMMIT_HASH --out "$PWD/checkout"
--json echoes the resolved commitHash, kind, and seq - record them in
your build metadata so an artifact names the exact tree it was built from.
Add --sparse <glob> (repeatable) to check out only the paths a job needs -
the same scoping as the runner's top-level sparse:,
with the same glob rules:
swarmfile materialize hash:$SWARMFILE_COMMIT_HASH --out "$PWD/checkout" --sparse "src/**"
A few things to know before relying on it:
- The executable bit is preserved. A file marked executable (with
chmod +xon the drive,swarmfile chmod +x, or a git push of mode100755) checks out executable on macOS and Linux, so./scripts/ci.shruns as-is. Windows has no executable bit. - The marker owns the directory. A new or empty destination is
materialized fresh; one this tool already wrote is updated in place,
pruning only the paths in its marker, so your job's own build outputs and
caches survive a re-checkout. A non-empty destination with no marker, or one
owned by a different project, is refused rather than merged or deleted -
there is no
--pruneflag. - One destination, one writer. The marker's written-paths list is read and
rewritten per run with no cross-process lock. Give each job its own
--outdirectory; two concurrent materializations into the same one can interleave. - Symlinks are written, not copied. On Windows that needs Developer Mode
(or elevation); a symlink the OS refuses is reported in the run's per-file
failuresrather than silently skipped, like any other entry that can't be written. - Not every commit has a hash. A commit finalized before hashing was available, and an archived commit, have none; address those by seq or tag. An unknown hash is the typed
hash_not_found, not a silent fallback.
swarmfile clone <project-id> <path> --checkout -b <ref> is the git-familiar
spelling for the same operation, cross-project included: -b takes the full
ref vocabulary (branch, tag, hash, seq), and --sparse <glob> (repeatable)
narrows the checkout exactly as it does for materialize. The mount form of
clone keeps -b branch-only, because a mount works one branch and a tag or
hash has no branch to live on; it refuses --sparse, since a mount always
shows the whole branch.
materialize still needs a running engine (the CLI is control-socket-only), so
on a driverless box it's usually a mounted engine elsewhere doing the checkout,
or the runner above doing the whole loop. The /runner-runs check-reporting
API below needs no engine at all.
Inspecting run history#
swarmfile runner runs # every branch in the project
swarmfile runner runs --branch main # narrowed to one branch
Each row carries the matched rule's name, the triggering commit, status
(running / success / failure / timed_out), exit code, and a
truncated log. There is no separate swarmfile runner runs show <id>
command - the list output already includes each run's log inline, so
there's nothing further to drill into.
The same history is browsable without the CLI: in the Desktop App, via the project left rail's CI runs row; on the web dashboard, under the project's Runs tab (open to every project member, like Files and History - guests are refused). Both surfaces also offer Run workflow and Cancel for a project's runs - those act on Swarmfile's hosted-CI service, never on a self-hosted runner, whose runs are only ever created and finished by the runner engine. The adjacent CI tab, which holds hosted-CI settings, secrets and runner sizes, is the admin/owner-only one (like Commits, RFIs and Permissions).
Reporting checks from your own CI system#
You don't have to run Swarmfile's own runner to gate a protected branch's --require-checks-to-pass - any CI system that can make two authenticated HTTP calls can report a check result the same way a runner does. The /runner-runs API is deliberately not tied to the runner binary: it only requires a project API key, the same credential a runner engine uses (see swarmfile api-key create).
The two calls are the whole contract:
# Start of your job - creates a `running` row so a crash reports as failure,
# not silence.
RUN_ID=$(curl -sf -X POST "$SWARMFILE_HUB_URL/runner-runs" \
-H "Authorization: Bearer $SWARMFILE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"projectId": "'"$SWARMFILE_PROJECT_ID"'",
"branchId": "'"$SWARMFILE_BRANCH_ID"'",
"branchName": "'"$SWARMFILE_BRANCH"'",
"changesetSeq": '"$SWARMFILE_CHANGESET_SEQ"',
"ruleName": "external-lint"
}' | jq -r '.run.id')
# ...your actual CI job runs here...
# End of your job - reports the real outcome.
curl -sf -X PATCH "$SWARMFILE_HUB_URL/runner-runs/$RUN_ID" \
-H "Authorization: Bearer $SWARMFILE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"status": "success"}' # or "failure" / "timed_out"
What each field means for gating, exactly as a Swarmfile-runner-produced row would be read: ruleName is an arbitrary label you choose - it's what shows up as one named check. requireChecksToPass first narrows to only the runs reported for the exact commit being merged, then requires the latest run of every distinct ruleName reported at that commit to be success. A rule that reported at an earlier commit but hasn't reported again at the new head isn't carried forward as still-required - it simply isn't part of the gate for this head, the same as a rule that's never run at all; a branch with zero matching runs at its head fails closed rather than passing vacuously. changesetSeq must be the commit number of the exact commit your job ran against (swarmfile status/mr status locally, or resolved from whatever your CI already checked out) - a check reported against an older commit doesn't satisfy a merge request whose branch has since moved past it. branchId is the branch's real id, not its name; resolve it once via swarmfile branch list or GET /branches.
A GitHub Actions job reporting into this looks like:
- name: Report check start
run: |
RUN_ID=$(curl -sf -X POST "$SWARMFILE_HUB_URL/runner-runs" \
-H "Authorization: Bearer ${{ secrets.SWARMFILE_API_KEY }}" \
-H "Content-Type: application/json" \
-d '{"projectId":"...","branchId":"...","branchName":"...","changesetSeq":...,"ruleName":"gh-actions-build"}' \
| jq -r '.run.id')
echo "RUN_ID=$RUN_ID" >> "$GITHUB_ENV"
# ...your build/test steps...
- name: Report check result
if: always()
run: |
STATUS=${{ job.status == 'success' && 'success' || 'failure' }}
curl -sf -X PATCH "$SWARMFILE_HUB_URL/runner-runs/$RUN_ID" \
-H "Authorization: Bearer ${{ secrets.SWARMFILE_API_KEY }}" \
-H "Content-Type: application/json" \
-d "{\"status\": \"$STATUS\"}"
This is exactly what Swarmfile's own runner does internally - there's no separate, more-privileged path only the runner binary can use. A project API key is the whole trust boundary, same as everywhere else in the API (see Security below).
Security#
A rule's run: command runs with the permissions of that machine's user -
the same trust model as a self-hosted GitHub Actions runner. The runner's API key
itself is narrowly scoped: it can
create and update its own run records, plus everything any other
project-scoped API key can already do (read metadata, download blocks,
read/write changesets and branches). It cannot post comments, touch ACLs,
or reach admin routes - the same restriction every API key has today.