# Operating Helix Foundry ## Installation profiles `./scripts/setup.sh --compose` starts the local-model profile and downloads the tested Qwen3 4B model. The app is available on `http://localhost:3001`. On Apple Silicon, native Ollama can use Metal: run the data services without `--profile local` and set `MODEL_ENDPOINT=http://host.docker.internal:11434` in `.env`. Keep `.env`, data volumes, and backups private. The application encrypts connector/provider credentials using `ENCRYPTION_KEY`; losing that key makes those credentials unrecoverable, and backups do not include it. Do not rotate it by simply replacing the environment value. Helix Foundry has no sign-in, so do not admit other devices. The API refuses a non-loopback `HOST` unless `ALLOW_NETWORK_ACCESS=true`, which Compose sets because it publishes the app on `127.0.0.1` only. Do not expose it to a network through a reverse proxy, tunnel, or port forward. Internal HelixDB, Ollama, and executor ports are not published by Compose. ## Diagnostics and schema migration `pnpm doctor` checks the database, executor, model endpoint, and local secrets. `/api/health` is a process health check. Startup applies the indexed version-1 metadata schema idempotently; `node --import tsx scripts/migrate.ts` applies it explicitly. Startup rejects database schema versions newer than the application. Jobs claim leases and periodically renew them. A worker restart reclaims expired leases. While setup's build from a confirmed ontology is queued or running, including while a restarted worker waits for its lease to expire, scheduled and CDC syncs of the sources whose datasets it reads stay queued, so the build publishes the input versions it validated. Each waits at most 15 minutes from when it was queued, however often the worker restarts, so a slow or crashing build cannot hold a source's data back for longer. A build still running when that time is up can therefore fail: the sync runs, and when the build publishes, it stops with "Inputs changed during publication", or with "Input data changed. Revalidate this proposal." if the sync finished after the build validated its inputs but before publishing started. Nothing is published, and the workspace stays as it was. Retry the build with **Retry** in setup: it builds again from the new data, validates it and publishes it. Syncs of other sources run as usual, and syncing a source by hand starts its waiting sync like a manual one. Other builds (chat requests, continuous repair, and setup from a goal) end with a proposal for review rather than publishing, so they hold no syncs; a proposal whose inputs a sync changed has to be revalidated before it is published. AI stages save definitions and usage after each completed stage; a failed/canceled run can be resumed through Activity or the API. Input changes invalidate cached stages. Failed connector/pipeline jobs retry three times with bounded backoff. A stopped worker does not stop published dashboard queries. ## Change capture Database synchronization prefers native CDC. PostgreSQL needs logical WAL, a replication-capable login, a primary key/usable replica identity, table publication permission, and available replication slots. Foundry creates a publication and slot scoped to each selected table. MySQL needs row-format binary logging, SELECT access, and replication client/slave privileges. Read-only logins still import through the scheduled fallback. Foundry never changes server-wide database configuration or writes business records back to sources. The capture service runs independently of inference, saves its dirty counter and log position atomically in HelixDB, and acknowledges PostgreSQL only after that durable write. One queued snapshot coalesces multiple commits; its starting counter prevents concurrent changes from being marked applied accidentally. Existing complete snapshots stay available during imports. Reconnects reconcile a new snapshot, including after restoration of old metadata; schema changes still need review. Transient stream failures retry from saved positions, then fall back to scheduled refreshes. This is native log capture with full-snapshot materialization, not row-at-a-time replication: existing two-million-row import and executor limits still apply. Monitor source-side WAL/binlog disk retention. Configure PostgreSQL's `max_slot_wal_keep_size` and a suitable MySQL binlog expiration policy for the installation; the application does not alter either server setting. Stop capture before manually removing Foundry-owned slots/publications. Keep their names from the source's sync metadata when decommissioning a source or changing its endpoint. Do not drop other applications' replication resources. ## Backup, restore, and upgrades Backups do not contain `ENCRYPTION_KEY` or `INTERNAL_KEY`. Keep a separate copy of `.env`, for example in a password manager; without it, a restored installation cannot decrypt its connector and AI credentials. Backups made by an older `backup.sh` do contain them: it copied `.env` into each backup as `config.env`, with both keys in plain text, by default into the repository's `backups/` folder, and older docs suggested `~/helix-foundry-backups/`. Anyone with a copy can decrypt every stored credential. `backup.sh` and `upgrade.sh` list the ones they find in those two folders. Once you have a new backup and a separate copy of `.env`, delete those old backups, or at least their `config.env`, and any copies of them elsewhere. An old backup without its `config.env` still restores with `--env-file` or `--accept-key-mismatch`. 1. Run `./scripts/backup.sh PATH` with a new folder outside the repository (the default is `~/helix-foundry-backups/TIME`). It stops writers and HelixDB, archives the data volumes without the executor's temporary `scratch` and `results` folders, starts again the services that were running and checks that the archives read back. It writes `settings.env` with the non-secret `.env` lines, including any executor limits, and `manifest.txt` with the commit the services run, its HelixDB image and metadata schema version, the image of the HelixDB container and a fingerprint of `ENCRYPTION_KEY`. Exit status 3 means the backup is complete but the services did not start again. Any other failure leaves no backup: it removes everything it wrote, and the folder if it created it, starts again the services it stopped, and says so if they did not start. The fingerprint is a SHA-256 of the key with a fixed prefix, so it identifies the key without revealing it. Every file is readable only by you. It also saves a copy of `restore.sh` and `backup-lib.sh` in the backup, so a checkout of an older release, whose own `scripts/restore.sh` cannot read this format, can still restore it: from the repository root, run `bash BACKUP/restore.sh BACKUP --replace`. 2. To keep the keys with the backup, run `./scripts/backup.sh --include-key PATH`. It stores the whole `.env` as `env.enc`, encrypted with AES-256-CBC and a key derived from your passphrase with PBKDF2-SHA256 (600,000 iterations). Type the passphrase when asked, or set `BACKUP_PASSPHRASE`; never put it on the command line. Anyone with the backup and the passphrase can decrypt every stored credential. 3. Copy the backup to private off-host storage. 4. To restore, run `./scripts/restore.sh PATH --replace`. It checks the backup before it stops anything, including that both archives read back. It keeps `.env` if its key matches the backup's fingerprint. Otherwise it uses `--env-file FILE` (your copy of `.env`), then `env.enc`, then a `config.env`, which backups made before this format hold, with the key in plain text; restore warns about it. `--env-file` reads the file once, so it can come straight from your password manager's command-line tool through process substitution, `--env-file <(COMMAND)`, without a plain-text copy on disk. It refuses a key that does not match. `--accept-key-mismatch` restores the data anyway, with the key in `--env-file` if you pass one, which then replaces `.env`, and otherwise with the current `.env`, even when the backup has `env.enc`. The stored credentials then cannot be decrypted, so you enter them again. An older backup whose `config.env` you removed does not record its key, so it needs `--env-file` or `--accept-key-mismatch`, and restore warns that it cannot check the key. It also refuses a backup written in a newer HelixDB storage format or with a newer metadata schema than the checkout, and names the commit to check out. If the backup or the checkout names a HelixDB image without a version, it cannot compare them and refuses unless you pass `--accept-helix-mismatch`, which also overrides the storage format check; HelixDB may then refuse the data. Then it replaces both data volumes, and only then replaces `.env`, saving the old one as `.env.pre-restore-TIME`; it holds the old keys, so delete it once the restore is confirmed. It makes `.env` readable only by you. Last, it rebuilds and starts the services. If a step fails, it says what state it left, and running it again finishes the restore. 5. The project name prefixes the volume names. The scripts find it as Compose does: `COMPOSE_PROJECT_NAME` from the shell, then from `.env`, then `name:` in `compose.yaml` (`helix-foundry`). If you set it in the shell, set the same name for every backup, restore and upgrade. `restore.sh` refuses a `.env` that names another project than the current one. 6. Upgrade with `./scripts/upgrade.sh`. The backup is its first step and cannot be skipped: it fetches the upstream branch, takes a fresh backup (`--backup-dir DIR` chooses the folder, `--include-key` passes through), checks it, and stops without changing anything if the backup fails. It records the previous commit and HelixDB image in the backup's `manifest.txt` and `upgrade.txt`, fast-forwards, and rebuilds and restarts with `docker compose up -d --build`. It records the commit the services run in the git directory (`helix-foundry-deployed`). Until the rebuild succeeds, it also keeps the unfinished upgrade there (`helix-foundry-upgrade`), so if the rebuild fails, the next run finishes the upgrade: it takes another backup, asks again about a HelixDB change, and names the backup from before the first attempt for going back, because HelixDB may already have opened the data with the new image. Without a record of the running commit, for example after a `git pull` by hand, it still backs up and rebuilds, but cannot name the commit to go back to; `git reflog` shows it. Afterwards, run diagnostics, sample ingestion, a saved query, and proposal validation before using the installation again. A checkout from before `upgrade.sh` gets it with one `git pull --ff-only`, without rebuilding; then run `./scripts/upgrade.sh`. Before that pull, if the checkout was not pulled since the stack last started, record the running commit with `git rev-parse HEAD > "$(git rev-parse --git-path helix-foundry-deployed)"`, so backups and the rollback advice name it. 7. HelixDB upgrades are one way. The first time HelixDB v0.0.10 or later opens a database, it upgrades the index storage, and v0.0.9 and earlier then refuse it with `unsupported_index_storage_version`. When a release changes the pinned HelixDB image in `compose.yaml`, `upgrade.sh` prints a warning and waits for `yes` (or `--yes`) before the backup and the new image; it refuses a release that moves HelixDB back to an older version. To revert an upgrade, run the two commands `upgrade.sh` printed, from the repository root: `git reset --keep PREVIOUS_COMMIT`, which keeps the branch so the next `upgrade.sh` fast-forwards again, then `bash BACKUP/restore.sh BACKUP --replace` with the backup taken before the upgrade. That is the copy of `restore.sh` saved in the backup: after the reset, `scripts/restore.sh` is the old release's, and one from before `upgrade.sh` needs a `config.env`, which backups no longer hold. Model weights are reproducible and excluded from application backups. Parquet versions and staged artifacts are immutable; unsuccessful previews can consume disk. Publication has the executor write each object type's and relationship's rows to a file in `DATA_DIR/results`, and deletes the file once it has read it, so one exists at a time per publication. The files are uncompressed JSON lines, often several times the size of the dataset's Parquet snapshot, and limited by `EXECUTOR_RESULT_LIMIT` (see [deployment limits](#deployment-limits)). A stopped publication leaves its file behind; the worker removes files there more than ten minutes old when it starts and every ten minutes. Backups leave out `results` and the executor's `scratch`, which hold nothing a restored installation needs. There is no automatic artifact garbage collector in this MVP. Remove unreferenced files only during offline maintenance, after checking proposal/run previews and version records. ## Deployment limits Two DuckDB jobs execute at a time, and up to 50 more wait for a turn. Each has a 512 MB per-process memory setting (`DUCKDB_MEMORY`) and a time limit, counted from when the executor receives it, so it includes time spent waiting for a turn: - `EXECUTOR_QUERY_TIMEOUT` (120 seconds by default) for queries: dashboards, saved queries, questions, and a build's checks of its widgets and relationships. - `EXECUTOR_JOB_TIMEOUT` (900 seconds, 15 minutes, by default) for imports, pipelines and publication queries, which read or write whole tables. Both take whole or decimal seconds from 1 to 86400 (one day), without a unit, such as `1800` or `90.5`. The API, the worker and the executor check them when they start, and refuse to start with any other value, such as `2m` or `100000`, with "Invalid EXECUTOR_JOB_TIMEOUT: 2m (whole or decimal seconds, from 1 to 86400)". In Compose, Docker then keeps restarting them and the health check fails until the value is fixed. A job that runs out of time is stopped with "Stopped at the executor's time limit", which names the setting. One whose time runs out while it waits never runs, and fails with "The executor was busy with other jobs". The API reads the same settings and waits 15 seconds longer, so it always receives the executor's answer; Compose and `pnpm dev` give both the values in `.env`. A job that needs more memory spills to `DATA_DIR/scratch`. Foundry sets no limit on that by default. DuckDB then stops a job whose spill files would pass 90% of the free disk space, but that bound applies to each job on its own: two large jobs at once, or one alongside other writes to the data volume, can still fill the disk. Importing a large file loads the whole table and profiles every column, and can spill several times the file's Parquet size. `DUCKDB_SPILL_LIMIT` caps each job's spill files, as a size with a unit such as `20GB` or `20GiB`. A job that needs more fails with DuckDB's "Out of Memory Error", which names the `max_temp_directory_size` setting; raise or remove the limit to import it. Without a limit, the same error means the job's spill files reached 90% of the free disk space. With a limit set, keep twice that much free for spill files, since two jobs can spill at once. A publication query also writes its result rows, as JSON lines, to the job's scratch directory, then moves them to `DATA_DIR/results` for the worker. They are not spill files, so `DUCKDB_SPILL_LIMIT` does not count them; `EXECUTOR_RESULT_LIMIT` (2 GB by default) caps each one. A larger result fails the publication with "Result exceeds EXECUTOR_RESULT_LIMIT". A relationship whose key matches many records on both sides is the usual cause, but a type with many records can reach it too; raise the limit in `.env` for one. Nothing limits the rest: the Parquet snapshots that imports and pipelines write to `DATA_DIR/datasets`, staged artifacts, and HelixDB's own volume all grow with your data. Each job's scratch directory is removed when the job ends, including when it is killed for its time or output limit or because the executor container is stopping. Once it is listening, the executor also removes the job directories that were in `scratch` when it started, which clears spill files left by an out-of-memory kill; an executor that cannot listen, because another one has the port, removes nothing. Nothing else in `scratch` is touched. In Compose, set `DUCKDB_SPILL_LIMIT`, `EXECUTOR_RESULT_LIMIT`, `EXECUTOR_QUERY_TIMEOUT` and `EXECUTOR_JOB_TIMEOUT` in `.env` and run the start command again (`docker compose --profile local up -d`, or without `--profile local` when the stack runs without it) to apply them; an empty value means the default. `pnpm dev` reads them from `.env` too. The executor container has a 2 GB memory cap and an internal-only network; see [Memory](#memory) for the other services. One worker performs one inference job at a time. Use one worker replica in this single-server release. The supported runtime is Node.js 24. There are no accounts: each installation serves one person on their own computer, who owns every workspace. API tokens are workspace scoped, optionally read-only, and revocable. See `VALIDATION.md` for benchmark scope and unverified deployment combinations. ## Memory HelixDB, the bundled model, the app, the worker and the executor share Docker's memory. Compose gives each service a budget and a place in the order the kernel stops them, so that when memory runs out the model or a single request fails with an error and HelixDB keeps running. | Service | Budget in `compose.yaml` | Override in `.env` | Measured (1.25 M-item workspace) | Stopped first when Docker runs out | | ---------- | ------------------------ | -------------------------- | ------------------------------------------------------------------------ | ---------------------------------- | | `ollama` | 6 GB limit | `OLLAMA_MEMORY_LIMIT` | 3.3 GB with `qwen3:4b` loaded | 1st (`oom_score_adj: 500`) | | `app` | No limit | — | 0.15–0.26 GB; 3.2 GB loading the graph of a 0.5 M-record workspace (#47) | 2nd (`300`, with the executor) | | `executor` | 2 GB limit | — | about 70 MB idle | 2nd (`300`) | | `worker` | No limit | — | about 0.55 GB | 3rd (`200`) | | `helix` | 2 GB reserved, no limit | `HELIX_MEMORY_RESERVATION` | 1.8–2.0 GB | Last (`0`) | - The app and worker have no memory limit on purpose. Node sizes its heap from the container limit, so a 2 GB limit caps the heap near 1.1 GB, and opening the graph explorer on a large workspace then crashes the app ("JavaScript heap out of memory"). Their `oom_score_adj` still puts them ahead of HelixDB when memory runs out. - `qwen3:4b` needs about 3.3 GB with an 8K context, and more for the 16K and 32K contexts used for long prompts. A prompt that does not fit in `OLLAMA_MEMORY_LIMIT` fails with an error; raise the limit on a machine with room to spare. - The model is unloaded 30 seconds after each answer instead of Ollama's default five minutes, so its memory comes back between questions. The next question reloads it in a few seconds. Set `MODEL_KEEP_ALIVE` (an Ollama duration such as `5m`, or `-1` to keep it loaded) to change this. - Give Docker at least 8 GB when you use local AI. Setup warns before Local is chosen on less, and the in-container doctor reports the memory. With less, use Claude or OpenAI, or native Ollama on macOS (`MODEL_ENDPOINT`), which runs outside Docker's memory. - If the kernel still stops HelixDB, Docker restarts it, and the app says "The database stopped because Docker ran out of memory" instead of "Cannot reach Helix". The app learns this from the kernel's `oom_kill` counter in `/proc/vmstat`, so it needs no access to Docker itself.