Operations reference

Last updated October 7, 2026

On this page

Draft for review

Everything about running an OrchKernel server, for looking things up: how big it can get, what data goes where, every setting, the command line, tokens and webhooks, backups and restore, upgrades, what survives a crash, watching it run and troubleshooting. To set a server up, start with Install in five minutes, then Production hardening.

OrchKernel is one program, orchkernel, that is both the command line and the server. The guide calls it ok: the install script and the Docker image make an ok link beside it, and alias ok=orchkernel does the same anywhere else. The server keeps everything for one company in one SQLite file, the state file.

       people (browser), the ok CLI, outside agents, webhook senders
                               |
                               | HTTPS (--domain, your certificate, or a proxy)
                               v
              ok serve
                 /api/*   the API and the event stream
                 /*       the web workspace, built into the program
                               |
                               | one writer
                               v
               state.db  (SQLite)  + secrets.key, backups/, tls/
                               |
          +--------------------+---------------------+
          v                    v                     v
   model providers     tool servers and      off-site backups
                       connections           (encrypted first)

How big it can get#

One server, one state file, one company. There are no replicas and no failover: availability comes from fast restarts (a service manager, or the Recreate Deployment with a health probe), and runs that wait for an approval survive a restart. If you need zero downtime, OrchKernel is not ready for you yet.

One company is comfortable at a million items on one server, as measured: saving, lists, dashboards and search, by words and by meaning, answer in a fraction of a second for 16 people at once (see At a million items, with its two caveats). Keep larger systems of record where they are: a CRM's 500,000 leads or an ERP's orders stay in that system as federated collections, which add nothing to this size (see Large CRM and ERP data: federate first). A federated list asks the vendor for the person's rows (their filter, the owner match of their access, the page they need) where the adapter declares those filters, so a rep's list over 500,000 vendor leads is one request; a filter the adapter does not declare is applied to at most 5,000 rows read.

The team's measurements on one Apple M2 laptop (8 cores), with 16 people working at once, on a 324,000-record file after the scale fixes:

What Measured
Saving a record (a governed write) 36 ms p50, 67 ms p95
A filtered list 35 ms p50, 61 ms p95
A sales rep's search, one person (common word / rare word) 74 / 13 ms p50
The same, 16 people at once 468 / 112 ms p50
The first server tick after the file grew; the first open 5.0 s; 3.0 s (0.33 s on later opens)
Memory during that load about 750 MB at its peak

Those searches matched words only. With search by meaning on (the local model), a rep with 4,000 notes of five windows each, among 24,000 notes and 38,000 knowledge entries, measured in process: 50 ms p50 for a common word and 32 ms for a rare one, one person at a time, and 630 ms and 340 ms with 16 people at once. Each such search scores every one of the rep's own vectors exactly; a rep with more chunks than OK_SEARCH_CANDIDATES goes through the vector index instead (31 and 18 ms, 400 and 240 ms with 16 at once). The index holds every vector there is: at a million vectors it answers in about a millisecond, from 533 MB of files it maps rather than reads into memory.

Captured mail is cheap to keep: the rules, archived history and the local embedding model cost nothing, and the local model embeds 57 to 63 chunks a second on three cores of such a server, about five hours for a million emails. ok serve runs it on a worker of its own (86 chunks a second on a million-item file, against 93 a second run directly), with requests as fast as before. What curating costs depends on how much mail reaches the model (see What capture costs).

At a million items#

The team measured one 100-person company at about a million items on the same laptop (500,000 leads, 250,000 interactions, 100,000 captured emails, 10,000 knowledge entries: a 3.3 GB file), with the machine busy and the file too large for its file cache, after the round 2 scale fixes:

What, 16 people at once unless noted Measured at 1M Before the fixes
Saving a record over HTTP 71 to 76 ms p50, 127 to 133 ms p95 the same
A filtered list 20 ms p50, 21 to 27 ms p95 the same
Dashboard tiles from precomputed totals under 6 ms p50 the same
A plain copy of the state (ok backup); opening the restored copy 27 s for 3.3 GB; 95 ms the same
A sales rep's search by words, one person (common word) 82 ms p50 (p95 372 ms) 18.0 s
The same, 16 people at once 127 to 217 ms p50 (p95 under 380 ms) 18 s
A sales lead's search, 16 people at once 80 to 181 ms p50 10.9 s
A rep's search by meaning, one person 0.12 s 8.7 s
The same, 16 people at once 0.85 s p50, 1.1 s p95; recall 0.93; no result the person may not read 20.8 s
The first open after the file grew 35 to 43 ms 20.9 s
A federated CRM of 500,000 vendor leads, 16 people for 10 s 47,000 to 80,000 answers from about 95 vendor requests, no row leaked no answer
The local model embedding inside the server 86 to 97 chunks a second 2.4 a second

Two soft spots, said plainly:

  • Words and meaning together have a p95 of 3.2 s, from each person's first (cold) search after a restart; later searches take the times above.
  • The first ticks after an upgrade that adds indexes hold writes for 8 to 12 s each, once per new index, about seven times.

A person's first search after a restart, with the file larger than the machine's file cache, took 1.4 to 2.5 s while their rows were read from disk; later ones took the times above. Search by meaning with 16 people at once is close to a second, and is the part to watch past this size.

Past these sizes the limits are known: the vector index does not merge its segments yet (measured to a million vectors), building a new dashboard total pauses writes for a few seconds on a collection of 500,000 rows, and 5 and 11 million items have not been measured. Postgres is not needed for any of the sizes above.

Disk. Keep free at least three times the state file: an encrypted backup is made from a copy of the state (about its size) into a bundle (smaller, compressed), plus the backups you keep in backups/. Admins are warned under OK_AUDIT_MIN_FREE_MB (1 GB) free.

Your data#

Where What
On your server, always The state file (people, records, threads, tasks, runs, the audit log, token hashes, sealed secrets), secrets.key, backups, certificates, the search index and the local search model. Every file readable only by the service user.
Sent to the model providers you connect What an agent's run needs: the skill's instructions, the records and messages it read, the conversation. Each call carries the most sensitive data class in it, and goes only to a model allowed that class (max_sensitivity); otherwise it is refused, never sent to a weaker model. Restricted data (personal data, captured mail) goes only to models you mark for it, such as a local one. See Models.
Sent to tool servers and connections you configure The arguments of each tool call an agent makes, after the gate allowed it.
Sent to your mail server Invites, password resets and notices, to the people they are for.
Sent to your backup bucket Encrypted backups (age, X25519) and a small plain listing file per backup with its id, time, size, versions and key fingerprints (no company name, nothing secret). The bucket's owner cannot read the backups.
Fetched by the server on its own The search-by-meaning model, once, from Hugging Face (OK_EMBED_MODEL_URL points elsewhere; the -embed images and ok embed fetch --dir avoid it); with --domain, certificates from Let's Encrypt.
Never Telemetry, usage reports, update checks.

Purged and deleted content stays in backups until they age out: keep backups only as long as your retention promises allow (Context capture).

Configuration#

Settings come from the environment, and many also from ok serve flags.

Setting Default What it does
OK_STATE or --state ./.orchkernel/state.db The state file. Set but empty, the program stops with "a value is required for '--state '".
OK_BIND or --bind 127.0.0.1:8080 Where the server listens without HTTPS. Use 0.0.0.0:8080 in a container. Not used with HTTPS.
OK_PUBLIC_URL http://<bind> on a loopback bind; with HTTPS, https://<first domain> The https address people open. Without a usable one, password sign-in is off. See The public address.
OK_DOMAIN or --domain unset Serve HTTPS for these hosts (comma list) with Let's Encrypt certificates. See Domain and HTTPS.
OK_ACME_EMAIL or --acme-email unset Contact address for certificate expiry mail.
OK_ACME_DIRECTORY or --acme-directory production production, staging (Let's Encrypt's test server) or an ACME directory URL.
OK_ACME or --acme off Under --tenants: a certificate for every tenant's host.
OK_TLS_CERT, OK_TLS_KEY or --tls-cert, --tls-key unset Your own PEM certificate chain and key instead of ACME.
OK_HTTPS_BIND or --https-bind 0.0.0.0:443 Where HTTPS listens.
OK_HTTP_BIND or --http-bind 0.0.0.0:80 Where the redirect, the ACME challenges and /api/health listen.
OK_ACME_START_WAIT 120 Seconds the start waits for the first certificates before serving placeholders.
OK_SHUTDOWN_WAIT 30 Seconds a stopping server waits for requests in flight.
OK_SECRETS_KEY the one in secrets.key Base64 of 32 random bytes. Seals connection secrets, two-step seeds, the mail password, model keys, the backup bucket's secret and every file's own key. Set, it wins over the key file. Refused under --tenants. See The key file.
OK_SECRETS_FILE secrets.key beside the state file Another place for the key file.
OK_WEBHOOK_SECRETS or --webhook-secrets unset One secret per webhook source: helpdesk=...,billing=....
OK_WEBHOOK_SECRET or --webhook-secret the one in secrets.key The secret for sources without their own.
OK_CORS_ORIGINS or --cors-origin unset Other web origins allowed to call the API from a browser.
OK_TRUSTED_PROXIES or --trusted-proxy unset Your proxy's address or range, so the server reads the client's address from X-Forwarded-For.
OK_BACKUP_SCHEDULE the page's (daily at 02:00 UTC) off, hourly or daily@HH:MM (UTC). Set, it wins and the page shows it read-only; the same for the next four.
OK_BACKUP_KEEP the page's (7d,4w,6m) How many daily, weekly and monthly backups retention keeps.
OK_BACKUP_DIR backups/ beside the state The local backup folder; set but empty turns it off.
OK_BACKUP_S3_ENDPOINT, _BUCKET, _PREFIX, _REGION, _ACCESS_KEY_ID, _SECRET_ACCESS_KEY unset An off-site destination set by the operator instead of the page. Allowed without a recovery key, with a warning at start and in ok doctor.
OK_BACKUP_RECIPIENTS unset Further age public keys (age1..., comma list) every backup is encrypted to: an operator's own key, kept elsewhere. Tenants' backups too.
OK_FILE_MAX_MB 25 The largest file anyone may attach, in MB. A company's admins may set a lower limit, never a higher one. See Files.
OK_FILES_QUOTA_GB 20 A single company's file storage, in GB, every preview counted.
OK_TENANT_FILES_QUOTA_GB 5 Each tenant's file storage under --tenants, unless its limits.files_quota_mb says otherwise (and limits.file_max_mb for its largest file).
OK_FILES_READ_IMAGES on off stops agents reading images, whatever admins set; the setting shows locked.
OK_FILES_STORE local local keeps file blobs in files/ beside the state; s3 in the bucket the next variables name. A store set wrong stops the start.
OK_FILES_S3_ENDPOINT, _BUCKET, _PREFIX, _REGION, _ACCESS_KEY_ID, _SECRET_ACCESS_KEY, _PATH_STYLE unset The bucket for OK_FILES_STORE=s3. Blobs are sealed before upload; keys are <prefix>files/<shard>/<id> (<prefix>tenants/<name>/files/... under --tenants). ok doctor tests it.
OK_TENANT_BACKUP_PRIVATE_ENDPOINTS off 1 lets tenants' S3 destinations be on this machine or a private network.
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY or GOOGLE_API_KEY, OLLAMA_HOST unset Turn on a model provider when there is no llm.yaml. See Models.
OK_LLM_CONFIG ./llm.yaml if present The model configuration file.
OK_LLM_CHECK on 0 skips the check of each provider key at start.
OK_MCP_CONFIG ./mcp.yaml if present Tool servers. Tools on servers it does not list are refused.
OK_EMBED_MODEL_DIR the models/ cache beside the state A folder already holding the search-by-meaning model (set in the -embed images); nothing is downloaded. ok embed fetch --dir DIR fills one.
OK_SQLITE_SYNC normal full makes every acknowledged write survive a power cut or an OS crash too, at the cost of slower writes. See What survives a crash.
OK_PASSWORD_DENYLIST unset A file of further refused passwords, one per line.
OK_SMTP_* unset A mail server for companies with none of their own. See Email.
OK_CONNECTION_HOSTS, OK_CONNECTION_PROXY unset Hosts connections may reach, and a proxy for their traffic.
OK_ARCHIVE_AFTER_HOURS 168 Finished tasks and runs untouched this long leave memory (they stay in the file). off never archives.
OK_AUDIT_MIN_FREE_MB 1024 Free disk space, in MB, under which admins are warned.
RUST_LOG info How much the server logs to standard output. ok::timing=debug adds where each open, tick and search spent its time.
--dev-auth off Trusts an x-ok-actor header instead of tokens. A flag only, never read from the environment. Never on a real server.
OK_DEV_ECHO_TOOLS, OK_DEV_AUTH off Development switches. Never on a real server.
--no-lock off Break glass: write without the state lock where the file system refuses locks. See Break glass.

The install script reads OK_INSTALL_DIR, OK_DATA_DIR, OK_VERSION, OK_REQUIRE_SIGNATURE, OK_INSTALL_FILE, OK_INSTALL_URL and OK_INSTALL_SHA256 (see Install). The image sets OK_STATE=/data/state.db and OK_BIND=0.0.0.0:8080, and its health check uses OK_HEALTH_PORT (8080; 80 with deploy/compose.https.yml). Tuning settings (OK_DB_READERS, OK_SEARCH_CANDIDATES, OK_ARCHIVE_TICK_MS, the audit checkpoint settings) are in docs/deployment.md in the repository.

Keep provider keys, webhook secrets, OK_SECRETS_KEY, bucket secrets and tokens in a secret store (a Kubernetes Secret, Docker secrets, a systemd EnvironmentFile of mode 0600, Vault). Never commit llm.yaml, .env, secrets.key or a recovery key.

ok doctor#

ok doctor checks the install in one run and never writes the state; it is safe beside a running server:

Section Says
Build The SQLite note, when the program was built without OrchKernel's settings
State The file, its size, storage version and journal mode, OK_SQLITE_SYNC, and a network file system
Lock Free, or held by the server (its pid and since when) or by a command
Files "private", or the files others can read; --fix-permissions fixes them
Secrets Where the key and the webhook secret came from and the key's id; what the state holds sealed, by key id; how many secrets do not open
HTTPS Each stored certificate with the days left (a staging one warned about), a configured domain without one, your own certificate files; or "not served by OrchKernel"
Models The server's and the workspace's models, each provider checked (no tokens spent), the default model, and whether agents can run
Backups The schedule and next run, each destination's last success and failure, the recovery key, and warnings (no recovery key, on this server only, an operator destination without recipients)

Each line is ok, warning or error, and the exit status is 1 when any is an error.

The command line#

The commands an operator uses most; CLI and API reference has them all.

Command Does While the server runs
ok serve [--bind A] [--domain H] [--tls-cert F --tls-key F] Runs the API and the workspace
ok init --admin ada --email ... Creates the company and the first admin's setup link refused
ok doctor [--fix-permissions] Checks the install works
ok backups list [--json] Backups in every destination, with their checks works
ok backups run Backs up now refused: use Back up now or POST /api/backups/run
ok backups verify <id> [--destination D] [--identity F] Opens a backup and checks it works
ok backups prune Removes what retention no longer keeps, never the last verified backup refused
ok backups recovery-key [--out F] Creates a recovery key; its secret printed once, or written 0600 refused: use the Backups page
ok backups decrypt <file> --identity F --out DIR Opens a backup into a folder without restoring it works
ok restore <file or id> [--identity F] [--yes] [--confirm NAME] [--replace-config] Restores, with the server stopped (below) refused
ok backup <path> A plain, unencrypted copy of the state file alone, for scripts works
ok embed fetch --dir DIR Downloads the search-by-meaning model and checks it, for offline use works
ok secrets rotate [--new-key-env V] Re-seals every stored secret under a new key refused
ok audit verify Checks the audit log's chain works

People, tokens and access#

People sign in to the workspace with their email address and a password (see Sign-in settings). Tokens are for the API, the CLI's --server mode, outside agents and CI, never for signing in to the workspace.

A token belongs to one person or agent. The server keeps only its hash, so a lost token cannot be shown again: issue a new one.

You want to In the workspace With the API With the CLI
A token for yourself Your profile, API tokens, New token (label, lifetime, scopes) POST /api/tokens {"label":"laptop","ttl":"7d"}, signed in to the workspace ok --as maya token issue --for maya --label laptop, server stopped
A kernel agent's token POST /api/tokens {"actor":"policy-ci","label":"ci","scopes":["policy"]}, signed in as an admin ok --as ada token issue --for policy-ci --label ci --scope policy, server stopped
An outside agent's token Directory, the agent, Tokens POST /api/gateway/agents/<id>/tokens ok agent register
See tokens Your profile (yours); the people panel (an admin, anyone's) GET /api/tokens ok --as ada token list
Revoke one Revoke on its row DELETE /api/tokens/id/<id> ok --as ada token revoke --id <id>
  • A token never issues a token, and never changes how anyone signs in: POST /api/tokens, passwords, two-step, email and sessions answer 403 session_required "Sign in to the workspace to do this; a token cannot." to a bearer.
  • Nobody issues a person's token for them. An admin asking for someone else's gets 403 own_tokens_only; each person issues their own.
  • Lifetimes. A person's token lives at most 90 days (an admin changes the ceiling in Governance, Sign-in); 30 days unless chosen. A kernel agent's lives at most 90 days and must carry the policy or changes scope.
  • Old sign-in tokens. Tokens people pasted into the old sign-in page keep working on the API until they expire; the people panel counts them and revokes them all with one action.

Send a token as Authorization: Bearer <token>. The one exception is the event stream, GET /api/events/stream?token=..., because a browser cannot set headers there (the workspace itself uses its cookie).

Someone leaves. An admin offboards them from the people panel (or POST /api/actors/<id>/offboard): their sessions end, their tokens are deleted, their running work and delegations are cancelled. Deactivating instead is reversible: sign-in is refused, sessions end, and tokens stop working until the person is reactivated.

Webhooks#

POST /api/webhooks/<source> turns an event from another system into tasks, through the triggers (rules in a module that start work) that listen to that source. The support module's helpdesk source, for example, creates a "Triage incoming ticket" task for its triage agent. A source with no secret is refused. The sender signs each delivery:

  • x-ok-timestamp: the time of sending in Unix seconds, within five minutes of the server's clock.
  • x-ok-signature: sha256= and the hex HMAC-SHA256, keyed with the source's secret, of the timestamp, a dot and the body exactly as sent.
  • x-ok-delivery (optional): a unique id. A repeat within 24 hours creates nothing and answers with the first delivery's tasks, so the sender can retry safely.

With the server started with OK_WEBHOOK_SECRETS=helpdesk=... and the same secret in HELPDESK_SECRET:

sh
body='{"external_id":"Z-100","subject":"Charged twice"}'
ts=$(date +%s)
sig=$(printf '%s.%s' "$ts" "$body" | openssl dgst -sha256 -hmac "$HELPDESK_SECRET" | sed 's/^.* //')
curl -sS -X POST https://work.acme.example/api/webhooks/helpdesk \
  -H 'content-type: application/json' \
  -H "x-ok-timestamp: $ts" -H "x-ok-signature: sha256=$sig" \
  -H 'x-ok-delivery: evt-0001' --data-raw "$body"
What the sender sees When
202 {"delivery":"evt-0001","replayed":false,"tasks":["..."]} Accepted. Agents run after the answer, so a slow model never holds the sender.
202 with "replayed":true and the same tasks The same x-ok-delivery again.
401 "bad webhook signature" Wrong secret or a changed body.
401 "x-ok-timestamp is more than five minutes from the server clock" A stale or future timestamp.
401 "no webhook secret is configured for this source" Neither OK_WEBHOOK_SECRETS nor OK_WEBHOOK_SECRET covers it.
400 bad_request The body is not JSON.

Senders that cannot sign may send the secret itself in x-ok-webhook-secret. It is accepted (202, with "delivery":null) but offers no protection against a replayed request; prefer signatures.

Backups and restore#

Setting backups up (the recovery key, an off-site bucket, checking Verified, testing a restore) is in Production hardening. This section is how they work.

What a backup holds#

Member What
manifest.json The backup's id, kind, time, OrchKernel and storage versions, company, the audit log's id and head, the key ids, and each file's size and SHA-256; signed (an HMAC) with a key derived from the secrets key
state.db The state, copied with SQLite's online backup while the server runs. Model, mail, sign-in and backup settings are in it
secrets.key The secrets key and webhook secret in use, also when they came from the environment
secrets.key.<kid> Keys kept after a rotation, so older sealed values still open
config/llm.yaml, config/mcp.yaml The configuration files the server uses, when there are any
packs/<name>/ Pack folders beside the state, when there are any
files/<shard>/<id>[.thumb|.view|.txt] Every live file's blobs, as sealed on disk (a removed file's are not). With OK_FILES_STORE=s3 the bucket is the store: the manifest lists its objects instead, and a restore says which are missing. Turn on the bucket's versioning

A file removed after a backup stays in that backup until retention prunes it. A restore puts back the state and files/ together, moving the folder as it was aside as pre-restore-<time>.files/ (removed with the pre-restore-<time>.db copy).

Left out, and rebuilt or fetched at start: the search index, the local model, certificates, the lock, -wal and -shm.

The format and the keys#

A backup is orchkernel-<id>.tar.gz.age, the id being its UTC time and six random hex digits (2026-10-07T020000Z-a1b2c3): a tar archive, gzipped, encrypted in the age format to X25519 keys. Beside it, orchkernel-<id>.json lists it without opening it (kind, time, versions, size, checksum, key fingerprints, when and how it was verified); it names no company and holds nothing secret, and a restore never trusts it.

Every backup is encrypted to:

Key Opens it Kept
This server's key Derived from the secrets key; a restore on this server needs nothing else Nowhere: derived when needed
The recovery key Anyone holding the recovery key file Only by you, offline. The server keeps its public half and fingerprint
OK_BACKUP_RECIPIENTS Each operator key listed By the operator

A backup made before a key rotation opens with the kept secrets.key.<kid> or the recovery key. A new recovery key applies to backups from then on.

Opening a backup without OrchKernel, with the age tool and the recovery key file:

sh
age -d -i orchkernel-recovery-key-9f3e1a2b.txt orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age | tar xz

or ok backups decrypt <file> --identity <key file> --out <folder>. Either gives the files above; state.db opens with any SQLite tool.

How a run goes#

  1. When: at the scheduled time, or five minutes after a start when the last good backup is older than one interval. A run whose state has not changed since the last good one records "unchanged" and writes nothing.
  2. Make: copy the state, pack, compress and encrypt into backups/.partial/.
  3. Verify here: open it with the server's key, check every file's checksum and the manifest's signature, PRAGMA quick_check and the audit chain of the state inside. Only a verified backup is kept.
  4. Upload to the off-site bucket (one upload up to 256 MiB, in 64 MiB parts above), then verify there: download and compare (the default). A copy that differs is removed from the bucket and the run fails.
  5. Record: BackupTaken and BackupVerified events, and the run in the backup log the page shows.
  6. Prune each destination by retention.
  7. Report: on a failure, an inbox item for every admin (once per failing streak per destination) and a BackupFailed event.

One run at a time: Back up now during a run answers 409 busy.

Retention#

Per destination, over verified backups only: the newest backup of each of the last 7 days that have one, of each of the last 4 ISO weeks and of each of the last 6 months (the union is kept). Always kept as well: the newest verified backup, whatever its age, and backups made before a restore for 7 days. Unverified or failed files older than a day are removed only when a newer verified backup exists there. Nothing is pruned after a failed run, and the last verified backup is never removed, from the page either.

Restore from the Backups page#

For a server that runs. On the backup's row, choose Restore:

  1. Fetching and checking. The backup is fetched (downloaded from the bucket if needed) and checked as in a run; nothing changes yet. A backup this server's keys cannot open asks for the recovery key it was made with (paste the key file's text).
  2. Where it came from. "Made by this workspace" (its signature checks with a key this server holds), "An earlier state of this workspace" (its audit log is an earlier point of this one), or "Not checked". A backup that is not checked is restored only from the command line.
  3. What will change: the time it goes back to, how many changes since are undone ("not counted" when the backup is not an earlier point of the workspace as it is now, such as a copy made before an earlier restore), people, records, tasks and threads now and in the backup, tokens issued since (revoked), the secrets key (kept; sealed values re-sealed to it), and configuration files that differ (kept; the backup's written beside them as llm.yaml.from-backup). A red line says what is lost: "Everything done after 7 Oct 2026, 02:00 is lost (42 changes). Everyone is signed out."
  4. Type the company's name and choose Restore. A sign-in within the last ten minutes is needed; otherwise the page asks for your password.

Then the server makes a backup of the state as it is now (kind "Before a restore", kept 7 days), answers every request but the health check with 503 restoring, stops running agents, and restarts itself with the same command line (HTTPS included), never letting go of its lock. The new process swaps the backup in before it serves, and its log says "restored from the backup of 2026-10-07 02:00 UTC; everyone signs in again". The page then sends you to sign in. Passwords and two-step stay as they are now (nobody gets back a password they changed); sessions and sign-in links end.

The server's backup setup is not rolled back: the schedule, the off-site destination, the recovery key and the record of past backups stay as they were just before the restore.

If the server does not come back (the restart failed, the process exits with status 3), the page says so after two minutes. Start the server again as you normally do (ok serve, or your service): it finishes the restore before it serves. A crash at any step leaves either the state as it was or the restored one, never a mix; the plain copy of the state as it was, pre-restore-<time>.db beside it, is removed after 7 days.

Undoing a restore. Restore the backup of kind Before a restore the same way; the sheet says "This undoes the restore of ...".

ok restore, with the server stopped#

For a server that will not start, a backup that is not checked, or a new machine:

sh
ok --state /var/lib/orchkernel/state.db restore 2026-10-07T020000Z-a1b2c3
ok --state /var/lib/orchkernel/state.db restore /path/to/orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age

A backup id is looked for in the local folder, then in the off-site bucket. It runs the same checks, prints the same summary, asks for the company's name (--confirm NAME gives it; --yes skips the question for a checked backup), and does the same swap. It is refused with exit status 75 while the server runs. A backup that is not checked needs --yes and the company's name, and says why first.

On a new machine (the server and its key are lost): put the backup file and the recovery key file on it, and restore into an empty folder, giving the full state path:

sh
mkdir -m 700 /var/lib/orchkernel
ok --state /var/lib/orchkernel/state.db restore orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age \
  --identity orchkernel-recovery-key-9f3e1a2b.txt

It writes state.db, secrets.key (and kept keys) from the backup, and llm.yaml, mcp.yaml and pack folders where none exist (--replace-config replaces existing ones), then prints the exact command to start the server on it:

text
restored 2026-10-07T020000Z-a1b2c3 into /var/lib/orchkernel/state.db; start the server on it with:
  orchkernel --state /var/lib/orchkernel/state.db serve

Use that command (or set OK_STATE in your service): a plain ok serve elsewhere uses ./.orchkernel/state.db and starts an empty workspace. People sign in with the passwords they had at the backup.

Versions#

The backup's storage version Result
Newer than this program Refused before anything changes, naming the version to install
The same Restored
Older Restored and converted at the first open; the backup is the way back

A plain copy for scripts: ok backup#

ok backup /backups/state-$(date +%F).db copies the state file alone with SQLite's online backup API, 0600, safe while the server runs. It is not encrypted and holds no key or configuration; restoring it means stopping the server, moving state.db (and any -wal, -shm) aside and copying the file in, with the secrets key it was sealed under. Prefer the encrypted backups above.

Tenants#

Under --tenants, each tenant's admins have their own Backups page: off until they turn it on, kept under the operator's root, encrypted to the tenant's key, its recovery key and OK_BACKUP_RECIPIENTS. A restore from it goes through the tenancy restore, without restarting the server. The operator's nightly copies keep running beside them. See Hosting several organisations.

Upgrades and rollback#

  1. Back up (Back up now, or ok backups run with the server stopped), and stop every scheduled job that opens the state file with the CLI.
  2. Replace the program (run the install script again) or the image tag.
  3. Start the server and check /api/health.

Some releases change how the file is stored, and the first open by the new program converts it:

  • It checks there is room first. Without it, the open fails with "upgrading the state file needs N bytes free beside it, and M are" and nothing changes.
  • It writes a copy beside the file, state.db.pre-upgrade-v<n>.db (0600), and keeps it. Delete it yourself once the new release has run well.
  • It converts in one transaction, so a stop half-way leaves the old file untouched.

An older program refuses a converted file: "the state file is newer than this binary". Rolling back after a storage change means restoring the pre-upgrade copy or a backup, and losing what changed since the upgrade. Releases without a storage change can be rolled back by swapping the program again.

If an old server kept writing after the conversion, the next open may refuse with "the state file was written by two versions of orchkernel". Stop everything and restore the pre-upgrade copy or a later backup.

The first open by a release that adds record indexes builds them before the server answers: a few seconds per index on a large collection, once.

Running commands while the server runs#

The server holds a lock on state.db.lock, beside the state file, from before it opens the file until it exits (and across its own restart after a restore). The operating system lets go of it when the process ends, however it ends, so there is no stale lock to clean up on a local disk.

  • Commands that only read open a snapshot of the file and never write it, so they work beside the server and see what it has saved: ok actor list, ok runs, ok tasks, ok inbox, ok log, ok replay, ok threads, ok search, ok brain query, ok module list, ok token list, ok user list, ok mail show, ok connection list, ok policy export, ok backups list and the other listings.

  • Commands that change the state (ok init, ok actor add-human, ok token issue, ok pack install, ok kill, ok tick, ok ask, ok secrets rotate, ok backups run, ok restore, every other one) are refused while the server runs, with exit status 75 and nothing written:

    text
    The server is running on /data/state.db (pid 4121, since 09:30 UTC).
    Use the workspace or the API (add --server https://work.acme.example), or stop the server, then run this again.
    

    Two such commands at once are refused the same way, and so is a second server.

  • ok backup, ok audit, ok doctor, ok backups decrypt and ok embed fetch load no kernel of their own and work beside the server.

  • With --server, the people, mail and policy commands go to the running server's API instead (see CLI reference).

  • Under tenancy the same holds per tenant.

In Docker, docker exec orchkernel orchkernel --as ada actor list works beside the server; a command that changes the state needs the container stopped and a one-off container on the same volume.

The state on a network disk (NFS, SMB, AFP, WebDAV, FUSE) is not supported, as SQLite's write-ahead log is not: the start and ok doctor say so.

Break glass#

Where the file system refuses locks altogether (ENOLCK, EOPNOTSUPP), a command that changes the state is refused with "the file system holding /data/state.db does not support locks". If you are certain no server runs on the file, --no-lock lets it through, with "warning: writing without the lock; make sure no server runs on this file". --no-lock is refused while a live process holds the lock: there is no override for a running server, because a write beside it is exactly what the lock prevents.

What survives a crash#

What happened What you find after a restart
The process died (SIGKILL, out of memory, a panic) Every write the server answered for. The file is never damaged. Runs that were mid-step end as stopped, with the reason "restart", their steps so far kept. Runs waiting for an approval or an answer resume when it comes. Sessions, links and passwords are kept
The machine lost power or the OS crashed The file is never damaged. With the default (synchronous = NORMAL in WAL), the writes of the last moments before the cut may be gone even though they were answered. Set OK_SQLITE_SYNC=full to keep those too, at the cost of slower writes
The disk filled Writes fail with 500 and nothing is half written; reads keep working. A backup run fails, removes its partial file and prunes nothing
A crash during an upgrade <state>.pre-upgrade-v<n>.db is the file before it
A crash during a backup The older backups are untouched; the partial file is removed by the next run
A crash during a restore The next start puts the state as it was back, or finishes the restore if it had recorded it: never a mix
A crash during key rotation Settled at the next start: every secret opens with the key on disk
The vector index Replayed from its change log, or rebuilt in the background
Model calls in flight Not retried; their run is stopped with "restart"
Email being sent The link stays queued; the people panel offers Resend after 15 minutes
Outside agents' actions being executed Marked interrupted by the start-up sweep

The tests behind this table kill the server with SIGKILL in the middle of 16 people's writes, in the middle of a run, while a run waits for an approval, right after a password change, in the middle of a backup of a 200 MB file, at both steps of a key rotation and at three points of a restore, then check with a fresh process that PRAGMA integrity_check answers ok, ok audit verify passes, every answered write is there and a new server starts (crates/ok-cli/tests/it/crash.rs, restore_crash.rs).

Watching it run#

What How
Is it up GET /api/health, no token. Use it for load balancer and container probes.
Server log Standard output and standard error, level set by RUST_LOG. With --service: journalctl -u orchkernel (Linux) or orchkernel.log in the data folder (macOS).
What happened Events in the sidebar, GET /api/events, or the live GET /api/events/stream. Someone who is not an admin sees only events about their own threads, tasks and runs.
Backups The Backups page, GET /api/backups/status, ok backups list, ok doctor
One run, step by step ok replay --run <id> (or --thread, --task) rebuilds it from the log alone.
The audit log is intact ok audit verify and ok audit head, safe beside a running server. See Events and the audit log.

The server advances its own clock every 30 seconds: scheduled triggers, overdue work, stalled threads, archiving. This is the tick. To drive it from your own scheduler, call POST /api/tick with an admin token (deploy/k8s/tick-cronjob.yaml does this every minute).

When the audit log cannot record a decision (a full disk, say), the server refuses the action with 503 audit_unavailable rather than act unrecorded.

Troubleshooting#

You see What to do
The first-admin link is lost or expired Restart the server: while the company has no admin with a password, each start prints a new one. With an admin already, ok user sign-in-link ada (server stopped).
"password sign-in is off" at start, or on the sign-in page People must open an https:// address: set OK_PUBLIC_URL to it (behind a proxy), or use --domain. On your own computer, open http://localhost:8080 or http://127.0.0.1:8080.
The link names 127.0.0.1, which your browser cannot open The server runs elsewhere and listens only on itself. Use --domain, a proxy, or a tunnel: ssh -L 8080:127.0.0.1:8080 <server>.
"cannot listen on 0.0.0.0:443 ... Ports below 1024 need a privilege" Run with --service (it grants exactly that), sudo setcap cap_net_bind_service=+ep "$(command -v orchkernel)", or higher ports with a forward.
"cannot listen on ...: another program uses it" Another web server holds 80 or 443 (sudo ss -ltnp 'sport = :443'). Stop it, or put OrchKernel behind it as a proxy.
"no certificate for work.acme.example yet (...)"; browsers warn about the certificate Let's Encrypt could not reach port 80 for that name: check the DNS record answers the server's address, port 80 is open to the internet, and nothing else answers it. The server retries on its own; ok doctor shows the state. Use --acme-directory staging while you fix it, to stay inside Let's Encrypt's limits.
The certificate is "not trusted" Still on staging: remove --acme-directory staging (or OK_ACME_DIRECTORY) and restart; a certificate from another ACME server is replaced at start.
"Backups to offsite are failing: uploading ...: the provider answered 403 (AccessDenied ...)" The key lacks a permission on the bucket or folder: it needs to put, get, delete and list. Test on the Backups page names what failed.
"... could not be reached" The endpoint is wrong or blocked: check it from the server with curl -I <endpoint>. A bucket on this machine or a private network needs http:// to be allowed there; under --tenants the operator allows it with OK_TENANT_BACKUP_PRIVATE_ENDPOINTS=1.
"its secret access key is not set" or a locked secret The bucket secret is sealed under a key the server no longer has: enter the secret again on the Backups page.
"No recovery key yet" Create one on the Backups page and keep it offline. Backups made before it do not open with it.
"N secrets sealed under key ... and no key was found" The secrets key is missing or different: put secrets.key back, or set OK_SECRETS_KEY to the key it names. A backup carries its key: restoring the latest one also works.
A restore says "made by OrchKernel 0.4 at storage version 4; install 0.4 or later" The backup is from a newer release: install that release, then restore.
"this server cannot tell where it came from" The backup is not this workspace's, or was made under a key this server no longer holds and its history does not match. Restore it only if you know where it came from, with ok restore ... --yes and the company's name, server stopped.
The Restore page says the server has not come back Look at the server's log. Start it again as you normally do; the next start finishes or undoes the restore.
A command says "The server is running on ..." (exit 75) It changes the state: use the workspace or the API, or stop the server first.
"the state file is newer than this binary" You went back to an older release after an upgrade converted the file: install the newer one, or restore a backup made before the upgrade.

Not possible yet#

  • Publishing: no release, image or install script has been published yet; build them yourself (Install).
  • More than one server or replica for a company, or failover. There is no Postgres backend.
  • Single sign-on (OIDC, SAML) and passkeys.
  • Certificates by DNS or TLS-ALPN challenge, and wildcard certificates from Let's Encrypt (give your own with --tls-cert).
  • More than one off-site destination, or destinations other than S3-compatible storage (SFTP, for example).
  • Restoring to a point between two backups.
  • Changing state with the CLI while the server runs.
  • Keeping the state on a network disk.
  • Exporting metrics or traces to OpenTelemetry; read standard output and the event log.
  • Apple-notarized macOS programs (the install script does not quarantine them); Windows.

License#

OrchKernel is source-available under the Business Source License 1.1. You may use, modify and run it in production for your own organization, or run a deployment for one client organization; separate deployments for separate clients are fine. Offering it as a hosted or managed service in which one deployment serves more than one organization needs a commercial license. Each version converts to Apache-2.0 four years after it is first published, and the client libraries under sdk/ are Apache-2.0 already. The LICENSE file in the repository is the authoritative text.

For developers#

  • ok serve and the tenancy host start every kernel through one path, crates/ok-server/src/startup.rs (prepare_kernel): a pending restore is finished first (ok_server::restart::finish_pending_restore), then the secrets key check, the context hooks, the process configuration, the module reconcile, the gateway's observer, the audit verifier, the gateway start-up sweep, the statistics refresh, the local indexer worker, the backup job (backup_jobs.rs), then the request state. Each 30-second tick runs tick_one.
  • HTTPS is crates/ok-tls (the ACME manager on instant-acme, the certificate store, the SNI resolver, the port 80 router) served by crates/ok-server/src/https.rs; the flags are crates/ok-cli/src/serve_tls.rs.
  • Backups are crates/ok-backup (bundle, age encryption, the local and S3 destinations, retention, the run, restore); the routes are crates/ok-server/src/backups_api.rs, the settings and log rows crates/ok-kernel/src/backups.rs. A restore's swap is shared with the tenancy restore (crates/ok-tenancy/src/swap.rs); the restart hands the lock over across exec (crates/ok-server/src/restart.rs, ok_db::lock).
  • ok backup loads no kernel: it copies through SQLite's online backup API (Kernel::backup_file) into a .partial file made 0600, renamed into place once complete.
  • The state lock is ok_db::lock (flock); crates/ok-cli/src/access.rs classes every command (no kernel, snapshot, write, serve).
  • The key file is ok_secrets::keyfile; its resolution, generation and rotation are crates/ok-cli/src/secrets_file.rs. Private files are ok_db::private.
  • deploy/install.sh is tested by crates/ok-cli/tests/it/install_script.rs; the release workflow is .github/workflows/release.yml.
  • docs/deployment.md is the full operator reference: the Kubernetes bootstrap pod, MCP servers, connection key rotation, archiving and indexes.