Operations reference
Last updated October 7, 2026
On this page
- How big it can get
- At a million items
- Your data
- Configuration
- ok doctor
- The command line
- People, tokens and access
- Webhooks
- Backups and restore
- What a backup holds
- The format and the keys
- How a run goes
- Retention
- Restore from the Backups page
- ok restore, with the server stopped
- Versions
- A plain copy for scripts: ok backup
- Tenants
- Upgrades and rollback
- Running commands while the server runs
- Break glass
- What survives a crash
- Watching it run
- Troubleshooting
- Not possible yet
- License
- For developers
Draft for review
Everything about running an OrchKernel server, for looking things up: how big it can get, what data goes where, every setting, the command line, tokens and webhooks, backups and restore, upgrades, what survives a crash, watching it run and troubleshooting. To set a server up, start with Install in five minutes, then Production hardening.
OrchKernel is one program, orchkernel, that is both the command line and
the server. The guide calls it ok: the install script and the Docker image
make an ok link beside it, and alias ok=orchkernel does the same
anywhere else. The server keeps everything for one company in one SQLite
file, the state file.
people (browser), the ok CLI, outside agents, webhook senders
|
| HTTPS (--domain, your certificate, or a proxy)
v
ok serve
/api/* the API and the event stream
/* the web workspace, built into the program
|
| one writer
v
state.db (SQLite) + secrets.key, backups/, tls/
|
+--------------------+---------------------+
v v v
model providers tool servers and off-site backups
connections (encrypted first)
How big it can get#
One server, one state file, one company. There are no replicas and no
failover: availability comes from fast restarts (a service manager, or the
Recreate Deployment with a health probe), and runs that wait for an
approval survive a restart. If you need zero downtime, OrchKernel is not
ready for you yet.
One company is comfortable at a million items on one server, as measured: saving, lists, dashboards and search, by words and by meaning, answer in a fraction of a second for 16 people at once (see At a million items, with its two caveats). Keep larger systems of record where they are: a CRM's 500,000 leads or an ERP's orders stay in that system as federated collections, which add nothing to this size (see Large CRM and ERP data: federate first). A federated list asks the vendor for the person's rows (their filter, the owner match of their access, the page they need) where the adapter declares those filters, so a rep's list over 500,000 vendor leads is one request; a filter the adapter does not declare is applied to at most 5,000 rows read.
The team's measurements on one Apple M2 laptop (8 cores), with 16 people working at once, on a 324,000-record file after the scale fixes:
| What | Measured |
|---|---|
| Saving a record (a governed write) | 36 ms p50, 67 ms p95 |
| A filtered list | 35 ms p50, 61 ms p95 |
| A sales rep's search, one person (common word / rare word) | 74 / 13 ms p50 |
| The same, 16 people at once | 468 / 112 ms p50 |
| The first server tick after the file grew; the first open | 5.0 s; 3.0 s (0.33 s on later opens) |
| Memory during that load | about 750 MB at its peak |
Those searches matched words only. With search by meaning on (the local
model), a rep with 4,000 notes of five windows each, among 24,000 notes and
38,000 knowledge entries, measured in process: 50 ms p50 for a common word
and 32 ms for a rare one, one person at a time, and 630 ms and 340 ms with
16 people at once. Each such search scores every one of the rep's own
vectors exactly; a rep with more chunks than OK_SEARCH_CANDIDATES goes
through the vector index instead (31 and 18 ms, 400 and 240 ms with 16 at
once). The index holds every vector there is: at a million vectors it
answers in about a millisecond, from 533 MB of files it maps rather than
reads into memory.
Captured mail is cheap to keep: the rules, archived history and the local
embedding model cost nothing, and the local model embeds 57 to 63 chunks a
second on three cores of such a server, about five hours for a million
emails. ok serve runs it on a worker of its own (86 chunks a second on a
million-item file, against 93 a second run directly), with requests as fast
as before. What curating costs depends on how much mail reaches the model
(see What capture costs).
At a million items#
The team measured one 100-person company at about a million items on the same laptop (500,000 leads, 250,000 interactions, 100,000 captured emails, 10,000 knowledge entries: a 3.3 GB file), with the machine busy and the file too large for its file cache, after the round 2 scale fixes:
| What, 16 people at once unless noted | Measured at 1M | Before the fixes |
|---|---|---|
| Saving a record over HTTP | 71 to 76 ms p50, 127 to 133 ms p95 | the same |
| A filtered list | 20 ms p50, 21 to 27 ms p95 | the same |
| Dashboard tiles from precomputed totals | under 6 ms p50 | the same |
A plain copy of the state (ok backup); opening the restored copy |
27 s for 3.3 GB; 95 ms | the same |
| A sales rep's search by words, one person (common word) | 82 ms p50 (p95 372 ms) | 18.0 s |
| The same, 16 people at once | 127 to 217 ms p50 (p95 under 380 ms) | 18 s |
| A sales lead's search, 16 people at once | 80 to 181 ms p50 | 10.9 s |
| A rep's search by meaning, one person | 0.12 s | 8.7 s |
| The same, 16 people at once | 0.85 s p50, 1.1 s p95; recall 0.93; no result the person may not read | 20.8 s |
| The first open after the file grew | 35 to 43 ms | 20.9 s |
| A federated CRM of 500,000 vendor leads, 16 people for 10 s | 47,000 to 80,000 answers from about 95 vendor requests, no row leaked | no answer |
| The local model embedding inside the server | 86 to 97 chunks a second | 2.4 a second |
Two soft spots, said plainly:
- Words and meaning together have a p95 of 3.2 s, from each person's first (cold) search after a restart; later searches take the times above.
- The first ticks after an upgrade that adds indexes hold writes for 8 to 12 s each, once per new index, about seven times.
A person's first search after a restart, with the file larger than the machine's file cache, took 1.4 to 2.5 s while their rows were read from disk; later ones took the times above. Search by meaning with 16 people at once is close to a second, and is the part to watch past this size.
Past these sizes the limits are known: the vector index does not merge its segments yet (measured to a million vectors), building a new dashboard total pauses writes for a few seconds on a collection of 500,000 rows, and 5 and 11 million items have not been measured. Postgres is not needed for any of the sizes above.
Disk. Keep free at least three times the state file: an encrypted backup
is made from a copy of the state (about its size) into a bundle (smaller,
compressed), plus the backups you keep in backups/. Admins are warned
under OK_AUDIT_MIN_FREE_MB (1 GB) free.
Your data#
| Where | What |
|---|---|
| On your server, always | The state file (people, records, threads, tasks, runs, the audit log, token hashes, sealed secrets), secrets.key, backups, certificates, the search index and the local search model. Every file readable only by the service user. |
| Sent to the model providers you connect | What an agent's run needs: the skill's instructions, the records and messages it read, the conversation. Each call carries the most sensitive data class in it, and goes only to a model allowed that class (max_sensitivity); otherwise it is refused, never sent to a weaker model. Restricted data (personal data, captured mail) goes only to models you mark for it, such as a local one. See Models. |
| Sent to tool servers and connections you configure | The arguments of each tool call an agent makes, after the gate allowed it. |
| Sent to your mail server | Invites, password resets and notices, to the people they are for. |
| Sent to your backup bucket | Encrypted backups (age, X25519) and a small plain listing file per backup with its id, time, size, versions and key fingerprints (no company name, nothing secret). The bucket's owner cannot read the backups. |
| Fetched by the server on its own | The search-by-meaning model, once, from Hugging Face (OK_EMBED_MODEL_URL points elsewhere; the -embed images and ok embed fetch --dir avoid it); with --domain, certificates from Let's Encrypt. |
| Never | Telemetry, usage reports, update checks. |
Purged and deleted content stays in backups until they age out: keep backups only as long as your retention promises allow (Context capture).
Configuration#
Settings come from the environment, and many also from ok serve flags.
| Setting | Default | What it does |
|---|---|---|
OK_STATE or --state |
./.orchkernel/state.db |
The state file. Set but empty, the program stops with "a value is required for '--state '". |
OK_BIND or --bind |
127.0.0.1:8080 |
Where the server listens without HTTPS. Use 0.0.0.0:8080 in a container. Not used with HTTPS. |
OK_PUBLIC_URL |
http://<bind> on a loopback bind; with HTTPS, https://<first domain> |
The https address people open. Without a usable one, password sign-in is off. See The public address. |
OK_DOMAIN or --domain |
unset | Serve HTTPS for these hosts (comma list) with Let's Encrypt certificates. See Domain and HTTPS. |
OK_ACME_EMAIL or --acme-email |
unset | Contact address for certificate expiry mail. |
OK_ACME_DIRECTORY or --acme-directory |
production |
production, staging (Let's Encrypt's test server) or an ACME directory URL. |
OK_ACME or --acme |
off | Under --tenants: a certificate for every tenant's host. |
OK_TLS_CERT, OK_TLS_KEY or --tls-cert, --tls-key |
unset | Your own PEM certificate chain and key instead of ACME. |
OK_HTTPS_BIND or --https-bind |
0.0.0.0:443 |
Where HTTPS listens. |
OK_HTTP_BIND or --http-bind |
0.0.0.0:80 |
Where the redirect, the ACME challenges and /api/health listen. |
OK_ACME_START_WAIT |
120 |
Seconds the start waits for the first certificates before serving placeholders. |
OK_SHUTDOWN_WAIT |
30 |
Seconds a stopping server waits for requests in flight. |
OK_SECRETS_KEY |
the one in secrets.key |
Base64 of 32 random bytes. Seals connection secrets, two-step seeds, the mail password, model keys, the backup bucket's secret and every file's own key. Set, it wins over the key file. Refused under --tenants. See The key file. |
OK_SECRETS_FILE |
secrets.key beside the state file |
Another place for the key file. |
OK_WEBHOOK_SECRETS or --webhook-secrets |
unset | One secret per webhook source: helpdesk=...,billing=.... |
OK_WEBHOOK_SECRET or --webhook-secret |
the one in secrets.key |
The secret for sources without their own. |
OK_CORS_ORIGINS or --cors-origin |
unset | Other web origins allowed to call the API from a browser. |
OK_TRUSTED_PROXIES or --trusted-proxy |
unset | Your proxy's address or range, so the server reads the client's address from X-Forwarded-For. |
OK_BACKUP_SCHEDULE |
the page's (daily at 02:00 UTC) | off, hourly or daily@HH:MM (UTC). Set, it wins and the page shows it read-only; the same for the next four. |
OK_BACKUP_KEEP |
the page's (7d,4w,6m) |
How many daily, weekly and monthly backups retention keeps. |
OK_BACKUP_DIR |
backups/ beside the state |
The local backup folder; set but empty turns it off. |
OK_BACKUP_S3_ENDPOINT, _BUCKET, _PREFIX, _REGION, _ACCESS_KEY_ID, _SECRET_ACCESS_KEY |
unset | An off-site destination set by the operator instead of the page. Allowed without a recovery key, with a warning at start and in ok doctor. |
OK_BACKUP_RECIPIENTS |
unset | Further age public keys (age1..., comma list) every backup is encrypted to: an operator's own key, kept elsewhere. Tenants' backups too. |
OK_FILE_MAX_MB |
25 |
The largest file anyone may attach, in MB. A company's admins may set a lower limit, never a higher one. See Files. |
OK_FILES_QUOTA_GB |
20 |
A single company's file storage, in GB, every preview counted. |
OK_TENANT_FILES_QUOTA_GB |
5 |
Each tenant's file storage under --tenants, unless its limits.files_quota_mb says otherwise (and limits.file_max_mb for its largest file). |
OK_FILES_READ_IMAGES |
on | off stops agents reading images, whatever admins set; the setting shows locked. |
OK_FILES_STORE |
local |
local keeps file blobs in files/ beside the state; s3 in the bucket the next variables name. A store set wrong stops the start. |
OK_FILES_S3_ENDPOINT, _BUCKET, _PREFIX, _REGION, _ACCESS_KEY_ID, _SECRET_ACCESS_KEY, _PATH_STYLE |
unset | The bucket for OK_FILES_STORE=s3. Blobs are sealed before upload; keys are <prefix>files/<shard>/<id> (<prefix>tenants/<name>/files/... under --tenants). ok doctor tests it. |
OK_TENANT_BACKUP_PRIVATE_ENDPOINTS |
off | 1 lets tenants' S3 destinations be on this machine or a private network. |
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY or GOOGLE_API_KEY, OLLAMA_HOST |
unset | Turn on a model provider when there is no llm.yaml. See Models. |
OK_LLM_CONFIG |
./llm.yaml if present |
The model configuration file. |
OK_LLM_CHECK |
on | 0 skips the check of each provider key at start. |
OK_MCP_CONFIG |
./mcp.yaml if present |
Tool servers. Tools on servers it does not list are refused. |
OK_EMBED_MODEL_DIR |
the models/ cache beside the state |
A folder already holding the search-by-meaning model (set in the -embed images); nothing is downloaded. ok embed fetch --dir DIR fills one. |
OK_SQLITE_SYNC |
normal |
full makes every acknowledged write survive a power cut or an OS crash too, at the cost of slower writes. See What survives a crash. |
OK_PASSWORD_DENYLIST |
unset | A file of further refused passwords, one per line. |
OK_SMTP_* |
unset | A mail server for companies with none of their own. See Email. |
OK_CONNECTION_HOSTS, OK_CONNECTION_PROXY |
unset | Hosts connections may reach, and a proxy for their traffic. |
OK_ARCHIVE_AFTER_HOURS |
168 |
Finished tasks and runs untouched this long leave memory (they stay in the file). off never archives. |
OK_AUDIT_MIN_FREE_MB |
1024 |
Free disk space, in MB, under which admins are warned. |
RUST_LOG |
info |
How much the server logs to standard output. ok::timing=debug adds where each open, tick and search spent its time. |
--dev-auth |
off | Trusts an x-ok-actor header instead of tokens. A flag only, never read from the environment. Never on a real server. |
OK_DEV_ECHO_TOOLS, OK_DEV_AUTH |
off | Development switches. Never on a real server. |
--no-lock |
off | Break glass: write without the state lock where the file system refuses locks. See Break glass. |
The install script reads OK_INSTALL_DIR, OK_DATA_DIR, OK_VERSION,
OK_REQUIRE_SIGNATURE, OK_INSTALL_FILE, OK_INSTALL_URL and
OK_INSTALL_SHA256 (see Install). The image sets
OK_STATE=/data/state.db and OK_BIND=0.0.0.0:8080, and its health check
uses OK_HEALTH_PORT (8080; 80 with deploy/compose.https.yml). Tuning
settings (OK_DB_READERS, OK_SEARCH_CANDIDATES, OK_ARCHIVE_TICK_MS, the
audit checkpoint settings) are in docs/deployment.md in the repository.
Keep provider keys, webhook secrets, OK_SECRETS_KEY, bucket secrets and
tokens in a secret store (a Kubernetes Secret, Docker secrets, a systemd
EnvironmentFile of mode 0600, Vault). Never commit llm.yaml, .env,
secrets.key or a recovery key.
ok doctor#
ok doctor checks the install in one run and never writes the state; it is
safe beside a running server:
| Section | Says |
|---|---|
| Build | The SQLite note, when the program was built without OrchKernel's settings |
| State | The file, its size, storage version and journal mode, OK_SQLITE_SYNC, and a network file system |
| Lock | Free, or held by the server (its pid and since when) or by a command |
| Files | "private", or the files others can read; --fix-permissions fixes them |
| Secrets | Where the key and the webhook secret came from and the key's id; what the state holds sealed, by key id; how many secrets do not open |
| HTTPS | Each stored certificate with the days left (a staging one warned about), a configured domain without one, your own certificate files; or "not served by OrchKernel" |
| Models | The server's and the workspace's models, each provider checked (no tokens spent), the default model, and whether agents can run |
| Backups | The schedule and next run, each destination's last success and failure, the recovery key, and warnings (no recovery key, on this server only, an operator destination without recipients) |
Each line is ok, warning or error, and the exit status is 1 when any is
an error.
The command line#
The commands an operator uses most; CLI and API reference has them all.
| Command | Does | While the server runs |
|---|---|---|
ok serve [--bind A] [--domain H] [--tls-cert F --tls-key F] |
Runs the API and the workspace | |
ok init --admin ada --email ... |
Creates the company and the first admin's setup link | refused |
ok doctor [--fix-permissions] |
Checks the install | works |
ok backups list [--json] |
Backups in every destination, with their checks | works |
ok backups run |
Backs up now | refused: use Back up now or POST /api/backups/run |
ok backups verify <id> [--destination D] [--identity F] |
Opens a backup and checks it | works |
ok backups prune |
Removes what retention no longer keeps, never the last verified backup | refused |
ok backups recovery-key [--out F] |
Creates a recovery key; its secret printed once, or written 0600 | refused: use the Backups page |
ok backups decrypt <file> --identity F --out DIR |
Opens a backup into a folder without restoring it | works |
ok restore <file or id> [--identity F] [--yes] [--confirm NAME] [--replace-config] |
Restores, with the server stopped (below) | refused |
ok backup <path> |
A plain, unencrypted copy of the state file alone, for scripts | works |
ok embed fetch --dir DIR |
Downloads the search-by-meaning model and checks it, for offline use | works |
ok secrets rotate [--new-key-env V] |
Re-seals every stored secret under a new key | refused |
ok audit verify |
Checks the audit log's chain | works |
People, tokens and access#
People sign in to the workspace with their email address and a password
(see Sign-in settings). Tokens
are for the API, the CLI's --server mode, outside agents and CI, never for
signing in to the workspace.
A token belongs to one person or agent. The server keeps only its hash, so a lost token cannot be shown again: issue a new one.
| You want to | In the workspace | With the API | With the CLI |
|---|---|---|---|
| A token for yourself | Your profile, API tokens, New token (label, lifetime, scopes) | POST /api/tokens {"label":"laptop","ttl":"7d"}, signed in to the workspace |
ok --as maya token issue --for maya --label laptop, server stopped |
| A kernel agent's token | POST /api/tokens {"actor":"policy-ci","label":"ci","scopes":["policy"]}, signed in as an admin |
ok --as ada token issue --for policy-ci --label ci --scope policy, server stopped |
|
| An outside agent's token | Directory, the agent, Tokens | POST /api/gateway/agents/<id>/tokens |
ok agent register |
| See tokens | Your profile (yours); the people panel (an admin, anyone's) | GET /api/tokens |
ok --as ada token list |
| Revoke one | Revoke on its row | DELETE /api/tokens/id/<id> |
ok --as ada token revoke --id <id> |
- A token never issues a token, and never changes how anyone signs in:
POST /api/tokens, passwords, two-step, email and sessions answer 403session_required"Sign in to the workspace to do this; a token cannot." to a bearer. - Nobody issues a person's token for them. An admin asking for someone
else's gets 403
own_tokens_only; each person issues their own. - Lifetimes. A person's token lives at most 90 days (an admin changes the
ceiling in Governance, Sign-in); 30 days unless chosen. A kernel
agent's lives at most 90 days and must carry the
policyorchangesscope. - Old sign-in tokens. Tokens people pasted into the old sign-in page keep working on the API until they expire; the people panel counts them and revokes them all with one action.
Send a token as Authorization: Bearer <token>. The one exception is the
event stream, GET /api/events/stream?token=..., because a browser cannot
set headers there (the workspace itself uses its cookie).
Someone leaves. An admin offboards them from the people panel (or POST /api/actors/<id>/offboard): their sessions end, their tokens are deleted,
their running work and delegations are cancelled. Deactivating instead is
reversible: sign-in is refused, sessions end, and tokens stop working until
the person is reactivated.
Webhooks#
POST /api/webhooks/<source> turns an event from another system into tasks,
through the triggers (rules in a module that start work) that listen to
that source. The support module's helpdesk source, for example, creates a
"Triage incoming ticket" task for its triage agent. A source with no secret
is refused. The sender signs each delivery:
x-ok-timestamp: the time of sending in Unix seconds, within five minutes of the server's clock.x-ok-signature:sha256=and the hex HMAC-SHA256, keyed with the source's secret, of the timestamp, a dot and the body exactly as sent.x-ok-delivery(optional): a unique id. A repeat within 24 hours creates nothing and answers with the first delivery's tasks, so the sender can retry safely.
With the server started with OK_WEBHOOK_SECRETS=helpdesk=... and the same
secret in HELPDESK_SECRET:
body='{"external_id":"Z-100","subject":"Charged twice"}'
ts=$(date +%s)
sig=$(printf '%s.%s' "$ts" "$body" | openssl dgst -sha256 -hmac "$HELPDESK_SECRET" | sed 's/^.* //')
curl -sS -X POST https://work.acme.example/api/webhooks/helpdesk \
-H 'content-type: application/json' \
-H "x-ok-timestamp: $ts" -H "x-ok-signature: sha256=$sig" \
-H 'x-ok-delivery: evt-0001' --data-raw "$body"
| What the sender sees | When |
|---|---|
202 {"delivery":"evt-0001","replayed":false,"tasks":["..."]} |
Accepted. Agents run after the answer, so a slow model never holds the sender. |
202 with "replayed":true and the same tasks |
The same x-ok-delivery again. |
| 401 "bad webhook signature" | Wrong secret or a changed body. |
| 401 "x-ok-timestamp is more than five minutes from the server clock" | A stale or future timestamp. |
| 401 "no webhook secret is configured for this source" | Neither OK_WEBHOOK_SECRETS nor OK_WEBHOOK_SECRET covers it. |
400 bad_request |
The body is not JSON. |
Senders that cannot sign may send the secret itself in x-ok-webhook-secret.
It is accepted (202, with "delivery":null) but offers no protection
against a replayed request; prefer signatures.
Backups and restore#
Setting backups up (the recovery key, an off-site bucket, checking Verified, testing a restore) is in Production hardening. This section is how they work.
What a backup holds#
| Member | What |
|---|---|
manifest.json |
The backup's id, kind, time, OrchKernel and storage versions, company, the audit log's id and head, the key ids, and each file's size and SHA-256; signed (an HMAC) with a key derived from the secrets key |
state.db |
The state, copied with SQLite's online backup while the server runs. Model, mail, sign-in and backup settings are in it |
secrets.key |
The secrets key and webhook secret in use, also when they came from the environment |
secrets.key.<kid> |
Keys kept after a rotation, so older sealed values still open |
config/llm.yaml, config/mcp.yaml |
The configuration files the server uses, when there are any |
packs/<name>/ |
Pack folders beside the state, when there are any |
files/<shard>/<id>[.thumb|.view|.txt] |
Every live file's blobs, as sealed on disk (a removed file's are not). With OK_FILES_STORE=s3 the bucket is the store: the manifest lists its objects instead, and a restore says which are missing. Turn on the bucket's versioning |
A file removed after a backup stays in that backup until retention
prunes it. A restore puts back the state and files/ together, moving
the folder as it was aside as pre-restore-<time>.files/ (removed with
the pre-restore-<time>.db copy).
Left out, and rebuilt or fetched at start: the search index, the local
model, certificates, the lock, -wal and -shm.
The format and the keys#
A backup is orchkernel-<id>.tar.gz.age, the id being its UTC time and six
random hex digits (2026-10-07T020000Z-a1b2c3): a tar archive, gzipped,
encrypted in the age format to X25519
keys. Beside it, orchkernel-<id>.json lists it without opening it (kind,
time, versions, size, checksum, key fingerprints, when and how it was
verified); it names no company and holds nothing secret, and a restore never
trusts it.
Every backup is encrypted to:
| Key | Opens it | Kept |
|---|---|---|
| This server's key | Derived from the secrets key; a restore on this server needs nothing else | Nowhere: derived when needed |
| The recovery key | Anyone holding the recovery key file | Only by you, offline. The server keeps its public half and fingerprint |
OK_BACKUP_RECIPIENTS |
Each operator key listed | By the operator |
A backup made before a key rotation opens with the kept secrets.key.<kid>
or the recovery key. A new recovery key applies to backups from then on.
Opening a backup without OrchKernel, with the age tool and the recovery key file:
age -d -i orchkernel-recovery-key-9f3e1a2b.txt orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age | tar xz
or ok backups decrypt <file> --identity <key file> --out <folder>. Either
gives the files above; state.db opens with any SQLite tool.
How a run goes#
- When: at the scheduled time, or five minutes after a start when the last good backup is older than one interval. A run whose state has not changed since the last good one records "unchanged" and writes nothing.
- Make: copy the state, pack, compress and encrypt into
backups/.partial/. - Verify here: open it with the server's key, check every file's
checksum and the manifest's signature,
PRAGMA quick_checkand the audit chain of the state inside. Only a verified backup is kept. - Upload to the off-site bucket (one upload up to 256 MiB, in 64 MiB parts above), then verify there: download and compare (the default). A copy that differs is removed from the bucket and the run fails.
- Record:
BackupTakenandBackupVerifiedevents, and the run in the backup log the page shows. - Prune each destination by retention.
- Report: on a failure, an inbox item for every admin (once per failing
streak per destination) and a
BackupFailedevent.
One run at a time: Back up now during a run answers 409 busy.
Retention#
Per destination, over verified backups only: the newest backup of each of the last 7 days that have one, of each of the last 4 ISO weeks and of each of the last 6 months (the union is kept). Always kept as well: the newest verified backup, whatever its age, and backups made before a restore for 7 days. Unverified or failed files older than a day are removed only when a newer verified backup exists there. Nothing is pruned after a failed run, and the last verified backup is never removed, from the page either.
Restore from the Backups page#
For a server that runs. On the backup's row, choose Restore:
- Fetching and checking. The backup is fetched (downloaded from the bucket if needed) and checked as in a run; nothing changes yet. A backup this server's keys cannot open asks for the recovery key it was made with (paste the key file's text).
- Where it came from. "Made by this workspace" (its signature checks with a key this server holds), "An earlier state of this workspace" (its audit log is an earlier point of this one), or "Not checked". A backup that is not checked is restored only from the command line.
- What will change: the time it goes back to, how many changes since
are undone ("not counted" when the backup is not an earlier point of the
workspace as it is now, such as a copy made before an earlier restore),
people, records, tasks and threads now and in the backup, tokens issued
since (revoked), the secrets key (kept; sealed values re-sealed to it),
and configuration files that differ (kept; the backup's written beside
them as
llm.yaml.from-backup). A red line says what is lost: "Everything done after 7 Oct 2026, 02:00 is lost (42 changes). Everyone is signed out." - Type the company's name and choose Restore. A sign-in within the last ten minutes is needed; otherwise the page asks for your password.
Then the server makes a backup of the state as it is now (kind "Before a
restore", kept 7 days), answers every request but the health check with 503
restoring, stops running agents, and restarts itself with the same
command line (HTTPS included), never letting go of its lock. The new process
swaps the backup in before it serves, and its log says "restored from the
backup of 2026-10-07 02:00 UTC; everyone signs in again". The page then
sends you to sign in. Passwords and two-step stay as they are now (nobody
gets back a password they changed); sessions and sign-in links end.
The server's backup setup is not rolled back: the schedule, the off-site destination, the recovery key and the record of past backups stay as they were just before the restore.
If the server does not come back (the restart failed, the process exits
with status 3), the page says so after two minutes. Start the server again
as you normally do (ok serve, or your service): it finishes the restore
before it serves. A crash at any step leaves either the state as it was or
the restored one, never a mix; the plain copy of the state as it was,
pre-restore-<time>.db beside it, is removed after 7 days.
Undoing a restore. Restore the backup of kind Before a restore the same way; the sheet says "This undoes the restore of ...".
ok restore, with the server stopped#
For a server that will not start, a backup that is not checked, or a new machine:
ok --state /var/lib/orchkernel/state.db restore 2026-10-07T020000Z-a1b2c3
ok --state /var/lib/orchkernel/state.db restore /path/to/orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age
A backup id is looked for in the local folder, then in the off-site bucket.
It runs the same checks, prints the same summary, asks for the company's
name (--confirm NAME gives it; --yes skips the question for a checked
backup), and does the same swap. It is refused with exit status 75 while the
server runs. A backup that is not checked needs --yes and the company's
name, and says why first.
On a new machine (the server and its key are lost): put the backup file and the recovery key file on it, and restore into an empty folder, giving the full state path:
mkdir -m 700 /var/lib/orchkernel
ok --state /var/lib/orchkernel/state.db restore orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age \
--identity orchkernel-recovery-key-9f3e1a2b.txt
It writes state.db, secrets.key (and kept keys) from the backup, and
llm.yaml, mcp.yaml and pack folders where none exist (--replace-config
replaces existing ones), then prints the exact command to start the server
on it:
restored 2026-10-07T020000Z-a1b2c3 into /var/lib/orchkernel/state.db; start the server on it with:
orchkernel --state /var/lib/orchkernel/state.db serve
Use that command (or set OK_STATE in your service): a plain ok serve
elsewhere uses ./.orchkernel/state.db and starts an empty workspace.
People sign in with the passwords they had at the backup.
Versions#
| The backup's storage version | Result |
|---|---|
| Newer than this program | Refused before anything changes, naming the version to install |
| The same | Restored |
| Older | Restored and converted at the first open; the backup is the way back |
A plain copy for scripts: ok backup#
ok backup /backups/state-$(date +%F).db copies the state file alone with
SQLite's online backup API, 0600, safe while the server runs. It is not
encrypted and holds no key or configuration; restoring it means stopping the
server, moving state.db (and any -wal, -shm) aside and copying the
file in, with the secrets key it was sealed under. Prefer the encrypted
backups above.
Tenants#
Under --tenants, each tenant's admins have their own Backups page: off
until they turn it on, kept under the operator's root, encrypted to the
tenant's key, its recovery key and OK_BACKUP_RECIPIENTS. A restore from it
goes through the tenancy restore, without restarting the server. The
operator's nightly copies keep running beside them. See
Hosting several organisations.
Upgrades and rollback#
- Back up (Back up now, or
ok backups runwith the server stopped), and stop every scheduled job that opens the state file with the CLI. - Replace the program (run the install script again) or the image tag.
- Start the server and check
/api/health.
Some releases change how the file is stored, and the first open by the new program converts it:
- It checks there is room first. Without it, the open fails with "upgrading the state file needs N bytes free beside it, and M are" and nothing changes.
- It writes a copy beside the file,
state.db.pre-upgrade-v<n>.db(0600), and keeps it. Delete it yourself once the new release has run well. - It converts in one transaction, so a stop half-way leaves the old file untouched.
An older program refuses a converted file: "the state file is newer than this binary". Rolling back after a storage change means restoring the pre-upgrade copy or a backup, and losing what changed since the upgrade. Releases without a storage change can be rolled back by swapping the program again.
If an old server kept writing after the conversion, the next open may refuse with "the state file was written by two versions of orchkernel". Stop everything and restore the pre-upgrade copy or a later backup.
The first open by a release that adds record indexes builds them before the server answers: a few seconds per index on a large collection, once.
Running commands while the server runs#
The server holds a lock on state.db.lock, beside the state file, from
before it opens the file until it exits (and across its own restart after a
restore). The operating system lets go of it when the process ends, however
it ends, so there is no stale lock to clean up on a local disk.
-
Commands that only read open a snapshot of the file and never write it, so they work beside the server and see what it has saved:
ok actor list,ok runs,ok tasks,ok inbox,ok log,ok replay,ok threads,ok search,ok brain query,ok module list,ok token list,ok user list,ok mail show,ok connection list,ok policy export,ok backups listand the other listings. -
Commands that change the state (
ok init,ok actor add-human,ok token issue,ok pack install,ok kill,ok tick,ok ask,ok secrets rotate,ok backups run,ok restore, every other one) are refused while the server runs, with exit status 75 and nothing written:The server is running on /data/state.db (pid 4121, since 09:30 UTC). Use the workspace or the API (add --server https://work.acme.example), or stop the server, then run this again.Two such commands at once are refused the same way, and so is a second server.
-
ok backup,ok audit,ok doctor,ok backups decryptandok embed fetchload no kernel of their own and work beside the server. -
With
--server, the people, mail and policy commands go to the running server's API instead (see CLI reference). -
Under tenancy the same holds per tenant.
In Docker, docker exec orchkernel orchkernel --as ada actor list works
beside the server; a command that changes the state needs the container
stopped and a one-off container on the same volume.
The state on a network disk (NFS, SMB, AFP, WebDAV, FUSE) is not supported,
as SQLite's write-ahead log is not: the start and ok doctor say so.
Break glass#
Where the file system refuses locks altogether (ENOLCK, EOPNOTSUPP), a
command that changes the state is refused with "the file system holding
/data/state.db does not support locks". If you are certain no server runs on
the file, --no-lock lets it through, with "warning: writing without the
lock; make sure no server runs on this file". --no-lock is refused while a
live process holds the lock: there is no override for a running server,
because a write beside it is exactly what the lock prevents.
What survives a crash#
| What happened | What you find after a restart |
|---|---|
| The process died (SIGKILL, out of memory, a panic) | Every write the server answered for. The file is never damaged. Runs that were mid-step end as stopped, with the reason "restart", their steps so far kept. Runs waiting for an approval or an answer resume when it comes. Sessions, links and passwords are kept |
| The machine lost power or the OS crashed | The file is never damaged. With the default (synchronous = NORMAL in WAL), the writes of the last moments before the cut may be gone even though they were answered. Set OK_SQLITE_SYNC=full to keep those too, at the cost of slower writes |
| The disk filled | Writes fail with 500 and nothing is half written; reads keep working. A backup run fails, removes its partial file and prunes nothing |
| A crash during an upgrade | <state>.pre-upgrade-v<n>.db is the file before it |
| A crash during a backup | The older backups are untouched; the partial file is removed by the next run |
| A crash during a restore | The next start puts the state as it was back, or finishes the restore if it had recorded it: never a mix |
| A crash during key rotation | Settled at the next start: every secret opens with the key on disk |
| The vector index | Replayed from its change log, or rebuilt in the background |
| Model calls in flight | Not retried; their run is stopped with "restart" |
| Email being sent | The link stays queued; the people panel offers Resend after 15 minutes |
| Outside agents' actions being executed | Marked interrupted by the start-up sweep |
The tests behind this table kill the server with SIGKILL in the middle of 16
people's writes, in the middle of a run, while a run waits for an approval,
right after a password change, in the middle of a backup of a 200 MB file,
at both steps of a key rotation and at three points of a restore, then check
with a fresh process that PRAGMA integrity_check answers ok, ok audit verify passes, every answered write is there and a new server starts
(crates/ok-cli/tests/it/crash.rs, restore_crash.rs).
Watching it run#
| What | How |
|---|---|
| Is it up | GET /api/health, no token. Use it for load balancer and container probes. |
| Server log | Standard output and standard error, level set by RUST_LOG. With --service: journalctl -u orchkernel (Linux) or orchkernel.log in the data folder (macOS). |
| What happened | Events in the sidebar, GET /api/events, or the live GET /api/events/stream. Someone who is not an admin sees only events about their own threads, tasks and runs. |
| Backups | The Backups page, GET /api/backups/status, ok backups list, ok doctor |
| One run, step by step | ok replay --run <id> (or --thread, --task) rebuilds it from the log alone. |
| The audit log is intact | ok audit verify and ok audit head, safe beside a running server. See Events and the audit log. |
The server advances its own clock every 30 seconds: scheduled triggers,
overdue work, stalled threads, archiving. This is the tick. To drive it
from your own scheduler, call POST /api/tick with an admin token
(deploy/k8s/tick-cronjob.yaml does this every minute).
When the audit log cannot record a decision (a full disk, say), the server
refuses the action with 503 audit_unavailable rather than act unrecorded.
Troubleshooting#
| You see | What to do |
|---|---|
| The first-admin link is lost or expired | Restart the server: while the company has no admin with a password, each start prints a new one. With an admin already, ok user sign-in-link ada (server stopped). |
| "password sign-in is off" at start, or on the sign-in page | People must open an https:// address: set OK_PUBLIC_URL to it (behind a proxy), or use --domain. On your own computer, open http://localhost:8080 or http://127.0.0.1:8080. |
The link names 127.0.0.1, which your browser cannot open |
The server runs elsewhere and listens only on itself. Use --domain, a proxy, or a tunnel: ssh -L 8080:127.0.0.1:8080 <server>. |
| "cannot listen on 0.0.0.0:443 ... Ports below 1024 need a privilege" | Run with --service (it grants exactly that), sudo setcap cap_net_bind_service=+ep "$(command -v orchkernel)", or higher ports with a forward. |
| "cannot listen on ...: another program uses it" | Another web server holds 80 or 443 (sudo ss -ltnp 'sport = :443'). Stop it, or put OrchKernel behind it as a proxy. |
| "no certificate for work.acme.example yet (...)"; browsers warn about the certificate | Let's Encrypt could not reach port 80 for that name: check the DNS record answers the server's address, port 80 is open to the internet, and nothing else answers it. The server retries on its own; ok doctor shows the state. Use --acme-directory staging while you fix it, to stay inside Let's Encrypt's limits. |
| The certificate is "not trusted" | Still on staging: remove --acme-directory staging (or OK_ACME_DIRECTORY) and restart; a certificate from another ACME server is replaced at start. |
| "Backups to offsite are failing: uploading ...: the provider answered 403 (AccessDenied ...)" | The key lacks a permission on the bucket or folder: it needs to put, get, delete and list. Test on the Backups page names what failed. |
| "... could not be reached" | The endpoint is wrong or blocked: check it from the server with curl -I <endpoint>. A bucket on this machine or a private network needs http:// to be allowed there; under --tenants the operator allows it with OK_TENANT_BACKUP_PRIVATE_ENDPOINTS=1. |
| "its secret access key is not set" or a locked secret | The bucket secret is sealed under a key the server no longer has: enter the secret again on the Backups page. |
| "No recovery key yet" | Create one on the Backups page and keep it offline. Backups made before it do not open with it. |
| "N secrets sealed under key ... and no key was found" | The secrets key is missing or different: put secrets.key back, or set OK_SECRETS_KEY to the key it names. A backup carries its key: restoring the latest one also works. |
| A restore says "made by OrchKernel 0.4 at storage version 4; install 0.4 or later" | The backup is from a newer release: install that release, then restore. |
| "this server cannot tell where it came from" | The backup is not this workspace's, or was made under a key this server no longer holds and its history does not match. Restore it only if you know where it came from, with ok restore ... --yes and the company's name, server stopped. |
| The Restore page says the server has not come back | Look at the server's log. Start it again as you normally do; the next start finishes or undoes the restore. |
| A command says "The server is running on ..." (exit 75) | It changes the state: use the workspace or the API, or stop the server first. |
| "the state file is newer than this binary" | You went back to an older release after an upgrade converted the file: install the newer one, or restore a backup made before the upgrade. |
Not possible yet#
- Publishing: no release, image or install script has been published yet; build them yourself (Install).
- More than one server or replica for a company, or failover. There is no Postgres backend.
- Single sign-on (OIDC, SAML) and passkeys.
- Certificates by DNS or TLS-ALPN challenge, and wildcard certificates from
Let's Encrypt (give your own with
--tls-cert). - More than one off-site destination, or destinations other than S3-compatible storage (SFTP, for example).
- Restoring to a point between two backups.
- Changing state with the CLI while the server runs.
- Keeping the state on a network disk.
- Exporting metrics or traces to OpenTelemetry; read standard output and the event log.
- Apple-notarized macOS programs (the install script does not quarantine them); Windows.
License#
OrchKernel is source-available under the Business Source License 1.1. You
may use, modify and run it in production for your own organization, or run a
deployment for one client organization; separate deployments for separate
clients are fine. Offering it as a hosted or managed service in which one
deployment serves more than one organization needs a commercial license.
Each version converts to Apache-2.0 four years after it is first published,
and the client libraries under sdk/ are Apache-2.0 already. The LICENSE
file in the repository is the authoritative text.
For developers#
ok serveand the tenancy host start every kernel through one path,crates/ok-server/src/startup.rs(prepare_kernel): a pending restore is finished first (ok_server::restart::finish_pending_restore), then the secrets key check, the context hooks, the process configuration, the module reconcile, the gateway's observer, the audit verifier, the gateway start-up sweep, the statistics refresh, the local indexer worker, the backup job (backup_jobs.rs), then the request state. Each 30-second tick runstick_one.- HTTPS is
crates/ok-tls(the ACME manager oninstant-acme, the certificate store, the SNI resolver, the port 80 router) served bycrates/ok-server/src/https.rs; the flags arecrates/ok-cli/src/serve_tls.rs. - Backups are
crates/ok-backup(bundle, age encryption, the local and S3 destinations, retention, the run, restore); the routes arecrates/ok-server/src/backups_api.rs, the settings and log rowscrates/ok-kernel/src/backups.rs. A restore's swap is shared with the tenancy restore (crates/ok-tenancy/src/swap.rs); the restart hands the lock over acrossexec(crates/ok-server/src/restart.rs,ok_db::lock). ok backuploads no kernel: it copies through SQLite's online backup API (Kernel::backup_file) into a.partialfile made 0600, renamed into place once complete.- The state lock is
ok_db::lock(flock);crates/ok-cli/src/access.rsclasses every command (no kernel, snapshot, write, serve). - The key file is
ok_secrets::keyfile; its resolution, generation and rotation arecrates/ok-cli/src/secrets_file.rs. Private files areok_db::private. deploy/install.shis tested bycrates/ok-cli/tests/it/install_script.rs; the release workflow is.github/workflows/release.yml.docs/deployment.mdis the full operator reference: the Kubernetes bootstrap pod, MCP servers, connection key rotation, archiving and indexes.