Production hardening

Last updated October 6, 2026

On this page

Draft for review

Install in five minutes gets a server running with its first admin. This page is what to do before people rely on it: a domain with HTTPS, email, off-site backups with a recovery key and a tested restore, private files, the secrets key, sign-in settings and updates. It ends with the security and production checklists. Every setting and command is in the Operations reference.

Three facts shape everything here:

  • One server per state file. Never run two servers on the same file, and never two copies of a container or pod that mount it. The server holds a lock beside the file (state.db.lock) to make sure.
  • The server is the only writer while it runs. Commands that would change the state are refused while it runs (exit status 75); use the workspace, the API, or stop it first. Commands that only read work beside it. See Running commands while the server runs.
  • People sign in with a password; scripts use tokens. A person's first password comes from a one-time link. A token (orchk_...) is for the API, the CLI's --server mode, outside agents and CI.

Domain and HTTPS#

People must reach the server at an https:// address: the session cookie is Secure, and password sign-in is off on any other address except this machine's own (localhost, 127.0.0.1). There are three ways.

OrchKernel gets the certificate: --domain#

  1. Point the name at the server. Create a DNS A (and AAAA, if the server has IPv6) record for work.acme.example with the server's public address, and wait until dig +short work.acme.example answers it.

  2. Open ports 80 and 443 to the internet in the server's firewall and your cloud's security group. Port 80 answers Let's Encrypt's check and redirects everything else to HTTPS; nothing else is served on it.

  3. Try it on Let's Encrypt's staging server first. Its certificates are not trusted by browsers, but its limits are generous, so a mistake in steps 1 and 2 costs nothing:

    sh
    ok serve --domain work.acme.example --acme-directory staging
    

    The start says "By using --domain you accept the Let's Encrypt Subscriber Agreement (...)" once, then obtained a certificate for work.acme.example. ok doctor shows it under HTTPS, warned as a staging certificate.

  4. Then the real one. Stop it and start without --acme-directory (with the install script: --service --domain work.acme.example). Add --acme-email you@acme.example to hear from Let's Encrypt if renewal ever fails for long.

What happens then:

  • HTTPS on 443 with HSTS, port 80 redirecting to it (308) for your domain only; any other name gets 421, never a redirect.
  • OK_PUBLIC_URL is set to https://work.acme.example for you. One set to another host is refused at start.
  • Renewal is checked every 12 hours, 30 days before expiry. A failed one is retried after an hour, doubling to a day; with 14 days or fewer left, every admin gets an inbox item. No restart is needed.
  • Certificates are kept in tls/ beside the state (0700). They are not in backups: a restored server gets new ones.
  • A failed order at start does not stop the server: it serves a placeholder certificate, keeps trying, and says why in its log.
  • Several names: repeat --domain (or OK_DOMAIN=a.example,b.example); the first is the address people use.
  • Ports below 1024 need a privilege. The install script's --service gives the service exactly that (CAP_NET_BIND_SERVICE); otherwise run sudo setcap cap_net_bind_service=+ep "$(command -v orchkernel)", or use --https-bind 0.0.0.0:8443 --http-bind 0.0.0.0:8080 and forward 443 and 80 to them. Docker allows them inside the container.

With Docker Compose, deploy/compose.https.yml does all of this; see Install.

Your own certificate#

sh
ok serve --tls-cert /etc/orchkernel/fullchain.pem --tls-key /etc/orchkernel/key.pem

PEM files, the chain with the server's certificate first. The address people use is the certificate's first name (or OK_PUBLIC_URL, which must be a name it covers). The files are read again when they change (checked every minute) or on SIGHUP; a file that does not load keeps the old certificate and says why in the log. --domain and --tls-cert are two ways to get a certificate: give one.

Behind a proxy you already run#

Caddy, nginx, Traefik or a cloud load balancer ends TLS and passes plain HTTP to OrchKernel:

sh
OK_PUBLIC_URL=https://work.acme.example ok serve --bind 127.0.0.1:8080 --trusted-proxy 127.0.0.1
  • OK_PUBLIC_URL is the https address people open. It decides the session cookie (__Host-ok_session, Secure) and is the only origin a signed-in write is accepted from. The server never builds an address from the Host header.
  • --trusted-proxy (OK_TRUSTED_PROXIES) names the proxy, so the server reads the client's address from X-Forwarded-For for its sign-in limits.
  • Keep event streams unbuffered (/api/events/stream is server-sent events): in nginx, proxy_buffering off for that path.

docker-compose.yml has a commented Caddy example.

The public address#

Without TLS flags, OK_PUBLIC_URL decides everything about the address. A server bound to loopback with none uses http://<bind> (the cookie is then ok_session, without Secure: how ok demo serve and a laptop work). A server bound to anything else with no OK_PUBLIC_URL, or with an http:// address that is not loopback, turns password sign-in off: the sign-in page says so, the start prints "set OK_PUBLIC_URL to the https address people use", and tokens keep working. /api/health shows "password_sign_in":true once it is right.

Several organizations on one server, each on its own host name, get their certificates with --tenants --acme: see Hosting several organisations.

Email: Resend or another SMTP server#

Invites, password resets and the "your password changed" notes go by email once it is set up; until then every link is shown once to the admin who made it, to send themselves, and Forgot password? is not offered. An admin sets it in Governance, Sign-in, section Email:

  • Resend: the host (smtp.resend.com), username (resend) and TLS are filled in; give the API key as the password, the port (465, or 587) and a From address on a domain you verified at Resend. Leave Resend's click tracking off: it would rewrite reset links.
  • Other SMTP server: host, port, security (TLS, or STARTTLS; plain SMTP only to a loopback host in development), username, password and the From address.

Send test email says where it stopped if it fails ("Resend refused the sender: the domain of noreply@acme.example is not verified (550)"). The password is sealed with the secrets key and never shown again. A server can instead give every company the same mail server through OK_SMTP_HOST, OK_SMTP_PORT, OK_SMTP_SECURITY, OK_SMTP_USERNAME, OK_SMTP_PASSWORD, OK_SMTP_FROM and OK_SMTP_FROM_NAME, used when the company has none of its own. From the command line (server stopped): RESEND_KEY=re_... ok --as ada mail set --preset resend --from noreply@acme.example --password-env RESEND_KEY, then ok --as ada mail test. See Setting up a company.

Backups#

Backups run from the first start: every day at 02:00 UTC, encrypted, into backups/ beside the state. Each one holds everything a restore needs: the state, the secrets key, llm.yaml and mcp.yaml, and pack folders beside the state. Two things are left to you, and the Backups page (sidebar, Company, Backups; admins only) walks you through both: a recovery key, and a copy off this server. Its status line says what is missing:

text
Last backup  today 02:00    Next  tomorrow 02:00    This server
No recovery key yet. Backups open only with this server's secrets key ...
Backups are on this server only: add an off-site destination.

1. Create the recovery key and keep it offline#

Every backup is encrypted to two keys:

  • This server's key, derived from secrets.key. It needs nothing saved and makes a one-click restore on this server possible, but it is lost with the server.
  • Your recovery key, made once and kept by you. It opens every backup made after it was created, even when the server and its secrets key are gone.

Choose Create recovery key. It needs a sign-in within the last ten minutes; otherwise the page asks for your password first. The key is shown once: choose Download orchkernel-recovery-key-.txt (or Copy), tick I saved it somewhere safe, off this server, and Done. The page then shows its fingerprint, and each backup in the list records the fingerprint it was made with.

Keep the file like a root password: offline, in a password manager or a safe, never on the server or next to the backups. Whoever holds it can read every backup made since. If you lose both the server's key and the recovery key, no backup can be opened, by anyone. Make a new recovery key later applies to backups from then on; older ones still open with the key they were made with.

From the command line (server stopped): ok backups recovery-key --out /safe/place/recovery.txt.

2. Add an off-site destination#

Under Where backups go, Off-site, pick the provider. Any S3-compatible storage works; the tabs fill in what they can:

Provider Endpoint Region The key
Cloudflare R2 https://<account id>.r2.cloudflarestorage.com (the account id is on the R2 overview page) auto An R2 API token with Object Read & Write on the bucket: its access key id and secret
Backblaze B2 https://s3.<region>.backblazeb2.com, shown on the bucket's page the bucket's, such as us-west-004 An application key for the bucket: its key id and application key
AWS S3 https://s3.<region>.amazonaws.com the bucket's An IAM user's access key, allowed s3:PutObject, GetObject, DeleteObject and ListBucket on the bucket
MinIO http://minio:9000 (plain http only on this machine or a private network) us-east-1 A MinIO access key

Create the bucket first, private, with versioning or object lock if your provider offers it (so nobody holding the key can delete history). Fill Bucket, Folder in the bucket (orchkernel/), the key id and secret, and a Name (how backups and alerts name it). Choose Test: it writes, reads back and removes a small object, and says what failed if anything does. Then Save. An off-site destination cannot be saved before a recovery key exists, so a backup off this server always opens without it.

The secret is sealed with the secrets key and never shown again; entering a new one needs a recent sign-in. The endpoint must be https, except on this machine or a private network, and link-local and cloud metadata addresses are refused.

3. Back up now and check "Verified"#

Choose Back up now. The list gains a row per destination; each must say Verified under Checked:

  • on this server, the backup was opened again with the server's key and every file, its checksum and the state's own integrity checked;
  • off-site, it was uploaded, then downloaded again and compared.

A backup that fails a check is not kept. When a scheduled one fails, every admin gets one inbox item per destination ("Backups to offsite are failing: ...", with the destination's name), updated rather than repeated, and another when it works again; with no good backup for two scheduled runs, "No backup since ...".

4. Schedule and retention#

Schedule and retention sets Every day at a time (shown in UTC and in your own time), Every hour, or Only when asked, and how many to keep: the newest backup of each of the last 7 days, 4 weeks and 6 months. The newest good backup is never removed, and copies made just before a restore are kept 7 days. A server's operator may fix any of these with OK_BACKUP_* variables; the page then shows them read-only.

5. Test a restore, on another machine#

A backup you have never restored is a hope. Once, and after big changes, restore the latest one on a spare machine or a laptop, with only the backup and the recovery key, as you would after losing the server:

sh
mkdir -m 700 ~/restore-test
ok --state ~/restore-test/state.db restore orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age \
  --identity /safe/place/recovery.txt

It prints what the backup holds, asks for the company's name, writes the state, the secrets key and the configuration into the folder, and prints the exact command to start it:

text
restored 2026-10-07T020000Z-a1b2c3 into /home/ada/restore-test/state.db; start the server on it with:
  orchkernel --state /home/ada/restore-test/state.db serve

Run it with another --bind if 8080 is taken, sign in with your usual password, look at recent work, then stop it and delete the folder. Restoring the live server (from the Backups page or with ok restore) is in Backups and restore.

Private files#

Whatever the umask, the program creates the state folder 0700 (when it makes it), and the state file, its -wal and -shm, state.db.lock, secrets.key, backups, certificates, the pre-upgrade copies, the search index and the model cache readable only by the service user (0600 files, 0700 folders). Files from an older release, or copied in by hand, keep their mode; the server lists them once at start:

text
warning: 3 files can be read by other users on this host:
  /data/state.db (0644), /data/state.db-wal (0644), /data/state.db-vectors (0755)
  Run `ok doctor --fix-permissions`, or chmod 600 the files and 700 the folders.

ok doctor --fix-permissions sets 0600 and 0700 on those the service user owns and names each change. It changes modes only, so it is safe while the server runs. Keep the state on a local disk: NFS, SMB and other network file systems are not supported, and the start says so.

The key file#

ok init, the first ok serve on an empty folder and ok demo create secrets.key beside the state file when no key is found, mode 0600:

text
# OrchKernel secrets for the state files in this directory.
# Created 2026-10-06T09:30:00Z by `ok init`. Keep it private (mode 0600) and
# back it up apart from the state: without it, connection secrets, two-step
# seeds, the mail password and model keys do not open.
OK_SECRETS_KEY=...
OK_WEBHOOK_SECRET=...

Encrypted backups carry it inside (sealed by the backup's own encryption), so a restore needs nothing else. A plain copy of the state made with ok backup does not.

  • The environment wins. OK_SECRETS_KEY or OK_WEBHOOK_SECRET set (and not blank) is used instead of the file's value. To keep the key off the data disk, set OK_SECRETS_KEY from your secret store before the first start. When both hold a key and they differ, the start says so and names the key the stored secrets need.
  • A key is never made for a file sealed under another one. When the state holds sealed secrets and no key is found, nothing is generated: the start says "error: /data/state.db holds 4 secrets sealed under key 1a2b3c4d (2 connections, 1 mail password, 1 two-step seed) and no key was found. Set OK_SECRETS_KEY, or put the key file back at /data/secrets.key." The server still starts, so an admin can sign in and see what is locked.
  • Rotation. ok secrets rotate (server stopped) re-seals every stored secret under a new key in one transaction. With the key file, the new key replaces it and the old one is kept beside it as secrets.key.<key id>; backups made before then open with it, or with the recovery key. With OK_SECRETS_KEY, pass --new-key-env VAR and set OK_SECRETS_KEY to the new key afterwards. A rotation stopped half-way is settled at the next start.
  • Under --tenants each tenant's key is sealed under the operator key instead (Hosting several organisations).

Sign-in settings#

People sign in at the public address with Email and Password. An admin sets the rules in Governance, Sign-in:

  • Passwords are at least 12 characters, not a common password, and not the person's email address or name; they never expire. The server keeps an argon2id hash only. OK_PASSWORD_DENYLIST names a file of further refused passwords, one per line.
  • Two-step sign-in. Each person may turn on an authenticator app (TOTP) on their profile and gets ten recovery codes. Require it for admins at least, or for everyone; people it covers set it up at their next sign-in. The seeds are sealed with the secrets key: if the key is lost, people sign in with a recovery code and an admin resets their two-step.
  • Sessions end 7 days after sign-in, or after 18 hours without use, whichever is first; both can be changed. Each person sees their sessions on their profile and can end any of them. A password change ends every other session.
  • Recent sign-in. Changing a password or two-step, making recovery codes, issuing a token, creating a recovery key, entering an S3 secret and restoring a backup need a sign-in within the last ten minutes; otherwise the workspace asks for the password (and code) again.
  • Failed sign-ins. Ten wrong passwords or codes for one address within 15 minutes lock that address for 15 minutes, whether or not anyone holds it. One client address gets 429 after 60 failed credentials a minute, and after 30 sign-in requests a minute. Every failure reads "Email or password is wrong.", so the page never tells anyone who has an account. Behind a proxy, set --trusted-proxy so the real client is counted.
  • Single sign-on (OIDC, SAML) is not available yet.

Tokens, who may issue them and their lifetimes are in People, tokens and access.

Updates#

  1. Back up: Back up now on the Backups page (or ok backups run with the server stopped), and check it says Verified.
  2. Install the new version:
    • with the install script, run it again: curl -fsSL .../install.sh | sh -s -- --service replaces the program and restarts the service (--version 0.4.0 picks one);
    • with Docker, pull the new tag and recreate the container (docker compose pull && docker compose up -d, or a new docker run with the same volume). Pin a version tag (0.4.0) rather than latest in production, so a restart never upgrades by surprise.
  3. Check /api/health and the start's log, and sign in.

Some releases change how the state is stored and convert it at the first open; the way back is then the backup. See Upgrades and rollback.

Security checklist#

  • HTTPS: --domain, your own certificate, or a proxy ending TLS with OrchKernel not reachable directly. /api/health says "password_sign_in":true and "dev_auth":false.
  • Two-step sign-in required at least for admins (Governance, Sign-in).
  • A recovery key created and kept offline, away from the server and the backups; an off-site destination saved and Verified.
  • secrets.key private (or OK_SECRETS_KEY from a secret store), and ok doctor says every sealed secret opens.
  • The state, its key, certificates and backups readable only by the service user (ok doctor); the state on a local disk.
  • Tokens with lifetimes that fit; a lost device's token revoked at once; a leaver offboarded (POST /api/actors/<id>/offboard), which ends their sessions, stops their tokens, revokes their delegations, cancels the runs acting for them and withdraws their waiting approvals.
  • Each webhook source with its own long secret; CORS off unless needed; --trusted-proxy naming your proxy.
  • Outbound traffic limited to your model providers, tool hosts, mail server and backup bucket; OK_CONNECTION_HOSTS lists the hosts connections may reach.
  • Restricted data only on local or contractually covered models (Models).
  • Kill switches: who engages the global one, and how the company hears of it, written down (Policy, rules and kill switches).
  • No development flags (--dev-auth, OK_DEV_*) anywhere in the service definition.

Production checklist#

Item Done
HTTPS on the address people use (--domain tried on staging first, your own certificate, or a proxy)
The first admin created from the first-admin link; the setup wizard finished
People invited with roles and reporting lines; two-step required for admins
Email set up and Send test email passing
A model connected (the wizard, or llm.yaml reviewed: sensitivities, prices, default model)
mcp.yaml lists every tool server the modules use (no "note: no MCP configuration" at start)
Recovery key created and stored offline
Off-site destination saved; latest backups Verified on both destinations
A restore tested on another machine with only the backup and the recovery key
ok doctor passes: build, state, lock, private files, secrets, models, HTTPS, backups
Run as a service (--service, a container with --restart unless-stopped, or a Recreate Deployment); health probe on /api/health
Logs collected from the server's standard output
Webhook secrets set; senders sign
Kill switch and budget procedure written down and tried
Update runbook: back up, install, check health, sign in