Production hardening
Last updated October 6, 2026
On this page
- Domain and HTTPS
- OrchKernel gets the certificate: --domain
- Your own certificate
- Behind a proxy you already run
- The public address
- Email: Resend or another SMTP server
- Backups
- 1. Create the recovery key and keep it offline
- 2. Add an off-site destination
- 3. Back up now and check "Verified"
- 4. Schedule and retention
- 5. Test a restore, on another machine
- Private files
- The key file
- Sign-in settings
- Updates
- Security checklist
- Production checklist
Draft for review
Install in five minutes gets a server running with its first admin. This page is what to do before people rely on it: a domain with HTTPS, email, off-site backups with a recovery key and a tested restore, private files, the secrets key, sign-in settings and updates. It ends with the security and production checklists. Every setting and command is in the Operations reference.
Three facts shape everything here:
- One server per state file. Never run two servers on the same file, and
never two copies of a container or pod that mount it. The server holds a
lock beside the file (
state.db.lock) to make sure. - The server is the only writer while it runs. Commands that would change the state are refused while it runs (exit status 75); use the workspace, the API, or stop it first. Commands that only read work beside it. See Running commands while the server runs.
- People sign in with a password; scripts use tokens. A person's first
password comes from a one-time link. A token (
orchk_...) is for the API, the CLI's--servermode, outside agents and CI.
Domain and HTTPS#
People must reach the server at an https:// address: the session cookie is
Secure, and password sign-in is off on any other address except this
machine's own (localhost, 127.0.0.1). There are three ways.
OrchKernel gets the certificate: --domain#
-
Point the name at the server. Create a DNS
A(andAAAA, if the server has IPv6) record forwork.acme.examplewith the server's public address, and wait untildig +short work.acme.exampleanswers it. -
Open ports 80 and 443 to the internet in the server's firewall and your cloud's security group. Port 80 answers Let's Encrypt's check and redirects everything else to HTTPS; nothing else is served on it.
-
Try it on Let's Encrypt's staging server first. Its certificates are not trusted by browsers, but its limits are generous, so a mistake in steps 1 and 2 costs nothing:
ok serve --domain work.acme.example --acme-directory stagingThe start says "By using --domain you accept the Let's Encrypt Subscriber Agreement (...)" once, then
obtained a certificate for work.acme.example.ok doctorshows it under HTTPS, warned as a staging certificate. -
Then the real one. Stop it and start without
--acme-directory(with the install script:--service --domain work.acme.example). Add--acme-email you@acme.exampleto hear from Let's Encrypt if renewal ever fails for long.
What happens then:
- HTTPS on 443 with HSTS, port 80 redirecting to it (
308) for your domain only; any other name gets421, never a redirect. OK_PUBLIC_URLis set tohttps://work.acme.examplefor you. One set to another host is refused at start.- Renewal is checked every 12 hours, 30 days before expiry. A failed one is retried after an hour, doubling to a day; with 14 days or fewer left, every admin gets an inbox item. No restart is needed.
- Certificates are kept in
tls/beside the state (0700). They are not in backups: a restored server gets new ones. - A failed order at start does not stop the server: it serves a placeholder certificate, keeps trying, and says why in its log.
- Several names: repeat
--domain(orOK_DOMAIN=a.example,b.example); the first is the address people use. - Ports below 1024 need a privilege. The install script's
--servicegives the service exactly that (CAP_NET_BIND_SERVICE); otherwise runsudo setcap cap_net_bind_service=+ep "$(command -v orchkernel)", or use--https-bind 0.0.0.0:8443 --http-bind 0.0.0.0:8080and forward 443 and 80 to them. Docker allows them inside the container.
With Docker Compose, deploy/compose.https.yml does all of this; see
Install.
Your own certificate#
ok serve --tls-cert /etc/orchkernel/fullchain.pem --tls-key /etc/orchkernel/key.pem
PEM files, the chain with the server's certificate first. The address
people use is the certificate's first name (or OK_PUBLIC_URL, which must
be a name it covers). The files are read again when they change (checked
every minute) or on SIGHUP; a file that does not load keeps the old
certificate and says why in the log. --domain and --tls-cert are two
ways to get a certificate: give one.
Behind a proxy you already run#
Caddy, nginx, Traefik or a cloud load balancer ends TLS and passes plain HTTP to OrchKernel:
OK_PUBLIC_URL=https://work.acme.example ok serve --bind 127.0.0.1:8080 --trusted-proxy 127.0.0.1
OK_PUBLIC_URLis the https address people open. It decides the session cookie (__Host-ok_session,Secure) and is the only origin a signed-in write is accepted from. The server never builds an address from theHostheader.--trusted-proxy(OK_TRUSTED_PROXIES) names the proxy, so the server reads the client's address fromX-Forwarded-Forfor its sign-in limits.- Keep event streams unbuffered (
/api/events/streamis server-sent events): in nginx,proxy_buffering offfor that path.
docker-compose.yml has a commented Caddy example.
The public address#
Without TLS flags, OK_PUBLIC_URL decides everything about the address. A
server bound to loopback with none uses http://<bind> (the cookie is then
ok_session, without Secure: how ok demo serve and a laptop work). A
server bound to anything else with no OK_PUBLIC_URL, or with an http://
address that is not loopback, turns password sign-in off: the sign-in page
says so, the start prints "set OK_PUBLIC_URL to the https address people
use", and tokens keep working. /api/health shows
"password_sign_in":true once it is right.
Several organizations on one server, each on its own host name, get their
certificates with --tenants --acme: see
Hosting several organisations.
Email: Resend or another SMTP server#
Invites, password resets and the "your password changed" notes go by email once it is set up; until then every link is shown once to the admin who made it, to send themselves, and Forgot password? is not offered. An admin sets it in Governance, Sign-in, section Email:
- Resend: the host (
smtp.resend.com), username (resend) and TLS are filled in; give the API key as the password, the port (465, or 587) and a From address on a domain you verified at Resend. Leave Resend's click tracking off: it would rewrite reset links. - Other SMTP server: host, port, security (TLS, or STARTTLS; plain SMTP only to a loopback host in development), username, password and the From address.
Send test email says where it stopped if it fails ("Resend refused the
sender: the domain of noreply@acme.example is not verified (550)"). The
password is sealed with the secrets key and never shown again. A server can
instead give every company the same mail server through OK_SMTP_HOST,
OK_SMTP_PORT, OK_SMTP_SECURITY, OK_SMTP_USERNAME, OK_SMTP_PASSWORD,
OK_SMTP_FROM and OK_SMTP_FROM_NAME, used when the company has none of
its own. From the command line (server stopped): RESEND_KEY=re_... ok --as ada mail set --preset resend --from noreply@acme.example --password-env RESEND_KEY, then ok --as ada mail test. See
Setting up a company.
Backups#
Backups run from the first start: every day at 02:00 UTC, encrypted, into
backups/ beside the state. Each one holds everything a restore needs: the
state, the secrets key, llm.yaml and mcp.yaml, and pack folders beside
the state. Two things are left to you, and the Backups page (sidebar,
Company, Backups; admins only) walks you through both: a recovery
key, and a copy off this server. Its status line says what is missing:
Last backup today 02:00 Next tomorrow 02:00 This server
No recovery key yet. Backups open only with this server's secrets key ...
Backups are on this server only: add an off-site destination.
1. Create the recovery key and keep it offline#
Every backup is encrypted to two keys:
- This server's key, derived from
secrets.key. It needs nothing saved and makes a one-click restore on this server possible, but it is lost with the server. - Your recovery key, made once and kept by you. It opens every backup made after it was created, even when the server and its secrets key are gone.
Choose Create recovery key. It needs a sign-in within the last ten minutes; otherwise the page asks for your password first. The key is shown once: choose Download orchkernel-recovery-key-.txt (or Copy), tick I saved it somewhere safe, off this server, and Done. The page then shows its fingerprint, and each backup in the list records the fingerprint it was made with.
Keep the file like a root password: offline, in a password manager or a safe, never on the server or next to the backups. Whoever holds it can read every backup made since. If you lose both the server's key and the recovery key, no backup can be opened, by anyone. Make a new recovery key later applies to backups from then on; older ones still open with the key they were made with.
From the command line (server stopped): ok backups recovery-key --out /safe/place/recovery.txt.
2. Add an off-site destination#
Under Where backups go, Off-site, pick the provider. Any S3-compatible storage works; the tabs fill in what they can:
| Provider | Endpoint | Region | The key |
|---|---|---|---|
| Cloudflare R2 | https://<account id>.r2.cloudflarestorage.com (the account id is on the R2 overview page) |
auto |
An R2 API token with Object Read & Write on the bucket: its access key id and secret |
| Backblaze B2 | https://s3.<region>.backblazeb2.com, shown on the bucket's page |
the bucket's, such as us-west-004 |
An application key for the bucket: its key id and application key |
| AWS S3 | https://s3.<region>.amazonaws.com |
the bucket's | An IAM user's access key, allowed s3:PutObject, GetObject, DeleteObject and ListBucket on the bucket |
| MinIO | http://minio:9000 (plain http only on this machine or a private network) |
us-east-1 |
A MinIO access key |
Create the bucket first, private, with versioning or object lock if your
provider offers it (so nobody holding the key can delete history). Fill
Bucket, Folder in the bucket (orchkernel/), the key id and secret,
and a Name (how backups and alerts name it). Choose Test: it writes,
reads back and removes a small object, and says what failed if anything
does. Then Save. An off-site destination cannot be saved before a
recovery key exists, so a backup off this server always opens without it.
The secret is sealed with the secrets key and never shown again; entering a
new one needs a recent sign-in. The endpoint must be https, except on this
machine or a private network, and link-local and cloud metadata addresses
are refused.
3. Back up now and check "Verified"#
Choose Back up now. The list gains a row per destination; each must say Verified under Checked:
- on this server, the backup was opened again with the server's key and every file, its checksum and the state's own integrity checked;
- off-site, it was uploaded, then downloaded again and compared.
A backup that fails a check is not kept. When a scheduled one fails, every admin gets one inbox item per destination ("Backups to offsite are failing: ...", with the destination's name), updated rather than repeated, and another when it works again; with no good backup for two scheduled runs, "No backup since ...".
4. Schedule and retention#
Schedule and retention sets Every day at a time (shown in UTC and in
your own time), Every hour, or Only when asked, and how many to keep:
the newest backup of each of the last 7 days, 4 weeks and 6
months. The newest good backup is never removed, and copies made just before
a restore are kept 7 days. A server's operator may fix any of these with
OK_BACKUP_* variables; the page then shows them read-only.
5. Test a restore, on another machine#
A backup you have never restored is a hope. Once, and after big changes, restore the latest one on a spare machine or a laptop, with only the backup and the recovery key, as you would after losing the server:
mkdir -m 700 ~/restore-test
ok --state ~/restore-test/state.db restore orchkernel-2026-10-07T020000Z-a1b2c3.tar.gz.age \
--identity /safe/place/recovery.txt
It prints what the backup holds, asks for the company's name, writes the state, the secrets key and the configuration into the folder, and prints the exact command to start it:
restored 2026-10-07T020000Z-a1b2c3 into /home/ada/restore-test/state.db; start the server on it with:
orchkernel --state /home/ada/restore-test/state.db serve
Run it with another --bind if 8080 is taken, sign in with your usual
password, look at recent work, then stop it and delete the folder. Restoring
the live server (from the Backups page or with ok restore) is in
Backups and restore.
Private files#
Whatever the umask, the program creates the state folder 0700 (when it makes
it), and the state file, its -wal and -shm, state.db.lock,
secrets.key, backups, certificates, the pre-upgrade copies, the search
index and the model cache readable only by the service user (0600 files,
0700 folders). Files from an older release, or copied in by hand, keep their
mode; the server lists them once at start:
warning: 3 files can be read by other users on this host:
/data/state.db (0644), /data/state.db-wal (0644), /data/state.db-vectors (0755)
Run `ok doctor --fix-permissions`, or chmod 600 the files and 700 the folders.
ok doctor --fix-permissions sets 0600 and 0700 on those the service user
owns and names each change. It changes modes only, so it is safe while the
server runs. Keep the state on a local disk: NFS, SMB and other network
file systems are not supported, and the start says so.
The key file#
ok init, the first ok serve on an empty folder and ok demo create
secrets.key beside the state file when no key is found, mode 0600:
# OrchKernel secrets for the state files in this directory.
# Created 2026-10-06T09:30:00Z by `ok init`. Keep it private (mode 0600) and
# back it up apart from the state: without it, connection secrets, two-step
# seeds, the mail password and model keys do not open.
OK_SECRETS_KEY=...
OK_WEBHOOK_SECRET=...
Encrypted backups carry it inside (sealed by the backup's own encryption),
so a restore needs nothing else. A plain copy of the state made with ok backup does not.
- The environment wins.
OK_SECRETS_KEYorOK_WEBHOOK_SECRETset (and not blank) is used instead of the file's value. To keep the key off the data disk, setOK_SECRETS_KEYfrom your secret store before the first start. When both hold a key and they differ, the start says so and names the key the stored secrets need. - A key is never made for a file sealed under another one. When the state holds sealed secrets and no key is found, nothing is generated: the start says "error: /data/state.db holds 4 secrets sealed under key 1a2b3c4d (2 connections, 1 mail password, 1 two-step seed) and no key was found. Set OK_SECRETS_KEY, or put the key file back at /data/secrets.key." The server still starts, so an admin can sign in and see what is locked.
- Rotation.
ok secrets rotate(server stopped) re-seals every stored secret under a new key in one transaction. With the key file, the new key replaces it and the old one is kept beside it assecrets.key.<key id>; backups made before then open with it, or with the recovery key. WithOK_SECRETS_KEY, pass--new-key-env VARand setOK_SECRETS_KEYto the new key afterwards. A rotation stopped half-way is settled at the next start. - Under
--tenantseach tenant's key is sealed under the operator key instead (Hosting several organisations).
Sign-in settings#
People sign in at the public address with Email and Password. An admin sets the rules in Governance, Sign-in:
- Passwords are at least 12 characters, not a common password, and not
the person's email address or name; they never expire. The server keeps an
argon2id hash only.
OK_PASSWORD_DENYLISTnames a file of further refused passwords, one per line. - Two-step sign-in. Each person may turn on an authenticator app (TOTP) on their profile and gets ten recovery codes. Require it for admins at least, or for everyone; people it covers set it up at their next sign-in. The seeds are sealed with the secrets key: if the key is lost, people sign in with a recovery code and an admin resets their two-step.
- Sessions end 7 days after sign-in, or after 18 hours without use, whichever is first; both can be changed. Each person sees their sessions on their profile and can end any of them. A password change ends every other session.
- Recent sign-in. Changing a password or two-step, making recovery codes, issuing a token, creating a recovery key, entering an S3 secret and restoring a backup need a sign-in within the last ten minutes; otherwise the workspace asks for the password (and code) again.
- Failed sign-ins. Ten wrong passwords or codes for one address within
15 minutes lock that address for 15 minutes, whether or not anyone holds
it. One client address gets 429 after 60 failed credentials a minute, and
after 30 sign-in requests a minute. Every failure reads "Email or password
is wrong.", so the page never tells anyone who has an account. Behind a
proxy, set
--trusted-proxyso the real client is counted. - Single sign-on (OIDC, SAML) is not available yet.
Tokens, who may issue them and their lifetimes are in People, tokens and access.
Updates#
- Back up: Back up now on the Backups page (or
ok backups runwith the server stopped), and check it says Verified. - Install the new version:
- with the install script, run it again:
curl -fsSL .../install.sh | sh -s -- --servicereplaces the program and restarts the service (--version 0.4.0picks one); - with Docker, pull the new tag and recreate the container
(
docker compose pull && docker compose up -d, or a newdocker runwith the same volume). Pin a version tag (0.4.0) rather thanlatestin production, so a restart never upgrades by surprise.
- with the install script, run it again:
- Check
/api/healthand the start's log, and sign in.
Some releases change how the state is stored and convert it at the first open; the way back is then the backup. See Upgrades and rollback.
Security checklist#
- HTTPS:
--domain, your own certificate, or a proxy ending TLS with OrchKernel not reachable directly./api/healthsays"password_sign_in":trueand"dev_auth":false. - Two-step sign-in required at least for admins (Governance, Sign-in).
- A recovery key created and kept offline, away from the server and the backups; an off-site destination saved and Verified.
secrets.keyprivate (orOK_SECRETS_KEYfrom a secret store), andok doctorsays every sealed secret opens.- The state, its key, certificates and backups readable only by the service
user (
ok doctor); the state on a local disk. - Tokens with lifetimes that fit; a lost device's token revoked at once; a
leaver offboarded (
POST /api/actors/<id>/offboard), which ends their sessions, stops their tokens, revokes their delegations, cancels the runs acting for them and withdraws their waiting approvals. - Each webhook source with its own long secret; CORS off unless needed;
--trusted-proxynaming your proxy. - Outbound traffic limited to your model providers, tool hosts, mail server
and backup bucket;
OK_CONNECTION_HOSTSlists the hosts connections may reach. - Restricted data only on local or contractually covered models (Models).
- Kill switches: who engages the global one, and how the company hears of it, written down (Policy, rules and kill switches).
- No development flags (
--dev-auth,OK_DEV_*) anywhere in the service definition.
Production checklist#
| Item | Done |
|---|---|
HTTPS on the address people use (--domain tried on staging first, your own certificate, or a proxy) |
|
| The first admin created from the first-admin link; the setup wizard finished | |
| People invited with roles and reporting lines; two-step required for admins | |
| Email set up and Send test email passing | |
A model connected (the wizard, or llm.yaml reviewed: sensitivities, prices, default model) |
|
mcp.yaml lists every tool server the modules use (no "note: no MCP configuration" at start) |
|
| Recovery key created and stored offline | |
| Off-site destination saved; latest backups Verified on both destinations | |
| A restore tested on another machine with only the backup and the recovery key | |
ok doctor passes: build, state, lock, private files, secrets, models, HTTPS, backups |
|
Run as a service (--service, a container with --restart unless-stopped, or a Recreate Deployment); health probe on /api/health |
|
| Logs collected from the server's standard output | |
| Webhook secrets set; senders sign | |
| Kill switch and budget procedure written down and tried | |
| Update runbook: back up, install, check health, sign in |