Self-hosting, backups and upgrades

Last updated October 5, 2026

On this page

Draft for review

This page is for the person who runs OrchKernel for a company: how to start a server, give people access, receive webhooks, back it up while it runs and upgrade it. To try the product first, use the demo company in Quickstart with the demo company. The demo is for development only; a real company starts from ok init, as below.

OrchKernel is one program, orchkernel, that is both the command line and the server. The server keeps everything for one company in one SQLite file, the state file. The guide calls the program ok: the Docker image ships an ok link beside orchkernel, step 1 below makes one, and anywhere else alias ok=orchkernel does the same. Serving several organizations from one server is covered in Hosting several organizations. What is not possible yet is listed near the end.

What you are running#

       people (browser), the ok CLI, outside agents, webhook senders
                               |
                               | HTTPS
                               v
              reverse proxy (TLS, single sign-on, IP rules)
                               |
                               | HTTP, Authorization: Bearer <token>
                               v
              ok serve --bind 0.0.0.0:8080
                 /api/*   the API and the event stream
                 /*       the web workspace, built into the program
                               |
                               | one writer
                               v
                      state.db  (SQLite)
                               |
             +-----------------+-----------------+
             v                                   v
     model providers                    tool servers and connections

Three facts shape everything below.

  • One server per state file. Never run two servers on the same file, and never two copies of a container or pod that mount it.
  • The server works from memory. It writes what changed to the file after every request that changes something, and again every 30 seconds. A CLI command that changes state (ok token issue, ok actor add-human, ok pack install, ok kill) writes the file directly, and a running server does not see it. For example, a token issued with ok token issue while the server runs gets 401 until the server restarts, and the server's own writes can overwrite such a change. Run those commands only while the server is stopped. While it runs, use the web workspace or the API.
  • There are no passwords. Everyone signs in with a token, a long secret string that starts with orchk_. The sign-in page asks for one, and the API takes it as a bearer token. For single sign-on, put a proxy in front (see Single sign-on).

Deploy a single-company server#

These steps run the program directly on a Linux or macOS host. Docker, Compose and Kubernetes follow the same order and are below. Ada is the first admin, as in the demo company.

  1. Get the program. No release has been published yet, so build it from a checkout of the repository, in the repository folder:

    sh
    (cd ui && npm ci && npm run build)      # the web workspace, first
    cargo build --release -p ok-cli
    install -m 0755 target/release/orchkernel /usr/local/bin/orchkernel
    ln -sf orchkernel /usr/local/bin/ok
    

    You need Node 22 and stable Rust. Build inside the repository: its .cargo/config.toml makes concurrent requests much faster. A build made elsewhere works, but ok serve then starts with "warning: this build's SQLite takes process-wide locks ...".

  2. Choose where the state lives and keep it private to the service user:

    sh
    umask 077
    export OK_STATE=/var/lib/orchkernel/state.db
    

    Without OK_STATE (or --state), the file is ./.orchkernel/state.db under the current folder. The folder is created for you. With the usual umask of 022 the file is created readable by every user on the host (mode 0644); with umask 077 the folder is 0700 and the file 0600. Use the same umask for the commands that make backups.

  3. Create the company and its first admin, with the server stopped:

    sh
    ok init --admin ada
    

    You see kernel ready; admin ada. Use --as ada for admin commands.

    To set up a whole company in the same step, add a blueprint (a company described as data: teams, modules and assistants): ok init --admin ada --blueprint saas-startup prints the same line and then installed blueprint saas-startup: 4 teams, 0 people, packs [crm, support, product], 0 delegations, 0 views. ok blueprint list shows the bundled blueprints, and Setting up a company explains them.

    To add one module instead, run this from the repository folder (the bundled modules are in packs/; in the Docker image they are in /opt/orchkernel/packs):

    sh
    ok --as ada pack install packs/crm
    

    You see installed pack crm (2 collections, 1 skills, 1 agents).

  4. Issue Ada's token:

    sh
    ok --as ada token issue --for ada
    

    You see two lines: the token itself (orchk_ and 64 hex digits), then token 99251b8a for ada expires 2026-11-04T..., 30 days from now. Here 99251b8a is the token's id, which you use to revoke it. The token is shown this once. Store it in your secret store. Add --ttl 7d for another lifetime, or --label "ada laptop" to tell it apart later.

  5. Set the model and secrets in the environment (see Configuration): at least one provider key such as ANTHROPIC_API_KEY, or an llm.yaml (Models), and the webhook secrets if anything will send webhooks. Without a model the server still starts, prints "warning: no model provider is configured: set ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY or OLLAMA_HOST, or write llm.yaml ...", and every agent run fails: its task shows as failed.

  6. Start the server:

    sh
    ok serve --bind 0.0.0.0:8080
    

    You see OrchKernel serving on http://0.0.0.0:8080, preceded by notes about anything still missing. With the support module installed and no mcp.yaml, for example: "note: no MCP configuration for billing, helpdesk; their tools are refused until they are in mcp.yaml". Run the server under a service manager: systemd with Restart=always, a dedicated user, and an EnvironmentFile of mode 0600 for the keys.

  7. Check it is up:

    sh
    curl -s http://127.0.0.1:8080/api/health
    

    You see {"dev_auth":false,"ok":true}. The health check needs no token. dev_auth must be false on any real server.

  8. Sign in. Open the address in a browser. The sign-in page says "Sign in with the token an admin gave you." Paste Ada's token in Token and choose Sign in. You land on Threads. A wrong or expired token shows "That token wasn't recognised or has expired."

  9. Give everyone else a token. Add people as described in Setting up a company. Then, as Ada:

    1. Open People in the sidebar, under Governance (admins only). Each person has a row with Issue token, and Revoke all once they hold a token.
    2. Choose Issue token beside a person. In the demo company, beside Maya Patel: the sheet "Issue a token for Maya Patel" opens with Label and Valid for (1 day, 7 days, 30 days or 90 days; 30 days is preselected).
    3. Choose Issue token. The sheet "Token for Maya Patel" shows the token once, with Copy. Send it to the person privately, then choose I have copied it. The new token appears under their name.

To stop the server, send it SIGTERM or press Ctrl-C. It takes no new connections, gives the requests in flight up to OK_SHUTDOWN_WAIT seconds (30 by default), writes its state once more and exits with status 0.

Docker, Compose and Kubernetes#

No release has been tagged yet, so there is no published image or release archive. Build the image from the repository (docker build -t orchkernel/orchkernel .) and push it to your own registry. Once a v* tag is pushed, the release workflow is set up to publish ghcr.io/arsp2020/orchkernel with the version and latest tags.

The image runs as the non-root user orchkernel (uid 10001), sets OK_STATE=/data/state.db and OK_BIND=0.0.0.0:8080, works in /data (so /data/llm.yaml and /data/mcp.yaml are picked up without setting anything), has a health check on /api/health, and carries the bundled modules at /opt/orchkernel/packs.

Docker#

sh
docker volume create orchkernel-data
# Bootstrap first, in one-off containers on the same volume.
docker run --rm -v orchkernel-data:/data orchkernel/orchkernel init --admin ada
docker run --rm -v orchkernel-data:/data orchkernel/orchkernel \
  --as ada pack install /opt/orchkernel/packs/crm
docker run --rm -v orchkernel-data:/data orchkernel/orchkernel \
  --as ada token issue --for ada        # prints the token once
# Then start the server.
docker run -d --name orchkernel -p 8080:8080 -v orchkernel-data:/data \
  -e ANTHROPIC_API_KEY=... -e OK_WEBHOOK_SECRET=$(openssl rand -hex 32) \
  orchkernel/orchkernel

Docker Compose#

sh
cp .env.example .env                    # every setting, commented out
printf 'OK_WEBHOOK_SECRET=%s\n' "$(openssl rand -hex 32)" >> .env
docker compose build
docker compose run --rm orchkernel init --admin ada
docker compose run --rm orchkernel --as ada token issue --for ada
docker compose up -d

The service reads .env and nothing else from your shell. Uncomment only the settings you set. The compose file sets OK_STATE itself, and empty provider keys and an empty webhook secret count as unset. Ignore the line in .env.example that says OK_SECRETS_KEY is not read: it is, and it seals the secrets of connections. docker-compose.yml also has commented examples of an Ollama service and a Caddy proxy.

Kubernetes#

deploy/k8s/orchkernel.yaml holds a namespace, a 5Gi volume claim, a one-replica Deployment with the Recreate strategy (required: two pods must never open the same file), a Service and an Ingress example. Its image: names ghcr.io/arsp2020/orchkernel:latest, which does not exist yet: set it to the image you pushed. Create the orchkernel-secrets Secret yourself from deploy/k8s/secrets.example.yaml. Bootstrap with the Deployment scaled to zero, in a one-off pod on the same claim, never with kubectl exec into the running pod; docs/deployment.md in the repository has the commands.

Configuration#

Settings come from the environment, and some also from ok serve flags.

Setting Default What it does
OK_STATE or --state ./.orchkernel/state.db The state file. Set but empty, the program stops with "a value is required for '--state '".
OK_BIND or --bind 127.0.0.1:8080 Where the server listens. Use 0.0.0.0:8080 in a container.
OK_SHUTDOWN_WAIT 30 Seconds a stopping server waits for requests in flight.
OK_WEBHOOK_SECRETS or --webhook-secrets unset One secret per webhook source: helpdesk=...,billing=... (the flag is repeated once per source).
OK_WEBHOOK_SECRET or --webhook-secret unset The secret for sources without their own.
OK_CORS_ORIGINS or --cors-origin unset Other web origins allowed to call the API from a browser. Unset means the workspace's own origin only, which is all it needs.
OK_TRUSTED_PROXIES or --trusted-proxy unset Your proxy's address or range, so the server reads the client's address from X-Forwarded-For.
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY or GOOGLE_API_KEY, OLLAMA_HOST unset Turn on a model provider when there is no llm.yaml. See Models.
OK_LLM_CONFIG ./llm.yaml if present The model configuration file.
OK_LLM_CHECK on 0 skips the check of each provider key at start.
OK_MCP_CONFIG ./mcp.yaml if present Tool servers. Tools on servers it does not list are refused.
OK_SECRETS_KEY unset Base64 of 32 random bytes. Seals the secrets of connections. Without it, connections with a secret are locked.
OK_CONNECTION_HOSTS, OK_CONNECTION_PROXY unset Hosts connections may reach, and a proxy for their traffic.
OK_ARCHIVE_AFTER_HOURS 168 Finished tasks and runs untouched this long leave memory (they stay in the file). off never archives.
OK_AUDIT_MIN_FREE_MB 1024 Free disk space, in MB, under which admins are warned.
RUST_LOG info How much the server logs to standard output.
--dev-auth off Trusts an x-ok-actor header instead of tokens. A flag only, never read from the environment. Never on a real server.
OK_DEV_ECHO_TOOLS, OK_DEV_AUTH off Development switches. Never on a real server.

Tuning settings (OK_DB_READERS, OK_SEARCH_CANDIDATES, the audit checkpoint settings) are in docs/deployment.md in the repository.

Keep provider keys, webhook secrets, OK_SECRETS_KEY and the admin token in a secret store (a Kubernetes Secret, Docker secrets, a systemd EnvironmentFile of mode 0600, Vault). Never commit llm.yaml or .env.

People, tokens and access#

A token belongs to one person and is valid for 30 days unless you choose otherwise. The server keeps only a hash of it, so a lost token cannot be shown again: issue a new one. Your own tokens are on your profile: choose your name at the bottom of the sidebar, then the Sessions & tokens section.

You want to On the screen With the API (server running) With the CLI (server stopped)
Give someone a token People, Issue token (admins) POST /api/tokens {"actor":"maya","ttl":"7d"} ok --as ada token issue --for maya --ttl 7d
Get another token for yourself Your profile, Sessions & tokens, New token POST /api/tokens {"actor":"<your id>"}
See tokens People (everyone's, admins) or your profile (yours) GET /api/tokens: everyone's for an admin, your own for anyone else ok --as ada token list
Revoke one token Revoke on its row, then Revoke to confirm DELETE /api/tokens/id/<id> ok --as ada token revoke --id <id>
Revoke all of a person's tokens Revoke all on People, or Sign out everywhere on your profile DELETE /api/tokens/<person> ok --as ada token revoke maya

Only an admin issues or revokes tokens for someone else. When Maya, the support lead, tries over the API to issue one for Alice, she gets 403 "policy denied token.issue: only admins issue tokens for others"; revoking Alice's gets "policy denied token.revoke: only admins revoke others' tokens". Each API answer that issues a token returns it once, in token, with its id and expires_at.

Send the token as Authorization: Bearer <token>. The one exception is the event stream, GET /api/events/stream?token=..., because a browser cannot set headers there. A token in the address of any other request gets 401 "missing bearer token".

Someone leaves. An admin calls POST /api/actors/<id>/offboard. The person is marked inactive, their tokens stop working at once (401 "invalid token"), and their running work and delegations are cancelled. When someone who is not an admin calls it, the answer is 202 with an approval for an admin to decide.

Failed sign-ins. After 60 failed tokens from one address within a minute, that address gets 429 "too many failed authentications from this address" with a Retry-After header until the minute is up. A valid token from the same address still works. Behind a proxy, set --trusted-proxy so the server counts the real client and not the proxy.

Single sign-on#

OrchKernel has no login of its own beyond tokens. For single sign-on, put an OIDC-aware proxy in front (oauth2-proxy, Caddy with forward_auth, nginx with auth_request). The proxy signs the person in and adds the bearer token that belongs to them. Keep one token per person: a shared token makes every action in the audit log look like the same person.

Webhooks#

POST /api/webhooks/<source> turns an event from another system into tasks, through the triggers (rules in a module that start work) that listen to that source. The support module's helpdesk source, for example, creates a "Triage incoming ticket" task for its triage agent. A source with no secret is refused. The sender signs each delivery:

  • x-ok-timestamp: the time of sending in Unix seconds, within five minutes of the server's clock.
  • x-ok-signature: sha256= and the hex HMAC-SHA256, keyed with the source's secret, of the timestamp, a dot and the body exactly as sent.
  • x-ok-delivery (optional): a unique id. A repeat within 24 hours creates nothing and answers with the first delivery's tasks, so the sender can retry safely.

With the server started with OK_WEBHOOK_SECRETS=helpdesk=... and the same secret in HELPDESK_SECRET:

sh
body='{"external_id":"Z-100","subject":"Charged twice"}'
ts=$(date +%s)
sig=$(printf '%s.%s' "$ts" "$body" | openssl dgst -sha256 -hmac "$HELPDESK_SECRET" | sed 's/^.* //')
curl -sS -X POST https://orchkernel.example.com/api/webhooks/helpdesk \
  -H 'content-type: application/json' \
  -H "x-ok-timestamp: $ts" -H "x-ok-signature: sha256=$sig" \
  -H 'x-ok-delivery: evt-0001' --data-raw "$body"
What the sender sees When
202 {"delivery":"evt-0001","replayed":false,"tasks":["..."]} Accepted. Agents run after the answer, so a slow model never holds the sender.
202 with "replayed":true and the same tasks The same x-ok-delivery again.
401 "bad webhook signature" Wrong secret or a changed body.
401 "x-ok-timestamp is more than five minutes from the server clock" A stale or future timestamp.
401 "no webhook secret is configured for this source" Neither OK_WEBHOOK_SECRETS nor OK_WEBHOOK_SECRET covers it.
400 bad_request The body is not JSON.

Senders that cannot sign may send the secret itself in x-ok-webhook-secret. It is accepted (202, with "delivery":null) but offers no protection against a replayed request; prefer signatures.

Backups#

Everything the company has is in the state file: people, policy, tasks, threads, the brain, the audit log, token hashes and connections (their secrets sealed, never in clear). Back it up with ok backup, which is safe while the server runs and never writes to the state file.

Back up while running#

  1. Take the backup, with umask 077 so only the service user can read it:

    sh
    umask 077
    ok --state /var/lib/orchkernel/state.db backup /backups/state-$(date +%F).db
    

    You see backed up <state file> to <backup>. Missing folders are created, an existing file of the same name is replaced, and naming the state file itself is refused ("... is the state file itself"). In Docker: docker exec orchkernel orchkernel backup /data/backups/state-$(date +%F).db, then copy it off the volume.

  2. Check it:

    sh
    ok --state /backups/state-2026-10-05.db audit verify
    

    The last line is no break; head #25 sha256:... (the number is the count of recorded events), and the exit status is 0. The check does not change the backup's contents, but it leaves two small files beside it, state-2026-10-05.db-wal and state-2026-10-05.db-shm. Copy only the .db file.

  3. Copy it off the host, and run steps 1 to 3 every day from cron or your scheduler.

Do not copy state.db with cp while the server runs: recent writes may still be in state.db-wal beside it. A backup made by ok backup is one self-contained file.

A backup does not hold what lives outside the file. Keep these too, apart from the backup: llm.yaml and mcp.yaml, the provider keys, the webhook secrets, OK_SECRETS_KEY, and the module folders you installed from. Keep OK_SECRETS_KEY away from the backups it protects; a backup restored under a different key has its connections locked until their secrets are entered again.

A backup keeps content the context module has since purged, so keep backups only as long as your retention promises allow (Context capture).

Test a restore#

Start a second server on a copy of the backup, on another port:

sh
cp /backups/state-2026-10-05.db /tmp/restore-test.db
ok --state /tmp/restore-test.db serve --bind 127.0.0.1:8081

Sign in with a token that existed when the backup was taken, or call curl -s http://127.0.0.1:8081/api/actors -H "Authorization: Bearer <Ada's token>", and check that recent people and work are there (GET /api/tasks lists the tasks). Use a copy: the test server writes to the file it opens.

Restore for real#

  1. Stop the server.
  2. Move state.db aside, with state.db-wal and state.db-shm if they are there. After a clean stop they usually are not.
  3. Copy the backup to state.db.
  4. Start the server and check /api/health.

Everything since the backup is gone, tokens included: a token issued after the backup gets 401 "invalid token" and must be issued again. Tokens that existed at the backup work again.

Upgrades and rollback#

  1. Stop the server, and every scheduled job that opens the state file with the CLI.
  2. Back it up with ok backup.
  3. Replace the program or the image tag.
  4. Start the server and check /api/health.

Some releases change how the file is stored, and the first open by the new program converts it. The current conversion, to storage version 2, works like this:

  • It checks there is room first. Without it, the open fails with "upgrading the state file needs N bytes free beside it, and M are" and nothing changes.
  • It writes a copy beside the file, state.db.pre-upgrade-v1.db, and keeps it. Delete it yourself once the new release has run well.
  • It converts in one transaction, so a stop half-way leaves the old file untouched.

An older program refuses a converted file: "the state file is newer than this binary". Rolling back after a storage change means restoring the pre-upgrade copy or your backup, and losing what changed since the upgrade. Releases without a storage change can be rolled back by swapping the program again.

If an old server kept writing after the conversion, the next open may refuse with "the state file was written by two versions of orchkernel". Stop everything and restore the pre-upgrade copy or a later backup.

The first open by a release that adds record indexes builds them before the server answers: a few seconds per index on a large collection, once.

Watching it run#

What How
Is it up GET /api/health, no token. Use it for load balancer and container probes.
Server log Standard output, level set by RUST_LOG. Ship it to your log system.
What happened Events in the sidebar, GET /api/events, or the live GET /api/events/stream. Someone who is not an admin sees only events about their own threads, tasks and runs (in the demo, Ada sees 378 events where Maya sees 7); use an admin token for operations.
One run, step by step ok replay --run <id> (or --thread, --task) rebuilds it from the log alone. ok --as ada runs and ok --as ada tasks list the ids.
The audit log is intact ok audit verify and ok audit head, safe beside a running server. See Events and the audit log.

The server advances its own clock every 30 seconds: scheduled triggers, overdue work, stalled threads, archiving. This is the tick. To drive it from your own scheduler, call POST /api/tick with an admin token (deploy/k8s/tick-cronjob.yaml does this every minute). The answer lists what it did, such as {"escalated":[],"fired":[],"ran":[],...}. Someone who is not an admin gets 403 "policy denied tick: admins only".

When the audit log cannot record a decision (a full disk, say), the server refuses the action with 503 audit_unavailable rather than act unrecorded.

How big it can get#

One server, one state file, one company. There are no replicas and no failover: availability comes from fast restarts (a service manager or the Recreate Deployment with a health probe), and runs that wait for an approval survive a restart. If you need zero downtime, OrchKernel is not ready for you yet.

Plan on about 10,000 records per company. That is the size at which everything, search included, has been measured fast. The team's measurements on one Apple M2 laptop with 16 people working at once:

Records Saving a record, p95 Filtered list, p95 Search under load, p50 (common word)
10,000 17 ms 18 ms 62 ms
100,000 32 ms 72 ms 449 ms
324,000 42 ms 72 ms 1.2 s

Above 10,000, saving and listing hold up well, but a sales rep's search slows down, memory reaches about 600 MB at 324,000 records, and the first open and first tick after the file has grown take several seconds. Tuning for larger companies is on the roadmap.

License#

OrchKernel is source-available under the Business Source License 1.1. You may use, modify and run it in production for your own organization, or run a deployment for one client organization; separate deployments for separate clients are fine. Offering it as a hosted or managed service in which one deployment serves more than one organization needs a commercial license. Each version converts to Apache-2.0 four years after it is first published, and the client libraries under sdk/ are Apache-2.0 already. The LICENSE file in the repository is the authoritative text.

Security checklist#

  • TLS at the proxy; OrchKernel speaks plain HTTP and is not exposed directly.
  • /api/health says "dev_auth":false.
  • Tokens have lifetimes that fit; a lost device's token is revoked at once, and a leaver is offboarded (POST /api/actors/<id>/offboard), which stops their tokens and cancels their running work and delegations.
  • Each webhook source has its own long secret; CORS stays off unless needed; --trusted-proxy names your proxy.
  • The state file and backups are readable only by the service user (umask 077, or chmod 600 afterwards): they hold token hashes, records and the audit log. The program does not restrict this for you.
  • Outbound traffic is limited to your model providers and tool hosts, and OK_CONNECTION_HOSTS lists the hosts connections may reach.
  • Restricted data stays on local or contractually covered models (Models).
  • Today any signed-in person, not only an admin, can engage or release any kill switch through the API (POST /api/kill). In the demo, Alice, a sales rep, can engage tool:helpdesk.reply. See Policy, rules and kill switches.

Production checklist#

Item Done
Proxy with TLS (and single sign-on if you use it) in front; OrchKernel not reachable directly
ok init run once before the first start; admin token in the secret store
People added with roles and reporting lines; one token each
No development flags in the service definition
Webhook secrets set; senders sign
llm.yaml reviewed: sensitivities, prices, default model; ok doctor passes
mcp.yaml lists every tool server the modules use (no "note: no MCP configuration" at start)
OK_SECRETS_KEY set and kept apart from backups; every connection tested
State file and backups mode 0600
Persistent storage for the state file; daily ok backup copied off the host; a restore tested
One server or replica per file (Recreate on Kubernetes); health probe on /api/health
Logs shipped from standard output
Kill switch and budget procedure written down and tried
Upgrade runbook: stop, back up, swap, start, health check

Not possible yet#

  • More than one server or replica for a company, or failover. There is no Postgres backend yet.
  • Downloading a published image or release archive: none has been tagged.
  • A built-in login with passwords or single sign-on; use a proxy.
  • Changing state with the CLI while the server runs. Use the workspace or the API.
  • Exporting metrics or traces to OpenTelemetry; read standard output and the event log.
  • Rolling back across a storage change without restoring a backup.
  • Restricting kill switches to admins at the API (see the security checklist).

For developers#

  • ok serve and the tenancy host start every kernel through one path, crates/ok-server/src/startup.rs (prepare_kernel): the secrets key check, the context hooks, the process configuration, the module reconcile, the gateway's observer, the audit verifier, the gateway start-up sweep, then the request state. Each 30-second tick runs tick_one: the clock, the audit tick, archiving, the daily statistics refresh, then a persist.
  • ok backup loads no kernel: it copies through SQLite's online backup API (Kernel::backup_file), so it never writes a snapshot over a running server's.
  • Failed credentials are counted in FailedAuthLimiter (crates/ok-server/src/auth.rs); the storage version 2 conversion is crates/ok-kernel/src/migrate_v2.rs.
  • docs/deployment.md is the full operator reference: the Kubernetes bootstrap pod, MCP servers, connection key rotation (ok secrets rotate), archiving and indexes.