Models

Last updated October 5, 2026

On this page

Agents plan with a language model. OrchKernel does not ship one: you connect your own provider with a key in the environment or a file called llm.yaml. For each model you set the most sensitive data it may see and what it costs. OrchKernel sends a call only to a model allowed to see the data in it, and charges the call to the run's budgets.

This page covers connecting a model, checking it with ok doctor, what goes wrong and how it looks, how calls are routed by data class, conversation history, and cost. What is not possible yet is listed at the end.

Where OrchKernel looks for models#

At start, OrchKernel takes the first of these that exists:

Order Source Notes
1 The file OK_LLM_CONFIG names A path relative to the folder you start from, or absolute. If the file does not exist, OrchKernel stops with "No such file or directory" instead of moving on.
2 llm.yaml in the folder you start from For the demo, that is the folder you run ok demo serve in, not the demo's own folder (orchkernel-demo).
3 Provider keys in the environment ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY (or GOOGLE_API_KEY), OLLAMA_HOST.

When a file is found, the environment keys are not picked up on their own; the file names each provider's key itself (api_key_env). A variable that is set but empty counts as not set.

Models are read once, at start. To change them, stop the server and start it again.

One key: the quickest start#

Set one variable and OrchKernel offers that provider's current models at list prices. Prices are in US dollars per million tokens, input / output.

Variable Models offered May see up to Price
ANTHROPIC_API_KEY claude-sonnet-5 (default) Restricted $2 / $10
claude-opus-5-5 Restricted $4 / $20
claude-fable-5-1 Restricted $10 / $50
claude-haiku-4-5-20251001 Confidential $1 / $5
OPENAI_API_KEY gpt-5 Confidential $1.25 / $10
gpt-5-mini Confidential $0.25 / $2
GEMINI_API_KEY or GOOGLE_API_KEY gemini-2.5-pro Confidential $1.25 / $10
gemini-2.5-flash Confidential $0.30 / $2.50
OLLAMA_HOST (for example http://localhost:11434) OLLAMA_MODEL, or llama3.1 Restricted Free

The default model is claude-sonnet-5 when the Anthropic key is set. Otherwise it is the first model of the first provider found, in the order of the table: gpt-5 with only an OpenAI key, gemini-2.5-pro with only a Gemini key. ANTHROPIC_BASE_URL and OPENAI_BASE_URL point those providers at another address, such as a proxy.

"May see up to" is the model's sensitivity cap, explained in Which model may see which data. In short: with only an OpenAI or Gemini key, no model may see restricted data such as the demo's customer tickets, and asks that read them are refused. In the demo that includes the Support Agent, the Churn-risk watcher and the Product Agent, whose skills all read tickets.

Switching the demo to a real model#

The demo starts with a stub planner when no model is configured: fixed plans for the walkthrough's asks and a note for anything else (see Asking an agent). Before you switch, the start banner reads Model none configured: a stub planner answers the walkthrough's asks.

  1. Stop the demo with Ctrl+C.
  2. Set a key in the same terminal, for example export ANTHROPIC_API_KEY=sk-ant-..., or put an llm.yaml in the folder you start from (see Writing llm.yaml).
  3. Check it: run ok doctor in the same folder. You should see a line starting ok for each provider and then default model: claude-sonnet-5 (or your file's default). See Checking with ok doctor.
  4. Start the demo again with ok demo serve. Before the banner, the terminal prints one line per provider from a key check, such as llm: house: reachable, or a warning: line if a key is missing or refused. In the banner, the Model line names the default model and where it came from, for example Model gpt-5, configured from the environment (ANTHROPIC_API_KEY, ...) or Model house-model, configured from llm.yaml.
  5. Sign in with the new tokens (each start revokes the old ones). As ada, open Governance, tab Policy. The Model governance card shows default: claude-sonnet-5 · configured providers serve 4 model(s) (your default and count) and a table with each Model and May see up to. With the stub it shows default: mock-default · configured providers serve 0 model(s) and one row, mock-default, Restricted.
  6. As pat (product manager), open Threads, click Product Agent under Talk to an agent and ask What should we build next quarter?. With the stub you get the note "No model is configured, so the demo's stub planner answered. ..." With a model allowed to see restricted data you get the model's answer and a green Done badge. Open Details on the run card: the Model step shows the tokens and the cost (see Cost). With a model capped at confidential the run fails instead, with "policy denied llm.complete: no allowed model for data class pii.customer".

Writing llm.yaml#

Copy llm.example.yaml from the repository to llm.yaml and keep the providers you want. A provider whose key is not set is skipped at start with a warning, so the copied file works with only one key set.

yaml
providers:
  - name: anthropic
    kind: anthropic
    api_key_env: ANTHROPIC_API_KEY
    models:
      - { id: claude-sonnet-5, max_sensitivity: restricted, input_per_mtok_cents: 200, output_per_mtok_cents: 1000, default: true }
      - { id: claude-haiku-4-5-20251001, max_sensitivity: confidential, input_per_mtok_cents: 100, output_per_mtok_cents: 500 }

Provider settings:

Field What it does
name Your name for the provider. It appears in messages ("house rejected the API key").
kind anthropic, openai, gemini, azure_openai or openai_compatible.
api_key_env The environment variable holding the key. Keeps the key out of the file.
api_key The key itself. Prefer api_key_env.
base_url The server's address. Needed for azure_openai and for most openai_compatible servers.
api_version Azure only, for example "2024-10-21".
models The models below.

Model settings:

Field What it does If left out
id The model id the provider knows. For Azure, the deployment name. Required
max_sensitivity The most sensitive data this model may see: public, internal, confidential or restricted. internal
input_per_mtok_cents, output_per_mtok_cents Price in cents per million tokens. 200 means $2.00. 0: calls cost nothing
default: true The company's default model. The first model in the file

Set max_sensitivity on purpose. Left out, a model may see internal data only, and every ask that reads customer data is refused.

OpenAI-compatible and local servers#

kind: openai_compatible covers any server that speaks the OpenAI chat completions API: Ollama, vLLM, LM Studio, Groq, Mistral, DeepSeek, Together, OpenRouter. It needs no key unless the server asks for one.

yaml
  - name: ollama
    kind: openai_compatible
    base_url: http://localhost:11434/v1
    models:
      - { id: llama3.1, max_sensitivity: restricted }

A model on your own machine or network is the usual place for restricted data: nothing leaves it, so you may mark it restricted. Keep hosted models at confidential or below unless your contract with the vendor says otherwise. Leave out the prices for a local model and its calls cost nothing.

Anthropic, OpenAI and Azure answer plans through native tool calling. OpenAI-compatible servers and Gemini are asked for JSON and their text is read as a plan. Every plan is checked before a single step runs, whichever way it arrived.

Checking with ok doctor#

ok doctor checks each provider without loading the company and without spending tokens: it asks the provider for its list of models. Run it in the folder you start the server from, so it reads the same llm.yaml.

sh
ok doctor

It prints where the configuration came from ("LLM configuration from llm.yaml"), one line per provider ending with the models you configured in brackets, and default model: .... It exits with 0 when no line says FAIL, 1 otherwise.

You see Meaning What to do
ok anthropic: reachable, key accepted [claude-fable-5-1, ...] Ready. Nothing.
ok house: reachable [house-model] Ready. A server that needs no key says only "reachable". Nothing.
ok house: reachable; not offered to this key: house-smal (check the model ids) [house-model, house-smal] The provider answers but does not list that model, usually a typo. Calls to it will fail, although the line says ok. Fix the id.
FAIL provider anthropic has no API key: ANTHROPIC_API_KEY is not set (export it, or set api_key in llm.yaml) The file names a key that is not set. The server skips this provider. Export the key, or remove the provider.
FAIL anthropic rejected the API key (HTTP 401: authentication_error: ...); check ANTHROPIC_API_KEY The key is wrong, revoked or for another account. With api_key in the file it ends "check api_key in llm.yaml". Replace the key.
FAIL local: cannot reach http://127.0.0.1:8896/v1: ... Connection refused (os error 61) The server is not running or the address is wrong. Start it or fix base_url.
FAIL no model provider is configured: set ANTHROPIC_API_KEY, ... Nothing to check. Set a key or write llm.yaml.

When a key is missing or rejected#

Errors name the provider and never contain your full key.

Situation ok serve / ok demo serve A run that needs a model
No provider at all ok serve warns "no model provider is configured: ..." at start. ok demo serve uses the stub planner instead. Fails with "llm: no model provider is configured: ..." (outside the demo).
A provider in llm.yaml without its key Warns at start: "warning: provider anthropic has no API key: ANTHROPIC_API_KEY is not set (...); its models are off". The others work. Goes to the providers that are left.
A key that is set but refused Starts normally and warns: "warning: house rejected the API key (HTTP 401: ...); check api_key in llm.yaml". Its models stay listed and may still be chosen. Fails with "llm: house rejected the API key (HTTP 401: ...); check api_key in llm.yaml".

The server checks every key once at start. Set OK_LLM_CHECK=0 to skip that check. The one-shot commands (ok ask, ok approve, ok answer, ok tick) do not check keys at start; they report a refused key on the first model call, for example run ... failed: llm: house rejected the API key (...). ok ask still exits with 0 when the run fails, so check its output in scripts.

A refused key, a rate limit that outlasts the retries, or a timeout fails the run. It shows as Failed with the message in the conversation's run card and on the Work page.

Retries and timeouts#

Each model call is retried when retrying can help.

Setting Default Change with
Attempts per call, the first included 4 OK_LLM_MAX_ATTEMPTS (at most 10)
Time one attempt may take 180 s OK_LLM_ATTEMPT_TIMEOUT_SECS
Time the whole call may take, retries included 300 s OK_LLM_DEADLINE_SECS
Longest retry-after honored 60 s
  • Retried: answers 429 (rate limit), 408, 409 and any 5xx (Anthropic's 529 overloaded among them), dropped connections and timeouts. The wait doubles from half a second up to 8 seconds, with some randomness, unless the provider says how long to wait.
  • Not retried: 401 and 403 (a refused key) and other 4xx answers.
  • When retries run out, the run fails with what happened, for example "llm: provider error: house HTTP 429: rate_limit: slow down (4 attempts)". A call that runs past its deadline says "no answer within the 300s deadline".

Which model may see which data#

Every collection has a data class: a label for how sensitive its data is, shown on the Brain page next to each collection, for example "Kept here · Customer PII" (see The brain). Each class has a sensitivity, from least to most sensitive:

Sensitivity Examples in the demo (Brain label, then the class label)
Public None in the demo (public)
Internal Internal (internal): Campaigns, Feature requests and most company records
Confidential Financial (financial): Forecasts; Confidential (confidential): Incidents, Investor updates
Restricted Customer PII (pii.customer): Tickets, Contacts; Candidate PII (pii.candidate): Candidates; Captured content (captured): Context items

A model call carries the most sensitive class of everything in it: the skill's collections, what the run has read, and any conversation history. The call goes to a model only if that model's max_sensitivity covers it.

How the model is chosen:

  1. The company's default model, if it may see this class.
  2. Otherwise the first model, in the order of your configuration, that may.
  3. If none may, the call is refused. It is never sent to a weaker model.

The refusal reads "no allowed model for data class pii.customer". In the demo, with only models that stop at confidential:

  • maya asking the Churn-risk watcher Northwind wants a refund for the double charge fails with "Round 1: policy denied llm.complete: no allowed model for data class pii.customer".
  • maya asking the Support Agent Summarize the open tickets fails with "Policy denied llm.complete: no allowed model for data class pii.customer".
  • pat asking the Product Agent What should we build next quarter? fails the same way.

None of these costs anything. Give a model max_sensitivity: restricted (Anthropic's from the environment, or a local model) and they reach it. With a confidential default and a second, restricted model, these asks go to the second model; the Model step's hover text names it.

To see the routing without running anything, as ada:

  1. Open Governance, tab Policy, card Try an action.
  2. Choose Who: Churn-risk watcher (agent), Acting for: Maya Patel, Action: Call a model, and type pii.customer under Data class.
  3. Choose Explain. With a model that may see restricted data you see Allowed, and the first Model line reads "The call goes to model claude-sonnet-5" (your model's name). Without one you see "Denied: no allowed model for data class pii.customer"; the later steps are marked "(after the decision)". Nothing is executed or recorded.

Type the class label, not the name the Brain page shows: pii.customer for Customer PII. Only pii. labels, financial, internal and public are read correctly today. confidential, captured, restricted and any other label are read as internal, so the answer can say a model may receive data it may not (see Not possible yet).

The Model governance card on the same tab lists each model and what it may see. It is read-only: it comes from your configuration at start.

Conversation history#

When you ask the same agent again in the same conversation, the model sees the earlier exchange as turns: your earlier asks and the agent's earlier answers ("Done: What should we build next quarter?" and its output), up to the 20 most recent, oldest first. In a loop run the model also sees its earlier rounds and their results.

  • Size. History is kept within 60,000 estimated tokens (OK_LLM_HISTORY_TOKENS changes it). When it is longer, the oldest turns are left out first; the first message says how many were left out. Estimates go by size, so compact data counts in full.
  • Data class. History counts toward the call's data class. Each of the agent's answers carries the class of what its run had read, and at least that of the collections its skills reach; an answer with no recorded class counts as restricted. If no allowed model may take the agent's earlier answers, they are left out and only your own messages go.
  • Where it applies. Only the asker's own messages and that agent's replies count. Other people's messages in a thread are not sent.
  • Starting fresh. New conversation at the top of the conversation starts an empty history.

Cost#

Each model call is charged at the price you configured, from the token counts the provider reports, rounded up to the next cent. A call to a priced model that used tokens is never recorded as free.

When a provider reports no token counts, OrchKernel estimates them from the size of the request and the answer (about four characters a token) and marks the Model step estimated.

Where you see it:

Where What it shows
A run card in a conversation, Details "5 steps · $0.05 · Open the run", then each step. The Model step reads "Planned the next step · 12,400 tokens" with its cost on the right; hover it for the model and the tokens in and out ("model house-small: 12,000 in, 400 out").
Work, tab Runs One row per run with Steps, Cost and Tokens.
Work, a run opened from that tab "Product Agent acting for Pat Kim · $0.05 · 12,400 tokens", then the trace and the output.
Directory, an agent's page, tab Scorecard Runs, Succeeded, Success rate, Cost and Tokens for the last 7 and 30 days.

For example, a call to a model priced at $3 / $15 per million tokens that reads 12,000 tokens and writes 400 costs 4.2 cents, recorded as $0.05.

Every call is also charged to the budgets in scope: the agent's token budget and money budget, the task's budget, and for a loop run its max_cost and max_tokens caps (see Agents). An agent past its token budget is refused before the call is made: "token budget exhausted (N used of M)". The stub planner costs nothing.

Not possible yet#

  • Choosing a model per agent or per skill. Every call goes to the default model when it may see the data, otherwise to the first model that may.
  • Changing models or what they may see from the screen. Edit llm.yaml and restart.
  • Adding or changing models without a restart.
  • Taking a model whose key was refused out of routing. It stays listed, and calls routed to it fail.
  • Testing routing in Try an action for the classes confidential, captured or a sensitivity level such as restricted: they are read as internal and the answer is wrong. pii.customer, financial and internal work.

For developers#

Environment#

Variable Default Purpose
OK_LLM_CONFIG ./llm.yaml The model configuration file.
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, GOOGLE_API_KEY Keys; used when there is no file.
ANTHROPIC_BASE_URL, OPENAI_BASE_URL The vendor's API Another address for a provider found from the environment.
OLLAMA_HOST, OLLAMA_MODEL llama3.1 A local Ollama server, used when there is no file.
OK_LLM_CHECK on 0 skips the key check at server start.
OK_LLM_MAX_ATTEMPTS 4 Attempts per call, at most 10.
OK_LLM_ATTEMPT_TIMEOUT_SECS 180 One attempt.
OK_LLM_DEADLINE_SECS 300 The whole call.
OK_LLM_HISTORY_TOKENS 60000 Conversation history per call, in estimated tokens.

Calls to remote providers honor HTTPS_PROXY and the other proxy variables; calls to localhost and 127.0.0.1 bypass the proxy.

API#

Method and path Returns
GET /api/llm/models { models: [model ids], governance: { allowed: [{ model, max_sensitivity, region }], default_model } }. With the demo's stub, models is empty and allowed holds mock-default.
GET /api/policy The policy, with the same governance under models.
POST /api/policy/explain { "do": { "op": "llm", "data_class": "pii.customer" }, "actor": "churn-watcher", "for": "maya" } explains a model call without making it: decision.kind and the steps, the first with check: "model".

A run's model step (kind: llm_call) carries model, input_tokens, output_tokens, cost in cents and estimated: true when the tokens were estimated. Each call is logged as an llm_called event with the model, the caller and the tokens; the Events page lists it under Runs as "Model called".

Live provider tests#

The provider tests run against a local stand-in. Two live tests call the real APIs only when asked:

sh
OK_LIVE_LLM_TEST=1 ANTHROPIC_API_KEY=... OPENAI_API_KEY=... \
  cargo test -p ok-llm --test it hardening::live_ -- --ignored