Models
Last updated October 5, 2026
On this page
- Where OrchKernel looks for models
- One key: the quickest start
- Switching the demo to a real model
- Writing llm.yaml
- OpenAI-compatible and local servers
- Checking with ok doctor
- When a key is missing or rejected
- Retries and timeouts
- Which model may see which data
- Conversation history
- Cost
- Not possible yet
- For developers
- Environment
- API
- Live provider tests
Agents plan with a language model. OrchKernel does not ship one: you connect
your own provider with a key in the environment or a file called llm.yaml.
For each model you set the most sensitive data it may see and what it costs.
OrchKernel sends a call only to a model allowed to see the data in it, and
charges the call to the run's budgets.
This page covers connecting a model, checking it with ok doctor, what goes
wrong and how it looks, how calls are routed by data class, conversation
history, and cost. What is not possible yet is listed
at the end.
Where OrchKernel looks for models#
At start, OrchKernel takes the first of these that exists:
| Order | Source | Notes |
|---|---|---|
| 1 | The file OK_LLM_CONFIG names |
A path relative to the folder you start from, or absolute. If the file does not exist, OrchKernel stops with "No such file or directory" instead of moving on. |
| 2 | llm.yaml in the folder you start from |
For the demo, that is the folder you run ok demo serve in, not the demo's own folder (orchkernel-demo). |
| 3 | Provider keys in the environment | ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY (or GOOGLE_API_KEY), OLLAMA_HOST. |
When a file is found, the environment keys are not picked up on their own;
the file names each provider's key itself (api_key_env). A variable that is
set but empty counts as not set.
Models are read once, at start. To change them, stop the server and start it again.
One key: the quickest start#
Set one variable and OrchKernel offers that provider's current models at list prices. Prices are in US dollars per million tokens, input / output.
| Variable | Models offered | May see up to | Price |
|---|---|---|---|
ANTHROPIC_API_KEY |
claude-sonnet-5 (default) |
Restricted | $2 / $10 |
claude-opus-5-5 |
Restricted | $4 / $20 | |
claude-fable-5-1 |
Restricted | $10 / $50 | |
claude-haiku-4-5-20251001 |
Confidential | $1 / $5 | |
OPENAI_API_KEY |
gpt-5 |
Confidential | $1.25 / $10 |
gpt-5-mini |
Confidential | $0.25 / $2 | |
GEMINI_API_KEY or GOOGLE_API_KEY |
gemini-2.5-pro |
Confidential | $1.25 / $10 |
gemini-2.5-flash |
Confidential | $0.30 / $2.50 | |
OLLAMA_HOST (for example http://localhost:11434) |
OLLAMA_MODEL, or llama3.1 |
Restricted | Free |
The default model is claude-sonnet-5 when the Anthropic key is set.
Otherwise it is the first model of the first provider found, in the order of
the table: gpt-5 with only an OpenAI key, gemini-2.5-pro with only a
Gemini key. ANTHROPIC_BASE_URL and OPENAI_BASE_URL point those providers
at another address, such as a proxy.
"May see up to" is the model's sensitivity cap, explained in Which model may see which data. In short: with only an OpenAI or Gemini key, no model may see restricted data such as the demo's customer tickets, and asks that read them are refused. In the demo that includes the Support Agent, the Churn-risk watcher and the Product Agent, whose skills all read tickets.
Switching the demo to a real model#
The demo starts with a stub planner when no model is configured: fixed
plans for the walkthrough's asks and a note for anything else (see
Asking an agent). Before
you switch, the start banner reads
Model none configured: a stub planner answers the walkthrough's asks.
- Stop the demo with Ctrl+C.
- Set a key in the same terminal, for example
export ANTHROPIC_API_KEY=sk-ant-..., or put anllm.yamlin the folder you start from (see Writing llm.yaml). - Check it: run
ok doctorin the same folder. You should see a line startingokfor each provider and thendefault model: claude-sonnet-5(or your file's default). See Checking with ok doctor. - Start the demo again with
ok demo serve. Before the banner, the terminal prints one line per provider from a key check, such asllm: house: reachable, or awarning:line if a key is missing or refused. In the banner, the Model line names the default model and where it came from, for exampleModel gpt-5, configured from the environment (ANTHROPIC_API_KEY, ...)orModel house-model, configured from llm.yaml. - Sign in with the new tokens (each start revokes the old ones). As ada,
open Governance, tab Policy. The Model governance card shows
default: claude-sonnet-5 · configured providers serve 4 model(s)(your default and count) and a table with each Model and May see up to. With the stub it showsdefault: mock-default · configured providers serve 0 model(s)and one row,mock-default, Restricted. - As pat (product manager), open Threads, click Product Agent under
Talk to an agent and ask
What should we build next quarter?. With the stub you get the note "No model is configured, so the demo's stub planner answered. ..." With a model allowed to see restricted data you get the model's answer and a green Done badge. Open Details on the run card: the Model step shows the tokens and the cost (see Cost). With a model capped at confidential the run fails instead, with "policy denied llm.complete: no allowed model for data class pii.customer".
Writing llm.yaml#
Copy llm.example.yaml from the repository to llm.yaml and keep the
providers you want. A provider whose key is not set is skipped at start with a
warning, so the copied file works with only one key set.
providers:
- name: anthropic
kind: anthropic
api_key_env: ANTHROPIC_API_KEY
models:
- { id: claude-sonnet-5, max_sensitivity: restricted, input_per_mtok_cents: 200, output_per_mtok_cents: 1000, default: true }
- { id: claude-haiku-4-5-20251001, max_sensitivity: confidential, input_per_mtok_cents: 100, output_per_mtok_cents: 500 }
Provider settings:
| Field | What it does |
|---|---|
name |
Your name for the provider. It appears in messages ("house rejected the API key"). |
kind |
anthropic, openai, gemini, azure_openai or openai_compatible. |
api_key_env |
The environment variable holding the key. Keeps the key out of the file. |
api_key |
The key itself. Prefer api_key_env. |
base_url |
The server's address. Needed for azure_openai and for most openai_compatible servers. |
api_version |
Azure only, for example "2024-10-21". |
models |
The models below. |
Model settings:
| Field | What it does | If left out |
|---|---|---|
id |
The model id the provider knows. For Azure, the deployment name. | Required |
max_sensitivity |
The most sensitive data this model may see: public, internal, confidential or restricted. |
internal |
input_per_mtok_cents, output_per_mtok_cents |
Price in cents per million tokens. 200 means $2.00. | 0: calls cost nothing |
default: true |
The company's default model. | The first model in the file |
Set max_sensitivity on purpose. Left out, a model may see internal data
only, and every ask that reads customer data is refused.
OpenAI-compatible and local servers#
kind: openai_compatible covers any server that speaks the OpenAI chat
completions API: Ollama, vLLM, LM Studio, Groq, Mistral, DeepSeek, Together,
OpenRouter. It needs no key unless the server asks for one.
- name: ollama
kind: openai_compatible
base_url: http://localhost:11434/v1
models:
- { id: llama3.1, max_sensitivity: restricted }
A model on your own machine or network is the usual place for restricted
data: nothing leaves it, so you may mark it restricted. Keep hosted models
at confidential or below unless your contract with the vendor says
otherwise. Leave out the prices for a local model and its calls cost nothing.
Anthropic, OpenAI and Azure answer plans through native tool calling. OpenAI-compatible servers and Gemini are asked for JSON and their text is read as a plan. Every plan is checked before a single step runs, whichever way it arrived.
Checking with ok doctor#
ok doctor checks each provider without loading the company and without
spending tokens: it asks the provider for its list of models. Run it in the
folder you start the server from, so it reads the same llm.yaml.
ok doctor
It prints where the configuration came from ("LLM configuration from
llm.yaml"), one line per provider ending with the models you configured in
brackets, and default model: .... It exits with 0 when no line says
FAIL, 1 otherwise.
| You see | Meaning | What to do |
|---|---|---|
ok anthropic: reachable, key accepted [claude-fable-5-1, ...] |
Ready. | Nothing. |
ok house: reachable [house-model] |
Ready. A server that needs no key says only "reachable". | Nothing. |
ok house: reachable; not offered to this key: house-smal (check the model ids) [house-model, house-smal] |
The provider answers but does not list that model, usually a typo. Calls to it will fail, although the line says ok. |
Fix the id. |
FAIL provider anthropic has no API key: ANTHROPIC_API_KEY is not set (export it, or set api_key in llm.yaml) |
The file names a key that is not set. The server skips this provider. | Export the key, or remove the provider. |
FAIL anthropic rejected the API key (HTTP 401: authentication_error: ...); check ANTHROPIC_API_KEY |
The key is wrong, revoked or for another account. With api_key in the file it ends "check api_key in llm.yaml". |
Replace the key. |
FAIL local: cannot reach http://127.0.0.1:8896/v1: ... Connection refused (os error 61) |
The server is not running or the address is wrong. | Start it or fix base_url. |
FAIL no model provider is configured: set ANTHROPIC_API_KEY, ... |
Nothing to check. | Set a key or write llm.yaml. |
When a key is missing or rejected#
Errors name the provider and never contain your full key.
| Situation | ok serve / ok demo serve |
A run that needs a model |
|---|---|---|
| No provider at all | ok serve warns "no model provider is configured: ..." at start. ok demo serve uses the stub planner instead. |
Fails with "llm: no model provider is configured: ..." (outside the demo). |
A provider in llm.yaml without its key |
Warns at start: "warning: provider anthropic has no API key: ANTHROPIC_API_KEY is not set (...); its models are off". The others work. | Goes to the providers that are left. |
| A key that is set but refused | Starts normally and warns: "warning: house rejected the API key (HTTP 401: ...); check api_key in llm.yaml". Its models stay listed and may still be chosen. | Fails with "llm: house rejected the API key (HTTP 401: ...); check api_key in llm.yaml". |
The server checks every key once at start. Set OK_LLM_CHECK=0 to skip that
check. The one-shot commands (ok ask, ok approve, ok answer,
ok tick) do not check keys at start; they report a refused key on the first
model call, for example run ... failed: llm: house rejected the API key (...). ok ask still exits with 0 when the run fails, so check its output
in scripts.
A refused key, a rate limit that outlasts the retries, or a timeout fails the run. It shows as Failed with the message in the conversation's run card and on the Work page.
Retries and timeouts#
Each model call is retried when retrying can help.
| Setting | Default | Change with |
|---|---|---|
| Attempts per call, the first included | 4 | OK_LLM_MAX_ATTEMPTS (at most 10) |
| Time one attempt may take | 180 s | OK_LLM_ATTEMPT_TIMEOUT_SECS |
| Time the whole call may take, retries included | 300 s | OK_LLM_DEADLINE_SECS |
Longest retry-after honored |
60 s |
- Retried: answers 429 (rate limit), 408, 409 and any 5xx (Anthropic's 529 overloaded among them), dropped connections and timeouts. The wait doubles from half a second up to 8 seconds, with some randomness, unless the provider says how long to wait.
- Not retried: 401 and 403 (a refused key) and other 4xx answers.
- When retries run out, the run fails with what happened, for example "llm: provider error: house HTTP 429: rate_limit: slow down (4 attempts)". A call that runs past its deadline says "no answer within the 300s deadline".
Which model may see which data#
Every collection has a data class: a label for how sensitive its data is, shown on the Brain page next to each collection, for example "Kept here · Customer PII" (see The brain). Each class has a sensitivity, from least to most sensitive:
| Sensitivity | Examples in the demo (Brain label, then the class label) |
|---|---|
| Public | None in the demo (public) |
| Internal | Internal (internal): Campaigns, Feature requests and most company records |
| Confidential | Financial (financial): Forecasts; Confidential (confidential): Incidents, Investor updates |
| Restricted | Customer PII (pii.customer): Tickets, Contacts; Candidate PII (pii.candidate): Candidates; Captured content (captured): Context items |
A model call carries the most sensitive class of everything in it: the
skill's collections, what the run has read, and any conversation history.
The call goes to a model only if that model's max_sensitivity covers it.
How the model is chosen:
- The company's default model, if it may see this class.
- Otherwise the first model, in the order of your configuration, that may.
- If none may, the call is refused. It is never sent to a weaker model.
The refusal reads "no allowed model for data class pii.customer". In the demo, with only models that stop at confidential:
- maya asking the Churn-risk watcher
Northwind wants a refund for the double chargefails with "Round 1: policy denied llm.complete: no allowed model for data class pii.customer". - maya asking the Support Agent
Summarize the open ticketsfails with "Policy denied llm.complete: no allowed model for data class pii.customer". - pat asking the Product Agent
What should we build next quarter?fails the same way.
None of these costs anything. Give a model max_sensitivity: restricted
(Anthropic's from the environment, or a local model) and they reach it. With
a confidential default and a second, restricted model, these asks go to the
second model; the Model step's hover text names it.
To see the routing without running anything, as ada:
- Open Governance, tab Policy, card Try an action.
- Choose Who: Churn-risk watcher (agent), Acting for: Maya Patel,
Action: Call a model, and type
pii.customerunder Data class. - Choose Explain. With a model that may see restricted data you see Allowed, and the first Model line reads "The call goes to model claude-sonnet-5" (your model's name). Without one you see "Denied: no allowed model for data class pii.customer"; the later steps are marked "(after the decision)". Nothing is executed or recorded.
Type the class label, not the name the Brain page shows: pii.customer for
Customer PII. Only pii. labels, financial, internal and public are
read correctly today. confidential, captured, restricted and any other
label are read as internal, so the answer can say a model may receive data it
may not (see Not possible yet).
The Model governance card on the same tab lists each model and what it may see. It is read-only: it comes from your configuration at start.
Conversation history#
When you ask the same agent again in the same conversation, the model sees the earlier exchange as turns: your earlier asks and the agent's earlier answers ("Done: What should we build next quarter?" and its output), up to the 20 most recent, oldest first. In a loop run the model also sees its earlier rounds and their results.
- Size. History is kept within 60,000 estimated tokens
(
OK_LLM_HISTORY_TOKENSchanges it). When it is longer, the oldest turns are left out first; the first message says how many were left out. Estimates go by size, so compact data counts in full. - Data class. History counts toward the call's data class. Each of the agent's answers carries the class of what its run had read, and at least that of the collections its skills reach; an answer with no recorded class counts as restricted. If no allowed model may take the agent's earlier answers, they are left out and only your own messages go.
- Where it applies. Only the asker's own messages and that agent's replies count. Other people's messages in a thread are not sent.
- Starting fresh. New conversation at the top of the conversation starts an empty history.
Cost#
Each model call is charged at the price you configured, from the token counts the provider reports, rounded up to the next cent. A call to a priced model that used tokens is never recorded as free.
When a provider reports no token counts, OrchKernel estimates them from the size of the request and the answer (about four characters a token) and marks the Model step estimated.
Where you see it:
| Where | What it shows |
|---|---|
| A run card in a conversation, Details | "5 steps · $0.05 · Open the run", then each step. The Model step reads "Planned the next step · 12,400 tokens" with its cost on the right; hover it for the model and the tokens in and out ("model house-small: 12,000 in, 400 out"). |
| Work, tab Runs | One row per run with Steps, Cost and Tokens. |
| Work, a run opened from that tab | "Product Agent acting for Pat Kim · $0.05 · 12,400 tokens", then the trace and the output. |
| Directory, an agent's page, tab Scorecard | Runs, Succeeded, Success rate, Cost and Tokens for the last 7 and 30 days. |
For example, a call to a model priced at $3 / $15 per million tokens that reads 12,000 tokens and writes 400 costs 4.2 cents, recorded as $0.05.
Every call is also charged to the budgets in scope: the agent's token budget
and money budget, the task's budget, and for a loop run its max_cost and
max_tokens caps (see Agents). An agent past its token budget
is refused before the call is made: "token budget exhausted (N used of M)".
The stub planner costs nothing.
Not possible yet#
- Choosing a model per agent or per skill. Every call goes to the default model when it may see the data, otherwise to the first model that may.
- Changing models or what they may see from the screen. Edit
llm.yamland restart. - Adding or changing models without a restart.
- Taking a model whose key was refused out of routing. It stays listed, and calls routed to it fail.
- Testing routing in Try an action for the classes
confidential,capturedor a sensitivity level such asrestricted: they are read as internal and the answer is wrong.pii.customer,financialandinternalwork.
For developers#
Environment#
| Variable | Default | Purpose |
|---|---|---|
OK_LLM_CONFIG |
./llm.yaml |
The model configuration file. |
ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY, GOOGLE_API_KEY |
Keys; used when there is no file. | |
ANTHROPIC_BASE_URL, OPENAI_BASE_URL |
The vendor's API | Another address for a provider found from the environment. |
OLLAMA_HOST, OLLAMA_MODEL |
llama3.1 |
A local Ollama server, used when there is no file. |
OK_LLM_CHECK |
on | 0 skips the key check at server start. |
OK_LLM_MAX_ATTEMPTS |
4 | Attempts per call, at most 10. |
OK_LLM_ATTEMPT_TIMEOUT_SECS |
180 | One attempt. |
OK_LLM_DEADLINE_SECS |
300 | The whole call. |
OK_LLM_HISTORY_TOKENS |
60000 | Conversation history per call, in estimated tokens. |
Calls to remote providers honor HTTPS_PROXY and the other proxy variables;
calls to localhost and 127.0.0.1 bypass the proxy.
API#
| Method and path | Returns |
|---|---|
GET /api/llm/models |
{ models: [model ids], governance: { allowed: [{ model, max_sensitivity, region }], default_model } }. With the demo's stub, models is empty and allowed holds mock-default. |
GET /api/policy |
The policy, with the same governance under models. |
POST /api/policy/explain |
{ "do": { "op": "llm", "data_class": "pii.customer" }, "actor": "churn-watcher", "for": "maya" } explains a model call without making it: decision.kind and the steps, the first with check: "model". |
A run's model step (kind: llm_call) carries model, input_tokens,
output_tokens, cost in cents and estimated: true when the tokens were
estimated. Each call is logged as an llm_called event with the model, the
caller and the tokens; the Events page lists it under Runs as "Model
called".
Live provider tests#
The provider tests run against a local stand-in. Two live tests call the real APIs only when asked:
OK_LIVE_LLM_TEST=1 ANTHROPIC_API_KEY=... OPENAI_API_KEY=... \
cargo test -p ok-llm --test it hardening::live_ -- --ignored