Changing a skill or playbook: the change pipeline and evals
Last updated October 5, 2026
On this page
- The pipeline at a glance
- Walkthrough: change the Support Agent's playbook
- 1. Draft the change
- 2. Propose it
- 3. Read the review card
- 4. Approve and promote
- 5. See the agent use the new version
- What evals are, and what their results mean
- Approving without evals
- When a change is approved automatically
- Rejecting a change
- Canary
- An agent proposing a change to its own playbook
- Who may propose and review what
- From the command line
- Not possible yet
- For developers
- API
- The propose_playbook step
- Events
Draft for review
Nothing that decides how an agent behaves is edited in place. A new playbook, a new skill version, a policy rule, an agent's definition, a blueprint or a company module is first proposed. Skill and playbook changes then run the skill's evals (short automatic checks, explained below), a person reviews the change, it may run as a canary on a share of the work, and finally it is promoted (it goes live) or rejected. Every step is recorded, so you can always tell who changed an agent, why, and what the checks said.
A playbook is the plain-language part of a skill: its goal, when it runs, the steps it follows, its guardrails and approval points. A skill also carries the tools it may use, its permissions, its risk tier and its evals. See Agents and Packs and skills.
This page walks through a playbook change in the demo company, explains what
eval results mean, and lists who may review which kind of change. It
describes the golive branch as it works today. What is not possible yet is
listed at the end.
The pipeline at a glance#
| State | What it means | What can happen next |
|---|---|---|
| Proposed | Filed, evals run (for skills and playbooks). Waiting for a reviewer. | Approve, Approve without evals (admin), Re-run evals, Reject |
| Reviewed | A person approved it, or it was approved automatically. Nothing is live yet. | Canary, Promote, Reject |
| Canary | The new version serves a share of the agent's tasks; the old one serves the rest. | Promote, Reject |
| Promoted | Live. For a skill or playbook, a new version is now the active one. | Nothing; propose another change to go back |
| Rejected | Closed with the reviewer's reason. | Nothing; propose a better one |
A playbook change replaces only the playbook. The skill's tools, permissions, risk tier and evals stay as they are. Evals always come from the version that is live now, so a proposal cannot make its own checks easier.
Walkthrough: change the Support Agent's playbook#
In the demo, Maya Patel leads the Support team, and the Support Agent
is the only agent that holds the Ticket support skill. That makes Maya the
lead of every agent holding the skill, so she may propose, review and promote
changes to it without an admin. Start the demo and sign in as maya with the
token ok demo serve printed (see Quickstart).
1. Draft the change#
- Open Directory and click Support Agent. A side panel opens; choose Open agent page ("Playbook, history and scorecard").
- Choose the Playbook tab. The heading reads "Ticket support ticket-support v1", then one card per section: Goal, Guardrails, Approval points, When, Inputs, Steps, Outputs, Measures, Run.
- On the Steps card, choose Edit. The Edit Steps sheet shows one
step per line, written as
name: what the agent does. Add a line at the end:Check the SLA: Before drafting, note how long the customer has waited and say so in the reply when it is over a day. - Choose Save to draft. The Steps card is outlined and marked edited, and a Draft bar appears at the top: "Steps changed", with Discard and Propose change.
Nothing has changed for the agent yet. The draft lives only in this browser tab; leaving the page drops it.
2. Propose it#
- Choose Propose change. A sheet opens: "Propose a playbook change to Ticket support", with each changed section side by side (before on the left, after on the right).
- Fill Why: "Customers who waited over a day want it acknowledged". Propose stays greyed out until Why has text.
- Choose Propose. The toast says "Proposed: proposed; it waits for review", and an Open proposals card appears under the playbook ("Proposed", "from v1", your reason) with Review in Governance.
At this moment the kernel ran the skill's evals against the new playbook and put an inbox item, "Change proposed: playbook of ticket-support (from v1) (Customers who waited over a day want it acknowledged)", in front of every reviewer: the admins (ada), and the leads of the teams whose agents hold the skill (maya). Clicking that inbox item today opens Governance, Schema tab ("No open schema proposals."), not Changes; choose the Changes tab yourself.
The sheet says "an admin's or the team lead's proposal is reviewed at once unless it weakens a guardrail or an approval point". In the demo without a model that never happens; see When a change is approved automatically.
If you remove or shorten a guardrail or an approval point, the sheet shows a red Weakens a guardrail: guardrails (or : approval points) badge before you propose. The toast then ends "(weakens a guardrail: guardrails removed or shortened)", and the review card carries "Weakens a guardrail: guardrails removed or shortened". Such a change is never approved automatically: a person must review it.
3. Read the review card#
Open Governance, then the Changes tab (or press Review in Governance). Changes are listed newest first. The card shows:
| Part | What it tells you |
|---|---|
| State badge and title | "Proposed", "playbook of ticket-support (from v1)": the skill and the version it starts from |
| Reason and proposer | The Why text, and "proposed by Maya Patel" (or an agent) |
| Warnings | Red badges such as "Weakens a guardrail: guardrails removed or shortened" |
| Evidence | For an agent's proposal, "evidence:" and links to the runs it points at |
| Playbook diff | Each changed section, before and after |
| Evals | A line of counts and one chip per eval case |
| Reviews | Who approved or rejected, with their note |
| content | The whole proposed playbook as JSON, folded |
| Actions | Approve, Approve without evals (admins only), Re-run evals, Reject; once reviewed, Promote and Reject |
In the demo the evals line reads "Evals: 1 passed (by the demo's stand-in, not a model)", with one chip, "support-reads-the-ticket-before-replying · passed (stand-in)". Press the chip to open it:
support-reads-the-ticket-before-replying · passed (stand-in) Asked: "Resolve ticket t-1001" Expected: the first step reads Tickets. The demo's scripted planner answered this, not a model, so the pass tests nothing about the change.
Details under it shows Expected (the ops the case asks for) and Got (the plan the planner actually produced).
The demo runs without a model unless you set a provider key, so its scripted planner (the stub planner) answers some eval cases. Its passes are marked passed (stand-in). They let you walk the pipeline but say nothing about whether the change is good. With a real model the chip says passed or failed.
Everyone who can open Governance sees these cards and their buttons. Someone who may not review, such as tom, finds out only after choosing Approve: the toast says "policy denied change.review: only an admin or the lead of every agent holding ticket-support may change.review".
4. Approve and promote#
- As maya, choose Approve. A sheet "Approve this change" asks for an optional Note ("Read the diff; one added step"). Choose Approve. The toast says "Reviewed". After a moment the card's badge changes to Reviewed, "Maya Patel approved: Read the diff; one added step" appears under the evals, and the actions become Promote and Reject.
- Choose Promote. The toast says "Promoted" and the card shows Promoted.
To try a canary first, see Canary below.
5. See the agent use the new version#
- On the Support Agent's page, the History tab lists
ticket-supportv2 Active, proposed by Maya Patel, with her reason, and v1 (proposed by "installed") with Propose rollback. - The Overview tab lists the skill as Ticket support v2.
- In a second terminal, in the folder where you ran
ok demo,ok playbook show ticket-support --state orchkernel-demo/state.dbprints"version": 2and, underprompt, the exact text the planner is given. Its Steps list now includes "6. Check the SLA: ...".
Every run of the agent from now on, single-step or loop, is planned from that prompt. With a real model you see the effect in its replies. With the demo's stub planner you do not: it answers the walkthrough's asks the same way whatever the playbook says.
Rolling back is a change like any other: Propose rollback on an older version files a proposal with that version's playbook. The toast says "Rollback proposed (proposed)", and the card on the Changes tab reads "playbook of ticket-support (from v2)", "Roll back to v1". It goes through evals and review again.
What evals are, and what their results mean#
Each skill carries eval cases. A case asks the skill something, for example
"Resolve ticket t-1001", and says what the plan must start with, for example
"the first step reads Tickets". When a skill or playbook change is proposed,
the kernel builds the same prompt a real run would get from the proposed
playbook, asks the model, and checks the first steps of its plan. By default
the expected steps must come first and in order; a case marked contains
passes when each expected step appears anywhere in the first round.
Every playbook needs at least one eval case. A skill with none cannot take a playbook change. In the demo, the Curator agent's Playbook tab says "This skill has no eval case; a playbook change needs at least one, so an admin must add evals through a skill proposal first", and proposing anyway fails with "skill context-curation has no eval cases; a playbook needs at least one".
| Result | What it means | What you can do |
|---|---|---|
| passed | The model planned what the case expects. | Approve. |
| passed (stand-in) | The demo's stub planner answered, not a model. | Approve if you have read the diff; it tests nothing. |
| failed | The model answered and the plan did not match, or the answer could not be read. | Nobody can approve it, not even an admin. Reject it, or propose a better change. |
| not run | The case could not run: "no model configured", a model error, a rate limit or a budget refusal. It says nothing about the change. | Re-run the evals once a model answers, or an admin approves without evals. |
When an eval did not pass, Approve is greyed out and the card says why, for example: "2 evals did not run (a-morning-run-starts-from-resolved-tickets, a-coverage-question-reads-the-knowledge-base). Re-run them once a model answers; an admin may approve without evals." A not run chip opens to "Not run: no model configured" and what to do about it.
Re-run evals runs them again and stores the results. The toast gives the counts, and why any are still not run: "Evals: 0 passed, 0 failed, 2 not run. 2 still not run: no model configured". A re-run that cannot run a case never turns an earlier failure into "not run": a failure stays a failure until a re-run passes it. Re-running is possible only while the change is still proposed.
Approving without evals#
When every eval that did not pass is not run (none failed), an admin sees Approve without evals. Try it in the demo with the Knowledge-base writer, whose evals the stub planner does not answer:
- As maya, open the Knowledge-base writer's page (Directory, Knowledge-base writer, Open agent page), then the Playbook tab. Edit the Goal (add "Keep each article under 300 words."), choose Save to draft, Propose change, give a reason ("Shorter articles get read") and Propose. On the Changes tab its card shows "Evals: 2 not run", and Maya's Approve is greyed out. She has no Approve without evals.
- Sign out and sign in as ada. On the same card choose Approve without evals. The sheet says "2 evals did not run (...). Approving without them is recorded on the change, in its history and in the event log with your reason."
- Fill Why may it go ahead without its evals? (required; the button stays greyed out while it is empty), for example "No model in the demo; I read the diff: only the goal text changed", and choose Approve without evals. The toast says "Approved without evals".
- The card shows a yellow Approved without evals line: "by Ada Park: No model in the demo; ... (not run: a-morning-run-starts-from-resolved-tickets, a-coverage-question-reads-the-knowledge-base)". Choose Promote.
The agent's History tab shows the same line next to v2, and Events lists "Ada Park approved a change without evals: ... (2 not run)".
Only a human admin can do this. Team leads, agents and the automatic review never can, and a failed eval still blocks it.
When a change is approved automatically#
A playbook change is approved the moment it is proposed (its state is Reviewed straight away, so promoting is the next step, and the toast says "Proposed and reviewed: promote it from Governance") only when all of these hold:
- an admin, or the lead of every team whose agents hold the skill, proposed it;
- it weakens no guardrail or approval point;
- every eval passed, answered by a real model.
The review is recorded as "auto-reviewed: proposed by an admin or the lead of every holder". In the demo without a model nothing is auto-reviewed, because stand-in passes never count. Promoting is always a separate, deliberate step.
Rejecting a change#
Choose Reject on a proposed, reviewed or canary change. The sheet "Reject this change" requires a reason (Why?). The toast says "Rejected", the card moves to Rejected, and the reason follows the proposer ("proposed by Maya Patel · Keep the SLA step"). A canary stops at once. A promoted change cannot be rejected; propose a rollback instead.
Canary#
A canary installs the new version next to the old one and sends a share of the agent's tasks to it. Each task is routed by its id, so the same task always sees the same version. The new version is not active until you promote.
On the Changes tab, Canary 10% appears only on a reviewed skill
change, not on a playbook change. For a playbook change use the API (or the
CLI, see From the command line). As maya, with her
token in $TOKEN and the change's id from GET /api/changes:
curl -X POST -H "Authorization: Bearer $TOKEN" -H 'content-type: application/json' \
http://127.0.0.1:8080/api/changes/<id>/canary -d '{"percent": 10}'
The card then shows Canary and 10% canary, with Promote and Reject. While the canary runs, the agent's History tab lists the new version (v3 in the walkthrough) as proposed by "installed", with Propose rollback, and it stays listed that way after you reject the canary. The kernel does not count canary runs or roll a canary back by itself today: watch the agent's runs and scorecard, then promote or reject.
An agent proposing a change to its own playbook#
An agent can suggest an improvement to its own playbook at the end of a run,
with the runs that show why. It does this with the propose_playbook step,
as the last step of its plan.
- It may do this only for a skill that it alone holds.
- It must give a reason. The run it is in is always attached as evidence; other runs it cites must be its own successful runs from the last 7 days.
- It may have one open proposal per skill at a time, and propose at most once a day per skill.
- The proposal replaces only the playbook, and it is never approved automatically: a person always reviews it.
On the Changes tab the card reads "proposed by Support Agent" and
"evidence: 4411e9d1", a link to the run on the Work page. Review it like any
other change. The run's result is {"proposal": "...", "state": "proposed"}.
A second proposal while the first is open fails the run with "op 0:
proposal_open: support-agent already has an open proposal for
ticket-support".
A model decides when to propose. The demo's stub planner never does, so to see one in the demo send an explicit plan (see For developers).
Who may propose and review what#
| Change | Who may propose | Who may review, canary and promote |
|---|---|---|
| Playbook | An admin; a team lead whose team holds the skill; an agent that alone holds the skill, from a run | An admin, or the lead of every team whose agents hold the skill. Never an agent. |
| Skill (a new version: tools, permissions, evals) | An admin. A lead may file one that changes only what a playbook could. | Same as a playbook |
| Policy rules, or a policy bundle from the repository | Rules: any person, through the API. A bundle: see Policy | A human admin who neither proposed nor wrote it. If there is only one active admin, they may review their own, and the change is marked self-reviewed. Rejecting: any admin, or the proposer. |
| Company module | See Modules | A human admin who did not propose it |
| Blueprint that adds company modules to the catalog | See Setting up a company | A human admin who did not propose it |
| Agent definition, or a blueprint that adds no modules | Any person, through the API | Not restricted today: see the note below |
What this means in the demo:
- Maya reviews and promotes the playbook changes of the Support team's agents: Support Agent, Triage, Knowledge-base writer, QA reviewer and Churn-risk watcher. In the demo each skill is held by one agent only.
- Tom (support agent, not a lead) sees the cards and their buttons, but Approve answers "policy denied change.review: only an admin or the lead of every agent holding ticket-support may change.review".
- A lead whose team holds a skill that another team's agents also hold may propose, but the change waits for an admin. No skill is shared that way in the demo.
- Pat may propose a policy rule change, but neither Pat nor Maya may review it ("only a human admin who neither proposed nor wrote a policy change may change.review it"); only Ada can. Ada may review her own rule proposals only because she is the demo's only admin. Pat may reject his own proposal.
Note: agent definition changes are not yet held to these rules. In the
demo, Tom (not an admin, not a lead) can propose a change to the Support
Agent's definition through the API, approve it himself and promote it, and
it goes live. Until this is fixed, check Governance › Changes for
agent ... cards you did not expect. The agent's actions still pass through
the gate and policy (a refund still waits for an admin).
From the command line#
ok playbook show support-agent # each skill's playbook, version, eval count and prompt
ok playbook show ticket-support # one skill
ok playbook propose ticket-support playbook.json --reason "Acknowledge long waits" --as maya
ok changes list # id, state, summary, number of evals, reason
ok changes review <id> --note "Read the diff" --as maya
ok changes review <id> --reject --note "Put it in the Draft step" --as maya
ok changes canary <id> --percent 10 --as maya
ok changes promote <id> --as maya
These commands work directly on a state file, not through a running server.
Pass it with --state (or OK_STATE); for the demo that is
--state orchkernel-demo/state.db. ok playbook show is safe while the
server runs. Make changes only with the server stopped, or use the screens or
the API instead.
ok playbook propose takes the whole playbook as YAML or JSON, and either a
skill id or an agent that holds exactly one skill. It prints the proposal.
From the CLI the demo's stub planner does not answer, so every eval of a CLI
proposal is not run ("no model configured"). ok changes review then
refuses with "evals not run: ...; re-run them once a model answers, or an
admin may approve without evals and say why", and the CLI can do neither. A
change proposed in the screens (whose evals passed) can be reviewed, canaried
and promoted from the CLI.
ok changes list is ordered by id, not by date.
Not possible yet#
- A canary from the Changes tab for a playbook change (only for a skill change); use the API or CLI.
- Counting canary runs, or rolling a canary back automatically when it fails.
- Seeing which version of a skill a given run used: runs do not record it.
- Proposing a whole skill version (new tools or eval cases) from the screen; it is an API call by an admin.
- From the CLI: re-running evals, approving without evals, or rejecting a change that is already reviewed or on canary.
ok changes listshows how many evals a change has, not their results.- Editing a proposal: reject it and propose again.
For developers#
API#
All under /api with Authorization: Bearer <token>.
| Method and path | Body | Notes |
|---|---|---|
GET /changes |
Every proposal the caller may see, newest first, paged. Each has kind, state, evals, reviews, warnings, evidence_runs, approved_without_evals, and for a canaried or promoted skill or playbook change installed_version. |
|
POST /changes |
{ kind, rationale, ... } |
kind: skill (manifest), agent (actor), rules (rules), blueprint, module, playbook. A policy bundle goes through POST /policy/proposals. |
GET /agents/:id/playbook |
The agent, each skill with its playbook and version, the diff from a fork's original, and the version history. | |
POST /agents/:id/playbook/propose |
{ skill?, playbook, reason, evidence? } |
skill may be left out when the agent holds one. |
POST /agents/:id/playbook/convert |
{ skill? } |
Turns plain instructions into a playbook proposal (409 has_playbook if it has one). The Playbook tab's Convert to a playbook button calls it. |
GET /changes/:id/diff |
Section-by-section diff; for a skill proposal, one eval:<name> entry per eval case added, changed or removed, with weakens. |
|
POST /changes/:id/review |
{ approve, note, without_evals? } |
without_evals: true with approve: true and a non-empty note is the admin's approval without evals. |
POST /changes/:id/evals |
Re-run evals; only while proposed. | |
POST /changes/:id/canary |
{ percent } |
A Policy proposal starts shadow instead. |
POST /changes/:id/promote |
||
POST /changes/:id/reject |
{ note } |
Errors you will meet:
| Status and code | When |
|---|---|
403 denied |
Someone who may not propose, review, canary or promote. |
403 not_held |
An agent proposing for a skill it does not alone hold. |
403 bad_evidence |
An agent citing a run that is not its own successful run from the last 7 days. |
403 same_person |
The proposer or author reviewing a policy change while another admin exists. |
409 evals_failing |
Approving, canarying or promoting with a failed eval. |
409 evals_not_run |
A plain approval while an eval did not run. |
409 nothing_to_waive |
Approving without evals when every eval passed. |
409 playbook_only |
A non-admin's skill proposal that changes tools, permissions, risk, code, inputs, outputs or evals. |
409 proposal_open |
An agent's second open proposal for the skill. |
422 evals_required |
A playbook change for a skill with no eval cases. |
422 reason_required |
Approving without evals with no reason. |
422 plan_eval, vacuous_eval |
A skill proposal adding an eval case that carries its own plan, or expects nothing. |
429 rate_limited |
An agent's second proposal for the skill within 24 hours. |
The propose_playbook step#
An agent's plan proposes with the propose_playbook op, allowed only as the
last op of a single-mode plan or in a loop's final done reply:
{ "op": "propose_playbook", "skill": "ticket-support",
"playbook": { "goal": "...", "when": { "prose": "...", "triggers": [] },
"inputs": [], "steps": [{ "name": "...", "do": "..." }],
"guardrails": [], "approval_points": [], "outputs": [],
"measures": [], "run": { "mode": "single" } },
"reason": "Replies that state the wait get fewer follow-ups",
"evidence": [] }
To try it in the demo, send it as an explicit plan: POST /api/ask with
{ "agent": "support-agent", "text": "Suggest a playbook improvement", "new_thread": true, "plan": { "ops": [ <the op> ] } } as maya, using the
current playbook from GET /api/agents/support-agent/playbook with one
change. In the screens, open a conversation with the Support Agent at
/threads/new?agent=support-agent&dev=1 and use Explicit plan.
Events#
| Event | When |
|---|---|
change_proposed |
A proposal is filed (kind, summary); "Change proposed" on the Events page |
change_evals_run |
Evals ran, at proposal or on a re-run: counts and each case's name and status ("Evals run by Maya Patel: 1 passed, 0 failed, 0 not run") |
change_state_changed |
Reviewed, canary, promoted or rejected ("Change moved on" on the Events page) |
change_approved_without_evals |
An admin approved without evals: reason and the cases not run |