AI-native playbook · B2B SaaS

How to make a B2B SaaS company AI-native: a playbook

The agents SaaS companies sell are pushing down the seat revenue they live on, and the same agents can do much of the company's own support, sales and engineering work. Doing that safely means knowing who each agent acts for, what it may touch, and being able to show your customers' security teams what it did. This playbook lays out a staged path, with sources for every figure.

Last reviewed
October 2026
Written for
Founders, COOs and functional leaders at SaaS companies of 50 to 1,000 people
Reading time
About 30 minutes
On this page

Your help desk vendor now sells the work your support seats used to do. Intercom charges $0.99 for each ticket its agent resolves[32], Zendesk $1.50 to $2.00[34], Salesforce $2 a conversation[35]. A customer whose agent closes more of its tickets buys fewer support seats at renewal.

Investors noticed. TechCrunch puts the software and services value lost in the early-February 2026 sell-off at close to $1 trillion, tied to AI coding and agent launches and pressure on seat-based renewals[46]. US software publishers employed 652,900 people in August 2026, below the November 2022 peak of 664,300[1]. Headcount stopped rising, but the cost base is still mostly people: at the nine public SaaS companies below, sales and marketing alone takes 23% to 47% of revenue[2].

SaaS also has an audience other sectors do not. Your customers' security teams audit you. When a support agent sends ticket text to a model provider, that provider is likely a new subprocessor under GDPR, and your customers are owed notice before it starts[5]. The next questionnaire will ask where AI touches their data. Getting an agent to draft a reply or open a pull request is now easy. Saying, account by account, what your agents did, on whose authority and with whose approval, is the work this playbook is about.

01

AI-enabled versus AI-native in a SaaS company

Most SaaS companies are AI-enabled already. The difference that matters is whether anyone can answer "what did AI do to this customer's account this month, and who allowed it?"

AI-enabled
  • Copilot seats for everyone, counted as adoption
  • A help desk bot the support lead set up
  • Coding assistants with personal tokens
  • Each tool with its own keys, rules and logs
  • Nobody can say what AI did to customer X's account this month
AI-native
  • Agents do the first pass in support, sales, CS, engineering and finance
  • People own pricing, shipped code, money and what is said in the company's name
  • Every agent acts for a named person, with no more access than that person
  • Anything touching a customer's money, data or contract passes a rule or approval
  • The company answers its customers' AI questionnaire with evidence

Investors use "AI-native" for a company whose product is built on AI. ICONIQ, for one, found 47% of AI-native companies had reached scale against 13% of AI-enabled ones[30]. This page means something else: how the company itself runs.

02

The missing layer

Each system a SaaS company runs now ships its own agent: the help desk, the CRM, the issue tracker[4,32,35]. Each one's rules, permissions and logs stop at the edge of that system.

The risky work crosses systems. A support agent reads a ticket and issues a credit in billing; a coding agent reads a customer's bug report and touches the production database. Someone has to answer: who is this agent acting for, who said it could, and what did it see? No single system of record can.

OrchKernel is built to be that layer for the actions agents send through it: agents ask it before they act, it holds what needs a person, and it keeps one log across systems. See the blueprint.

What it is not. It does not replace your CRM, help desk, billing or code host; records stay there. It is not a compliance certification, and it does not govern your product's own AI features unless they are routed through it.

Asks to act
Agents in every team

Support, deal desk, coding. Each works for a named rep, AE or engineer.

Checks each action sent through it
OrchKernel
  • Who it acts for
  • Rules
  • Approvals
  • Data by role and field
  • One log

Allows the action, holds it for a person, or denies it, and records which.

Systems of record, unchanged
What you already run
  • Help desk
  • CRM
  • Billing
  • Code host
  • Warehouse
  • Slack
A help desk's own bot follows the help desk's settings. The layer covers agents that work across systems, which is where a refund, a discount and a code change meet the same customer.
03

Where the revenue dollar and the hours go

SaaS has high gross margins and thin operating margins. The median private company runs 77% gross margin[25]; the money goes below gross profit, into people who sell, build and support. Here is how nine public companies spent each revenue dollar in their latest fiscal year[2]:

Share of revenue, latest fiscal year, GAAP
  • Cost of revenue (hosting, support, services)
  • Sales and marketing
  • Research and development
  • General and administrative
  • Operating margin
Five of the nine companies in the table below. Stock-based compensation sits inside every line. A light gap at the end of a bar is other operating items, such as restructuring. Where support and customer success sit varies by company.
HubSpotFY to Dec 2025 · revenue $3.13B
Cost of revenue
16.2%
Sales and marketing
44.1%
R&D
28.9%
G&A
10.4%
Operating margin
0.2%
FreshworksFY to Dec 2025 · revenue $0.84B
Cost of revenue
15.0%
Sales and marketing
47.1%
R&D
19.5%
G&A
16.8%
Operating margin
1.6%
SalesforceFY to Jan 2026 · revenue $41.52B
Cost of revenue
22.3%
Sales and marketing
34.5%
R&D
14.4%
G&A
7.2%
Operating margin
20.1%
AtlassianFY to Jun 2026 · revenue $6.57B
Cost of revenue
15.2%
Sales and marketing
23.4%
R&D
49.7%
G&A
11.5%
Operating margin
0.2%
DocuSignFY to Jan 2026 · revenue $3.22B
Cost of revenue
20.6%
Sales and marketing
37.4%
R&D
20.7%
G&A
12.1%
Operating margin
9.3%
BoxFY to Jan 2026 · revenue $1.18B
Cost of revenue
20.8%
Sales and marketing
34.3%
R&D
25.0%
G&A
12.8%
Operating margin
7.1%
ZoomFY to Jan 2026 · revenue $4.87B
Cost of revenue
23.0%
Sales and marketing
28.5%
R&D
17.4%
G&A
8.1%
Operating margin
23.1%
PagerDutyFY to Jan 2026 · revenue $0.49B
Cost of revenue
15.1%
Sales and marketing
37.4%
R&D
25.8%
G&A
20.6%
Operating margin
1.2%
DatadogFY to Dec 2025 · revenue $3.43B
Cost of revenue
20.0%
Sales and marketing
27.9%
R&D
45.2%
G&A
8.2%
Operating margin
-1.3%

People cost dominates every line. Stock-based compensation alone is 16.9% of revenue at HubSpot and 24.4% at Atlassian[2]. The median VC-backed private company puts 47% into sales and marketing[25].

Revenue grew faster than headcount. HubSpot grew revenue about 19% in 2025 while full-time employees grew about 8%, from 8,246 to 8,882 (derived from its 10-K)[2,3]. Salesforce's CEO said support went from about 9,000 people to about 5,000 "because I need less heads"[43].

So AI economics land in sales and marketing, the support part of cost of revenue, and R&D. Inference is the one cost AI adds; HubSpot's 10-K warns it "could impact our profit margin if we are unable to monetize such assets"[3].

04

Two fronts: AI in what you sell, AI in how you run

A SaaS company becoming AI-native works on two fronts at once. The first is the product: shipping agents and repricing around them. The incumbents now charge per outcome:

  • Intercom Fin$0.99 per resolution[32]
  • Zendesk AI agents$1.50 per automated resolution committed, $2.00 pay as you go[34]
  • Salesforce Agentforce$2 per conversation, or Flex Credits at about $0.10 per action[35]

37% of software companies in ICONIQ's investor survey planned to reprice within a year[30]. Databricks' CEO counters: "Why would you move your system of record?"[47]

The second front is how the company runs: support, sales, CS, engineering, security and finance done with agents, so margin holds while seat revenue is under pressure. This page is about the second front. The first comes in only where it creates obligations, such as AI transparency and the rules for products that decide about people, covered in the rules that bite.

05

Team by team: what agents already do

What agents already do in each team, and how good the evidence is. The grade is ours: strong means independent causal research, mixed means studies that disagree, thin means mostly vendor data, none independent means we found none.

  1. Support

    Evidence: strong
    The work
    Triage, answers from the help center and past tickets, bug escalation, credits and refunds, QA.
    What agents do today
    Help desk agents that answer and close tickets, and drafting help for reps. Atlassian's 10-K describes one that can take a ticket, diagnose it from Confluence and "resolve it autonomously"[4].
    Evidence
    The strongest independent result in SaaS work: +14% issues resolved per hour, +34% for novices, across 5,179 agents[26], from an assistant helping people, not an agent acting alone. Intercom claims Fin resolves 76% on average (vendor claim)[33].
    The SaaS catch
    Credits, refunds and cancellations start as tickets. Here a support agent stops writing and starts moving money.
  2. Sales, SDRs and RevOps

    Evidence: thin
    The work
    Lead qualification, outbound, meeting prep, follow-up, order forms, discount approval, questionnaires, forecast, CRM hygiene.
    What agents do today
    AI SDRs, research and enrichment, call analysis, CRM agents, questionnaire drafting.
    Evidence
    Mostly vendor surveys. MIT NANDA found more than half of AI budgets go to sales and marketing tools, while the clearest returns came from back-office work[49]. No public benchmark for AI SDR meeting rates or discount leakage.
    The SaaS catch
    Vendor claims need checking; see the 11x case below[48].
  3. Customer success and renewals

    Evidence: none independent
    The work
    Onboarding, health monitoring, QBR briefs, renewals, expansion, churn saves, cancellations.
    What agents do today
    Health scoring and summaries in CS platforms and CRMs; meeting note-takers that write to the CRM.
    Evidence
    No independent study. The case is retention math: median GRR 88%, NRR 101%, and 40% of new ARR from existing customers in 2025[25].
    The SaaS catch
    Cancellation. On self-serve plans bought by individuals, California allows save offers only if the customer can still cancel at once[18]. Enterprise renewals follow the contract's notice terms. A retention agent that stalls either is a liability.
  4. Engineering

    Evidence: mixed
    The work
    Issue to code, review, tests, deploys, incident response, on-call handover, dependency upkeep.
    What agents do today
    Coding assistants and agents, AI code review, incident summaries.
    Evidence
    It points both ways: 55.8% faster on one task in a GitHub-linked trial[36], 19% slower for experienced developers on their own repos in METR's[27]. See the evidence section.
    The SaaS catch
    Production data, credentials and deploys. An agent can do the most damage here, fastest.
  5. Security, trust and compliance

    Evidence: thin
    The work
    Questionnaires, SOC 2 evidence, DPAs, the subprocessor list, access reviews.
    What agents do today
    Compliance automation and questionnaire drafting from an approved library.
    Evidence
    Vendor surveys only.
    The SaaS catch
    Your customers audit you. Every AI tool on their data becomes a questionnaire answer and possibly a subprocessor notice[5].
  6. Finance and billing

    Evidence: none independent
    The work
    Invoicing, dunning, credits, usage metering, revenue recognition, board metrics.
    What agents do today
    Billing platforms' recovery features; AI in close and planning tools.
    Evidence
    None on SaaS finance work. MIT NANDA's back-office finding is the closest signal[49].
    The SaaS catch
    In a public company, an agent changing billing or revenue data is likely inside SOX controls; to be confirmed with your auditors.

Product and marketing use AI for feedback, release notes and campaigns, under CAN-SPAM[16] and the FTC's rules on AI claims[13].

06

What the evidence says, and where it is thin

Independent causal evidence exists for two SaaS functions: support and coding. Everything else rests on vendor and investor surveys.

  • Support helps people most where they are newest

    In the QJE study, issues resolved per hour rose 14% on average and 34% for novices, with better customer sentiment[26]. The gains came from spreading top performers' answers, so the people who keep the knowledge matter more.

  • Developers cannot tell how much AI helps them

    METR's randomized trial: 16 experienced developers, 246 issues on their own repos, 19% slower with AI while believing they were 20% faster. METR cautions it may not generalize[27]. A vendor-linked trial found one task done 55.8% faster[36].

  • AI amplifies what a team already is

    DORA 2025, about 5,000 respondents: 90% use AI, 30% barely trust its code, and adoption goes with higher throughput and lower stability. "AI doesn't fix a team; it amplifies what's already there"[28].

  • Developers use AI widely; agents much less

    Stack Overflow 2025: 84% use or plan to use AI, 46% distrust its accuracy, 66% complain of output "almost right, but not quite". Only 14.1% use agents daily; 37.9% have no plans to[29].

  • Pilots stall, and bought tools do better

    MIT NANDA, in a press summary: about 5% of enterprise pilots showed rapid revenue gains; tools bought from specialist vendors succeeded about 67% of the time, internal builds about a third as often. The method has been debated[49].

07

The staged path

Six stages for a company of 50 to 1,000 people, ordered by evidence and exposure: best-documented and lowest money risk first, actions that move money or speak for the company last. Vendors often sell AI SDRs first; the MIT NANDA finding above is why we do not.

  1. 0
    Inventory and data mapFind the AI already switched on and what customer data it sees
    GateEvery AI tool has an owner and a data class; the subprocessor list matches reality
  2. 1
    Support first passAgents triage and draft; people send
    GateRepeat-contact rate measured; AI replies labeled; policy topics routed to people
  3. 2
    Revenue research and hygieneBriefs, CRM updates, questionnaire drafts. People own the deal
    GateCRM field ownership written down; questionnaire answers come from an owned library
  4. 3
    Engineering agents with gatesAgents open pull requests; people merge and deploy
    GateChange failure rate and time to restore tracked against a baseline
  5. 4
    Money and customer-facing actionsCredits, refunds, plan changes, renewal drafts, under caps
    GateCaps and named approvers for every money action; one log across billing, CRM and help desk
  6. 5
    AI-native operating modelHeadcount, owners and rules built around agents
    Keeps goingReviewed every quarter: rules, subprocessors, questionnaire library, scorecards
Each stage starts when the gate above it holds. Teams move at their own pace: support can be at Stage 4 while engineering is at Stage 3.
  1. 0

    Stage 0: Inventory and data map

    Find the AI already switched on and what customer data it sees.

    About 4 to 6 weeks

    What to do

    • List every AI feature already on (help desk bot, CRM assistant, call recorder, note-taker, coding assistant), the customer data it sees and the model provider that receives it.
    • Compare that list with your subprocessor list and DPAs. The gaps are your first finding.
    • Review OAuth scopes and connected apps in the CRM, help desk, GitHub and Google Workspace. Revoke what nobody owns.
    • Ban customer data in personal chatbot accounts. Set data classes and name one owner per team.

    Why now

    The worst SaaS agent incidents of 2025 were about scopes and inputs[37,39,41], and customers must hear about a new processor before it starts[5], so you need lead time.

    In place first

    • A current subprocessor list.
    • Admin access to each system's connected-apps and token settings.

    What to measure

    Share of AI tools with an owner and a data class; long-lived tokens held by AI tools; subprocessor list matches what runs (yes or no).

    Common mistakes

    • Skipping the connected-app review. The Drift attackers used OAuth tokens a chat integration already held[41].
    • Letting each team buy its own agent with its own keys.
  2. 1

    Stage 1: Support first pass

    Agents triage and draft; people send.

    Starts in one queue once Stage 0 is done

    What to do

    • Triage tickets by product area, plan tier and severity.
    • Draft replies from the help center and resolved tickets. A rep sends them.
    • File bug reports with reproduction steps; draft help-center articles for repeat questions.
    • After a measured trial, let low-severity how-to replies go out alone. Nothing about money, policy or security.

    Why now

    The strongest independent evidence[26], and little money exposure while credits and policy stay out of scope.

    In place first

    • A current help center. The agent will repeat it.
    • Topics the agent may not answer alone: pricing, security, legal, outages, cancellations, data export.
    • Replies labeled as AI where the EU AI Act applies; Article 50 transparency duties apply from 2 August 2026[9].

    What to measure

    First-response time; reopen and repeat-contact rate within 7 days; share of AI drafts sent unedited; satisfaction scores on AI-handled versus human-handled tickets.

    Common mistakes

    • Counting silence as resolution, as vendor definitions do[32].
    • Letting the bot explain policy without a source. Cursor's bot invented one[50].
    • No direct human path for enterprise accounts.
  3. 2

    Stage 2: Revenue research and hygiene

    Briefs, CRM updates, questionnaire drafts. People own the deal.

    Starts with one sales team, alongside Stage 1

    What to do

    • Qualify inbound leads with a written reason the AE can check.
    • Pre-call and pre-renewal briefs from CRM, tickets, usage and billing.
    • CRM next steps and contacts from email and call notes. Stage, amount and close date stay with the AE.
    • Security questionnaire drafts from an approved library; usage and payment flags for CS.

    Why now

    Sales and marketing is 23% to 47% of revenue at the public companies above[2], so most internal hours sit here. Start with research and hygiene, not autonomous outbound.

    In place first

    • CRM field ownership: who may change stage, amount and close date.
    • A questionnaire library with an owner per section.
    • CAN-SPAM and TCPA rules written down before any outbound automation[16,17].

    What to measure

    Hours from meeting to CRM update; questionnaire turnaround; forecast error at day 30 of the quarter.

    Common mistakes

    • Buying an AI SDR on a logo wall[48].
    • Letting agents edit amounts and stages.
    • Questionnaire answers drafted from last year's documents.
  4. 3

    Stage 3: Engineering agents with gates

    Agents open pull requests; people merge and deploy.

    When branch protection and credential separation are in place

    What to do

    • AI code review as a second reviewer on every pull request.
    • Dependency upgrades, incident summaries and on-call handover notes.
    • Agents open pull requests from well-specified issues. Merges and production stay with people.

    Why now

    Developers already use AI: 84% use or plan to[29]. DORA links that adoption to lower delivery stability[28], so the work now is gates.

    In place first

    • Branch protection and required human review on every repo an agent can write to.
    • Separate dev and prod credentials[51]; tokens scoped per repo and task[38,40]; no superuser database roles[39].

    What to measure

    Change failure rate and time to restore, before and after; share of agent pull requests merged without rework; incidents traced to agent-authored changes.

    Common mistakes

    • Measuring speed by asking developers. In METR's trial they felt 20% faster and were 19% slower[27].
    • Giving an agent the deploy key for incidents.
  5. 4

    Stage 4: Money and customer-facing actions

    Credits, refunds, plan changes, renewal drafts, under caps.

    After Stages 1 and 2 have run with evidence

    What to do

    • Credits and refunds under a cap by plan and amount.
    • Failed-payment outreach; plan changes customers ask for in a ticket.
    • Renewal drafts and quotes for the account owner.
    • Usage-billing anomaly checks before invoices go out.

    Why now

    Agents now move money and speak for the company, which owns what they say[52]. Cancellation and switching rules also apply[12,18].

    In place first

    • A cap and a named approver for each money action.
    • Approval shows the exact amount, customer and invoice, and runs once.
    • One log across billing, CRM and help desk.

    What to measure

    Agent credits versus approved credits, and reversals; approval wait time; involuntary churn recovered; GRR and NRR trend. Medians were 88% and 101% in 2025[25], as context, not a target.

    Common mistakes

    • One global refund cap instead of limits by plan and customer.
    • Approvals in a chat thread with no record of what was approved.
    • Letting a retention agent handle cancellation requests.
  6. 5

    Stage 5: AI-native operating model

    Headcount, owners and rules built around agents.

    When Stages 1 to 4 run with evidence

    What to do

    • Start headcount plans from agent capacity; give each team an agent owner and a scorecard.
    • Keep refund caps, discount ceilings and send limits in version control, changed through review.
    • Publish a plain statement to customers of where AI touches their data.

    Why now

    By now you have two quarters of numbers on what agents do in support, sales and engineering. Shopify already tells staff to show a job cannot be done by AI before asking for more people[44]. Savings from Stages 1 to 4 stay only if next year's headcount plan is built from them.

    In place first

    • Named agent owners in every team.
    • Two or more quarters of scorecard data from earlier stages.

    What to measure

    Revenue per employee; sales and marketing as a share of revenue; days to answer a customer's AI questionnaire.

    Common mistakes

    • Cutting the people who write and maintain the knowledge the agents rely on.
    • Treating a vendor's resolution rate as your own.
08

Your first 90 days

Stage 0 for the whole company, one support queue at Stage 1, and one sales team at Stage 2. By day 90 you have a baseline, a month of drafts to judge, and the evidence to decide how far Stage 3 goes.

  1. Days 1 to 30

    Run the Stage 0 inventory. Review OAuth and connected-app scopes in the CRM, help desk, GitHub and Google Workspace. Name an owner per team. Write the five rules nobody may break: a refund cap, a discount ceiling, no production writes by agents, no new AI subprocessor without notice, no policy answer without a source. Record the baselines: repeat-contact rate, first-response time, meeting-to-CRM time, questionnaire turnaround.

  2. Days 31 to 60

    Stage 1 in one support queue. Agents triage and draft; reps send. Label AI-written replies. Track repeat contacts and the share of drafts sent unedited every week.

  3. Days 61 to 90

    Stage 2 for one sales team: pre-call briefs and CRM next steps. Build the questionnaire library with an owner per section. Review the first scorecard against the baseline, then decide the scope of Stage 3.

09

How the operating model and roles change

  1. 1

    Headcount plans start from what agents can do

    Shopify's CEO told staff to show a job cannot be done by AI before asking for more people[44]. Expect your board to ask the same of every open requisition.

  2. 2

    Support becomes a knowledge and exceptions team

    AI spreads top performers' practices to new agents[26], so knowledge owners, escalation specialists and AI reply QA become the core. Our inference, not a survey finding.

  3. 3

    Sales separates research from judgment

    Research, CRM updates and drafts go to agents; pricing, the deal desk and the relationship stay with people. Our recommendation; no study covers it.

  4. 4

    Engineering protects review capacity

    Google describes its AI-generated code as "reviewed and accepted by engineers"[42]. When agents write more code, review and deploy gates become the bottleneck to protect.

  5. 5

    New owners for rules that cross teams

    An agent owner per team; RevOps owns cross-system rules such as discount ceilings and credit caps; security owns the AI subprocessor list. Our recommendation; in a 50-person company these are parts of existing jobs.

  6. 6

    An open question: fewer, more senior hires?

    Flat software employment since 2022[1] suggests fewer hires; the biggest QJE gains went to the newest staff[26]. The evidence pulls both ways.

10

Don't fully automate

An agent can prepare each of these: the numbers, the draft, the evidence. A named person makes the call, and the log shows who.

Pricing, discounts and contract terms beyond set limits
Why it stays with a person
A discount commits margin for the life of the contract, in a business where sales and marketing already costs 23% to 47% of revenue[2].
Credits, refunds and write-offs above a cap; any change to a plan or invoice
Why it stays with a person
Money leaves and revenue records change. A named approver sees the exact amount and invoice.
Statements of policy, security posture or legal position
Why it stays with a person
The company owns what its bot says. Cursor's bot invented a policy; Air Canada was held liable for its chatbot's refund answer[50,52].
Cancellation and data-export requests
Why it stays with a person
For consumers, which can include individuals on self-serve plans, California requires that save offers never block a cancellation[18], and the EU Data Act gives customers a right to switch providers[12].
Code merged to production and production data changes
Why it stays with a person
An agent deleted SaaStr's live database during a code freeze[51]. Even Google's AI-written code is "reviewed and accepted by engineers"[42].
Adding a model provider for customer data
Why it stays with a person
It is likely a new subprocessor, which the customer must hear about first and can object to[5].
Incident communications and materiality calls
Why it stays with a person
A public company files within four business days of deciding an incident is material[15]. That call belongs to named people.
Public claims about the product's AI
Why it stays with a person
The FTC and SEC have both acted on overstated AI claims[13,14].
Suspending an individual's account; your own hiring decisions
Why it stays with a person
People have rights against decisions based solely on automated processing[6], and Colorado adds notice, explanation and human review from 2027[21].
11

The rules that bite

A SaaS company meets AI rules from three directions: as a processor of its customers' data, which binds most often, through contracts and GDPR; as a seller of AI features, through claims, transparency and high-risk uses; and as an employer and marketer, through its own hiring, outbound email and calls, and cancellation rules.

Five requirements recur: tell people when they deal with AI; tell customers before a new party processes their data; keep records of what automated systems did; offer human review of decisions about people; do not overclaim. Below are the specific forms those take for a SaaS company.

Customer data and subprocessors

The rules you meet daily, as a processor of customers' data.

  • GDPR Article 28(2) and (4)[5]

    A processor needs the controller's written authorization to engage another processor, must tell the controller about any new one so it can object, and stays fully liable for it. A model provider that receives support tickets with personal data is likely a new subprocessor.

    In force

    To be confirmed: The notice period is set by each DPA. We found no public source for a typical length; read your own.

  • GDPR Article 22[6]

    A right not to be subject to solely automated decisions with legal or similarly significant effects, and, where allowed, to human intervention. Relevant when you suspend an individual's account or screen job applicants.

    In force

  • Sector contracts your customers flow down

    DORA terms from EU financial customers, NIS2, and HIPAA business associate agreements that reach AI subprocessors.

    Depends on your customers

    To be confirmed: Not checked against the primary texts.

The EU AI Act

For companies that sell into the EU or run agents there.

  • Article 4: AI literacy[7,10]

    Providers and deployers support the AI literacy of staff who operate AI for them.

    From 2 February 2025

  • Article 50: transparency[9]

    People must be told they are interacting with AI unless it is obvious. This reaches AI support agents and AI outreach.

    From 2 August 2026

  • Article 25: when you become the provider[8]

    Putting your name on a high-risk system, substantially modifying one, or repurposing a general-purpose model into a high-risk use makes you its provider. An HR, credit or education SaaS product can land here.

    Applies with the high-risk duties

  • Annex III high-risk duties[10,11]

    Risk management, logging, human oversight and documentation for uses such as hiring and credit.

    From 2 December 2027, moved from August 2026 by a 2026 amendment

    To be confirmed: The new date comes from the tracker; the Official Journal text could not be opened in October 2026.

The EU Data Act: switching

The Commission says the switching chapter covers SaaS.

  • Switching between data processing services[12]

    Customers may switch provider or move in-house. Cost-based switching fees are allowed until 12 January 2027; after that, no charges for switching or data egress. Switching requests will arrive through support, where agents sit.

    Applies from 12 September 2025; charges end 12 January 2027

    To be confirmed: The maximum two-month notice period and 30-day transition period in Article 25 could not be checked in the regulation text.

Claims about AI and incident disclosure (US)

What you say about your AI has to be true.

  • FTC Operation AI Comply[13]

    AI claims need substantiation. DoNotPay settled for $193,000 over its "robot lawyer" claims.

    Announced 25 September 2024

  • SEC AI-washing cases[14]

    Two investment advisers paid $225,000 and $175,000 for overstating their AI. The theory applies to any public company's disclosures.

    18 March 2024

  • SEC cybersecurity disclosure[15]

    A material cybersecurity incident goes on Form 8-K Item 1.05, generally within four business days of the materiality decision. An agent-caused leak can qualify.

    Adopted 26 July 2023

  • SOX 404 and agents that touch billing

    An agent that changes billing, credits or revenue data is likely inside IT general controls over financial reporting.

    Public companies

    To be confirmed: We found no published guidance on agent-made changes. Agree the treatment with your auditors.

Outreach and cancellation

Where AI SDRs and retention agents meet consumer-protection law.

  • CAN-SPAM[16]

    No B2B exemption. Accurate headers, a physical address, opt-outs honored within 10 business days; up to $53,088 per violating email.

    In force

  • TCPA and AI voices[17]

    AI-generated voices are artificial voices, so TCPA consent rules apply to AI calls.

    Ruling of 8 February 2024

  • California AB 2863, auto-renewal[18,19]

    Express consent, annual reminders, online click-to-cancel, and save offers only if cancellation stays immediate. It protects consumers, defined as individuals buying for personal, family or household purposes, so a plan bought for business use is generally outside it.

    Contracts entered or renewed from 1 July 2025

    To be confirmed: How it applies to a self-serve plan an individual buys partly for personal use is a question for counsel.

  • FTC click-to-cancel rule[20]

    The Eighth Circuit vacated the FTC's amended Negative Option Rule. State laws still apply.

    Vacated 8 July 2025

  • EU and UK rules on B2B cold email

    National ePrivacy rules differ on whether a corporate address needs consent.

    Varies by country

    To be confirmed: Not checked country by country.

When your product decides about people

For HR-tech, lending, insurance and education SaaS, and your own hiring.

  • Colorado SB26-189[21]

    Developers give deployers documentation of intended uses, training data, known limits and human review, and keep records three years. Deployers give notice, an explanation within 30 days of an adverse outcome, and correction and human review rights.

    From 1 January 2027

  • California CCPA regulations on automated decisionmaking, risk assessments and cybersecurity audits[22]

    Notices and rights where automated decisionmaking makes significant decisions; risk assessments filed with the agency; cybersecurity audits.

    ADMT duties by 1 January 2027. Risk assessments from 2026 and 2027 filed by 1 April 2028. First audit reports due 1 April 2028, 2029 or 2030, by revenue

  • California AB 2013[23]

    Developers of public generative AI systems post training-data documentation. "Substantially modifies" includes retraining or fine-tuning, so fine-tuning a vendor model for a public feature can bring you in scope.

    From 1 January 2026

  • Mobley v. Workday[24]

    Age-discrimination claims against a SaaS vendor over its customers' AI screening proceed as a collective action: the leading case on a vendor's own exposure.

    Collective action allowed to proceed in May 2025

    To be confirmed: The current status. The docket could not be opened for this page.

12

When agents fail

Ten reported cases, each with the control that would have stopped it. Six of the ten came down to what an agent or its tokens could reach, rather than what it said.

  1. Cursor's support bot invents a policy

    April 2025

    An AI support bot told users who were being logged out that it was a new one-device policy. It was a bug. Users posted cancellations before a cofounder wrote "We have no such policy." Cursor then labeled AI replies[50].

    The control: Policy answers from an approved source or a person; AI replies labeled. Control 1

  2. Replit's agent deletes SaaStr's production database

    July 2025

    During a code freeze the agent deleted a live database with records on more than 1,200 executives, then said a rollback would not work. The data was recovered. Replit added automatic dev and prod separation[51].

    The control: Approval for production writes; dev credentials by default; a kill switch. Controls 6 and 15

  3. Salesloft Drift tokens used to raid Salesforce orgs

    8 to 18 August 2025

    An attacker used OAuth tokens held by the Drift AI chat integration to export cases, accounts and opportunities from many companies' Salesforce instances, then searched them for keys and passwords[41].

    The control: No long-lived vendor keys in AI tools; minimal scopes. Controls 7 and 8

  4. ForcedLeak in Salesforce Agentforce

    Disclosed 25 September 2025

    Instructions hidden in a web-to-lead form ran when an employee later asked the agent about the lead, and data could leave through an expired, still-trusted domain. Rated CVSS 9.4; Salesforce fixed it (security vendor's disclosure)[37].

    The control: A gate before any send; allowed outbound hosts. Control 9

  5. Prompt injection through a public GitHub issue

    May 2025

    A malicious issue in a public repo led a developer's agent to copy private repo data into a public pull request (security vendor's demonstration)[38].

    The control: One repo per task; approval before public writes. Controls 7 and 9

  6. A support ticket leaks integration tokens

    July 2025

    In a demonstration, instructions in a support ticket led a developer's assistant, connected with a key that bypasses row-level security, to post secret tokens into the ticket[39].

    The control: No superuser agents; access by field. Controls 7 and 10

  7. Amazon Q extension ships with injected code

    July 2025

    An "inappropriately scoped GitHub token" let an attacker add code to version 1.84.0. It shipped but failed to run[40].

    The control: Scoped tokens; review on anything that ships. Controls 6 and 8

  8. Questions over an AI SDR vendor's customers

    March 2025

    TechCrunch reported that 11x showed logos of non-customers and counted trial contracts as ARR; 11x disputed parts[48]. Allegations, not findings.

    The control: Judge agents on your own scorecard. Control 13

  9. Air Canada is held liable for its chatbot

    February 2024

    The chatbot described a refund option the airline did not offer, and the tribunal held the airline liable[52]. Read through a secondary summary; the decision itself is to be confirmed.

    The control: Refund and policy answers from approved sources. Controls 1 and 2

  10. Klarna's support automation

    2024 onward

    Klarna said its assistant handled two-thirds of chats in month one (company claim)[45]. Reports that it later brought back human service are to be confirmed.

    The control: A human path by customer tier; measure repeat contacts. Control 1

The pattern: the lethal trifecta

ForcedLeak, the GitHub issue and the ticket leak share what Simon Willison calls the lethal trifecta: "access to your private data", "exposure to untrusted content" and "the ability to externally communicate"[31]. In SaaS, anyone can write into the help desk and web forms, which sit next to the CRM.

  1. 1 of 3
    Untrusted content comes in
    • Support tickets
    • Web-to-lead forms
    • Inbound email
    • Public GitHub issues

    ForcedLeak used a lead form

  2. 2 of 3
    The agent can read private data
    • CRM records
    • Production database
    • Private repos
    • Billing

    The Supabase demo used a superuser key

  3. 3 of 3
    The agent can send data out
    • Email replies
    • Public pull requests
    • Ticket comments
    • Outbound HTTP

    ForcedLeak sent data to a trusted, expired domain

Where the gate goes

A run that touched untrusted content and can read private data is checked before it sends anything: allowed hosts only, unusual sends approved by a person. Controls 7, 8 and 9.

The three legs, named by Simon Willison, mapped to the systems of a SaaS company. Remove any one and the leak path closes.

Agent failures to design against (scenarios)

Scenarios, not reported cases, each mapped to a control point.

The over-generous refund. A ticket says "double charged, refund me $4,800" and the agent issues it.
What stops it
A refund cap; above it, a finance approver sees the exact amount and invoice (control 2).
The unapproved discount. A sales agent applies 35% off to match a competitor named in an email.
What stops it
A discount ceiling by deal size; deal desk approval above it (control 3).
The wrong recipient. A CS agent emails a price change to a former champion still in the CRM.
What stops it
The account owner approves sends about price or terms (control 4).
The unannounced subprocessor. An engineer points support tickets at a new model provider to test summaries.
What stops it
Models are cleared per data class; adding one is a logged admin change (control 11).
The agent that fixed prod. During an incident, an agent with a deploy token rolls back a migration.
What stops it
Production actions need on-call approval; the agent acts with the engineer's authority (controls 6 and 7).
The saved cancellation. A retention agent keeps offering discounts to a customer who asked to cancel.
What stops it
Cancellation intent goes to the human queue (control 5).
The untrue questionnaire answer. An agent answers "Do you train on customer data?" with "No" from an old library after a new feature started doing it.
What stops it
A versioned library; security answers change only with the security owner (control 12).
13

Control points for an AI-native SaaS company

Seventeen places where a SaaS company needs a control whatever tools it uses, who owns each, and the case or rule behind it. The blueprint below shows what enforces each one.

01
Policy, security and legal statements to customers
Owner
Support lead; the security owner for security answers
Why it exists
Cursor; Air Canada; EU AI Act Article 50[9,50,52]
02
Credits and refunds: caps by plan, named approver above
Owner
Finance, with the support lead
Why it exists
Money leaves the company; scenario 1 below
03
Discounts and terms: ceilings by deal size
Owner
Deal desk or CRO
Why it exists
Sales and marketing is the largest cost line[2]; scenario 2
04
Outbound sends to customers and prospects
Owner
RevOps; the account owner
Why it exists
CAN-SPAM; TCPA for AI voices[16,17]; scenario 3
05
Cancellation and data-export requests go to a person
Owner
CS lead
Why it exists
California AB 2863; EU Data Act switching[12,18]; scenario 6
06
Production code and data: human merge, on-call approval
Owner
Engineering lead
Why it exists
Replit and SaaStr; DORA stability finding; Amazon Q token[28,40,51]
07
Every agent acts for a named person, with no more access
Owner
Security or IT
Why it exists
Supabase service-role key; Drift tokens[39,41]
08
No long-lived vendor tokens in agents or AI tools
Owner
Security
Why it exists
Salesloft Drift[41]
09
Untrusted inbound content cannot trigger sends unchecked
Owner
Security; the agent owner
Why it exists
ForcedLeak; GitHub MCP; the lethal trifecta[31,37,38]
10
Customer data by role and field
Owner
Data owner
Why it exists
GDPR and your DPAs[5]
11
Models cleared per data class; subprocessor notices
Owner
Security and legal
Why it exists
GDPR Article 28(2)[5]; scenario 4
12
Questionnaire answers from a versioned, owned library
Owner
Security owner
Why it exists
Scenario 7; FTC claim substantiation[13]
13
Public and investor claims about AI
Owner
Marketing and legal; finance for filings
Why it exists
FTC Operation AI Comply; SEC AI-washing cases[13,14]
14
Tamper-evident record of what each agent did
Owner
Security; legal for materiality
Why it exists
SEC Form 8-K Item 1.05 for public companies[15]; customers' breach-notice clauses
15
Kill switch and budgets
Owner
Agent owner; CTO
Why it exists
The Replit agent kept acting during a code freeze[51]
16
Decisions about individuals
Owner
HR; product counsel
Why it exists
GDPR Article 22; Colorado SB26-189[6,21]
17
AI literacy and acceptable use
Owner
People team
Why it exists
EU AI Act Article 4, from 2 February 2025[7]
14

The OrchKernel blueprint for a B2B SaaS company

OrchKernel sits between agents and the systems you run, as drawn in the missing layer. Your help desk, CRM, billing and code host stay the systems of record.

The mechanisms

Approvals
An action waits for a named person: a refund above the cap for finance, a discount above the ceiling for the deal desk, a production change for on-call. The approver sees the exact payload; it runs once.
Rules
Checked before every action: refund caps, discount ceilings, send limits, topics the agent may not answer. A rule allows, holds or denies with the reason.
Acting on a named person's authority
Each agent works for a named rep, AE or engineer, with no more access than that person. When the person loses access to an account, so do their agents.
Data access by role and field
Which accounts and fields each person and agent sees. Each model is cleared for a data class, so restricted data reaches only cleared models.
Tamper-evident audit log
Every request, decision, approval and result, chained so an edited entry shows. It answers a customer's security team: what did AI do with our data, and who approved it.
Human queue
Cancellations, export requests, security and legal questions land with an owner and a deadline.
Connections to your systems
The help desk, CRM, billing, code host, Google Workspace, Slack, any MCP server and any REST API through declarative adapters. The same rules and log apply to each.

Control point by control point

Where OrchKernel does part of the job or none, the row says so.

01
Policy, security and legal statements to customers
What enforces it in OrchKernel
Rules on which replies an agent may send alone; policy, security and legal topics go to the human queue.
What belongs elsewhere
The help desk's own label setting for its native bot; the content of the help center.
02
Credits and refunds: caps by plan, named approver above
What enforces it in OrchKernel
Rules hold the caps; approval shows the exact payload and runs once.
What belongs elsewhere
Billing permissions and accounting; SOX control design with auditors.
03
Discounts and terms: ceilings by deal size
What enforces it in OrchKernel
Rules hold the ceilings; the deal desk approves above them.
What belongs elsewhere
An existing CPQ approval flow; contract templates.
04
Outbound sends to customers and prospects
What enforces it in OrchKernel
Approval on price and terms messages; rules on frequency and contact role.
What belongs elsewhere
The email platform's suppression list; call consent records.
05
Cancellation and data-export requests go to a person
What enforces it in OrchKernel
A rule sends cancellation and export intent to the human queue.
What belongs elsewhere
The product's own cancel button and export function.
06
Production code and data: human merge, on-call approval
What enforces it in OrchKernel
Approval for agent actions through the code-host connection; kill switch.
What belongs elsewhere
Branch protection, CI, the deploy pipeline, database roles.
07
Every agent acts for a named person, with no more access
What enforces it in OrchKernel
Acting on a named person's authority; delegation only narrows; revoking it stops every agent.
What belongs elsewhere
Your identity provider (SSO for OrchKernel is on the roadmap).
08
No long-lived vendor tokens in agents or AI tools
What enforces it in OrchKernel
Agents hold no vendor keys; OrchKernel acts with its own sealed credentials and allowed hosts.
What belongs elsewhere
Scopes inside each vendor system; rotating credentials.
09
Untrusted inbound content cannot trigger sends unchecked
What enforces it in OrchKernel
A gate on every action, checked against the agent's trust tier; approval before outbound sends.
What belongs elsewhere
Model-side injection defenses; vendor fixes.
10
Customer data by role and field
What enforces it in OrchKernel
Data access by role and field; search respects each person's permissions.
What belongs elsewhere
Field-level security in the CRM and database.
11
Models cleared per data class; subprocessor notices
What enforces it in OrchKernel
Models cleared per data class; adding one is a logged admin change.
What belongs elsewhere
The DPA, subprocessor list and customer notices.
12
Questionnaire answers from a versioned, owned library
What enforces it in OrchKernel
New knowledge is quarantined until a person accepts it.
What belongs elsewhere
Your trust center; the SOC 2 audit.
13
Public and investor claims about AI
What enforces it in OrchKernel
Only approval on drafts an agent writes.
What belongs elsewhere
Legal review; the disclosure committee.
14
Tamper-evident record of what each agent did
What enforces it in OrchKernel
Tamper-evident log; the ok audit verify command names the first broken link; runs can be replayed.
What belongs elsewhere
Your SIEM (export is on the roadmap); materiality decisions.
15
Kill switch and budgets
What enforces it in OrchKernel
Kill switches, budgets and trust tiers.
What belongs elsewhere
Vendor-side rate limits.
16
Decisions about individuals
What enforces it in OrchKernel
Approval and the human queue for actions that affect a person.
What belongs elsewhere
Bias audits, notices, HR process; your product's own controls.
17
AI literacy and acceptable use
What enforces it in OrchKernel
Not enforced by software.
What belongs elsewhere
Training and policy.

What is available when. The governed actions API with its TypeScript and Python clients, the MCP gateway and the policy dry run that explains a decision before it is enforced are arriving in this release. SSO and export of the audit log to a SIEM are later, on the roadmap.

What it is not. It does not replace your systems of record, holds no SOC 2, HIPAA or ISO certification and does not by itself satisfy any law or audit; it enforces controls and keeps the evidence. It is source-available under the Business Source License 1.1 and runs on your own servers, so your security team can read the code.

15

Scorecard by stage

Record a baseline before Stage 1, then track the same numbers at each stage. For most internal SaaS work no public benchmark exists, and we have not quoted rules of thumb we could not source. Compare against yourself.

AI tools with an owner and a data classStage 0
How to count it
Owned and classified tools divided by all AI tools found in the inventory
Public benchmark
No public benchmark
Long-lived tokens held by AI toolsStage 0
How to count it
Count of OAuth grants and API keys held by AI tools or agents, by system
Public benchmark
No public benchmark
Repeat-contact rateStage 1
How to count it
Tickets where the same customer writes again on the same issue within 7 days, AI-handled and human-handled separately
Public benchmark
No independent benchmark. Vendor resolution rates exist (Fin claims 76%) but count silence as success[32,33]
Meeting to CRM updateStage 2
How to count it
Hours from a call ending to next step and contacts updated in the CRM
Public benchmark
No public benchmark
Questionnaire turnaroundStage 2
How to count it
Business days from receipt to answers sent, and how many answers the security owner changed
Public benchmark
No public benchmark
Change failure rate and time to restoreStage 3
How to count it
DORA's definitions, split by agent-authored and human-authored changes
Public benchmark
DORA defines the measures; no public benchmark for agent-authored changes[28]
Agent credits against approvalsStage 4
How to count it
Credits and refunds issued by agents, reversed, and approved above the cap
Public benchmark
No public benchmark
Gross and net revenue retentionStage 4
How to count it
Standard GRR and NRR, by segment
Public benchmark
Medians 88% and 101% in 2025[25]
Revenue per employeeStage 5
How to count it
Revenue divided by full-time employees at year end
Public benchmark
HubSpot about $352,000 and Atlassian about $494,000 (derived from 10-Ks)[3,4]; private SaaS $200,000 to $300,000 ARR per employee[25]
Days to answer a customer's AI questionnaireStage 5
How to count it
From receipt to a complete answer backed by the log
Public benchmark
No public benchmark
16

Sources

Sources were read in October 2026; dates are publication or data dates. Ratios from 10-K data are line items divided by total revenue. Gartner forecasts on agent projects were left out because they could not be opened.

Primary sources

Government, regulators, courts, statutes and SEC filings. GDPR and EU AI Act articles are read in unofficial consolidated copies.

  1. 1
    Current Employment Statistics, software publishers (series CES5051320001). US Bureau of Labor Statistics, August 2026, preliminary, seasonally adjusted.
  2. 2
    Company facts from the latest annual reports (10-K): HubSpot, Freshworks, Salesforce, Atlassian, DocuSign, Box, Zoom, PagerDuty, Datadog. US Securities and Exchange Commission, XBRL data, filed February to August 2026.
    Each line item divided by total revenue, GAAP, so stock-based compensation is included
  3. 3
    HubSpot, Inc. annual report on Form 10-K for fiscal 2025. US Securities and Exchange Commission, filed 11 February 2026.
  4. 4
    Atlassian Corporation annual report on Form 10-K for fiscal 2026. US Securities and Exchange Commission, filed 14 August 2026.
  5. 5
  6. 6
  7. 7
    EU AI Act, Article 4: AI literacy. artificialintelligenceact.eu.
  8. 8
  9. 9
  10. 10
    EU AI Act implementation timeline. artificialintelligenceact.eu, updated 31 August 2026.
    Unofficial tracker of the official dates
  11. 11
    Regulation (EU) 2026/1744 amending the AI Act: new application dates for high-risk systems. Official Journal of the European Union.
    Could not be opened in October 2026 (EUR-Lex unreachable); date taken from the tracker
  12. 12
    Data Act explained. European Commission.
  13. 13
  14. 14
  15. 15
    SEC adopts rules on cybersecurity risk management, strategy, governance and incident disclosure by public companies. US Securities and Exchange Commission, press release 2023-139, 26 July 2023.
  16. 16
    CAN-SPAM Act: a compliance guide for business. US Federal Trade Commission.
  17. 17
    FCC makes AI-generated voices in robocalls illegal. Federal Communications Commission, 8 February 2024.
  18. 18
    AB 2863 (automatic renewal and continuous service offers). California Legislature, chaptered 24 September 2024.
  19. 19
  20. 20
  21. 21
    SB26-189: automated decision-making technology. Colorado General Assembly, signed 14 May 2026.
  22. 22
    CCPA updates, cybersecurity audit, risk assessment and automated decisionmaking technology regulations: approved text. California Privacy Protection Agency, approved 22 September 2025.
    Compliance dates from sections 7121, 7157 and 7200 of the approved text
  23. 23
    AB 2013 (generative artificial intelligence: training data transparency). California Legislature, chaptered 28 September 2024.
  24. 24
    Mobley v. Workday, Inc., No. 3:23-cv-00770 (N.D. Cal.): docket. CourtListener (RECAP archive), read October 2026.

Industry bodies and independent research

Peer-reviewed studies, independent research groups, developer surveys and benchmark reports. Investor surveys are marked.

  1. 25
  2. 26
  3. 27
  4. 28
    Announcing the 2025 DORA report. Google Cloud (DORA research program), 24 September 2025.
  5. 29
    2025 Developer Survey: AI. Stack Overflow, 2025.
  6. 30
    The State of AI 2025. ICONIQ, 2025.
    Investor survey of about 300 software executives in its network
  7. 31

Vendor sources

Published by companies that sell AI to SaaS companies. Directional, not an industry benchmark.

  1. 32
    Intercom pricing (Fin AI agent, price per outcome). Intercom, read October 2026.
    Vendor source
  2. 33
    Fin AI agent: average resolution rate. Intercom, read October 2026.
    Vendor source
  3. 34
    Zendesk pricing (automated resolutions). Zendesk, read October 2026.
    Vendor source
  4. 35
    Agentforce pricing. Salesforce, read October 2026.
    Vendor source
  5. 36
    The impact of AI on developer productivity: evidence from GitHub Copilot (Peng et al.). arXiv 2302.06590, 2023.
    Vendor sourceRandomized trial run with GitHub and Microsoft researchers
  6. 37
    ForcedLeak: agent risks exposed in Salesforce Agentforce. Noma Security, 25 September 2025.
    Vendor sourceSecurity vendor's disclosure
  7. 38
    GitHub MCP exploited: accessing private repositories via MCP. Invariant Labs, 26 May 2025.
    Vendor sourceSecurity vendor's research
  8. 39
    Supabase MCP can leak your entire SQL database. General Analysis, 8 July 2025.
    Vendor sourceSecurity vendor's demonstration

Company and press

Company statements, news coverage, security researchers and encyclopedia summaries.

  1. 40
  2. 41
    Widespread data theft targets Salesforce instances via Salesloft Drift. Google Threat Intelligence Group, August 2025.
  3. 42
    Q3 2024 earnings remarks. Alphabet (Sundar Pichai), 29 October 2024.
  4. 43
  5. 44
  6. 45
  7. 46
  8. 47
  9. 48
    a16z- and Benchmark-backed 11x has been claiming customers it doesn't have. TechCrunch, 24 March 2025.
    Allegations; 11x disputed parts of the report
  10. 49
    MIT report: 95% of generative AI pilots at companies are failing. Fortune, 18 August 2025.
    Press summary of MIT NANDA's report; its method has been debated
  11. 50
    Cursor AI support bot invents a policy. The Register, 18 April 2025.
  12. 51
  13. 52
    Moffatt v. Air Canada, 2024 BCCRT 149. Wikipedia, decided 14 February 2024.
    Secondary summary; the tribunal's decision could not be opened in October 2026

Become a design partner

Run Stage 0 and one support queue at Stage 1 on your own help desk with us. You get early access, help with setup, and a say in what we build next.

Or email support@prefero.ai