On this page
Your help desk vendor now sells the work your support seats used to do. Intercom charges $0.99 for each ticket its agent resolves[32], Zendesk $1.50 to $2.00[34], Salesforce $2 a conversation[35]. A customer whose agent closes more of its tickets buys fewer support seats at renewal.
Investors noticed. TechCrunch puts the software and services value lost in the early-February 2026 sell-off at close to $1 trillion, tied to AI coding and agent launches and pressure on seat-based renewals[46]. US software publishers employed 652,900 people in August 2026, below the November 2022 peak of 664,300[1]. Headcount stopped rising, but the cost base is still mostly people: at the nine public SaaS companies below, sales and marketing alone takes 23% to 47% of revenue[2].
SaaS also has an audience other sectors do not. Your customers' security teams audit you. When a support agent sends ticket text to a model provider, that provider is likely a new subprocessor under GDPR, and your customers are owed notice before it starts[5]. The next questionnaire will ask where AI touches their data. Getting an agent to draft a reply or open a pull request is now easy. Saying, account by account, what your agents did, on whose authority and with whose approval, is the work this playbook is about.
AI-enabled versus AI-native in a SaaS company
Most SaaS companies are AI-enabled already. The difference that matters is whether anyone can answer "what did AI do to this customer's account this month, and who allowed it?"
- Copilot seats for everyone, counted as adoption
- A help desk bot the support lead set up
- Coding assistants with personal tokens
- Each tool with its own keys, rules and logs
- Nobody can say what AI did to customer X's account this month
- Agents do the first pass in support, sales, CS, engineering and finance
- People own pricing, shipped code, money and what is said in the company's name
- Every agent acts for a named person, with no more access than that person
- Anything touching a customer's money, data or contract passes a rule or approval
- The company answers its customers' AI questionnaire with evidence
Investors use "AI-native" for a company whose product is built on AI. ICONIQ, for one, found 47% of AI-native companies had reached scale against 13% of AI-enabled ones[30]. This page means something else: how the company itself runs.
The missing layer
Each system a SaaS company runs now ships its own agent: the help desk, the CRM, the issue tracker[4,32,35]. Each one's rules, permissions and logs stop at the edge of that system.
The risky work crosses systems. A support agent reads a ticket and issues a credit in billing; a coding agent reads a customer's bug report and touches the production database. Someone has to answer: who is this agent acting for, who said it could, and what did it see? No single system of record can.
OrchKernel is built to be that layer for the actions agents send through it: agents ask it before they act, it holds what needs a person, and it keeps one log across systems. See the blueprint.
What it is not. It does not replace your CRM, help desk, billing or code host; records stay there. It is not a compliance certification, and it does not govern your product's own AI features unless they are routed through it.
Support, deal desk, coding. Each works for a named rep, AE or engineer.
- Who it acts for
- Rules
- Approvals
- Data by role and field
- One log
Allows the action, holds it for a person, or denies it, and records which.
- Help desk
- CRM
- Billing
- Code host
- Warehouse
- Slack
Where the revenue dollar and the hours go
SaaS has high gross margins and thin operating margins. The median private company runs 77% gross margin[25]; the money goes below gross profit, into people who sell, build and support. Here is how nine public companies spent each revenue dollar in their latest fiscal year[2]:
- Cost of revenue (hosting, support, services)
- Sales and marketing
- Research and development
- General and administrative
- Operating margin
People cost dominates every line. Stock-based compensation alone is 16.9% of revenue at HubSpot and 24.4% at Atlassian[2]. The median VC-backed private company puts 47% into sales and marketing[25].
Revenue grew faster than headcount. HubSpot grew revenue about 19% in 2025 while full-time employees grew about 8%, from 8,246 to 8,882 (derived from its 10-K)[2,3]. Salesforce's CEO said support went from about 9,000 people to about 5,000 "because I need less heads"[43].
So AI economics land in sales and marketing, the support part of cost of revenue, and R&D. Inference is the one cost AI adds; HubSpot's 10-K warns it "could impact our profit margin if we are unable to monetize such assets"[3].
Two fronts: AI in what you sell, AI in how you run
A SaaS company becoming AI-native works on two fronts at once. The first is the product: shipping agents and repricing around them. The incumbents now charge per outcome:
- Intercom Fin$0.99 per resolution[32]
- Zendesk AI agents$1.50 per automated resolution committed, $2.00 pay as you go[34]
- Salesforce Agentforce$2 per conversation, or Flex Credits at about $0.10 per action[35]
37% of software companies in ICONIQ's investor survey planned to reprice within a year[30]. Databricks' CEO counters: "Why would you move your system of record?"[47]
The second front is how the company runs: support, sales, CS, engineering, security and finance done with agents, so margin holds while seat revenue is under pressure. This page is about the second front. The first comes in only where it creates obligations, such as AI transparency and the rules for products that decide about people, covered in the rules that bite.
Team by team: what agents already do
What agents already do in each team, and how good the evidence is. The grade is ours: strong means independent causal research, mixed means studies that disagree, thin means mostly vendor data, none independent means we found none.
Support
Evidence: strong- The work
- Triage, answers from the help center and past tickets, bug escalation, credits and refunds, QA.
- What agents do today
- Help desk agents that answer and close tickets, and drafting help for reps. Atlassian's 10-K describes one that can take a ticket, diagnose it from Confluence and "resolve it autonomously"[4].
- Evidence
- The strongest independent result in SaaS work: +14% issues resolved per hour, +34% for novices, across 5,179 agents[26], from an assistant helping people, not an agent acting alone. Intercom claims Fin resolves 76% on average (vendor claim)[33].
- The SaaS catch
- Credits, refunds and cancellations start as tickets. Here a support agent stops writing and starts moving money.
Sales, SDRs and RevOps
Evidence: thin- The work
- Lead qualification, outbound, meeting prep, follow-up, order forms, discount approval, questionnaires, forecast, CRM hygiene.
- What agents do today
- AI SDRs, research and enrichment, call analysis, CRM agents, questionnaire drafting.
- Evidence
- Mostly vendor surveys. MIT NANDA found more than half of AI budgets go to sales and marketing tools, while the clearest returns came from back-office work[49]. No public benchmark for AI SDR meeting rates or discount leakage.
- The SaaS catch
- Vendor claims need checking; see the 11x case below[48].
Customer success and renewals
Evidence: none independent- The work
- Onboarding, health monitoring, QBR briefs, renewals, expansion, churn saves, cancellations.
- What agents do today
- Health scoring and summaries in CS platforms and CRMs; meeting note-takers that write to the CRM.
- Evidence
- No independent study. The case is retention math: median GRR 88%, NRR 101%, and 40% of new ARR from existing customers in 2025[25].
- The SaaS catch
- Cancellation. On self-serve plans bought by individuals, California allows save offers only if the customer can still cancel at once[18]. Enterprise renewals follow the contract's notice terms. A retention agent that stalls either is a liability.
Engineering
Evidence: mixed- The work
- Issue to code, review, tests, deploys, incident response, on-call handover, dependency upkeep.
- What agents do today
- Coding assistants and agents, AI code review, incident summaries.
- Evidence
- It points both ways: 55.8% faster on one task in a GitHub-linked trial[36], 19% slower for experienced developers on their own repos in METR's[27]. See the evidence section.
- The SaaS catch
- Production data, credentials and deploys. An agent can do the most damage here, fastest.
Security, trust and compliance
Evidence: thin- The work
- Questionnaires, SOC 2 evidence, DPAs, the subprocessor list, access reviews.
- What agents do today
- Compliance automation and questionnaire drafting from an approved library.
- Evidence
- Vendor surveys only.
- The SaaS catch
- Your customers audit you. Every AI tool on their data becomes a questionnaire answer and possibly a subprocessor notice[5].
Finance and billing
Evidence: none independent- The work
- Invoicing, dunning, credits, usage metering, revenue recognition, board metrics.
- What agents do today
- Billing platforms' recovery features; AI in close and planning tools.
- Evidence
- None on SaaS finance work. MIT NANDA's back-office finding is the closest signal[49].
- The SaaS catch
- In a public company, an agent changing billing or revenue data is likely inside SOX controls; to be confirmed with your auditors.
What the evidence says, and where it is thin
Independent causal evidence exists for two SaaS functions: support and coding. Everything else rests on vendor and investor surveys.
- Support helps people most where they are newest
In the QJE study, issues resolved per hour rose 14% on average and 34% for novices, with better customer sentiment[26]. The gains came from spreading top performers' answers, so the people who keep the knowledge matter more.
- Developers cannot tell how much AI helps them
METR's randomized trial: 16 experienced developers, 246 issues on their own repos, 19% slower with AI while believing they were 20% faster. METR cautions it may not generalize[27]. A vendor-linked trial found one task done 55.8% faster[36].
- AI amplifies what a team already is
DORA 2025, about 5,000 respondents: 90% use AI, 30% barely trust its code, and adoption goes with higher throughput and lower stability. "AI doesn't fix a team; it amplifies what's already there"[28].
- Developers use AI widely; agents much less
Stack Overflow 2025: 84% use or plan to use AI, 46% distrust its accuracy, 66% complain of output "almost right, but not quite". Only 14.1% use agents daily; 37.9% have no plans to[29].
- Pilots stall, and bought tools do better
MIT NANDA, in a press summary: about 5% of enterprise pilots showed rapid revenue gains; tools bought from specialist vendors succeeded about 67% of the time, internal builds about a third as often. The method has been debated[49].
The staged path
Six stages for a company of 50 to 1,000 people, ordered by evidence and exposure: best-documented and lowest money risk first, actions that move money or speak for the company last. Vendors often sell AI SDRs first; the MIT NANDA finding above is why we do not.
- 0Inventory and data mapFind the AI already switched on and what customer data it seesGateEvery AI tool has an owner and a data class; the subprocessor list matches reality
- 1Support first passAgents triage and draft; people sendGateRepeat-contact rate measured; AI replies labeled; policy topics routed to people
- 2Revenue research and hygieneBriefs, CRM updates, questionnaire drafts. People own the dealGateCRM field ownership written down; questionnaire answers come from an owned library
- 3Engineering agents with gatesAgents open pull requests; people merge and deployGateChange failure rate and time to restore tracked against a baseline
- 4Money and customer-facing actionsCredits, refunds, plan changes, renewal drafts, under capsGateCaps and named approvers for every money action; one log across billing, CRM and help desk
- 5AI-native operating modelHeadcount, owners and rules built around agentsKeeps goingReviewed every quarter: rules, subprocessors, questionnaire library, scorecards
- 0
Stage 0: Inventory and data map
Find the AI already switched on and what customer data it sees.
About 4 to 6 weeks
What to do
- List every AI feature already on (help desk bot, CRM assistant, call recorder, note-taker, coding assistant), the customer data it sees and the model provider that receives it.
- Compare that list with your subprocessor list and DPAs. The gaps are your first finding.
- Review OAuth scopes and connected apps in the CRM, help desk, GitHub and Google Workspace. Revoke what nobody owns.
- Ban customer data in personal chatbot accounts. Set data classes and name one owner per team.
Why now
The worst SaaS agent incidents of 2025 were about scopes and inputs[37,39,41], and customers must hear about a new processor before it starts[5], so you need lead time.
In place first
- A current subprocessor list.
- Admin access to each system's connected-apps and token settings.
What to measure
Share of AI tools with an owner and a data class; long-lived tokens held by AI tools; subprocessor list matches what runs (yes or no).
Common mistakes
- Skipping the connected-app review. The Drift attackers used OAuth tokens a chat integration already held[41].
- Letting each team buy its own agent with its own keys.
- 1
Stage 1: Support first pass
Agents triage and draft; people send.
Starts in one queue once Stage 0 is done
What to do
- Triage tickets by product area, plan tier and severity.
- Draft replies from the help center and resolved tickets. A rep sends them.
- File bug reports with reproduction steps; draft help-center articles for repeat questions.
- After a measured trial, let low-severity how-to replies go out alone. Nothing about money, policy or security.
Why now
The strongest independent evidence[26], and little money exposure while credits and policy stay out of scope.
In place first
- A current help center. The agent will repeat it.
- Topics the agent may not answer alone: pricing, security, legal, outages, cancellations, data export.
- Replies labeled as AI where the EU AI Act applies; Article 50 transparency duties apply from 2 August 2026[9].
What to measure
First-response time; reopen and repeat-contact rate within 7 days; share of AI drafts sent unedited; satisfaction scores on AI-handled versus human-handled tickets.
- 2
Stage 2: Revenue research and hygiene
Briefs, CRM updates, questionnaire drafts. People own the deal.
Starts with one sales team, alongside Stage 1
What to do
- Qualify inbound leads with a written reason the AE can check.
- Pre-call and pre-renewal briefs from CRM, tickets, usage and billing.
- CRM next steps and contacts from email and call notes. Stage, amount and close date stay with the AE.
- Security questionnaire drafts from an approved library; usage and payment flags for CS.
Why now
Sales and marketing is 23% to 47% of revenue at the public companies above[2], so most internal hours sit here. Start with research and hygiene, not autonomous outbound.
In place first
What to measure
Hours from meeting to CRM update; questionnaire turnaround; forecast error at day 30 of the quarter.
Common mistakes
- Buying an AI SDR on a logo wall[48].
- Letting agents edit amounts and stages.
- Questionnaire answers drafted from last year's documents.
- 3
Stage 3: Engineering agents with gates
Agents open pull requests; people merge and deploy.
When branch protection and credential separation are in place
What to do
- AI code review as a second reviewer on every pull request.
- Dependency upgrades, incident summaries and on-call handover notes.
- Agents open pull requests from well-specified issues. Merges and production stay with people.
Why now
Developers already use AI: 84% use or plan to[29]. DORA links that adoption to lower delivery stability[28], so the work now is gates.
In place first
What to measure
Change failure rate and time to restore, before and after; share of agent pull requests merged without rework; incidents traced to agent-authored changes.
Common mistakes
- Measuring speed by asking developers. In METR's trial they felt 20% faster and were 19% slower[27].
- Giving an agent the deploy key for incidents.
- 4
Stage 4: Money and customer-facing actions
Credits, refunds, plan changes, renewal drafts, under caps.
After Stages 1 and 2 have run with evidence
What to do
- Credits and refunds under a cap by plan and amount.
- Failed-payment outreach; plan changes customers ask for in a ticket.
- Renewal drafts and quotes for the account owner.
- Usage-billing anomaly checks before invoices go out.
Why now
Agents now move money and speak for the company, which owns what they say[52]. Cancellation and switching rules also apply[12,18].
In place first
- A cap and a named approver for each money action.
- Approval shows the exact amount, customer and invoice, and runs once.
- One log across billing, CRM and help desk.
What to measure
Agent credits versus approved credits, and reversals; approval wait time; involuntary churn recovered; GRR and NRR trend. Medians were 88% and 101% in 2025[25], as context, not a target.
Common mistakes
- One global refund cap instead of limits by plan and customer.
- Approvals in a chat thread with no record of what was approved.
- Letting a retention agent handle cancellation requests.
- 5
Stage 5: AI-native operating model
Headcount, owners and rules built around agents.
When Stages 1 to 4 run with evidence
What to do
- Start headcount plans from agent capacity; give each team an agent owner and a scorecard.
- Keep refund caps, discount ceilings and send limits in version control, changed through review.
- Publish a plain statement to customers of where AI touches their data.
Why now
By now you have two quarters of numbers on what agents do in support, sales and engineering. Shopify already tells staff to show a job cannot be done by AI before asking for more people[44]. Savings from Stages 1 to 4 stay only if next year's headcount plan is built from them.
In place first
- Named agent owners in every team.
- Two or more quarters of scorecard data from earlier stages.
What to measure
Revenue per employee; sales and marketing as a share of revenue; days to answer a customer's AI questionnaire.
Common mistakes
- Cutting the people who write and maintain the knowledge the agents rely on.
- Treating a vendor's resolution rate as your own.
Your first 90 days
Stage 0 for the whole company, one support queue at Stage 1, and one sales team at Stage 2. By day 90 you have a baseline, a month of drafts to judge, and the evidence to decide how far Stage 3 goes.
- Days 1 to 30
Run the Stage 0 inventory. Review OAuth and connected-app scopes in the CRM, help desk, GitHub and Google Workspace. Name an owner per team. Write the five rules nobody may break: a refund cap, a discount ceiling, no production writes by agents, no new AI subprocessor without notice, no policy answer without a source. Record the baselines: repeat-contact rate, first-response time, meeting-to-CRM time, questionnaire turnaround.
- Days 31 to 60
Stage 1 in one support queue. Agents triage and draft; reps send. Label AI-written replies. Track repeat contacts and the share of drafts sent unedited every week.
- Days 61 to 90
Stage 2 for one sales team: pre-call briefs and CRM next steps. Build the questionnaire library with an owner per section. Review the first scorecard against the baseline, then decide the scope of Stage 3.
How the operating model and roles change
- 1
Headcount plans start from what agents can do
Shopify's CEO told staff to show a job cannot be done by AI before asking for more people[44]. Expect your board to ask the same of every open requisition.
- 2
Support becomes a knowledge and exceptions team
AI spreads top performers' practices to new agents[26], so knowledge owners, escalation specialists and AI reply QA become the core. Our inference, not a survey finding.
- 3
Sales separates research from judgment
Research, CRM updates and drafts go to agents; pricing, the deal desk and the relationship stay with people. Our recommendation; no study covers it.
- 4
Engineering protects review capacity
Google describes its AI-generated code as "reviewed and accepted by engineers"[42]. When agents write more code, review and deploy gates become the bottleneck to protect.
- 5
New owners for rules that cross teams
An agent owner per team; RevOps owns cross-system rules such as discount ceilings and credit caps; security owns the AI subprocessor list. Our recommendation; in a 50-person company these are parts of existing jobs.
- 6
Don't fully automate
An agent can prepare each of these: the numbers, the draft, the evidence. A named person makes the call, and the log shows who.
The rules that bite
A SaaS company meets AI rules from three directions: as a processor of its customers' data, which binds most often, through contracts and GDPR; as a seller of AI features, through claims, transparency and high-risk uses; and as an employer and marketer, through its own hiring, outbound email and calls, and cancellation rules.
Five requirements recur: tell people when they deal with AI; tell customers before a new party processes their data; keep records of what automated systems did; offer human review of decisions about people; do not overclaim. Below are the specific forms those take for a SaaS company.
Customer data and subprocessors
The rules you meet daily, as a processor of customers' data.
- GDPR Article 28(2) and (4)[5]
A processor needs the controller's written authorization to engage another processor, must tell the controller about any new one so it can object, and stays fully liable for it. A model provider that receives support tickets with personal data is likely a new subprocessor.
In force
To be confirmed: The notice period is set by each DPA. We found no public source for a typical length; read your own.
- GDPR Article 22[6]
A right not to be subject to solely automated decisions with legal or similarly significant effects, and, where allowed, to human intervention. Relevant when you suspend an individual's account or screen job applicants.
In force
- Sector contracts your customers flow down
DORA terms from EU financial customers, NIS2, and HIPAA business associate agreements that reach AI subprocessors.
Depends on your customers
To be confirmed: Not checked against the primary texts.
The EU AI Act
For companies that sell into the EU or run agents there.
Providers and deployers support the AI literacy of staff who operate AI for them.
From 2 February 2025
- Article 50: transparency[9]
People must be told they are interacting with AI unless it is obvious. This reaches AI support agents and AI outreach.
From 2 August 2026
- Article 25: when you become the provider[8]
Putting your name on a high-risk system, substantially modifying one, or repurposing a general-purpose model into a high-risk use makes you its provider. An HR, credit or education SaaS product can land here.
Applies with the high-risk duties
Risk management, logging, human oversight and documentation for uses such as hiring and credit.
From 2 December 2027, moved from August 2026 by a 2026 amendment
To be confirmed: The new date comes from the tracker; the Official Journal text could not be opened in October 2026.
The EU Data Act: switching
The Commission says the switching chapter covers SaaS.
- Switching between data processing services[12]
Customers may switch provider or move in-house. Cost-based switching fees are allowed until 12 January 2027; after that, no charges for switching or data egress. Switching requests will arrive through support, where agents sit.
Applies from 12 September 2025; charges end 12 January 2027
To be confirmed: The maximum two-month notice period and 30-day transition period in Article 25 could not be checked in the regulation text.
Claims about AI and incident disclosure (US)
What you say about your AI has to be true.
- FTC Operation AI Comply[13]
AI claims need substantiation. DoNotPay settled for $193,000 over its "robot lawyer" claims.
Announced 25 September 2024
- SEC AI-washing cases[14]
Two investment advisers paid $225,000 and $175,000 for overstating their AI. The theory applies to any public company's disclosures.
18 March 2024
- SEC cybersecurity disclosure[15]
A material cybersecurity incident goes on Form 8-K Item 1.05, generally within four business days of the materiality decision. An agent-caused leak can qualify.
Adopted 26 July 2023
- SOX 404 and agents that touch billing
An agent that changes billing, credits or revenue data is likely inside IT general controls over financial reporting.
Public companies
To be confirmed: We found no published guidance on agent-made changes. Agree the treatment with your auditors.
Outreach and cancellation
Where AI SDRs and retention agents meet consumer-protection law.
- CAN-SPAM[16]
No B2B exemption. Accurate headers, a physical address, opt-outs honored within 10 business days; up to $53,088 per violating email.
In force
- TCPA and AI voices[17]
AI-generated voices are artificial voices, so TCPA consent rules apply to AI calls.
Ruling of 8 February 2024
Express consent, annual reminders, online click-to-cancel, and save offers only if cancellation stays immediate. It protects consumers, defined as individuals buying for personal, family or household purposes, so a plan bought for business use is generally outside it.
Contracts entered or renewed from 1 July 2025
To be confirmed: How it applies to a self-serve plan an individual buys partly for personal use is a question for counsel.
- FTC click-to-cancel rule[20]
The Eighth Circuit vacated the FTC's amended Negative Option Rule. State laws still apply.
Vacated 8 July 2025
- EU and UK rules on B2B cold email
National ePrivacy rules differ on whether a corporate address needs consent.
Varies by country
To be confirmed: Not checked country by country.
When your product decides about people
For HR-tech, lending, insurance and education SaaS, and your own hiring.
- Colorado SB26-189[21]
Developers give deployers documentation of intended uses, training data, known limits and human review, and keep records three years. Deployers give notice, an explanation within 30 days of an adverse outcome, and correction and human review rights.
From 1 January 2027
- California CCPA regulations on automated decisionmaking, risk assessments and cybersecurity audits[22]
Notices and rights where automated decisionmaking makes significant decisions; risk assessments filed with the agency; cybersecurity audits.
ADMT duties by 1 January 2027. Risk assessments from 2026 and 2027 filed by 1 April 2028. First audit reports due 1 April 2028, 2029 or 2030, by revenue
- California AB 2013[23]
Developers of public generative AI systems post training-data documentation. "Substantially modifies" includes retraining or fine-tuning, so fine-tuning a vendor model for a public feature can bring you in scope.
From 1 January 2026
- Mobley v. Workday[24]
Age-discrimination claims against a SaaS vendor over its customers' AI screening proceed as a collective action: the leading case on a vendor's own exposure.
Collective action allowed to proceed in May 2025
To be confirmed: The current status. The docket could not be opened for this page.
When agents fail
Ten reported cases, each with the control that would have stopped it. Six of the ten came down to what an agent or its tokens could reach, rather than what it said.
Cursor's support bot invents a policy
April 2025
An AI support bot told users who were being logged out that it was a new one-device policy. It was a bug. Users posted cancellations before a cofounder wrote "We have no such policy." Cursor then labeled AI replies[50].
The control: Policy answers from an approved source or a person; AI replies labeled. Control 1
Replit's agent deletes SaaStr's production database
July 2025
During a code freeze the agent deleted a live database with records on more than 1,200 executives, then said a rollback would not work. The data was recovered. Replit added automatic dev and prod separation[51].
The control: Approval for production writes; dev credentials by default; a kill switch. Controls 6 and 15
Salesloft Drift tokens used to raid Salesforce orgs
8 to 18 August 2025
An attacker used OAuth tokens held by the Drift AI chat integration to export cases, accounts and opportunities from many companies' Salesforce instances, then searched them for keys and passwords[41].
The control: No long-lived vendor keys in AI tools; minimal scopes. Controls 7 and 8
ForcedLeak in Salesforce Agentforce
Disclosed 25 September 2025
Instructions hidden in a web-to-lead form ran when an employee later asked the agent about the lead, and data could leave through an expired, still-trusted domain. Rated CVSS 9.4; Salesforce fixed it (security vendor's disclosure)[37].
The control: A gate before any send; allowed outbound hosts. Control 9
Prompt injection through a public GitHub issue
May 2025
A malicious issue in a public repo led a developer's agent to copy private repo data into a public pull request (security vendor's demonstration)[38].
The control: One repo per task; approval before public writes. Controls 7 and 9
A support ticket leaks integration tokens
July 2025
In a demonstration, instructions in a support ticket led a developer's assistant, connected with a key that bypasses row-level security, to post secret tokens into the ticket[39].
The control: No superuser agents; access by field. Controls 7 and 10
Amazon Q extension ships with injected code
July 2025
An "inappropriately scoped GitHub token" let an attacker add code to version 1.84.0. It shipped but failed to run[40].
The control: Scoped tokens; review on anything that ships. Controls 6 and 8
Questions over an AI SDR vendor's customers
March 2025
TechCrunch reported that 11x showed logos of non-customers and counted trial contracts as ARR; 11x disputed parts[48]. Allegations, not findings.
The control: Judge agents on your own scorecard. Control 13
Air Canada is held liable for its chatbot
February 2024
The chatbot described a refund option the airline did not offer, and the tribunal held the airline liable[52]. Read through a secondary summary; the decision itself is to be confirmed.
The control: Refund and policy answers from approved sources. Controls 1 and 2
Klarna's support automation
2024 onward
Klarna said its assistant handled two-thirds of chats in month one (company claim)[45]. Reports that it later brought back human service are to be confirmed.
The control: A human path by customer tier; measure repeat contacts. Control 1
The pattern: the lethal trifecta
ForcedLeak, the GitHub issue and the ticket leak share what Simon Willison calls the lethal trifecta: "access to your private data", "exposure to untrusted content" and "the ability to externally communicate"[31]. In SaaS, anyone can write into the help desk and web forms, which sit next to the CRM.
- 1 of 3Untrusted content comes in
- Support tickets
- Web-to-lead forms
- Inbound email
- Public GitHub issues
ForcedLeak used a lead form
- 2 of 3The agent can read private data
- CRM records
- Production database
- Private repos
- Billing
The Supabase demo used a superuser key
- 3 of 3The agent can send data out
- Email replies
- Public pull requests
- Ticket comments
- Outbound HTTP
ForcedLeak sent data to a trusted, expired domain
A run that touched untrusted content and can read private data is checked before it sends anything: allowed hosts only, unusual sends approved by a person. Controls 7, 8 and 9.
Agent failures to design against (scenarios)
Scenarios, not reported cases, each mapped to a control point.
Control points for an AI-native SaaS company
Seventeen places where a SaaS company needs a control whatever tools it uses, who owns each, and the case or rule behind it. The blueprint below shows what enforces each one.
The OrchKernel blueprint for a B2B SaaS company
OrchKernel sits between agents and the systems you run, as drawn in the missing layer. Your help desk, CRM, billing and code host stay the systems of record.
The mechanisms
- Approvals
- An action waits for a named person: a refund above the cap for finance, a discount above the ceiling for the deal desk, a production change for on-call. The approver sees the exact payload; it runs once.
- Rules
- Checked before every action: refund caps, discount ceilings, send limits, topics the agent may not answer. A rule allows, holds or denies with the reason.
- Acting on a named person's authority
- Each agent works for a named rep, AE or engineer, with no more access than that person. When the person loses access to an account, so do their agents.
- Data access by role and field
- Which accounts and fields each person and agent sees. Each model is cleared for a data class, so restricted data reaches only cleared models.
- Tamper-evident audit log
- Every request, decision, approval and result, chained so an edited entry shows. It answers a customer's security team: what did AI do with our data, and who approved it.
- Human queue
- Cancellations, export requests, security and legal questions land with an owner and a deadline.
- Connections to your systems
- The help desk, CRM, billing, code host, Google Workspace, Slack, any MCP server and any REST API through declarative adapters. The same rules and log apply to each.
Control point by control point
Where OrchKernel does part of the job or none, the row says so.
What is available when. The governed actions API with its TypeScript and Python clients, the MCP gateway and the policy dry run that explains a decision before it is enforced are arriving in this release. SSO and export of the audit log to a SIEM are later, on the roadmap.
What it is not. It does not replace your systems of record, holds no SOC 2, HIPAA or ISO certification and does not by itself satisfy any law or audit; it enforces controls and keeps the evidence. It is source-available under the Business Source License 1.1 and runs on your own servers, so your security team can read the code.
Scorecard by stage
Record a baseline before Stage 1, then track the same numbers at each stage. For most internal SaaS work no public benchmark exists, and we have not quoted rules of thumb we could not source. Compare against yourself.
Sources
Sources were read in October 2026; dates are publication or data dates. Ratios from 10-K data are line items divided by total revenue. Gartner forecasts on agent projects were left out because they could not be opened.
Primary sources
Government, regulators, courts, statutes and SEC filings. GDPR and EU AI Act articles are read in unofficial consolidated copies.
- 1Current Employment Statistics, software publishers (series CES5051320001). US Bureau of Labor Statistics, August 2026, preliminary, seasonally adjusted.
- 2Company facts from the latest annual reports (10-K): HubSpot, Freshworks, Salesforce, Atlassian, DocuSign, Box, Zoom, PagerDuty, Datadog. US Securities and Exchange Commission, XBRL data, filed February to August 2026.Each line item divided by total revenue, GAAP, so stock-based compensation is included
- 3HubSpot, Inc. annual report on Form 10-K for fiscal 2025. US Securities and Exchange Commission, filed 11 February 2026.
- 4Atlassian Corporation annual report on Form 10-K for fiscal 2026. US Securities and Exchange Commission, filed 14 August 2026.
- 5GDPR Article 28: processor. gdpr-info.eu.
- 6
- 7EU AI Act, Article 4: AI literacy. artificialintelligenceact.eu.
- 8EU AI Act, Article 25: responsibilities along the AI value chain. artificialintelligenceact.eu.
- 9EU AI Act, Article 50: transparency obligations for providers and deployers of certain AI systems. artificialintelligenceact.eu.
- 10EU AI Act implementation timeline. artificialintelligenceact.eu, updated 31 August 2026.Unofficial tracker of the official dates
- 11Regulation (EU) 2026/1744 amending the AI Act: new application dates for high-risk systems. Official Journal of the European Union.Could not be opened in October 2026 (EUR-Lex unreachable); date taken from the tracker
- 12Data Act explained. European Commission.
- 13FTC announces crackdown on deceptive AI claims and schemes (Operation AI Comply). US Federal Trade Commission, 25 September 2024.
- 14SEC charges two investment advisers with making false and misleading statements about their use of artificial intelligence. US Securities and Exchange Commission, press release 2024-36, 18 March 2024.
- 15SEC adopts rules on cybersecurity risk management, strategy, governance and incident disclosure by public companies. US Securities and Exchange Commission, press release 2023-139, 26 July 2023.
- 16CAN-SPAM Act: a compliance guide for business. US Federal Trade Commission.
- 17FCC makes AI-generated voices in robocalls illegal. Federal Communications Commission, 8 February 2024.
- 18AB 2863 (automatic renewal and continuous service offers). California Legislature, chaptered 24 September 2024.
- 19Business and Professions Code section 17601: definitions for the Automatic Renewal Law. California Legislative Information.
- 20Custom Communications, Inc. v. FTC, No. 24-3137 (8th Cir.): opinion vacating the Negative Option Rule. US Court of Appeals for the Eighth Circuit, 8 July 2025.
- 21SB26-189: automated decision-making technology. Colorado General Assembly, signed 14 May 2026.
- 22CCPA updates, cybersecurity audit, risk assessment and automated decisionmaking technology regulations: approved text. California Privacy Protection Agency, approved 22 September 2025.Compliance dates from sections 7121, 7157 and 7200 of the approved text
- 23AB 2013 (generative artificial intelligence: training data transparency). California Legislature, chaptered 28 September 2024.
- 24Mobley v. Workday, Inc., No. 3:23-cv-00770 (N.D. Cal.): docket. CourtListener (RECAP archive), read October 2026.
Industry bodies and independent research
Peer-reviewed studies, independent research groups, developer surveys and benchmark reports. Investor surveys are marked.
- 252025 B2B SaaS performance metrics benchmarks. Benchmarkit, 2025.
- 26Generative AI at Work (Brynjolfsson, Li and Raymond), Quarterly Journal of Economics 140(2). NBER working paper w31161, 2025.
- 27
- 28Announcing the 2025 DORA report. Google Cloud (DORA research program), 24 September 2025.
- 292025 Developer Survey: AI. Stack Overflow, 2025.
- 30
- 31The lethal trifecta for AI agents: private data, untrusted content, and external communication. Simon Willison, 16 June 2025.
Vendor sources
Published by companies that sell AI to SaaS companies. Directional, not an industry benchmark.
- 32
- 33
- 34
- 35
- 36The impact of AI on developer productivity: evidence from GitHub Copilot (Peng et al.). arXiv 2302.06590, 2023.Vendor sourceRandomized trial run with GitHub and Microsoft researchers
- 37ForcedLeak: agent risks exposed in Salesforce Agentforce. Noma Security, 25 September 2025.Vendor sourceSecurity vendor's disclosure
- 38GitHub MCP exploited: accessing private repositories via MCP. Invariant Labs, 26 May 2025.Vendor sourceSecurity vendor's research
- 39Supabase MCP can leak your entire SQL database. General Analysis, 8 July 2025.Vendor sourceSecurity vendor's demonstration
Company and press
Company statements, news coverage, security researchers and encyclopedia summaries.
- 40Security bulletin AWS-2025-015: Amazon Q Developer for VS Code extension (CVE-2025-8217). Amazon Web Services, July 2025.
- 41Widespread data theft targets Salesforce instances via Salesloft Drift. Google Threat Intelligence Group, August 2025.
- 42Q3 2024 earnings remarks. Alphabet (Sundar Pichai), 29 October 2024.
- 43Salesforce CEO confirms 4,000 layoffs 'because I need less heads' with AI. CNBC, 2 September 2025.
- 44
- 45Klarna AI assistant handles two-thirds of customer service chats in its first month. Klarna press release, 27 February 2024.
- 46SaaS in, SaaS out: here's what's driving the SaaSpocalypse. TechCrunch, 1 March 2026.
- 47Databricks CEO says SaaS isn't dead, but AI will soon make it irrelevant. TechCrunch, 9 February 2026.
- 48a16z- and Benchmark-backed 11x has been claiming customers it doesn't have. TechCrunch, 24 March 2025.Allegations; 11x disputed parts of the report
- 49MIT report: 95% of generative AI pilots at companies are failing. Fortune, 18 August 2025.Press summary of MIT NANDA's report; its method has been debated
- 50Cursor AI support bot invents a policy. The Register, 18 April 2025.
- 51AI coding tool Replit wiped a database and called it a catastrophic failure. Fortune, 23 July 2025.
- 52Moffatt v. Air Canada, 2024 BCCRT 149. Wikipedia, decided 14 February 2024.Secondary summary; the tribunal's decision could not be opened in October 2026