Spending caps are a confession

Spending caps are a confession

Cohere just launched North 2, its workplace agent platform, "now with memory and spending caps." Memory makes sense. Agents that forget everything between runs are annoying. The spending caps are the part I keep thinking about.

Here is my take, and I know some people will disagree: if an agent platform needs a spending cap, it is telling you it can't predict what its agents will do. A cap is a guess about the worst case, written as a dollar amount.

That's fine for a side project. For a bank, an insurer, a hospital, or a government agency, it misses the point. Those buyers need to know what an agent will do before it runs. How much it cost afterward is a question for finance.

A spending cap is a billing feature

Let me be fair. Spending caps are useful. Nobody wants a runaway loop burning through tokens over a long weekend. Every cloud account has a budget alert for a reason, and I'm not knocking Cohere for shipping one. My argument is about the category, not one product.

But look at what a cap actually controls. It controls cost. It does not control which systems the agent touched, which records it read, which emails it sent, or which decisions it made on the way to the limit. An agent that wires money to the wrong account for 40 cents of compute is well under budget.

Governance answers different questions. What is this agent allowed to do? What data can it see? Who signs off before an action is final? Where is the record of every step? A budget ceiling answers none of these. It tells you when to stop worrying about the invoice. It tells you nothing about when to start worrying about the agent.

Why agents get unpredictable

Agents become hard to predict for a simple reason: we give them too much. One agent, one giant prompt, access to every tool in the company, and a goal like "handle onboarding." Then we act surprised when it takes a path nobody planned.

When the scope is wide open, the number of possible actions is huge, and cost becomes the only thing you can easily measure. So you cap it. The cap is a symptom. The real problem is the design.

The fix is boring, which is why it works. Break the work into small tasks. Give each agent one job, the minimum tools to do it, and a clear handoff to the next step. Put a person at the points where judgment or liability lives. Log everything. Do this and cost becomes predictable as a side effect, because each step does a known, bounded amount of work.

Example: KYC at a bank

Take a know-your-customer workflow for opening a business account. The "one big agent" version reads the application, pulls documents, searches the web, decides whether the customer is risky, and maybe opens the account. Good luck explaining that to an examiner.

Here is the version I would build:

  • Intake agent: reads the application and checks it for completeness. It can read the form and write to a case record. Nothing else.

  • Document extraction agent: pulls entity names, beneficial owners, addresses, and registration numbers from the uploaded documents into a fixed schema. Read-only access to the document store.

  • Sanctions check agent: runs the extracted names against the sanctions and PEP lists the bank already uses. It calls one screening service and returns matches with confidence scores. It cannot approve or reject anyone.

  • Analyst review: a human compliance analyst sees the case, the extracted fields, the screening results, and the source documents side by side, then approves, rejects, or sends it back.

Each agent has a narrow job, so you know what it will do before it runs. If the sanctions agent ever tried to write to the core banking system, it would fail, because nobody gave it that tool. The audit trail shows every input, output, and decision. Cost per case is easy to forecast because every step does the same bounded thing each time. We go deeper on this pattern in our guide to AI agents for banks and financial services.

Example: prior authorization at a hospital

Prior auth is a good test because it is high volume, heavy on rules, and full of protected health information. A scoped version looks like this:

  • An intake agent pulls the order and the payer from the EHR.

  • A policy agent looks up that payer's criteria for that procedure.

  • A clinical evidence agent extracts the relevant notes, labs, and imaging results, only for that patient and that encounter.

  • A drafting agent assembles the request packet and flags any criteria that are missing.

  • A nurse or utilization reviewer approves the packet before anything goes to the payer.

None of these agents needs the whole chart, and none can submit anything on its own. That is least privilege in practice. When an auditor asks why a request went out, you can show the exact evidence the agent used and the person who approved it. Try doing that with a spending cap.

How we build this at StackAI

This is how we think about StackAI. Teams build agentic workflows where each agent owns one atomized task inside a larger org process. The process is the thing you are governing. The agents are workers inside it, each with a job description.

A few things make that hold up in regulated settings:

  • Scoped tools. Agents connect to the systems they need through 300+ integrations and MCP servers, with access set per agent and per step. We wrote about why this matters in MCP for regulated enterprises.

  • Contained execution. When an agent needs to run code or work in a real environment, it gets a sandbox, a computer, or a terminal that is isolated from everything else. More on that in sandboxes and terminals for enterprise agents.

  • Human approval steps placed exactly where the risk is, as part of the workflow itself.

  • Deploy where your data lives: HIPAA and GDPR ready, on-prem, in your VPC or private cloud, or on StackAI cloud. Details are on our security page.

  • People who have done this before. Forward-deployed engineers and AI strategists work with your team to map the process before anyone writes a prompt.

It also stays low-code, so the compliance team can read the workflow without a translator.

Do we track cost? Of course. But on a well-scoped workflow, cost is a line item you can forecast, not a fence you hope holds.

What to ask your vendor

Next time an agent platform walks you through its spending caps, ask these instead:

  • Can I see, before it runs, every tool and data source each agent can access?

  • Can I limit an agent to read-only on a specific system?

  • Where can I require a human to approve before an action is final?

  • Is there a step-by-step audit log of inputs, outputs, tool calls, and approvals?

  • Can I split one process into several agents with separate permissions?

  • Can I deploy on-prem or in my own VPC?

  • What happens when an agent tries something outside its scope?

If the answers are vague and the demo keeps drifting back to the budget dashboard, you have your answer.

Predict the work and the bill follows

I'm not against spending caps. I'm against selling them as safety. A cap tells you a platform has a ceiling. It tells you nothing about what happens underneath it.

Regulated teams don't get to tell a regulator "the agent stayed under budget." They have to explain what the agent did, why, and who approved it. Build for that, and the bill becomes the least interesting part of the conversation.

If you want to see what a scoped, auditable workflow looks like on one of your own processes, book a demo. Bring your messiest workflow. We like those.

Bernard Aceituno – Co-Founder and President at StackAI
Bernard Aceituno

Co-Founder at StackAI

Table of Contents

Make your organization smarter with AI.

Deploy custom AI Assistants, Chatbots, and Workflow Automations to make your company 10x more efficient.