Hello Reader
Uber burned through its entire 2026 AI budget in four months. Open AI CEO Sam Altman told reporters recently that enterprise customers were calling him to say their annual AI budget was gone by the end of Q1.
If that sounds familiar, you're not alone. Budgets set in late 2025 weren't built for AI agents.
A growing AI invoice is usually a sign your team is doing real work with the technology, running agents, building workflows and producing more than your headcount would normally allow. Most organizations don't have a framework for knowing whether that bill is worth paying, and that's where the spend goes sideways.
This edition gives you the framework for knowing what your spend is producing and the specific moves for keeping it under control without cutting off the workflows your team depends on.
What’s Driving the Spike
When I first signed up for Claude and started building workflows in Cowork, I let things run unattended. I hit my limits quickly and added overages, but I kept going over budget because I wasn’t watching what Claude was doing in the background.
You type a question into your AI tool, it answers and the whole thing wraps up in about four seconds. An agent doing the same kind of work is a completely different operation. It makes dozens of background requests per task, sometimes hundreds, drafting, checking, cross-referencing and reformatting without you lifting a finger. Every one of those steps consumes tokens, and the meter doesn't care whether you're watching.
—
A token is the unit an AI model uses to process text, everything you send in and everything that comes back. Think of it as a meter that never stops running. A quick question and answer might use a few hundred. An agent doing complex, multi-step work can burn through tens of thousands before it's done. So when your team hits their limit mid-task and comes to you about it, the meter just ran out.
—
Goldman Sachs projected in May that token consumption will multiply 24 times between 2026 and 2030.
For marketers and agency owners, you can already feel that coming.
Content creation, personalization, campaign analysis, email sequences and SEO research are all agent-heavy workflows, and every one of them will show up on your invoice in the next three months.
The harder problem is visibility.
You’re looking at a growing invoice, fielding complaints from your team about running out of tokens and trying to figure out what to do without cutting access to the tools.
Most leaders can’t account for which workflows are consuming the most tokens or which ones are worth the spend. Which means every budget decision becomes a guess.
Why a Rising AI Bill Is Worth Paying
Michael Domanic, Head of AI at Section AI, expected to spend $20,000 on inference last month. His team landed closer to $30,000, and his reaction was that this was a good sign.
If your AI line item is flat, it probably means your team is using AI for individual tasks like drafting and research. The teams spending more are the ones where AI is embedded in recurring workflows.
If your AI budget is running into the thousands or tens of thousands, and that spend is enabling your team to do the work of two or three people, you’re still saving money.
Domanic ran the numbers with his CTO and found that across a 50-person company, annual inference costs came to roughly the equivalent of two full-time employees, but the output of that spend was distributed across every person in the organization.
What the Teams Pulling Ahead Are Doing Differently
The organizations further along in their AI programs share three things in common, and none of them are training programs or standard tool lists (which are now considered table stakes).
The three things that separate mature AI programs from early-stage ones are integration, governance and measurement.
Here’s what that means:
Wiring AI into Existing Systems
A standalone AI tool your team opens in a separate tab is a productivity aid. AI connected to your platforms and daily work is a system, and it’s the system that allows you to 10x outputs.
Let’s look at a real life example.
A BDR using Claude might ask it to help draft a follow up email they paste into Gmail. This might save them a few minutes but it’s not transforming their work.
Turning this into a system might mean connecting Claude with Zoom and Hubspot, and designing a workflow that drafts follow-up emails based on the transcript, creates content based on prospect pain points and sends you call insights via Slack so you can keep up with buyer trends.
Building Governance and Oversight
A recent Notion survey of 6,000 working professionals found that half of AI decision makers can’t account for which tools their employees are using.
Get ahead of it: an approved tool list, documented use cases and a clear review process for what goes out the door.
Measuring the Right Things
Mature AI programs track quality metrics (error rates, rework), workflow metrics (cycle time, throughput) and financial metrics tied to cost and revenue. Early-stage teams rely on self-reported time saved, and while that metric is useful for internal momentum, it doesn’t hold up in a budget conversation with a CFO or a client.
The teams that can defend AI spend are the ones measuring outputs and outcomes.
How to Manage Your AI Spend
Managing rising AI costs well requires an intentional strategy.
Here’s how to tackle it:
- Audit before you restrict. Before setting new limits, understand where your tokens are going. Which workflows are consuming the most? Which ones are producing outcomes you can point to? A blanket cap will cut productive workflows along with wasteful ones, so let the picture drive your decisions.
- Match the model to the task. The most capable model is also the most expensive, and it’s overkill for most of what marketing teams do day to day. Running a frontier reasoning model on your social captions costs more and produces the same output that a lighter model would generate. Build a habit of deliberate model selection: save the expensive models for complex, high-stakes work and use lighter ones for routine output.
- Set caps to prevent runaway automations. When I let Cowork run unattended, it was running in pointless loops. Those limits exist to catch misconfigured workflows and keep a runaway automation from draining your budget overnight. If someone on your team is consistently hitting their ceiling, have a conversation before making any decisions. They might be the person keeping more work afloat than anyone realizes.
- Track spend at the workflow level. Individual agent costs don’t tell you much in isolation. What matters is the total cost of a workflow relative to the business outcome it’s producing, and you’ll want to give it at least a quarter before drawing conclusions. The meaningful indicators take time to show up.
Resist the urge to cut spending too early. If your costs are rising because your team is building real workflows and running agents on real work, that line item should be growing.