The Anatomy of Token Overspend
Agentic AI workflows are structurally different from a standard chatbot exchange. A user doesn’t ask a question and receive an answer. An agent plans, executes, iterates, and often backtracks — generating tokens at every step, including the wrong ones.
Jenkins identified two primary drivers of his own overruns: model selection and conversational drift. The first is a routing problem. Not every task requires the most capable — and most expensive — model in a provider’s lineup. Sending a basic arithmetic check to a high-reasoning model is the AI equivalent of hiring a senior architect to hang a picture frame.
The second is subtler. Drift happens when a session wanders. The user engages, the agent elaborates, the user encourages more, and the conversation migrates far from the original intent. “You kind of find yourself just chatting, and things getting away from you,” Jenkins said. “All of a sudden you end up in who-knows-where, and you’re like, ‘No, I don’t want that at all.’” Every token spent reaching that dead end is waste with no refund mechanism.
Industry-level data suggests this is not an isolated pattern. Estimates from research firms indicate agentic workflows can require significantly more tokens per task than standard interactions — in some cases by an order of magnitude. Reports of major enterprises burning through annual AI budgets within months, or spending nine figures in a single month after deploying access without usage caps, point to the same structural vulnerability at scale.
Model Routing as a Cost Control Discipline
Jenkins’s response was practical and, notably, self-taught — largely from optimization content circulating on social platforms rather than from internal engineering guidance. He began routing tasks by complexity:
- Lightweight models for simple operations like basic math or formatting
- Mid-tier models for routine coding and standard drafts
- High-reasoning models reserved for genuine strategic or architectural decisions
He also adopted orchestration layers — tools designed to compress AI output and reduce token volume. One approach he described forces the AI to respond in short, blunt sentences rather than elaborated prose, which he estimated reduces token consumption by roughly 70%. Another compresses replies to near-minimal length for tasks where brevity is sufficient.
The tools exist. The problem is distribution. Jenkins was candid: these are techniques that require a degree of technical curiosity and time investment that most employees — in sales, in marketing, in customer success — neither have nor should be expected to develop. “Do we need sales leaders and service leaders and marketers finding this stuff?” he asked. The answer, implicitly, is no. Which means the efficiency gap between power users and the rest of the workforce compounds over time.
The Governance Gap: Vibe-Coded Tools and Continuity Risk
A related problem is emerging quietly in the background of many organizations: informally built, employee-created AI tools that have become operationally significant without ever going through a formal development or review process.
Jenkins refers to these as “vibe-coded” tools — software assembled quickly, often on weekends, by technically capable employees who saw a need and filled it. The risk is not that these tools don’t work. The risk is that they work well enough to become load-bearing, and then the person who built them leaves, gets sick, or simply moves on.
His response has been structural: expanding the DevOps organization specifically to govern this category of internal tooling, applying continuity, security, and scalability standards to software that originated as side projects. “It can’t just be with Tim that vibe-coded it on the weekend,” he said.
This is a governance problem that most enterprise AI cost discussions underweight. Token spend is visible on a dashboard. A critical internal automation with no documentation, no owner, and no failover is not.
The Human Risk Layer: Insecurity as an Organizational Force
When asked to rank the problems he’s encountered rolling out AI across several hundred employees, Jenkins didn’t lead with cost. He named three: inefficiency, inequality, and insecurity — and he returned to the third repeatedly.
The insecurity he describes is not abstract. It manifests in specific, observable behaviors. When Jenkins built automations inside a leadership team member’s department — unprompted, because he’d learned how — the reaction was unease rather than appreciation. “This put me on edge. I should be coming to you with these things. I’ve got to catch up. I feel so behind.”
The same dynamic repeated at multiple management layers: Am I doing my job? Am I keeping up? Will this replace my team?
Unequal Access Compounds the Problem
Jenkins’s company initially deployed ChatGPT organization-wide, then licensed a more capable model for a smaller group concentrated in sales and marketing. The reaction from teams left out was immediate. “People were like, wait a minute, why don’t I have Claude? Why do they get that and we don’t get that?”
This is a real organizational dynamic. When AI capability becomes a visible differentiator between departments — not just between companies — it introduces a new axis of internal inequality. Employees who lack access to better tools aren’t just less productive; they’re aware of the gap, and that awareness shapes behavior.
Jenkins’s argument is that this anxiety is more corrosive to adoption than any single runaway invoice. An employee who feels threatened by the technology, or who perceives it as unevenly distributed, is less likely to engage with it seriously — and more likely to treat cost concerns as a convenient reason to disengage.
Hybrid Org Design: Mapping Humans and Agents Together
Jenkins’s structural response to the human-agent operating model has been to make it explicit. His executive team went through a formal org-design exercise in which each leader mapped their department to include not just the people who report to them, but the AI agents those people now manage directly.
The output is a hybrid org chart — human roles and agent functions layered together — that the company treats as a living management document rather than a one-time exercise.
This is a meaningful operational shift. It acknowledges that agents are not just tools; they are, in a functional sense, participants in workflows that require oversight, continuity planning, and accountability structures. Treating them as invisible infrastructure is how organizations end up with both runaway token bills and undocumented automations running critical processes.
The Underlying Labor Math
The most significant financial signal Jenkins points to isn’t the token bill. It’s headcount growth relative to revenue.
His company’s ratio of annual recurring revenue per employee has been rising — not because of layoffs, but because the organization is no longer scaling headcount at the rate it once did to support revenue growth. “I’m not arguing that we want to reduce a whole bunch of headcount because of AI, but we should not be growing the headcount at the same rate that we were before,” he said. “That’s a big change in our business, and it’s all attributed to AI.”
This is the cost story that matters at the enterprise level. A $1,000 weekend token overrun is, as Jenkins acknowledges, a rounding error. The structural shift in the relationship between labor and output is not.
What This Actually Requires
The practical takeaway from Jenkins’s experience is not that agentic AI is too expensive to use. It’s that using it without deliberate cost architecture, governance structures, and organizational change management produces predictable and avoidable problems.
Specifically:
- Model routing is not optional. Defaulting to the most capable model for every task is a cost decision, not a quality decision. Organizations need routing logic — whether human, automated, or both.
- Prompt drift requires session discipline. Agentic workflows need defined scope boundaries. Open-ended sessions with no exit criteria burn tokens and produce outputs that require rework.
- Vibe-coded tools need a governance path. Informal internal automation is happening whether organizations acknowledge it or not. The question is whether it gets managed before it becomes a continuity risk.
- Access inequality is a change management problem. Uneven tool distribution doesn’t just affect productivity — it affects trust, morale, and adoption rates across the organization.
- The org chart needs to reflect reality. If agents are doing meaningful work, they belong in the operational model — with owners, oversight, and accountability — not in the background as invisible infrastructure.
The token bill at dinner is a symptom. The underlying condition is an operating model that hasn’t yet caught up with the technology it’s running.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!