From Tokenmaxxing to Token Panic
Rippling, best known as an HR and workforce management platform, went deep on AI adoption at the start of the year. The result: token spend was tracking toward 40% of its entire R&D headcount budget, growing at 80% month-over-month. A small slice of employees — roughly 10–15% — were driving about 60% of total spend.
The instinct wasn’t to shut it down. It was to figure out what was actually happening, and why.
What they found was predictable in hindsight: employees defaulted to the newest, most expensive frontier models for everything. Grammar fixes. Routine lookups. Tasks that a cheaper model handles just fine.
The Inference Provider Problem
Rippling’s Chief Product Officer Matt MacInnis put it plainly: the inference providers — OpenAI, Anthropic — have no incentive to help you control your spend. They don’t offer great usage visibility, and they don’t coordinate with each other. Every dollar you burn is a dollar they keep.
That’s not a criticism so much as a structural reality. Enterprises that treat AI spend like a utility bill are going to get surprised.
What AI Spend Console Actually Does
The product has two core components working together:
- A spend dashboard that tracks token usage per employee, team, and tool — scoring attributes like prompts per day alongside actual work output (lines of code, pull requests, onboarded customers).
- An AI gateway that routes prompts to the most cost-effective model for the task at hand, rather than defaulting to whatever’s newest and priciest.
The routing piece is where the real savings live. Rippling’s own benchmarks found that models like Z.ai’s GLM 5.2 delivered near-identical performance to frontier models for coding tasks at 85% lower cost. Grok performed well across the board. The point isn’t loyalty to any one lab — it’s matching the model to the job.
The Results
Rippling says it brought token spend down from 40% of its R&D headcount budget to around 15%. In April, at peak panic, the company consumed 605 billion tokens. In July, it hit 600 billion tokens again — but the cost was 37% of what April’s bill had been.
Same usage. Dramatically lower cost. Just smarter routing.
The Harder Problem: Beyond Engineering
Software engineers were the easy case. They have measurable outputs — commits, pull requests, code shipped. The dashboard can connect token spend to productivity without much ambiguity.
The harder challenge is everyone else. Rippling is working on extending this to customer onboarding teams, where productivity might be measured in customers onboarded or data reconciliation tasks completed. But MacInnis is candid: if the company can’t link token consumption to real output in non-engineering functions, broader employee access may not be justified.
That’s a meaningful shift in how enterprises might think about AI access. It may no longer be a default perk like email or Slack. It could become something you earn — or something that gets metered based on demonstrable productivity.
Who It’s For
AI Spend Console is included for existing Rippling HR subscribers (with usage-based AI costs on top) and available as a standalone product that can integrate with other HR systems of record.
If you’re already running an AI gateway elsewhere, you can still use the spend analytics layer — but the spend governance features require Rippling’s own gateway.
The Useful Takeaway
The tokenmaxxing era had a predictable ending: uncapped spend, no visibility, and a CFO with a very uncomfortable slide deck. The correction isn’t to restrict AI — it’s to route it intelligently and measure what it actually produces.
Rippling built this tool for itself first. That’s usually a decent sign it solves a real problem.
Comments (0) No comments yet
Want to join this discussion? Login or Register.
No comments yet. Be the first to share your thoughts!