Founder Insights

The Smarter Way to Manage AI Spend

By Aneesh Sachdeva, Founder & CEO of Glia Intelligence

Last week one of our clients asked whether they should start migrating to open-source models. Their owners had seen the numbers everyone is seeing. Claude Fable 5, the most intelligent model on the market, runs about $20 per million tokens at list price. GLM-5.2, an open-source model with credible benchmark scores, runs about $2.15 per million tokens. When the gap looks like 10x, two questions become impossible to ignore: are we spending too much on AI? And should we be moving off of Claude or ChatGPT, or investing less in them? Underneath both sits the same fear of becoming over-reliant on a single provider.

The cost concern is both everywhere and real.

Everywhere: When we hosted a dinner for 40+ C-suite and VP finance operators in June, most said they feel little AI cost pressure today, and almost all of them expect it soon. We call this the calm before the bill.

Real: Open source absolutely has a role to play. But switching on the hype won't solve the concern, because cost anxiety is usually a measurement problem wearing a costume. If you knew the ROI on every dollar of token spend, you wouldn't be capping spend; you'd be budgeting it the way you budget labor.

Since 2024 we've transformed 52+ teams across 20+ businesses with a team of 7, and this article is the playbook we run: where AI costs actually come from, why this is a measurement problem, a framework for deciding when and how to move work to cheaper models, and the usage policy we recommend to clients.

1. Token spend is labor spend

Software spend is fixed. A seat costs the same whether someone logs in every day or never. Token spend scales with the amount of work being done, and that difference changes the management philosophy entirely. AI elevates your people from doers of work to directors of work, and the tokens they direct are output, like the hours of a team. You're investing in a workforce rather than buying seats, so manage it with the attention you give your people: know who is directing the work, what they're producing, and how they're performing.

If you've ever outsourced, you already know this decision. When a chunk of back-office or analyst work got well-defined enough, you moved it to an outsourced team because the spend was more efficient there. Tokens are the same decision with a different cost measure.

The number to watch is token spend as a percentage of labor spend. Ours runs about 21%, a 4.7-to-1 ratio of people spend to token spend for a team of 7, and based on how productive the team has been, we expect roughly 40% by the end of the year. For contrast, most companies that come to us worried about AI costs are spending less than 0.1% of their labor budget on tokens. If you run a business where every person needs a computer to do their job, a healthy target to strive toward is 10%, provided you follow the framework in section 4.

A note on the ratio: it compares operating labor spend to operating token spend. If AI runs inside your product, those tokens are cost of goods sold and belong in a different bucket. Margins differ by business model, and so will your target; the 10% here covers the work your people direct.

2. Your AI spend follows a power law

In every organization we monitor, roughly 20% of users generate 80% of the token spend. We have yet to find an exception.

Our proprietary tool Watchtower tracks usage across our own team and our clients, and the pattern is the same: the top 20% of users account for ~80% of spend, the top 10% alone account for 61%, and the bottom half of the company combined accounts for about 4%. The shape barely changes with company size, whether it's 200, 2,000 or 20,000 FTEs: SemiAnalysis reports the same concentration across its enterprise conversations, from tech-forward Fortune 500s down, and we see it in every organization Watchtower tracks.

The top decile of users accounts for ~61% of AI spend

Average across 6 orgs; each org's users ranked and split into tenths

Average across six client organizations tracked by Watchtower, October 2025 through June 2026. Each org's users are ranked by spend and split into tenths; bars show the average share of org spend per tenth. SemiAnalysis reports the same shape across its enterprise conversations.

This means your organization doesn't have an average user to write policy for. It has a spectrum, from people who open Claude a few times a week to people directing it like a team of direct reports, and the right way to manage spend is different at each end. That's why the broad policies spreading across the industry misfire in both directions. Some real examples of poor policies, from SemiAnalysis's conversations with enterprises: an aerospace manufacturer caps everyone at $250 a month and disabled its most capable model org-wide, and a travel company switched 1,500 employees' default to a cheaper model to control spend most of those employees were never generating. Policies like these regulate the 80% who barely spend while throttling the 20% doing the productive spending.

The good news is that the power law makes cost management easier. If spend were spread evenly, managing it would mean managing everybody. Instead it concentrates in a small group of people you can name, sit down with, and understand.

When I was a data scientist at Uber, we had a saying: nobody opens the Uber app just for fun. It's a utility: you open it because you need to go somewhere. Your company's Claude or ChatGPT works the same way, and nobody opens it at work unless they're trying to get something done. So when someone's spend starts climbing, treat it with curiosity instead of alarm. Go sit with them, ask what they're working on, what they've learned, and what they're experimenting with. You'll usually find your next power user, and sometimes your next big use case.

For everyone else, the math says to relax. The bottom half of your company drives 4% of your spend. There are no meaningful savings there, but there is plenty to lose: a newcomer pushed onto a cheaper model has a worse first experience, trusts the tool less, and adopts slower.

“Trust, once lost, costs far more than tokens to win back.”

3. The price of intelligence

Before reading the cost-cutting headlines as advice, look at where they come from. Uber's caps apply to AI coding tools, imposed after roughly 5,000 engineers burned through the year's budget by April. Microsoft's evaluation of Kimi K3 is for the models inside Copilot, where tokens are part of the cost of goods sold. SemiAnalysis estimates that over 70% of Anthropic and OpenAI revenue today traces to coding. Engineering was always going to optimize first: the work is verifiable, the measurements already exist (tests, velocity, bugs), and open-source models are trained heavily toward engineering workloads. None of that transfers automatically to the rest of your company. An engineer switching to GLM-5.2 does not mean your lawyer can, because outside their specialty these models are jagged, and the frontier models remain the far better generalists.

So what do the price gaps actually look like? Group the leading models by tier of intelligence and the picture gets much less dramatic than the headlines. The near-10x differential compares Fable, the most intelligent model on the market, against models two tiers below it. Within a tier, the gap between the Anthropic model and the cheapest credible alternative runs 40 to 67%. That is still real money at volume; the question is what you give up to capture it, and that's where the value of the intelligence comes in.

Within a tier of intelligence, the gap is 40–67%, not 10x

Standard list prices, July 2026, blended 3:1 input:output.−% shown vs the Anthropic model in each tier

AnthropicOpenAIOther providers
Standard list prices as of July 2026, blended at a 3:1 input-to-output ratio. Tiers follow the Artificial Analysis Intelligence Index v4.1 rankings; percentages are relative to the Anthropic model in each tier.

Now plot intelligence against price and you can see exactly how much cost-performance flexibility each platform buys you. Moving down Anthropic's own ladder saves 50% (Fable to Opus) or 70% (Fable to Sonnet). Add OpenAI and the span widens to 89% (Fable to Luna). Access to all 16 providers behind today's leaderboard buys you a maximum span of 90%. In other words, two platforms deliver about 99% of the full market's cost-performance flexibility, without new security reviews, new procurement cycles, or your people juggling unfamiliar tools to chase models.

Two platforms span nearly the entire cost-performance curve

Intelligence Index (AA v4.1, July 2026) vs standard list price, blended 3:1.

AnthropicOpenAIOther providers
Artificial Analysis Intelligence Index v4.1 (July 2026) against standard list price, blended 3:1, log scale. The 16-provider count and tier groupings come from the same leaderboard.

There are two legitimate ways to capture those savings, and neither is forcing everyone onto a dumber model.

Individual task

Can an equally intelligent but cheaper model do it?

If yes, switch: Fable to Sol, Opus to Terra, Sonnet to Luna. Verify with evals; same tier does not mean same behavior.

Workflow

Can you restructure the work so cheaper models carry it, within your approved tools?

If yes, optimize the workflow so cheaper models carry the load. Evals tell you when quality holds.

For workflow optimization, the smartest model orchestrates and delegates execution downward (Anthropic publishes guidance on exactly this pattern), or you tighten the prompts, skills, and scaffolding until a cheaper model passes your evals. You can even hand the optimization itself to the smart model: give Fable your evals and ask it to get the workflow running on Sonnet without a quality drop.

The order of operations matters. Build with the smartest model first and prove the outcome is worth having. Put evals in place to set the quality bar. Then optimize toward the cheapest model that clears the bar. Evals are non-negotiable even for lateral swaps: Claude and ChatGPT sit in the same tiers on aggregate, but they behave differently on specific tasks, the way two people with the same SAT score aren't equally good at every job. We'll publish a separate guide on building and maintaining evals for non-engineering teams; subscribe to Glia to get it when it ships.

One instinct to retire: “this task is small, so I should use the cheapest model.” Writing an email, summarizing a meeting, or formatting a document consumes a negligible number of tokens; the difference between running it on a frontier model and a budget one is pennies. Default to an Opus-class model and don't sweat it. What actually drives cost is true delegation: whole tasks handed to agents end to end, with a person acting as director across multiple sessions. By the same description, that's also what drives productivity.

As for the open-source models and the budget challengers, their most important role in your buying decision is the pricing pressure they put on Anthropic and OpenAI. Both labs need your market share, and intelligence per dollar keeps improving because of that competition. At scale both platforms bill by usage, so adding the second one costs you nearly nothing: whatever you run on one, you aren't paying the other for.

4. The decision framework

The moment “are we spending too much on AI?” pops into your head, walk this tree.

Gate 1. Are tokens at least 1% of your labor spend? If not, you don't have a cost problem, you have an adoption problem, and the correct move is to promote use rather than cut it. A good test for what crossing 1% looks like: your people direct AI like a coworker daily, delegating real work end to end rather than asking the occasional question. Push proficiency up that ladder, from AI as a search engine, to thought partner, to coworker. Premature optimization at this stage teaches newcomers to distrust the tools, and it saves almost nothing; remember that the bottom half of your company generates about 4% of spend.

Gate 2. Can you see where spend goes and what it produces? You should be able to answer: where are the 20% of users driving 80% of our spend? What are they doing with the tokens, and what work is coming out? What outcomes have changed? That implies per-user and per-team visibility into both cost and production. If you can't answer these, put the measurement in place before touching anything, because you can't optimize what you can't measure.

Measuring doesn't have to be clinical, and this is where most companies get stuck: the barriers are about people rather than technology. Much of the measurement is conversation. Sit with the people driving spend and the wins are usually obvious. Work you used to outsource now done in-house. Marketers writing copy and building emails without handing off to copywriters and designers. An FP&A team running its own analytics without filing tickets to the data team. A CEO seeing progress and blockers across every team without waiting on directs.

Gate 3. Within your top spend, where is ROI weak, and why? If a use case is negative or underwhelming, remember who you're dealing with: your most proficient users. Don't punish the innovation; redirect it. Sit down together, figure out whether the use case is simply unproductive, and point that talent, and the tokens they know how to drive, at something better. If a use case started strong and returns are thinning as it scales, that's the actual signal to optimize. Now run the playbook from section 3, lateral swaps and workflow optimization gated by evals, until you find the best cost-performance point.

Notice what's not in the tree: switching platforms.

Notice also what walking the tree requires: use cases cataloged, spend and outcomes visible per person, evals in place, and someone who knows the strengths and weaknesses of the models. That is the job description of an AI ops function. We're not telling you to stand one up tomorrow, but companies are already experimenting with AI ops teams, and once adoption is company-wide, the value of owning this framework centrally becomes obvious.

5. The usage policy

Here's the policy shape we recommend to clients, and the philosophy behind it: caps exist for predictability and visibility, never for rationing.

  1. 01

    If someone hits their cap doing productive work, raise it fast. The review is the point, because it's where you learn what your best people are doing.

  2. 02

    Tier users by demonstrated proficiency. The four tiers map to the four levels of AI proficiency, a framing we borrowed from Notion: AI as a search engine, as a coworker, as a team of coworkers, and as a system the organization builds around. Tier moves follow demonstrated proficiency rather than seniority.

1. General

$100
Who
Everyone by default. Building the habit and finding use cases.
Models
Opus-class default; cheaper models available

2. Active users

$400
Who
Daily users who direct AI like a coworker.
Models
Opus-class default; cheaper models introduced deliberately

3. Power users

$1,200
Who
AI as a team of coworkers. Losing it would visibly cut their output.
Models
Frontier models (Fable-class) unlocked

4. Pioneers

~$3,000
Who
People redesigning how the company works around AI.
Models
All models. The cap is a circuit breaker, and spend can arguably be booked as R&D.
  1. 03

    Three model rules do most of the work. Default everyone to an Opus-class model: $100 buys a lot of Opus mileage, most people never change defaults, and a strong first experience is what turns a skeptic into a daily user. At the early tiers, trust is worth more than cost. Introduce cheaper models at upgrade time, framed as a speed benefit for lighter work rather than a demotion. And unlock the frontier tier for Tiers 3 and 4: on decisions with five-figure consequences you'd rather have more intelligence than less, and your pioneers should be actively hunting for what this level of intelligence makes possible.

  2. 04

    Upgrades should be fast and human. Alert users at 75% of cap so the review happens before anyone sits idle. The review is 15 to 45 minutes with one designated reviewer, and it's really an interview: how are you using it, what's the impact so far, what do you want to do next. Log it in a shared sheet and you get, for free, a living catalog of use cases, a census of your power users, and wins worth celebrating company-wide. Watch for the opposite failure too: unused seats are the more expensive problem.

Take this with you

If you own an AI budget: token spend is labor spend, so manage the people directing it. The power law means your costs concentrate in a small group you can know personally; optimize with them, and promote adoption everywhere else, where trust matters more than cost. Measure impact, quality, and cost per person before optimizing anything, and capture savings inside the two major platforms before shopping for a third. Adoption first, measurement second, optimization last.

We nerd out on this topic. If you want help wrapping your head around your own AI spend and policy, we'll do a free 30-minute workshop with you: send a note to [email protected]. And subscribe to Glia for the evals guide and everything else we publish next.

Sources

  • Artificial Analysis Intelligence Index v4.1 and model list prices, July 2026. artificialanalysis.ai/models
  • Glia Watchtower usage data: 182 users across 6 client organizations, Oct 2025 to Jun 2026 (internal). gliai.com
  • SemiAnalysis, “TokenBudgeting: Our Conversations with Enterprises on Token Spend,” June 2026. newsletter.semianalysis.com
  • SemiAnalysis on subscription token multiples, via X. x.com/SemiAnalysis_
  • TechCrunch, “Uber caps employee AI spending after blowing through budget in four months,” June 2026. techcrunch.com
  • DigiTimes, “Microsoft evaluates Kimi K3 for Copilot to lower inference costs,” July 2026. digitimes.com
  • Anthropic (@ClaudeDevs), model orchestration and delegation guidance for Claude Fable 5, via X. x.com/ClaudeDevs