Skip to content
Back to Resources
Published Sep 24, 2026

Token Management Is an Engineering Skill

Better token management starts with engineering skill. Learn how stronger AI practices can make spend more efficient, predictable, and defensible.


Are you a “Token Legend”? How about a “Cache Wizard”? Those were two of the rankings offered by Claudenomics, a dashboard a Meta employee built to track how many tokens each person at the company used. Meta has since shut it down, but it remains a memorable artifact of our whiplash-inducing feelings about tokenmaxxing.

First, we couldn’t spend tokens fast enough. AI was magic, then the future, then a paradigm shift we couldn’t afford to miss. The bills came in and everyone had second thoughts. Maybe we should slow down. Let’s be practical about this.

Then the models kept advancing. Every pragmatic decision looked foolish six months later. Why worry about a budgetary rounding error when the next model might eclipse the last one in intelligence, cost efficiency, or both? And so we tokenmaxxed. AI spend became proof of commitment to the future.

This back-and-forth suggests we’re thinking about the bill all wrong. Token spend is a lagging indicator of engineering skill. The best way to control it is to make token management one AI skill among many.

When the bill comes due

AI budgets differ, but in our conversations with engineering organizations, they tend to follow the same arc: $5,000 a month today, a projected $10,000 in six months, then something closer to $150,000 six months after that. The starting point doesn’t matter much. It’s runaway spend either way, and it usually triggers one of two reflexes: cap it or ignore it.

Engineering leaders in the cap-it camp see themselves as pragmatic. They have the CFO or CEO in their ear, and they can’t prove the ROI of all this token spending yet. Why not put a ceiling on it?

A cap feels like insurance against a spectacular bill. In practice, the problem moves. Engineers route around the cap with different tools or cheaper models. Some stop using tools that would have paid off. The spend becomes harder to see, and the team gets slower.

The ignore-it camp also has a practical-sounding case. Setting a sensible token budget is difficult. Halfway through the exercise, OpenAI drops a model that costs 80% less and the spreadsheet is obsolete.

Ignoring the bill can even feel visionary. Model improvements will eventually swallow the cost, so limiting tokens is as foolish as limiting engineers. There may be some truth in that. You still won’t be able to defend it in a budget meeting, and the budget meeting is inevitable.

Both approaches assume tokens are the thing we should manage. How did we end up there?



How tokens became the default metric

Nobody sat down, considered every option, and chose tokens. AI arrived. Teams used it, started measuring it, and climbed the metrics ladder one rung at a time. Token spend is the rung we’re stuck on.

The first rung was seats. With any new tool, the first question is whether anyone is using it at all. Then you notice you’re paying for hundreds of licenses and only a small fraction of people are active, so you move to active seats. Eventually, everybody is active. That’s the sign of a good tool, but it’s not much of a metric.

Each rung requires more work. Measuring seats is automatic. At the other end of the ladder, some poor team has to fill out a budget request for every prompt. That’s obviously silly, but there is a point where measurement becomes more work than it’s worth.

Teams stop climbing when they find a comfortable rung. Right now, that rung is token spend. Tokens are precise and already instrumented. They also map directly to dollars. Almost every other candidate metric requires more work.

For a while, teams could keep experimenting while watching for exceptional costs that suggested something had gone wrong. Now leadership, having spent a year pushing engineers up the usage leaderboard to prove the organization was AI-native, is telling them to get spending under control.

The AI bubble hasn’t popped. Nobody is rushing to decommission Copilot licenses. We’ve simply asked a cost metric to do the job of a value metric. It can’t.

Treat the bill as a skill report

Most runaway bills trace back to a skill gap. Spend is the symptom. Manage only the symptom and you end up with a cap set by finance and an engineering organization irritated enough to find workarounds.

Token spend becomes useful when you connect it to the behavior that produced it.

What you see The skill to build Better practice
Enormous single sessions Context management and task decomposition Plan in one session, execute in shorter ones, and checkpoint before the context gets noisy
Agents running unbounded overnight Scoping, guardrails, and stopping conditions Set a finish line and a hard ceiling before the run starts
Frontier models used for trivial tasks Model selection Start with the cheaper model and move up when the task requires it
The same context sent with every call Caching architecture Cache stable context; retrieve only what the current task needs
Multiple teams building the same agents Shared skills and tooling Reusable harness templates, evals, and documented patterns

Every row in that table describes a technique, and techniques are teachable.

The bill can also reveal patterns that defy the first instinct to economize. Efficiency and quality often move together. Stuffing a context window burns more tokens and tends to produce worse results. Chroma’s research on context rot found that models use their context unevenly and become less reliable as input length grows.

context-rot

The same pattern applies to unbounded agents. Giving an agent an endless task without stopping conditions is tokenmaxxing, but the tokens aren’t doing useful work. It’s like driving a gas-guzzling SUV across the street. You’ll get there. You’ll just pay more than the trip deserved.

Governance should look a lot like coaching. Track the big spenders, but don’t bring the hammer down. Find out what they were doing. Help engineers get better at using AI and the bill will stabilize as a side effect.

Too much stability is also a bad sign. Teams should still experiment. An expensive failed experiment can be cheap if it keeps five other teams from repeating it. A surprise bill discovered by the CFO is different. Engineering leaders should understand token spend—and what caused it—well enough to handle that conversation before finance starts it.

Three skills the organization needs

Individual engineers can’t fix this alone. The organization needs a few new habits of its own.

Define the outcome before the budget

An engineering leader who sees the job as running a factory floor will inevitably end up in a token-cost conversation. Cost is one of the few variables that framing gives you to control. It’s also much easier to measure than value.

Take a page from Amazon and work backward. Define the customer experience, get clear about what the team needs to build, then measure whether AI changes how quickly the work reaches customers.

 

Time to ship cuts through the “Well, I feel like I’m coding faster” noise. Are we actually shipping value? Capability maturity gives that number context.

We’ve seen companies introduce AI metrics and wind up with historically high PR-opening rates while PR-merging rates stay flat. That is a maximally efficient way to produce work nobody ships.

Focus and stack-rank

As we’ve written before, AI can speed you up. It can definitely speed you up. A team with no direction can accelerate in the wrong direction just as easily.

Stack-rank candidate AI use cases against the OKRs you already have, then fund the top few. Put AI against the slowest or most painful parts of software delivery. If the work matters, the improvement should show up somewhere beyond the token bill.

Sometimes the best use cases are unglamorous: automating Jira ticket creation or turning a ticket into a first-draft PR. Sometimes the right move is to finish one important thing quickly instead of doing five things slightly faster. AI has also blurred plenty of role boundaries, so get explicit about who owns what.

Land and expand

AI hype naturally produces ambition. The organization has to channel that ambition without smothering it or paying for every idea at once.

Start with an opportunity that lends itself to measurement. Expand once you have an approach another team can repeat. Ask engineers to contribute what they learn to a shared skills repository so every team doesn’t pay for the same lesson. Each win should make the next one easier. That’s the point of building AI capability in stages.

Prove the workflow before optimizing it. Caching and spend management are useful once you have something worth spending on. Optimizing an unproven workflow is just a cheaper way to do something you may not need.

The larger risk is wasted potential. AI makes previously impractical work possible, including agent runs that continue overnight and give engineers something useful to review in the morning. If the return holds up, keep expanding. Starting small is how you get big.

Aim for a predictable number

Token spend creates sticker shock. AI still feels new, even as the costs rise like a mature budget category. That curve is what finance sees and reacts to.

Most CFOs would rather have a predictable number owned by someone credible than an artificially low number nobody can explain. Marketing, comms, and every other function bought AI tools that have become costly and difficult to forecast. Engineering can be the function that walks into the room already knowing what happened.

The goal is to outgrow token spend as a management tool. Its usefulness is a measure of organizational immaturity. Mature teams know which work deserves expensive models, where agents need limits, and when a big bill bought something valuable.

Then the token bill becomes what it should have been all along: a receipt.


icon-stackup@2x

Not sure whether your token bill reflects skill or drift? Book a free 45-minute StackUp session with an Uplevel expert to see where your AI spend is going and how it compares to other engineering orgs. 

Start with StackUp →


FAQs

What is token management?

Token management means controlling AI costs through engineering technique: tighter context, smaller tasks, sensible model choices, caching, and limits on agent runs. Better technique usually brings the bill down without an organization-wide cap.

Should we cap our AI token budget?

An organization-wide cap is usually a poor first move. It encourages engineers to use cheaper models where they don’t belong, move work to personal accounts, or abandon useful tools. A spending limit can make sense for a single agent run. That’s a guardrail attached to a specific risk, not a substitute for knowing how your teams use AI.

What causes runaway AI token spend?

Common causes include oversized context windows, agents without stopping conditions, expensive models used for routine work, repeated context that should be cached, and several teams building the same internal tool. Each pattern points to a skill the organization can teach.

Is token spend a good measure of AI ROI?

No. It tells you the cost of AI use. Pair it with delivery measures such as time to ship and PR merge rates to see whether that use is producing anything customers receive. More on why token spend is the wrong ROI metric.

How much should an engineering organization spend on AI tokens?

There is no useful universal benchmark. The right number depends on what the spend produces. A predictable bill you can explain is healthier than a low bill you can’t.

Who should own AI token spend?

Engineering leadership, with finance as a partner. The owner needs visibility into delivery, not just cost, so they can explain the number and the engineering behavior behind it.

Does reducing token usage hurt output quality?

Often, the opposite happens. Tighter context, smaller tasks, and clear stopping conditions can reduce spend while making results more reliable.

Table of Contents

    Amy Carrillo Cotten is Director of Client Transformation at Uplevel. With 12+ years of technology industry experience as a change consultant and program manager, she works directly with engineering leaders and their teams to increase growth, reduce risk, and maximize innovation.

    StackUp Hero (1)

    Where's your AI spend going?

    Six liabilities drive up cost in large engineering orgs. Get 45 minutes with an Uplevel expert and know where to focus first — for free.

    Related Resources

    Valuemaxxing Strikes Back
    AI Engineering

    Valuemaxxing Strikes Back

    Discover how engineering leaders can shift focus from token spend to value delivery, unlocking AI's true potential for meaningful outcomes and better ROI.

    What Is Developer Productivity in the AI Era?
    AI Engineering

    What Is Developer Productivity in the AI Era?

    Code generation has exploded. Here's what developer productivity actually means now — and why your metrics may not be showing the full picture.

    Top Engineering Intelligence Platforms [2026]
    AI Engineering

    Top Engineering Intelligence Platforms [2026]

    Boards want a number on AI ROI, but most platforms only show what happened last week. Compare the top engineering intelligence platforms for 2026.