Want to become a partner or sponsor, click here to schedule time to talk with the team.
FinOps & Beyond is what engineering, finance, and IT leaders read to understand FinOps, and what it means for operating models, accountability, and spend decisions.

No project you've ever run was staffed entirely by senior members. Not a client engagement, not an internal build, not a product team. You'd never do it. Experienced senior members are expensive, and most of the work on any project doesn't need one to execute. So you mix the team. Senior members take the judgment calls, everyone else takes the work that matches their experience, ultimately translating to their rate. That's not a cost trick. It's just how work gets done.

Now look at how we run AI. One model on everything, usually the most powerful one, pointed at every task whether it needs the horsepower or not. Nobody staffs a project that way. So why are we running AI that way?

Here's the perspective that I think fixes it. AI is a tool, and tokens are how you measure the work it did. Same as hours measure the work a person did. The model you pick is the person you staffed. Once you see it like that, AI cost runs on rules you already know and understand.

When you staff a project you don't invent a new unit to think in. You think in hours and skill levels. You keep a timesheet, total it against the project, and decide whether the outcome was worth what it took. Tokens drop straight into that. The tokens are the hours. The model you pick is the person you staffed, and right now most teams are putting a principal engineer on data entry and wondering why the bill is high.

Match the tier to the task

The mapping is simple. A token is a unit of work, the same as a billed hour. You track it against a project and you end up with a ledger of what the work took. The model you run is the skill tier you staffed. A smaller, cheaper model is your junior. Fast, cheap, fine for the grunt work. A frontier model is your senior. Expensive, and you bring it in for the calls where being wrong is costly and you need the knowledge and the experience.

You would never put your most expensive person on data entry. Running your most expensive model on a task a cheap one handles is the same waste, and it's everywhere right now, because almost nobody treats model choice as a staffing decision. They pick the biggest model once, wire it into everything, and never look again.

So the first move is one you already know how to make. Match the tier of work to the tier of worker. Cheap model for the cheap work. Escalate to the expensive model only when the work actually needs the judgment. On a real project that one decision moves the bill more than any discount a vendor will hand you.

The catch is the tooling. The harness to switch cleanly between models mid-workflow is still young, nowhere near fully baked. So for now it's a manual call, and here's how I make it:

  • For complex work that needs research and knowledge that isn't sitting there ready, I use the newer, slower, more token-heavy models. In FinOps that's pointing one, with the right data, at finding anomalies or working out why the bill won't reconcile. The deep digging.

  • For simpler work I switch to a smaller, faster model that burns fewer tokens. Reporting, writing, project management. That's where it wins.

The ceiling you just lost

If it stopped there it would just be a tidy metaphor. But there's one place the comparison misleads, and you should be made aware of it.

Human hours are scarce. A person bills their hours for a week, get tired, takes vacation, and human has a ceiling that was quietly doing the prioritizing for you. You can’t afford to chase every idea, so scarcity kills the weak ones before they ever got staffed.

AI has no ceiling even with tokens, except for cost (if that’s a limit, and it should be). You can burn a consultant-year of them in an afternoon and never feel the friction that used to stop you (and many a company have recently). The governor is gone and that changes what the timesheet means. When work was scarce, a full timesheet was a decent proxy for value, because nobody wasted a scarce, expensive person on junk. When work is cheap and unlimited, a full timesheet proves nothing. A million tokens of output that nobody needed is still a lot of "hours," and it produced nothing.

I wrote a whole issue on this a few weeks back. Cheap to make was never a reason to make it. The token version of the timesheet has the same trap baked in. It will happily rack up a huge, precise, professional-looking total for work that should never have run. Precision is not the same as worth.

Why this breaks the bill

Here's the part that should make anyone who prices work sit up, me included.

Why did we ever track hours? Not for the hours. We tracked them to price and defend value. The hour was a stand-in for effort, and effort was a stand-in for what the work was worth to whoever paid for it. A consultant bills the client by the hour. An internal team justifies its headcount and its budget by what all those hours produced. Same logic, different invoice.

If tokens are the new hours, and you still price the work by the effort, then AI just quietly collapsed the thing you're selling or the case you're making. The deliverable that was a hundred hours is now ten hours of a person plus a pile of tokens. Bill a client by effort and your revenue falls off a cliff. Justify a team by hours logged and you've made the argument for your own cuts. You removed most of the effort, so effort stopped being the thing to point at.

The way out is the same shift I keep pushing FinOps vendors toward. Stop pricing the effort. Price the outcome. When production gets cheap, the value was never in the production, and any pricing or headcount case that still points at production is about to break. Tokens don't just give you a cleaner cost lens. They push the whole model off effort and onto results, whether you're a services firm or a team defending a budget.

The people didn't disappear

If you read this the wrong way, it sounds like the people become the line you cut just that’s happening across the entire industry. That's the layoff pitch, and I’m convinced this is backwards. The models are staff you added, not staff you replaced. Someone still has to run the project. Someone assigns the work, points the right model at it, and judges whether what comes back is any good. That someone is a person, and the whole thing falls apart without them.

And here's the part that I think gets missed. The person has to match the tier they're guiding. A junior staff member can direct a cheap model through grunt work and check the results, because the work is what a junior already knows. Point that same junior staff member at a hard problem with a frontier model and they will get stuck or a suboptimal solution. Why? Well, they can run the task through the model, but they can't tell whether the answer is right, because judging a the answer takes senior-level knowledge. I’m arguing that the person guiding the work still has to be knowledgable enough to catch it when it's confidently wrong.

So AI doesn't flatten the skill ladder. It raises the stakes at the top of it. You still need people at every level, each guiding the tier of work they're actually qualified to judge. The junior member keeps the cheap model honest on the easy stuff. The senior member keeps the expensive model honest on the calls that carry risk. Take the people out and you don't have a leaner project. You have a pile of output nobody qualified ever checked.

The job

If you run AI work on any real project, here's where I suggest you start on Monday.

Keep the timesheet. Track tokens per project the way you'd track hours. If you can't say which engagement or task burned the tokens, you can't manage any of it.

Then, staff it like a project. Map your model tiers to your work tiers on purpose. Cheap model for the grunt work, expensive model only where judgment earns it, and check that split the way you'd check to make sure you haven’t overstaffed an engagement.

Finally, measure the output, not the tokens. The token total tells you what the work consumed, not whether it was worth doing. Track cost per finished thing that mattered instead: a deliverable shipped, a ticket resolved, a report someone acted on. Put the spend against what it produced, and that ratio is the number you take to a client or a CFO. The token count is just the fuel it burned to get there.

Tokens are the new timesheet. Track them like hours, staff them like people, and price them like a consultant who remembers the hours were never the point. The work still has to be worth doing. That part didn't get cheaper.

Written with the help of AI. All the ideas expressed are mine and mine alone.

FinOps Company Spotlight

If you would like your company included in the Spotlight, contact the CloudXray AI Team

Company: CloudXray AI

Category: Managed Services & Consulting

What They Do: FinOps Consulting & Advisory Services; Owners & maintainers of the single largest FinOps company directory (finops.cloudxray.ai)

Why It Matters: Companies still need guidance on implementing FinOps and understanding the landscape of companies that exist

Reply

Avatar

or to participate

Keep Reading