Want to become a partner or sponsor, click here to schedule time to talk with the team.
FinOps & Beyond is what engineering, finance, and IT leaders read to understand FinOps, and what it means for operating models, accountability, and spend decisions.
After last week’s issue, one of our reader’s (a FinOps vendor) sent me a customer story to review. One of their enterprise customers saw AI spend jump. They went looking, and what they found was that an unnecessarily expensive model had been set as the default across the team's local development environments. Hundreds of laptops, all quietly pointed at the wrong thing, all billing at the higher amount.
While I did not verify, it is the right story for what I want to talk about in this week’s issue, and the interesting part is not the part they thought was interesting.
The Detection Worked. So What.
AI costs increased and somebody got told. This is great, as some companies cannot even do this until its too late. Of the 300 companies I have gone through first-hand for the directory, 87% do reporting. Reporting is solved and everyone is still selling it. Great.
Between "spend is up" and "an expensive default is set on hundreds of dev laptops", there is still real work to do, and that work is one of two things. Either it is metadata you already had to use for reporting, or it is a person walking the floor asking people what they changed. Only one of those scales, and I’ll let you decide which one. 🤣
Here is the part I want you to think about. The customer in this story got lucky.
One wrong default across a team is a win. Big, uniform, and one cause. Change the shape of the problem, even slightly, and detection may give you nothing. A dozen people drifting up 15% each is not an anomaly, it is a Tuesday. Or, one agent run that fires for six hours and stops is over before your alert threshold has a chance to fire. Same dollars, with no spike and nothing to trace.
Detection did not work in the examples I gave, but detection worked this time for this customer.
Two Kinds of Cost Data
I have been writing about this topic for a month without really giving it a name that is identifiable.
The useful thought here is not cloud versus AI. That is just what it looks like right now. The thought is whether the item/resource/thing you want to attribute still exists when you go back looking for it.
Some of it does. An EC2 instance, a storage bucket, a subscription, a licensed seat. It’s there. If the metadata is wrong, you can “sweep” it up by mapping values after the fact. It is tedious and it is a real project, but the information was never actually lost. Call this retroactive.
And some of it does not. A model call, an API request (sometimes), ephemeral workloads, a per-invocation job, an agent run. It exists for a few hundred milliseconds and then there is nothing left but an amount. Whatever was attached at the instant it fired is all you will ever have, and if nothing was attached, no “sweep” will ever attribute it, because there is nothing to review and attribute. Call this run time.
You cannot “sweep” a model call. That is the whole point.
Which Means Governance Has to Change Shape
This is the part I think most teams are going to get wrong, and it is not their fault.
Cloud trained the entire industry on detective control. Find the drift, flag it, remediate it. Quarterly cleanups. Compliance dashboards. Tagging audits. Every one of those works because the resource is there after the fact, and after fifteen years of that being the only kind of governance anyone needed, it is what "governance" means to most people.
That is retroactive governance, and it does nothing for a run time signal. There is nothing left to inspect except an amount. By the time you have found the problem, the window where you could have attributed it has closed permanently, and finding it faster does not help.
Run time data needs run time governance. The control has to sit in the path the request travels and put the metadata on the call as it goes by, or require it and refuse the call without it. The data and the control have to happen at the same instant or they never meet at all.
That is a different design and a much harder conversation with engineering, because you are now standing in the request path instead of standing behind it “with a clipboard”.
I have discussed these in previous issues already and did not notice what the two had in common. The governance gate asks its cost questions at design time, before there is anything to audit. The rail sits in the request path, because an agent can ignore a policy document but not the path its requests travel through. Neither one waits for the spend to happen. That is the whole trick and I totally missed it.
And it puts tagging at the core, which I am sure you are tired of hearing (and I’m tired of discussing). The design gate cannot ask about a thing it cannot identify. The rail cannot attribute a request with nothing attached to it. When the control runs at call time, the schema stops being paperwork and becomes the enforcement layer. What you require at the gateway is your policy. There is no second document.
The Job
What steps can you take to help resolve some of the challenges. I think there are 3 moves that you can make.
Move One. Generate a visual for yourself. Draw two columns. Take your cost signals and sort them into retroactive and run time. Compute, storage, accounts and most SaaS seats (retroactive) go on the left. Model calls, per-invocation jobs, ephemeral environments and agent runs (run time) go on the right. I would guess you have never drawn this line, and I would also guess every control you own is sitting in the left column.
Move Two. Count what reaches a gateway, if you even have one. What share of your model calls actually pass through a proxy you control? Everything else is unattributable, and no tool is going to report that number to you, because no tool can see what does not pass through it. This one you have to go ask people.
(Note - yes, I know API calls can be attributed to teams or individuals, if there are multiple APIs being configured this way to support. My argument is that’s not enough for the future of AI costs and Tokenomics)
Move Three. Go find out what your standard developer environment sets as the default model, and who chose it. My bet is nobody chose it. It came with the template.
Some of your cost data is waiting for you to come back and fix it. The rest of it is gone the second you look away. Go find out which is which, because they do not take the same kind of governance, and almost everybody is running one kind against both types.
PS - The Customer Resolution
If you have read this far and you are wondering how the vendor solved the local laptop issue that was using the most expensive model. Well, they used their own tooling to deploy a change to all devices that changed the default model to a less expensive one.
Written with the help of AI. All the ideas expressed are mine and mine alone.
FinOps Company Spotlight
If you would like your company included in the Spotlight, contact the CloudXray AI Team

Category: Managed Services & Consulting
What They Do: FinOps Consulting & Advisory Services; Owners & maintainers of the single largest FinOps company directory (finops.cloudxray.ai)
Why It Matters: Companies still need guidance on implementing FinOps and understanding the landscape of companies that exist

