Want to become a partner or sponsor, click here to schedule time to talk with the team.
FinOps & Beyond is what engineering, finance, and IT leaders read to understand FinOps, and what it means for operating models, accountability, and spend decisions.

Two weeks ago, in this newsletter, I told you I hate tagging but you should do it anyway. I still don’t like it, but I think I’m convincing myself more and more that it is extremely critical for determining the success of AI’s return on investment and unit economics.

Where I struggle is around decay and how do you have an enforcement model. And it’s less about how does the tag actually get on the thing but how can you properly point to a resource and say … “We are using it for fill in the blank

So, I have come to the conclusion that tagging is not metadata. Tagging is configuration. We have just never treated it that way (keeping the format and design in documents and in human readable format) versus moving to how we treat and manage configuration. And until we do, I think we will struggle.

(If you already do this or I’m late to the game, I have no defensive.)

The Part of My Background That Keeps Coming Back

Before I focused on FinOps exclusively, I spent years in platform engineering, DevOps, and cloud operations. In that world, nobody asks a person to remember a setting. You do not send an email telling forty engineers to set the log level correctly. You set it as configuration, use the configuration in the pipeline, and the setting is used whether anyone was paying attention or not.

Configuration has a lifecycle, and it is a boring one. You declare the desired state and then something applies it. The real world drifts or changes, because the real world always does. And then the next apply pulls it back. And if its time to change it, that can be done.

Now look at how tagging actually gets done well. Default tags at the provider level. A tag block in the Terraform module. Inheritance from the resource group or the folder or the account. A policy that denies the deployment when a required key is missing. Every one of those is configuration management wearing a FinOps hat. In Issue 27 I wrote "inherit and don't ask" as a tagging rule. That is not a tagging rule. That is just how config has always worked.

Which means the tagging programs that fail are the ones that failed to implement in config and pipelines and only kept the wiki page. A document is a declared state with no apply step. And therefore, we are asking humans to be the reconciliation loop, which we are terrible at, myself very much included.

There Is No Next Apply For AI

Here is where it breaks for AI.

Everything I just discussed above depends on the thing still being there. The instance, the bucket, the subscription, the seat. It persists, so drift or decay is survivable. You can be wrong at provisioning and correct at reconciliation. That is exactly “the sweep” I wrote about last week, and it is the whole reason cloud tag debt is a cleanup project rather than a loss.

However, an AI model call does not persist. It fires, it bills, and it is gone in a few hundred milliseconds. There is no resource sitting there waiting for your next apply. There may be API calls or token meters capturing a running token total, but unfortunately there is no drift to detect, because there is no current state to compare against a desired one. There is an amount, and whatever was attached at the instant it left.

So the reconciliation loop, the part that makes configuration management actually work, does not exist for AI spend. You get one apply. It happens at call time. That is it.

Which is not a reason to give up on treating tagging as config. It is a reason to move the apply.

Three Places to Consider

If the only moment that counts is the moment the call fires, then the real question is not "what should we tag." It is how close to that moment you can attach the identity, and who (or what) has to remember to do it.

I believe there are three answers in the market right now, and most teams are arguing about them without noticing they are arguing about the same thing.

1. Cloud tags and provider-native config.

Project-scoped keys in OpenAI. Workspaces in Anthropic. Application inference profiles, Projects, or per-request metadata in Bedrock. Labels on a Vertex call. This is the closest thing to what you already do, the cheapest to start, and genuinely useful.

The limit is that provider billing only aggregates along a boundary you declared before the call. Sometimes that boundary rides on the credential, like an OpenAI project or an Anthropic workspace. Sometimes it is an ARN or a header that somebody has to remember to set on every single request, like a Bedrock inference profile or a Vertex label. Either way the boundary is whatever was wired up. Good luck on consistency.

And there are some caveats which are interesting … For example, AWS says request metadata is not enforced, requests that omit it still succeed, and there is no service-side policy to require it. That is the untagged resource problem rebuilt, on something you cannot go back and fix.

Granularity is uneven and mostly worse than you would like. Anthropic gives you minute-level buckets for tokens and daily-only for dollars, grouped by workspace or description and nothing else. OpenAI is the same, daily dollars grouped by project, key, or line item. Bedrock's finest grain is per usage type per day, and AWS indicates that none of the native methods produce a per-request row. Google is the outlier in that they have hourly usage windows in the billing export and up to 64 arbitrary labels per call on top of that. Although if you are on Provisioned Throughput, Vertex silently ignores those labels, which is a fun thing to find out later.

2. OpenTelemetry and identity carried in the application.

Another way to approach is to attach team, product, feature, and/or customer in the code, propagate it through the trace, and read cost off the spans. This is the only one of the three that can see the call graph, which matters enormously once an agent fans one user action out into forty calls and a dozen tool invocations. It is also the only one that can tell you cost per completed task rather than cost per call.

Two problems with it.

First, the GenAI semantic conventions standardize tokens but do not standardize dollars, so there is no gen_ai.cost attribute. In theory this is fine, since everyone will have a different rate card. I think it would make sense to have the field so teams can populate.

Second, is the one that should worry you more. For any of this to be in the trace, a developer has to put it there. Every call site, every new service, every code path somebody adds next quarter. Miss one and that spend has no identity, and as I said last week, there is nothing left to go back and fix.

We have run this experiment before. Tagging was optional at resource creation for a decade and coverage was never finished anywhere. The difference is that cloud let you sweep up afterward. This does not.

3. API management and key issuance.

This is what the FinOps Foundation now recommends, and I think people have skipped past how strong the wording is. Their Tokenomics working paper from June says establish API key governance immediately, with every key mapped to one team, app, or use case and provisioning requiring a named owner and a cost center. Then deploy a proxy or observability layer to inject the metadata the providers do not give you natively. In the paper, they name LiteLLM, Portkey, and Helicone.

The part worth noticing is the sequencing. In their maturity table, "deploy a lightweight proxy" sits in the Crawl column. Not Run. Crawl. A year ago the gateway was an advanced move. The Foundation now treats it as table stakes.

Their own framework page is honest about why. It names as an open problem the "lack of generally accepted frameworks for cost allocation across multi-agent workloads." That is the standards body telling you we need to solve this and find solutions in the meantime.

The One That Survives Contact With a Developer

So which of the three wins. I do not think that is the right question to be asking, because any good implementations runs at least two. But there is a test that separates them, and it is not technical.

So, here is the test. Does the person making the call have to remember anything?

Tetrate published their own numbers on this and I respect them for it. About thirty engineers, roughly four thousand dollars of model spend in a week, and about half of it on keys with no team, no workload, and no owner beyond an email address. Four keys were half the spend. Ten keys were eighty percent. Their conclusion is that attribution is a property of key issuance, not a property of usage.

Read that again next to everything I have written this year about tagging. The tag does not belong on the call. It belongs on the credential that made the call possible, because a credential is issued once by someone whose job it is, and a call is made ten million times by people who are not thinking about you.

Uber is the proof at scale. Their Model Gateway had three requirements at design time, and attribution was one of the three, not something bolted on after a bill surprised somebody. Everything routes through one endpoint. Eight hundred plus internal projects, more than a hundred million model requests a day. And the contract they offer their own engineers is one sentence: take the vanilla client, set the project ID, and we take care of everything else.

One field. Not a taxonomy. Not a standard anyone has to read. One field, set once, on the client.

What that bought them is in their own engineering write-up. Cost per thousand model requests down 34% from peak. Cost per session down 52% from its June peak. Roughly 40% of fleet-wide token consumption removed. You do not get to run those levers unless you can already see which project is spending what.

Now, here’s the honest counterweight, because I do not want to sell you all rainbows and butterflies.

Wealthsimple has run an internal LLM gateway since 2023, and worth noting, they did not build it for cost. They built it so they would know what data was leaving and who sent it. Cost attribution rode along afterward, not surprisingly. They made it free to employees, faster than going direct, with pooled rate limits and retry logic, so using it was the path of least resistance. They also built a nudge that pinged anyone who bypassed it, and then removed the nudge because it barely changed behavior.

Their coverage is about eighty percent. That is a gateway with the economics deliberately rigged in its favor, at a company where this is taken seriously. Eighty percent.

I also went looking for a credible survey telling us how many organizations run an AI gateway at all, and I couldn’t find one. Everything I found was vendor content or a self-selected sample. So when someone tells you the gateway is the industry standard now, including me, ask them for the number. Nobody has it.

The Job

Three moves. The first one costs you an hour.

Move One. Write down where your identity gets attached today, for each of the three routes.

Cloud tags and provider config. Application and OpenTelemetry. Key issuance and gateway. For most of you the first column is full, the second is partial and undocumented, and the third is empty. That gap is your answer and you do not need a tool to find it.

Move Two. Look at your key issuance process, if you have one.

Every key should have a named owner, a team, contact info, and a workload attached at the moment it is created, before anyone can call anything with it. If keys get handed out in a Slack thread, that is the single highest leverage fix available to you this quarter, and it is process work, not procurement.

Move Three. Find out what your managed device config sets for AI tooling, and who owns it.

Default model, telemetry endpoint, allowed model list. If the answer is that nobody sets it and it came with the template, you now know where your next uncontrolled default is going to come from.

Tagging was always configuration. Cloud let us be sloppy about it, because there was always another apply coming to clean up after us. AI does not offer that. The name goes on the call as it leaves, or it does not exist. Put it in the key, put it in the config, and stop asking people to remember.

Go Read These Yourself

I have embedded the links in the issue, but if you have reached this point and want to just have the list, I’m including them here.

The provider docs

The guidance

The people who actually built it

PS - The Laptop I Owed You

Last week I ended with a PS about how that vendor actually fixed their customer's problem. Hundreds of dev laptops pointed at an expensive default model, and the fix was pushing a config change to every device with their own tooling.

I left something out of that issue that I should not have. A laptop does not pass through your gateway. Neither does a CI runner with an environment variable, an IDE plugin, or a personal API key someone put in a config file eight months ago. Those are run time events on a path with no enforcement point, which is the worst combination available.

And, yet, the answer was sitting there the whole time. It is device configuration management. MDM. One of the oldest tool in the IT box.

Claude Code reads managed settings from OS policy at startup and re-checks every thirty minutes. Copilot and Codex both ship equivalents. You can pin the model list, pin the telemetry endpoint, and push a default the way you would push any other setting.

It is not airtight and I will not pretend otherwise. Setting a different base URL bypasses the managed path. Copilot's server-managed settings do not reach people using their own license. Codex continues without the managed layer if the config fetch times out. On an unmanaged device with admin rights, all bets are off anyway.

But it is the same move as everything else in this issue. Stop asking a person to remember. Put it in the config and let it arrive.

Written with the help of AI. All the ideas expressed are mine and mine alone.

FinOps Company Spotlight

If you would like your company included in the Spotlight, contact the CloudXray AI Team

Company: CloudXray AI

Category: Managed Services & Consulting

What They Do: FinOps Consulting & Advisory Services; Owners & maintainers of the single largest FinOps company directory (finops.cloudxray.ai)

Why It Matters: Companies still need guidance on implementing FinOps and understanding the landscape of companies that exist

Reply

Avatar

or to participate