Want to become a partner or sponsor, click here to schedule time to talk with the team.
FinOps & Beyond is what engineering, finance, and IT leaders read to understand FinOps, and what it means for operating models, accountability, and spend decisions.

This week's newsletter is going out later than normal. Why? Well, that's an interesting story.

I recently purchased a flat top grill at one of the big box retailers. The recommendations from friends and the ratings online made it look like a no brainer. This retailer also had a little bit of an early Labor Day Weekend sale, so I was saving $100 off the original price. Sounds great.

I picked it up and brought it home, and started to put it together. As I was going through the process, I got to the point where a piece that had 4 spot welds in the corners of this piece of metal had failed. Not sure why, and honestly, don't care. I had purchased new, and this clearly was a quality control issue.

So I spent the better part of the next hour taking pictures and emailing with support to figure out how to solve it. I got so wrapped up that when it came to write this week's newsletter, I totally forgot.

But whenever we have a challenge like this, I try and stay curious to figure out what this is attempting to teach me. The one that is meaningful for the purposes of this newsletter is tied to cost savings, tagging, and the context needed to make smart decisions when evaluating and improving technology usage.

And here's the part that really stuck with me. Every number and value, independently were in my favor. From the star rating, to the recommendations, to the $100. Unfortunately, none of them measured the thing that decided whether the value of the purchase. And I couldn't have known that from the numbers, because the failure was inside the box. It only showed up when I tried to use the thing.

That's most tagging programs.

The Same Subject

I've spent most of this year writing this newsletter, discussing governance. Whether its the gate at design time, the rail in the request path, or the tradeoff nobody wrote down.

While all super meaningful, in order to drive good governance and controls, you need to have a foundation that will help drive that and tagging is the work I'm actually doing for a client right now in order to help drive the controls.

Every control I've described in the last ten issues runs on a tag. The gate has to know which system. The rail can't attribute a request without something to attribute it to. Cost per unit of anything needs a denominator that resolves to a team, a product, or a customer.

The Part I Got Wrong

When I analyzed this project before the customer signed, I knew that their FinOps practice was immature or nonexistent. And I assumed, going in, that one of the major problems would be tagging coverage. Untagged resources, orphaned spend, the usual.

It wasn't. Coverage was actually pretty good. There were tags on nearly everything, applied consistently enough to pass any audit you'd point at it.

The tags just weren't the ones providing significant value and they were overly verbose.

Some of the tags were built for an org structure that had since changed. Some were technically correct and operationally useless or a key:value pair that was accurate but nobody was using. A few were genuinely important and applied correctly but values had changed over time and reporting was stale.

So while nothing broke and the bill could be paid, there was no crisis. That's the point. This kind of failure never announces itself. It's the spot weld. It holds right up until the moment you put weight on it, and everything before that moment looks fine.

Compliance Is a Vanity Metric

Tag coverage is a vanity metric, and it's the more dangerous kind, because it doesn't look like vanity. It looks like discipline.

A coverage percentage measures whether tags exist. It has nothing to say about whether they map to a question anyone is asking. You can hit 95% and still not be able to tell finance what a product line costs, because product line was never a tag key, or it was, and half the values are freeform text somebody typed.

That's Visibility Theater (sound familiar) in its purest form. Way back in Issue 02 I defined it as dashboards substituting for action. This is a step earlier and quieter: the metadata layer itself performing correctness while answering nothing.

Why This Happens to Everybody

Rarely does someone get the opportunity to design a tagging taxonomy from scratch. They typically inherit one.

Somebody stood up the first accounts, borrowed a tag list from a reference architecture or a previous employer, and it worked well enough at the time, because at the time the questions were simple. Then the organization changed. New business units, new products, a reorg, an acquisition, a cost center map that doesn't look like the one from three years ago. The questions people ask the bill changed completely.

The tags didn't. Nobody re-derived them, because there was no trigger. Tagging gets treated as a project with a finish line, and once the coverage number is green the project closes and the taxonomy quietly goes out of date while the compliance report keeps saying everything is fine.

That's the reframe I'd offer. Tagging isn't a state you reach. It's alignment against a moving target, with a decay rate. Two ways to fail: you don't have tags, or you have tags that answer last year's questions. The industry measures the first one obsessively and the second one not at all.

And this is why remediation is real work rather than a script. Analysis first, to find out what the business actually needs to ask, then alignment. The scripting part is the easy end.

The Job

Three moves to consider and should be followed in the order listed.

Identify what is important to you for reporting and governance. Regardless of what you currently have or do not have, determine what is valuable to you. And start simply. If environment, business unit, and owner or critical, then stick with those. KISS as I say (Keep it Simple, Silly)

Score your tags on use, not coverage. For each key, name the report, the chargeback, or the decision it fed. Keys that can't answer that aren't neutral. They cost effort to enforce, they add noise to every query, and they inflate a compliance number that's telling you something false about your visibility. Retire them and say out loud that you're retiring them.

Give the taxonomy an owner and a trigger. Someone owns the tag standard the way someone owns a schema. And it gets re-read on events, not on a calendar: a reorg, a new product line, a new business unit, an acquisition, a new cost center map. If your last taxonomy review predates your last reorg, you already know the answer.

Provide flexibility. Allow teams to manage their own to help them understand and use the resources. But they must follow the standard, such as lower case, 3 or 4 letter, etc.

I’ll end by telling you that the grill is getting sorted out. Its not final, but I’ll either be returning for a new one, or hopefully they can send me a replacement part with a suggested fix. But I'd made that purchase based on a set of numbers that all looked good and none of which were measuring the thing that mattered, and I didn't find out until ... well, you know.

Coverage tells you the tags are there. It doesn't tell you if they answer anything. Go find out what your organization actually needs to ask, and then check whether your tags can answer it. Most of the time nobody knows.

Written with the help of AI. All the ideas expressed are mine and mine alone.

FinOps Company Spotlight

If you would like your company included in the Spotlight, contact the CloudXray AI Team

Company: CloudXray AI

Category: Managed Services & Consulting

What They Do: FinOps Consulting & Advisory Services; Owners & maintainers of the single largest FinOps company directory (finops.cloudxray.ai)

Why It Matters: Companies still need guidance on implementing FinOps and understanding the landscape of companies that exist

Reply

Avatar

or to participate