Want to become a partner or sponsor, click here to schedule time to talk with the team.
FinOps & Beyond is what engineering, finance, and IT leaders read to understand FinOps, and what it means for operating models, accountability, and spend decisions.

I spent some time over the last 2 weeks chatting with a company that's still in the domain, but still in stealth. What is interesting is how they are thinking about the cloud, and it's had me stuck on a question all week.

They're building a cloud operations solution, and they treat performance, reliability, and cost as one problem instead of three. Not three distinct tools with a human in the middle trying to correlate across them. Just one system.

To me, this is the Iron Triangle. Same concept as the project management one where you pick two of fast, good, and cheap. You can't max all three, so every workload you run is sitting at some point between them.

Now compare that with what I wrote last week. Stripe is paying more than $7 billion for OpenRouter, which puts 400+ models behind one interface and picks one per request based on the task, the need, and the budget. You declare what you want and the system decides where it runs.

So here's the question I can't seem to shake.

Why can't a team declare what a workload actually needs, on speed, on resilience, and on what it's worth to the business, and have it built and run wherever that lands, on any hyperscaler? And when the answer changes, because the workload matters more this year than it did last year, why is that a migration project instead of a setting?

The first half of that question sounds like an architecture one. The second half is the one that has the wheels upstairs spinning. But before getting to why we haven’t solve this problem, it's worth being honest about what we do instead.

The Decision You Didn't Make

Pick any service in your estate and ask what tradeoff it's running between performance, reliability, and cost.

You may get an answer for one or two systems, the ones with a real SLA and an owner. For everything else, the position on the triangle is an accident. It's whatever the reference architecture had, whatever the last person copied from the service next to it, whatever the provider's default was on the day it got stood up. Multi-AZ because that's the template. Provisioned rather than on-demand because someone got burned once. An instance family chosen by someone who has since left.

Those are all reliability and performance decisions, and each one set a cost. Nobody wrote down what they were buying or why.

Then, a year or two later, a cost optimization program shows up and asks an engineer to turn one of those dials down. And the engineer says No. And can you blame them?

Why the Rightsizing Ticket Gets Ignored

This is the part I think the core FinOps industry keeps getting wrong.

A cost recommendation with no performance or reliability context isn't a recommendation. It's a request to make somebody's system potentially worse in exchange for money that doesn't necessarily show up in their OKRs. Does this sound familiar? “Rightsize this node group.” “Drop this to one AZ.” “Move this to a cheaper storage class.” Every one of those is a single-variable request dropped into, really, a three-variable decision, possibly made by a person who can't see the other two variables (or doesn’t understand them) and won't be the one paged at 3am when the tradeoff turns out to have been wrong.

So the ticket sits. And FinOps concludes engineering doesn't care about cost, and engineering concludes FinOps doesn't understand the system, and both are sort of right and neither is going to win that argument.

You don't fix that with better dashboards or more accurate recommendations. You fix it by changing what gets asked. Not "can this be cheaper," which is a question engineers are right to resist, but "is this the tradeoff we want," which is a question they're actually equipped to answer and usually want to be asked.

The Waste No Dashboard Can See

Here's what the single-variable approach misses.

Somewhere in your estate is an internal reporting tool (or pick your app of choice) that has a handful of people using, but has been built running multi-region with automated failover and a recovery objective measured in seconds. It's extremely well architected. It's well utilized. And yet, it will never appear on a waste report, because nothing about it is idle, oversized, or untagged. Every cost tool you own will look at it and see a healthy workload.

Yet, it's still waste. It bought four nines for a system where four hours of downtime would have inconvenienced a Slack channel. Nobody asked for that reliability, nobody priced it, and nobody will ever find it, because utilization can't see a requirement that was never written down.

This example is different than the idle instance. Idle resources are a hygiene problem and we have tools to help solve that problem. Over-reliability is a decision problem, and the only way to find it is to ask what is the value of a workload and compare that against what it's actually providing. Which brings me back to the router from last week’s issue.

What OpenRouter Actually Does

OpenRouter is the iron triangle built as a system. In other words, it's a tradeoff engine. Grunt work is routed to a small, fast model, where as judgment calls go to a frontier model. And the routing rule is where the team, department, or organization wrote down what that particular request was worth. Same idea I was arguing in Issue No 22 about matching model tier to work tier. The value isn't the savings. It's that the tradeoff got made deliberately, per request, by a rule someone can inspect and change.

Now ask what the equivalent is for cloud, and the honest answer is there isn't one. You make your tradeoff once, at architecture time, and then you live inside it for five years. Changing it means a migration project. So of course nobody revisits it. The cost of changing your mind is so high that the original decision (or maybe guess) calcifies into a permanent architecture.

This is the problem. Not that we can't find the cheapest hyperscaler. That teams have no way to move a workload's position on the triangle without a quarter of work, so the position never moves, even when the workload's value to the business has changed completely.

Why It Doesn't Exist Yet

I believe there are 3 reasons this doesn’t exist yet (and probably never will), and only one of them is technical.

First, the interface never converged. Model providers all copied one API shape, which is why a router is even buildable. The hyperscalers spent twenty years diverging, because being different at the API was the moat. There is no common shape for a VPC or an IAM policy.

Second, the comparison unit doesn't exist. Dollars per million tokens compares pretty cleanly. A vCPU-hour does not, across different silicon and different network behavior.

And third, the Enterprise Agreement (or similar) commitment kills it. You can't optimize across providers when your marginal cost isn't marginal. Sign a large EA and the cheapest place to run the next thing is the place you already committed to, every time, until the commit burns down. Routing assumes a spot market. Enterprise cloud is a futures contract.

Two Things Just Changed That Can Help

The comparison became possible. The FOCUS Steering Committee ratified version 1.4 on June 4, and the headline addition is 17 new columns for comparing commitment structures across providers, plus invoice and billing period datasets. The Foundation's own phrase is "expose the anatomy of the deal." Read as a reporting upgrade, it's incremental. Read as the input a tradeoff engine would need, it's the first normalized price sheet that includes the contract and not just the line items.

The exit tax is on a schedule. Google dropped egress fees for departing customers in January 2024, AWS and Microsoft followed within weeks, with conditions most people have never read (Azure wants a full exit with subscriptions cancelled, AWS excludes CloudFront and Direct Connect). And under the EU Data Act, switching fees are prohibited outright from January 12, 2027 for anyone serving EU customers.

You can already see the pattern working where the conditions are right. GPU jobs are usually stateless, finite, expensive enough to be worth moving, and frequently bought outside the big commitment (although one could argue this too is changing), and a whole tier of brokers and marketplaces has grown up routing that capacity across providers on exactly this basis. Teams running training jobs already choose their triangle position per run: interruptible and cheap, or guaranteed and expensive. That's the normal way to work with GPUs. It's an exotic idea for everything else you own.

The Job

Three moves for this week.

  1. Write down the triangle for your top ten workloads. One line each: which of the three is the release valve. If this system got slower, would anyone notice. If it was down for four hours on a Saturday, who would care and what would it cost. What are we paying for the answer.

    You will not have this written down for most of them, and the gap between what teams assume and what the business would actually accept is where the conversation needs to happen
    .

  2. Stop sending single-variable cost recommendations. Every cost improvement request goes out with the other two dials (performance and reliability) attached.

    Here's the change,
    Here's what it does to latency
    Here's what it does to your failure domain
    Here's what we're proposing to buy with the difference.

    It's more work per ticket yet it's the difference between a request engineers ignore and a decision they'll engage with. To be clear, this is a change management fix, not a tooling fix, which is why no vendor will sell it to you.

  3. Go hunt over-reliability. Find every workload carrying more nines than anyone asked for. Sort the list by cost of the redundancy. Then take the top few to the appropriate business owner and ask directly what they'd accept. Some will say keep it, and that's a fair answer, because now it's a decision instead of it just happening.

For a decade and more we've treated cost as its own discipline with its own tools, its own team, and its own tickets. That was always felt a little fake, and it produced a function that can describe spend precisely and change it slowly. Yet, performance, reliability, and cost were one decision the entire time. Somebody made it, a long time ago, usually by copying the service next to it, and then everyone stopped looking. Value isn't a number you drive down. It's a position you picked on purpose and can defend when someone asks why. Governance is what keeps it from rotting the moment you stop looking.

Written with the help of AI. All the ideas expressed are mine and mine alone.

FinOps Company Spotlight

If you would like your company included in the Spotlight, contact the CloudXray AI Team

Company: CloudXray AI

Category: Managed Services & Consulting

What They Do: FinOps Consulting & Advisory Services; Owners & maintainers of the single largest FinOps company directory (finops.cloudxray.ai)

Why It Matters: Companies still need guidance on implementing FinOps and understanding the landscape of companies that exist

Reply

Avatar

or to participate