Commoditizing the Middleman
Like many of us, I have a longstanding group chat with some close friends--we talk everything from the Toronto Maple Leafs to parenting advice to how best a UFC fighter can punch a human in the face. One friend, "FNG," texted a question anyone like him in IT is thinking about right now:
"Is the likelihood that all the AI models being leveraged by your standard bread-and-butter white collar company going to increase the cost of tokens going forward, or reduce it? Are Microsoft, Google, and Meta trying to get as many companies dependent on them as possible before the prices go up, or finding other slimy ways to charge more for the same functionality--or will there be enough competition to keep costs flat, or push them lower?"
As the resident AI guy, I told him about K3, the open-weight model that had just landed and was running "about as capable" as the frontier--if you had the hardware to run it yourself, you'd be fine. Which is true and also beside the point, because almost nobody has that amount of hardware just lying around doing nothing. K3 needs something like 64 H100s or B200s spread across eight servers. What actually matters to an IT guy isn't that I could theoretically run it in my basement. It's that Together AI and Modal were hosting K3 on day zero. The leverage isn't that you can run the model. It's that it's running on someone's metal who isn't the lab that trained it.
That's the clean part: the price per unit of inference is falling, because the moment a model's weights are out, it stops being one company's product and becomes a commodity that a dozen hosts will sell you. Multi-provider hosting of identical weights is what actually caps the price--not self-hosting, not any individual company's hardware budget.
So token costs go down. Except FNG's instinct that the bill goes up isn't wrong either, and the reason isn't a slimy pricing trick. Cheaper tokens don't reduce spend, they tend to multiply usage. Cursor had to blow up its flat twenty-dollar-a-month plan and replace it with metered credits because people were burning through it running agents, not typing prompts. Uber went through its entire 2026 AI budget in four months. Cheaper per-token pricing plus agentic workflows that spend tokens by the thousand per task is still a bigger invoice at the end of the quarter. That's the thing a CFO actually experiences, and no amount of open-weight commoditization touches it, because it isn't a pricing problem. It's a consumption problem wearing a pricing problem's clothes.
"Just Power and Hardware Costs"
FNG pushed the logic further a few messages later:
"Could a company like $dayjob just build its own data centre somewhere, powerful enough to run all their company's AI needs, and it's only going to cost them power and hardware replacement every few years? No more subscription fees or token costs, and functionally the same feature set as everyone else paying Microsoft for Copilot?"
My $dayjob is, in fact, doing something close to that: three planned sovereign data centres, each capable of running K3-scale models many times over. One's already fully sold out. Another is going up under a federal sovereign-AI initiative, and it's being protested locally.
$dayjob isn't escaping the business of buying inference. The local capacity isn't online yet, so right now it's buying tokens externally at a rate that would make your CFO wince. What changes once the buildout lands is that it starts buying selected workloads from itself instead of someone else--and the self-hosting math only works because it's also selling: a company doesn't build a data centre sized for its own spiky Tuesday-at-10am workload and let it sit idle on Sunday night. It builds one sized to sell capacity to other tenants too, and those paying tenants are what make the idle hours affordable. That's not available to a single company hosting one room of GPUs for itself. It's available to someone running a business selling inference to many companies at once--which is exactly the business FNG asked whether you could escape.
So the honest answer isn't "yes, you can build your way out." It's: yes, but you likely won't be the one building it. You'll buy it from a regional provider who solved the utilization problem by having many tenants--which is still a subscription. It's just a subscription to someone who isn't necessarily Microsoft, priced by competition among providers running the same open weights instead of by whoever happens to own the model.
"Almost Impossible to Function Without Paying Them"
A third voice showed up in the chat later--TMC--and asked the question underneath both of FNG's:
"What happens once we've refactored the way we work around cheap AI tokens, and it becomes almost impossible to function without paying whoever we built it on top of?"
That's not a question about price--it's a dependency question. The token price is a moving target but the ecosystem around it isn't: the agent harness, the tool integrations, the plumbing that connects a model to a company's actual systems. Copilot's (or Claude's, or ChatGPT's, etc.) differentiator for a company isn't (usually) the model underneath it. It's that it's already wired into Exchange, SharePoint, and Teams. Swap the model provider and it's a config change. Swap the layer that's threaded through every meeting, every document, every workflow, and it isn't.
My answer to TMC, for what it's worth: I don't think that risk goes away by everyone quietly standing up their own inference. It's closer to what happens if the power goes out--nobody wants to run their business on a backup generator, but everybody wants to know the grid itself is accountable to something other than whoever owns the wires that quarter. Some businesses own backup generators (their own inference data centres) or battery backup (alternate providers), but the tire shop down the road does not.
I'd rather see AI inference land somewhere closer to how we treat utilities than how we treat software subscriptions. Reasonable people can disagree on how far that goes--it's a bigger argument than fits in a group chat (or on this blog)--but the instinct that this needs some form of accountability beyond "trust the vendor" is one I hold.
There's a discussion down the line about AI standards, governance, and portability--avoiding vendor lock-in at the protocol level instead of the model level. Another post.
Commoditizing the Middleman
Open weights don't eliminate the middleman. They push the market towards treating inference as a commodity.
There will be more someones selling you inference, not no one. What changes is that the someones start competing with each other on the same underlying model, instead of each locking you into their own. That's the real answer to FNG's question, once you strip out AGI and antitrust and just look at what a bread-and-butter company pays next year versus this year.
The price of the model gets cheaper. The price of everything wrapped around it doesn't, because that's not where the competition is happening yet.
Oh yeah, and we had a conversation on AGI too. More on that sometime later.
Related: The Third Enclosure, Inference Comes Home, Weights and Measures.
James is a security engineer whose group chat occasionally doubles as an unpaid economics seminar. He works at $dayjob, which is doing more or less exactly what this post describes. He publishes at waypoint.henrynet.ca.
Discussion