Weights and Measures

July 21, 2026 — James Henry


Researching this post I tried to read OpenAI's own page announcing Stargate, the half-trillion-dollar data center project. 403 Forbidden.

This is an amusing echo of a previous post of mine where my $dayjob's egress got flagged by Cloudflare attempting to go to Dairyqueen.com.

OpenAI's site serves you fine in a browser and refuses anything that looks like a script--generic bot-blocking, the same default a million sites behind the same CDN use, defeated in my case by lying about a header. It means nothing about OpenAI specifically.

But it tickles my fancy that one of the companies scraping the entire Internet for content and monetizing it is blocking bot access to their own sites. I find it apropos of this article: they're gatekeeping knowledge.

The Gap That's Closing

If you're paying attention to the LLM newswire, you may have heard some version of this by now: open-weight models are catching the closed frontier.

Take the numbers the labs report about themselves with a grain of salt--especially lately. The 2026 search results for AI benchmarks are a swamp of content-farm blogs reposting and occasionally inventing vendor figures, and every number a lab publishes about its own model is a number a lab published about its own model. The signal worth trusting is the independent kind. Epoch AI puts the lag between the best open model and the best closed one at roughly four months. Stanford's AI Index clocked the top US and Chinese models at about 2.7% apart on head-to-head ratings--down from a seventeen-to-thirty-point gap three years ago, and closing despite the US spending something like twenty-three times more. On OpenRouter, Chinese open-weight models now account for around 61% of all the tokens people actually route through it. Four of the five most-used models are Chinese and open. Meta's Llama fell off the list.

Where the open models have genuinely caught up is on the benchmarks that make the press releases: contest math, graduate-level science questions. DeepSeek, Kimi, Zhipu's GLM--all within a point or two of the closed flagships on those. Inkling--an open-weight model from Thinking Machines (the lab some of OpenAI's leadership left to found), released July 15th with the weights on Hugging Face for anyone to download--matches or beats many closed models, and trails the frontier.

The frontier still clearly leads on two things. One is factuality--or more precisely, knowing what it doesn't know. Inkling's raw recall isn't really the problem (its 43.9 on a standard fact test is roughly Claude-Opus-class); the problem is that when it doesn't know, it bluffs rather than begs off. Artificial Analysis's Omniscience benchmark, which rewards a model for admitting ignorance instead of inventing an answer, gave Inkling a 63% hallucination rate and parked it near the bottom of the field. The other is the messy agentic work: long-horizon coding, using tools, staying coherent across a fifty-step task. A model you download and run on your own hardware is a real thing now. It is not yet the same thing as this month's frontier model. My last post said exactly that, and a month of newer numbers hasn't changed it.

The loudest headline of the month came when Axios declared China had "erased America's AI lead"--Moonshot's Kimi K3 topping Arena's front-end coding board above both Fable 5 and GPT-5.6 Sol, at 40% less cost. One grain of salt worth keeping: Anthropic has accused Moonshot and other Chinese labs of distilling its models--training on millions of exchanges siphoned from Claude, an accusation Beijing calls groundless. K3's weights are supposed to open up later this month.

The Moat and the Lock

So if the capability gap is narrowing to a few specific, stubborn places, the obvious question is what the closed frontier still has that the open models don't. The lazy answer is "not much, just secrecy."

The labs are ahead. The real question is how far. And they're ahead not only on the benchmarks--on the whole craft. They understand how to train these things, how to align them, how to take them apart and see what's inside, in ways the rest of us are behind on. Aside from being the incumbents, their real advantage is two things, and neither one is a file you can download.

The first is compute. The frontier is either a hyperscaler or married to one. Google and Meta own their data centers outright. OpenAI and Anthropic own almost none--they're fused at the balance sheet to Microsoft, Oracle, Amazon, and Google for sums with twelve digits in them. OpenAI has effectively stopped building its own; that Stargate page I couldn't read is about capacity another company owns and operates and leases to them. The concentration underneath all of it is the part that should give you pause: the big four cloud players are spending somewhere north of six hundred billion dollars on this in 2026, and Nvidia's own filings show four unnamed customers making up around 60% of its revenue. The entire AI frontier draws from one chip supplier and a handful of clouds.

The second is craft--the accumulated, tacit institutional knowledge of how to do this at all. By its very nature we don't actually know how deep it runs. Are we at their heels, or are they miles ahead and letting us see the last mile? One take is that each lab publishes a result about when open work starts closing on the same idea, and that behind every paper there's a backlog we never get to see. Another take is that they're scrambling to stay ahead. I'm not sure which of those two would bother me more. You can't measure that reservoir any more than you can hook the model.

Now hold those two next to each other and ask what open weights would actually cost the labs. Not the craft--a decade of alignment know-how doesn't evaporate because someone downloads a checkpoint. Not the data centers--Google's don't get smaller if Gemini's weights leak. What open frontier weights would cost is the revenue. Hand the weights over and anyone can distill them, fine-tune on top of them, and stand up a competing service at a fraction of the price--which is the very thing Anthropic just accused Moonshot of doing through nothing but an API. Open weights would make it trivial. And the API income is exactly what justifies the outlay for the next data center (although, as I argued last post, serving inference to the rest of us was never what these labs are actually after).

So the weights are load-bearing after all--just not for the reason on the label. The lock is real, but it guards the business model, not safety and not the capability lead. Closing the frontier weights protects the revenue that funds the frontier. What it does for safety is nothing anyone outside can check--and what it guarantees is that the most capable models are the ones no outsider can examine.

What You Can Only Do to a Model You Can Hold

There's a field called mechanistic interpretability whose entire job is to answer what is this model actually doing in there? Not what it outputs--what it computes, layer by layer, as inference moves through it. I've spent the better part of a year on a slice of it. Every technique in MI--probing the internal activations, watching a concept sharpen into a direction, ablating that direction to see what breaks--requires the same thing: the actual numbers flowing through the actual network, at every layer, which you can only get by running the weights yourself. No API gives you that.

So independent interpretability--the kind anyone outside the lab can do and check--exists only on open weights.

The obvious objection is that you don't need the weights on everyone's disk--you need access. Let vetted researchers in under NDA, hand approved auditors an activation-level API, run what Toby Shevlane named structured access: controlled, arm's-length interaction with a model you never actually hold. It's a serious proposal and I'd take it over nothing. But read what that same paper says structured access is for--it is explicitly designed to stop users from modifying or reverse-engineering the model, which is precisely the poking-around that interpretability is. It's built to prevent the thing I'm describing, not to deliver it. And even the friendly version has two holes: nobody offers it at frontier scale, and the moment a lab chooses who may look and on what terms, you are back to trusting the lab. A lab-selected auditor working under a lab-set NDA is a more polite "trust us," not independent verification. Independent means anyone can check--not that someone credentialed was permitted to.

I can show you it's not theoretical. In my own work, a concept auditor read a social-engineering prompt as malicious before the model wrote a single word of its reply--urgency spiking, source_credibility collapsing, deception crossing a threshold, all read straight off the internal state. That can run across 33 different open models, ten different families, from tiny ones up to 72 billion parameters. We could not run it on Fable 5, or Mythos 5, or GPT-5.6. Not because we lack the hardware but because you cannot hook a model you cannot download.

My work is basically one person's research line (with help from $dayjob!), self-published, and not (yet!) peer-reviewed, and I'll be the first to say the labs could almost certainly do it better. That's exactly the point I'm building to. Their interpretability is private. They look inside their closed models, for themselves, and share neither the weights nor, mostly, the methods. From where you stand, a closed model with world-class internal oversight and a closed model with none look identical. You can see neither the model nor the audit. It's "we looked, trust me bro," from the people who happen to be the best in the world at looking.

Trust Us, We Looked

Take the labs' stated reason at face value. They keep the frontier closed for safety--the June pullbacks came wrapped in exactly that language, a government cyber-review, an abundance of caution.

But closed is the precise thing that makes a safety claim impossible to check. You're asked to trust that the model is safe, about the one model we are structurally forbidden from inspecting.

To be fair to the labs, that isn't quite their argument. They don't claim that looking inside is dangerous. They claim that above some capability line, open weights are too risky to release, because anyone who holds them can fine-tune the safety training back out. And they're not wrong that this is real--the same openness that lets me read a model for deception lets someone else strip its guardrails off. As a place to draw a line, that's coherent.

The problem is who holds the ruler. The threshold--this model is safe to open, that one is not--is set by the labs, unfalsifiable from outside, and it happens to fall right at the edge of their most valuable models. You're asked to trust that the frontier is too dangerous to open, about the same models you are not allowed to inspect to check whether that's true. It's the identical trust-me-bro from a section ago, moved up one floor. Maybe the line sits where the danger actually is. Maybe it sits where the revenue is. From out here it's hard to tell the difference.

Who's Allowed in the Room

The MI work I have done, on rented/borrowed GPUs and a home server full of downloaded models, mapped a handful of concepts across 33 open-weight models. Even the small group of concepts I mapped out explains only a sliver--call it 6%--of the organized structure in there: the directions in a model's activations that sit above the statistical-noise floor and are plainly doing something, whether or not I can put a name to them. That sliver is my ceiling--a limit of my time and my budget, not of the field and not of the models. There is so much more in there to discover and understand.

It seems to me that, for anyone interested in what this technology actually does, inspectability is starting to feel like a right--and, from a security point of view, a need. It's not just the weights, it's the methods, and a publishing backlog nobody outside will ever see. The capable models are converging. The permission to look inside them is diverging. That's yet another enclosure--and it fences off nothing the labs were ever actually keeping.


Related: A Decent Subset of Human Knowledge, The Third Enclosure, Provenance, Not Enclosure, Reading the Flow, Using CAZ to See What the Models Are Thinking.


James is a security engineer who got 403'd trying to read a public web page about a half-trillion-dollar data center, and drew the obvious conclusions. This post is signed and verifiable.

Discussion