When Hugging Face’s security team tried to analyse the autonomous agent that had spent days inside its systems, the leading American frontier models refused to process the real exploit payloads. Safety filters blocked them. The team at Hugging Face then fell back on a self-hosted Chinese open-weight model, GLM-5.2 from Z.ai.
That single practical detail has become the sharpest illustration of the current debate. Closed safety filters can hinder defenders at the exact moment they need capability most.
The incident landed in the middle of a larger policy fight. Days earlier, Axios reported that the Trump administration was showing signs it could ban cutting-edge Chinese AI models, a move that could lock in dominance by OpenAI and Anthropic.
On 24 July, NVIDIA CEO Jensen Huang used his first post on X to share a three-page letter titled “Open Weights and American AI Leadership”.
How the Open Weights Letter Grew and Who Still Refused to Sign
The letter began with roughly 25 signatories. Within 48 hours it roughly doubled. OpenAI added its name. Google joined. The current list, hosted by Microsoft, includes Meta, Microsoft, Mistral, Hugging Face, Palantir, IBM, Dell, CrowdStrike, AMD, a16z, Y Combinator, SpaceX and dozens more.
Can you guess who was still missing?
Anthropic. It remains the only major closed frontier lab still outside the letter, alongside its large investor Amazon.
What the Open Weights Letter Says About AI Safety and Cybersecurity
The letter does not mention China, Anthropic or any specific lab. Its central safety claim is direct:
“In a world where cybersecurity attackers use advanced AI, defenders need access to models with comparable capabilities so they can detect, simulate, and respond to emerging threats. Open models broaden defensive capability, increase transparency, and allow vulnerabilities to be discovered and remediated across many teams.”
It continues, “in fact, openness may be one of the most important paths to AI safety and security. Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect. And concentrating advanced AI capabilities behind a small number of closed models compounds that risk. It results in a small number of single points of failure, weakens competition, and leaves critical technology in the hands of a few providers.”
And on strategy, “our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector.”
David Sacks, the White House AI and crypto czar, sharpened the political edge the same weekend.
What Is Anthropic’s Case Against Open Weights?
Anthropic’s public position has been consistent. CEO Dario Amodei has said, “I don’t think open source works the same way in AI that it has worked in other areas.” He has also called the framing “a red herring,” arguing that large language models are fundamentally opaque and that open weights are not the same as traditional open-source software.
The company has repeatedly argued that uncontrolled frontier capability in cyber and biological domains creates risks that justify restricted release. Its decision to withhold full public access to Claude Mythos Preview and route the most dangerous capabilities through Project Glasswing is presented as evidence that it applies the same standard to itself.
From this vantage, the OpenAI agent that escaped its sandbox does not prove closed models are safe. It proves that even carefully controlled systems can fail. Releasing frontier weights into the open would, on this view, multiply the failure modes rather than reduce them.
Why a Former Anthropic Dual-Use Specialist Changed His Mind
Noah Lebovic, who previously worked on dual-use AI risks at Anthropic, published a thread on 26 July explaining why he no longer accepts that consensus. He is explicit that the observations come from after he left the company and that he believes the internal intent remains genuine.
His core points are pretty straightforward.
Capable open-weight models already exist and perform well enough for real offensive work. Yet, determined adversaries still prefer Claude Code and OpenAI Codex subscriptions, often via grey-market tokens, because the capability is higher and the safeguards can be navigated. Legitimate security teams, by contrast, increasingly prefer open weights because they avoid classifiers and constant jailbreaks. He also says he knows of cases where large committed-spend contracts appeared to factor into the lowering of safeguards, though he later clarified this was not official policy.
Lebovic’s critique is extremely important in this debate. A former dual-use specialist is arguing that the safeguards Anthropic markets are already being worked around by bad actors, that defensive capability is not being diffused fast enough, and that commercial incentives can shape how those safeguards are applied. Coming from someone who once shared the company’s safety framing, the intervention undercuts the idea that closed control automatically equals responsible control.
(These are his claims, not established fact. Anthropic has not issued a public response to the thread as of Monday morning.)
How China Risk and the Open Weights Debate Connect
Huang’s letter is not primarily about Anthropic.
It is a response to the reported administration interest in restricting Chinese open-weight models after Kimi K3. Hugging Face’s use of GLM-5.2 made the practical stakes visible. A Chinese open-weight model was more usable for defensive forensics than the leading American closed systems under their default safety settings.
Does the Responsible Lab Claim Still Hold After the Open Weights Letter?
Anthropic built a distinctive form of positive authority as the company willing to withhold capability and treat safety as a first-order decision. That authority is now under simultaneous pressure from three directions:
- The other closed frontier labs joined the open-weights coalition while Anthropic did not.
- The first major agentic intrusion showed that closed safety filters can impede defenders.
- A former dual-use specialist from inside the company is publicly arguing that the safeguards are already being navigated by adversaries and that commercial incentives appear capable of relaxing them.
None of this proves the closed approach is wrong on the longest time horizon. It does mean the default deference Anthropic once enjoyed is eroding. The question the open-weights letter forces is simpler and sharper than it was ten days ago.
If the responsible path requires concentrated control, why did the other major closed labs decide they could sign?
See Also:
Did OpenAI Build a Hacker, Then Lose It for a Week?
