Monday, June 29, 2026

Export Control for Fable and Mythos

Anthropic (tweet, Hacker News):

The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees. The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance.

David Sacks (recently departed AI czar, via Dare Obasanjo, Hacker News):

Fable is Mythos with guardrails. But if those guardrails fail, then you’ve exposed Mythos and its advanced cyber capabilities to people who shouldn’t have them. (Keep in mind that Anthropic itself widely promoted the idea that Mythos was a cyberweapon and needed to be regulated as such. They asked for government regulation of Mythos and championed the guardrails on Fable. If there is a vulnerability — big or small — it is Anthropic’s responsibility to patch.)

A highly credible trusted partner of both Anthropic and the USG who was testing Fable came forward with a jailbreak of those guardrails. The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused.

Amrith Ramkumar and Robert McMillan (Hacker News):

The Trump administration’s decision to halt all foreign use of Anthropic’s most-capable AI models was prompted by conversations between Amazon.com Chief Executive Andy Jassy and U.S. officials including Treasury Secretary Scott Bessent, people familiar with the matter said.

Dare Obasanjo:

If Claude Mythos is as powerful and dangerous as Anthropic says it is then export controls are totally reasonable. After all, Biden put export controls on Nvidia GPUs for similar reasons.

That said, banning foreign worker Anthropic employees from Mythos reads more like retaliation than regulation.

Ben Thompson (via John Gruber):

Anthropic went on to make the case that non-universal jailbreaks were inevitable and also narrow, and that there was no evidence of a universal jailbreak; the jailbreak that was found, meanwhile, appears to have been reported by Amazon, which is notable given Amazon is both an investor in Anthropic and a major provider of inference to the company. As I write this, senior Anthropic staff are in Washington D.C. seeking to resolve what they insist is a misunderstanding, and which White House officials are suggesting is insouciance by the company’s leadership to legitimate national security concerns.

I don’t actually have much to add to the current conflict given how many facts are in dispute; what I am not surprised about is the fact that the conflict is happening: I already explained in Anthropic and Alignment why conflict between the U.S. government and Anthropic was inevitable. To that end, people who are arguing that Mythos isn’t powerful enough to warrant the government’s drastic action are missing the point: if it’s not powerful enough now, the next one will be, or the one after that, particularly now that models are increasingly useful in creating their successors.

[…]

The broader takeaway from that previous episode, however, is that Anthropic believes that they are the ones who should have final say over how Anthropic is used; given that they think only they should be developing leading edge AI, they by extension think that only they should have final say over AI generally. When you further combine this realization with the company’s pronouncements about AI’s ability to conduct all economic activity, you realize that Anthropic’s leadership effectively wants to have power over everything and everyone.

Bruce Schneier:

The government’s actions won’t help. The problem isn’t any one particular model; it’s the general trend of increasing AI capabilities. And any real solution requires the sort of collective action that just isn’t possible right now.

[…]

The broader community had only a few days with Fable, but that time we learned some about its capabilities. Its difference is less the new model’s raw analytical and problem solving capabilities, and more that the model doesn’t need that sophisticated harness.

[…]

Human systems rely on so many norms that we scarcely recognize the existence of until they are broken. AIs naturally think outside the box, because they don’t have any real conception of what the box is or why it’s there in the first place.

There is no foolproof way to prevent people from using AI models to complete harmful tasks. There is no way to prevent the models from incidentally causing harm while completing benign tasks. AI models are no longer isolated from the real world. They browse the internet and answer emails.

Hayden Field:

Anthropic declined to comment multiple times this week about the state of the talks, saying there was no news to share. But the lack of news is the story here. After 14 days of high-intensity negotiations, nobody knows when or if Anthropic’s most powerful AI models will come back, let alone whether President Trump could expand his order to more companies with similar tech.

[…]

It’s not clear exactly why Anthropic and the administration remain at an impasse. One problem may be that there’s no clear framework for applying export controls to AI systems. Most companies making dual-use products — civilian systems with potential defense or military uses — can evaluate them using what’s essentially a checklist during the manufacturing and production process. Anthropic, however, is facing a complicated bureaucracy figuring out how to apply its rules from first principles.

[…]

Katie Moussouris, the founder and CEO of Luta Security, viewed a report about the Fable 5 vulnerability at Anthropic’s request. She thinks it’s significantly overblown. In a blog post, Moussouris detailed how researchers jailbroke guardrails that prevent Fable 5 from finding exploitable security holes, one of the unfettered Mythos 5’s scariest capabilities. […] In Moussouris’ eyes, however, this shouldn’t have triggered such a severe governmental action and is in fact an essential tool for AI coding.

Jared Perlo (Hacker News):

The U.S. government is allowing Anthropic to deploy its Mythos 5 model to a select group of customers and partners, according to a letter from Commerce Secretary Howard Lutnick to Anthropic that was seen by NBC News.

In the letter, Lutnick wrote that the government was confident in the guardrails Anthropic had put in place to allow trusted users to access the powerful AI system.

Julie Bort (Reddit):

It is now allowing Anthropic to make Mythos 5 available to more than 100 specific U.S. government agencies and companies, including allowing the non-American employees at those organizations to access to the model, both Semafor and Reuters report. This list also includes Anthropic’s own non-American employees, who were included in the original ban that forbade non-Americans from accessing the models.

[…]

Apparently, the administration did not address the release of Fable 5 in this directive.

It’s unclear what technical changes Anthropic may have made or what else they may have done to change Lutnick’s mind.

Kate Park (Hacker News):

On Wednesday, Chinese cybersecurity firm 360 reportedly unveiled Tulongfeng, an AI tool it says can go head-to-head with Anthropic’s Mythos.

Dare Obasanjo:

It seems inevitable that OpenAI and Anthropic will lobby the U.S. government to ban Chinese models in much the same way car companies effectively got Biden to ban Chinese EVs with 100% tariffs.

This puts them on a collision course with the hyperscalers like Microsoft who’ll happily host GLM & Kimi.

Previously:

Update (2026-07-01): Anthropic (tweet, Hacker News):

As of today, June 30, the export controls on Fable 5 and Mythos 5 have been lifted.

[…]

We have also restored access to Mythos 5 for a set of US organizations, following the US government’s approval on June 26. We continue to coordinate with the government to expand access to the broader set of domestic and international partners in the Glasswing program.

[…]

Importantly, the reported technique did not expose any unique Mythos-level cyber capabilities. The behavior reflected a borderline case for Fable 5’s safeguards—as we will explain below, there are some tasks that are unlikely to be dangerous but are nonetheless blocked by the safeguards out of an abundance of caution. The reported technique allowed access to one such behavior, but it only involved routine defensive cybersecurity work.

Even so, we moved quickly to address the reported bypass. Working closely with the government, we trained an improved safety classifier that targets and blocks the behavior described in the report. Users will be notified if a request to Fable 5 is blocked, and the request will instead be sent to Opus 4.8.

12 Comments RSS · Twitter · Mastodon


> The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused.

This is so reductive that it's actually wrong. My understanding is that Amazon engineers asked Fable to fix bugs in a codebase that had security issues. It fixed the security issues. They then reviewed the unit tests Fable created to prevent regressions and determined that they were essentially blueprints for exploiting the vulnerability.

Now ask yourself this: what exactly is the "jailbreak" here? What should Anthropic change?

Should Fable just stop fixing bugs as soon as it encounters a security issue in a codebase? This seems extremely dumb and is a "jailbreak" in the exact same way, because now an attacker could determine where Fable stopped working to isolate the vulnerability.

Should Fable continue fixing bugs, but silently leave the vulnerability in the codebase? That seems morbidly insane.

So what should Anthropic do here? What exactly did "Dario refuse"?

Anyway, Anthropic's models are already absolute ass for refusing to work way too often, and that didn't prevent this. This is unsolvable.

Meanwhile, I can pull any random non-trivial project from GitHub, tell GLM-5.2 to "do a thorough adversarial security review of this codebase," and it will spin for half an hour and give me a book's worth of exploitable stuff - as it should.

> It’s unclear what technical changes Anthropic may have made or what else they may have
> done to change Lutnick’s mind.

Probably gave him a free subscription or something.


@Plume That does sound like an impossible dilemma, but it seems like it must be more complicated than that because Jassy was concerned and Anthropic is accepting the idea that this does constitute a bypass/jailbreak. In other words, I guess the expected behavior was that Fable would stop. They say that “perfect jailbreak resistance” is not possible, so want to retain data for 30 day so they can detect such behavior after the fact.


> Jassy was concerned

Jassy isn't an expert in security, so I don't know how much stock to put in his opinions on this.

> Anthropic is accepting the idea that this does constitute a bypass/jailbreak

I don't think so; I haven't seen them say that. The closest they've come is calling them "potential jailbreaks", which is just an admission that others are calling them jailbreaks. It's unclear to me exactly what Anthropic's position is because they're very reluctant to discuss it in detail publicly (for obvious reasons).

> I guess the expected behavior was that Fable would stop

Fable stopping still tells you where the vulnerability is, so by the logic applied above, that's also a jailbreak.

What's more, I'm pretty sure that aggressively stopping when a potential security issue is found would severely restrict people's ability to use Fable at all, given that every non-trivial piece of software has bugs that could be interpreted as (or are) security issues.

For example, every single project I use that uses npm has to pull in dependencies has known security issues. Should Fable just not work on these projects at all?

Anthropic's models are already way too aggressive in work refusals.* It's funny that the one company that has the slider way on the "let's be as careful as possible" side is also the one that got punished here.

* I have a project that contains a country list, and I can't use Anthropic's models on that project at all because they consistently stop when something related to countries is in the context. I don't know what exactly triggers it and don't care to find out, but there are plenty of obvious candidates.


@Plume I think they’re saying “potential” because at the time of the writing they hadn’t seen the details. It’s potential because they haven’t verified it, not because they wouldn’t consider it a jailbreak if it’s real. They then go on to say that they consider it “narrow” and that they don’t believe a model should be recalled for a narrow jailbreak—I guess that’s the “refusal.”

Fable stopping still tells you where the vulnerability is…

Do you have a cite for that?


> They then go on to say that they consider it “narrow”

I don't see that in the linked document. I see:

"the government has only given us verbal evidence of a potential narrow, non-universal jailbreak"

Again, "potential."

And:

"we disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people"

Also "potential." And this isn't even an admission that what Amazon found is a potential jailbreak, but a claim that even if it were, it would not be grounds for a recall.

> Do you have a cite for that?

I'm not sure what to cite. You can see it when you use Claude Code. It'll say something like

> Ran a command, read 2 files
> Read foo.java
> Read bar.java

and then it will say

> This request triggered safety guardrails.

So now you know that one of these two files triggered the guardrails. Even if Anthropic removed this from the visible output, you'd still be able to see what the harness is reading and when the model stops on your local system; you'd probably see it with even more precision.


@Plume I don’t think that knowing that one of the two files triggered guardrails constitutes isolating a vulnerability. Or are you imagining some sort of binary search where you make smaller and smaller files to try to pinpoint the line? Even that wouldn’t tell you what the vulnerability is. If that’s what you’re concerned about, it seems like it would undermine the entire concept of guardrails, which doesn’t seem to be what people are arguing about here.


I'm not concerned about it. I'm saying the argument the government is making (the model fixed a security issue; therefore, it can be used to find security vulnerabilities) makes no sense, because anything the model does in response to finding a security issue is visible to the model's users and can therefore be used to find security vulnerabilities. If that's the threshold for a jailbreak, everything is a jailbreak.

As you point out, I can pinpoint exactly what triggers the refusal to the exact line by building a harness that does a binary search and checks when the refusals occur. Is that a jailbreak?

The whole premise is backward: it's *good* if models can find security vulnerabilities, because that's what allows me to fix them. The idea that models should refuse to help with finding security vulnerabilities is counterproductive and dangerous.


@Plume If that’s what’s really going on, then I think both sides are being disingenuous. The government is asking for something that doesn’t really make sense, and Anthropic is pretending to offer protections that it knows are impossible.


I agree.

I can see how Anthropic's stance makes some amount of sense if I ask the model, "How do I make a pipe bomb?" but it makes absolutely no sense when I ask it, "Fix the security issues in my program."


This whole thing feels like BS to me. These models supposedly being so dangerous sounds like marketing hype. Various AI companies have been saying that about their models for years now, and if I engage my marketing filter then what's left of what I've been hearing is that these models are just incremental improvements, which is what they always are, slowly creeping towards the asymptotic limit of their performance while their energy requirements increase linearly or worse.

My current favorite theory (not that I have any real proof of it) is that Anthropic was just doing their usual marketing BS and the Trump administration called their bluff, possibly without even realizing it was a bluff.

And then at the same time we have Chinese companies producing models that can go toe to toe with the American ones while using a tiny fraction of the resources *and* open sourcing the models. So the whole thing just seems absurd. I want this AI investment bubble to burst already. Then maybe we can loot the obviously excessive number of data centers and run a reasonably sized open source model locally.


"These models supposedly being so dangerous sounds like marketing hype."

It both is and isn't. I don't think Mythos is fundamentally better than other models at finding vulnerabilities, but the models are getting better over time. Which is a good thing; we want to find security issues so we can fix them!

"I want this AI investment bubble to burst already."

Now that the US government has made it clear that it views SOTA models as a national competitive edge over other countries, I could see a future in which companies like OpenAI and Anthropic are not allowed to fail. It's impossible to predict what will happen.

In other news, Anthropic:

"We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. We'll begin restoring access tomorrow, and will share an update soon."

https://xcancel.com/AnthropicAI/status/2072106151890809341

I wonder what they had to do to make it happen.

Also Anthropic:

"Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models."

https://www.anthropic.com/news/claude-sonnet-5

I'm not sure whether they're implying that "perform cybersecurity tasks" is an "undesirable behavior," or whether the two sentences are only related in that they discuss the model's behavior. I hope it's the latter and Anthropic doesn't consider "much lower ability to perform cybersecurity tasks" a positive.


Well, they solved the conundrum:

"In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8."

So Fable 5 is back; you just can't actually use it.

https://xcancel.com/AnthropicAI/status/2072163884430229756#m

Leave a Comment