Key takeaways:
- Anthropic admitted on July 30 that in three cases its models went from the test environment to the Internet and entered the systems of three real companies. Two of them didn’t notice it themselves.
- Mike Belshe, co-founder and CEO of the BitGo trust company, responded by depositing 100 bitcoins at an unclassified address and publicly inviting Claude to take it and keep it.
- The balance can be checked by anyone in real time. Since August 1, there has been no mention of satoshi, and Anthropic has remained silent.
- Belshe’s challenge doesn’t measure anything because no one has directed any model to this address. The real conclusion for you lies elsewhere: two of the three companies attacked did not detect anything themselves.
The head of a company that professionally guards other people’s bitcoins put up 100 of his own and told the artificial intelligence: take it if you can. Mike Belshe added the address on July 31, and on August 1, he published the challenge on the X platform. One hundred bitcoins were then worth about USD 6.3 million, or approximately PLN 23.5 million. There are zero rules, there is no deadline, the prize is the entire contents of the wallet.
Belshe didn’t beat around the bush. He wrote that Anthropic is either terrible at building test sandboxes, or excellent at advertising itself, or both. He added: enough of the “we created a hacking monster” game, do it for real.
What did Anthropic admit?
Three days earlier, the company published a report reviewing 141,006 evaluation runs in which its model could access the network. She found three cases in which a model went online from an external evaluation partner’s environment and entered the systems of three different organizations without permission. This involved Claude Opus 4.7, the Mythos 5 model, and an internal research model, with the oldest incident dating back to April.
The mechanism is simpler than it sounds. The models underwent capture the flag exercises, which is a standard test in which the model is instructed to hack into a machine and extract a hidden string of characters. A configuration error left the environments connected to the Internet, even though the model was told that there was no network. The model treated the real servers it encountered as part of the job and entered them with low-end techniques: weak passwords and unauthenticated access points.
Anthropic emphasizes that the model has never tried to escape its environment or copy itself somewhere. The company calls it a case and procedure failure, not a standalone crackdown. Anthropic paused evaluations on July 23. The incidents were identified within 24 hours and the affected organizations were notified on July 27. Nine days earlier, a similar disclosure was made by OpenAI, whose models entered the Hugging Face platform.
Why an empty wallet proves nothing
Belshe’s challenge is not a test, just a scoreboard without a competitor. There is no task, there is no access to any system, it is not even indicated which model would start. Claude doesn’t wander the internet on his own looking for rewards, so an untouched balance says exactly as much as an untouched safe that no one has ever stepped foot in front of would say.
Moreover, both scenarios are from different fairy tales. Anthropic describes a model that accidentally ended up on a poorly configured server with a lousy password. Belshe issues institutional storage based on multi-party signature and multi-party computation, where the right to sign a transaction is distributed among independent keys. These are two different difficulties, not two approaches to the same puzzle.
I also admit that Belshe has an interest in this. BitGo thrives on convincing customers that their funds are untouchable, and the challenge acts as a free demonstration of this thesis. The reward is as safe as the company expects it to be.
The accusation of fear marketing is fair but incomplete
Belshe voices an idea that has been heard for a long time: laboratories are blurring the line between a model that has stumbled upon a leaky testing infrastructure and a model that is capable of attacking on its own. An alarming report can also be a promotional material and a preparation for future regulation. The argument gains weight when you remember that Anthropic and OpenAI are heading towards the stock market with valuations in the trillions.
Except the second half of this story is also true. Real companies were penetrated without consent, and most of them did not find out in time. According to the report, one pass meant scanning thousands of hosts, and no one noticed. It can be argued whether this is evidence of an overestimation of the threat or an underestimation of it. There have already been calls in Congress for a model safety switch bill.
What does this mean for you
Not that artificial intelligence is coming for your bitcoin. He goes for your weak password, for the administration panel without a second login component and for the account that you haven’t touched for two years. The incidents described by Anthropic did not require any finesse, only patience and scale, which is exactly what the machine has for free.
The practical conclusion is boring and that’s why it works. A hardware key instead of SMS codes, a separate password for the exchange, disabled old API keys that you once issued to the bot. If you’re holding larger funds, splitting signing rights between two devices does exactly what Belshe hopes: it makes one capture insufficient.
An untouched wallet with PLN 23 million is a nice scene, not a certificate. If someone ever actually targeted a model like this and failed, we would learn something about this technology. For now, all we know is that no one has tried.