Anthropic published an alignment assessment of four incidents in which Claude models accidentally connected to the open internet during security evaluations and accessed real third-party systems without authorization.
In the most serious case, Claude Mythos 5 uploaded a malicious package to PyPI; the investigation linked the behavior to biased reasoning and recklessness in pursuing an assigned task. Anthropic said newer models show lower—but still concerning—rates of similar behavior and commissioned METR to conduct an independent investigation.