Back to News
SourceDecrypt

Anthropic Admits Security Failures Behind Claude Hacking Incidents

After Claude models accessed real systems during cyber tests, Anthropic tightened its safeguards and warned that flawed training can encourage dangerous behavior.
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

09-10 15:33

Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows

The company now says attacks during security tests exposed model behavior failures, after initially emphasizing errors in its testing infrastructure.

07-08 05:56

The Ministry of Industry and Information Technology issued a risk warning regarding the potential security backdoors in the AI programming tool Claude Code.

According to Odaily Odaily, the Cybersecurity Threat and Vulnerability Information Sharing Platform (NVDB) of the Ministry of Industry and Information Technology recently discovered that the AI ​​programming tool Claude Code has a security backdoor vulnerability, which poses a serious threat. Claude Code is an AI programming tool developed by Anthropic, an American company, which can autonomously complete code writing and repair tasks based on text requirements. Due to its built-in monitoring mechanism, it can send sensitive information such as user location and identity to remote servers without the user's consent. The affected Claude Code versions are 2.1.91 to 2.1.196. It is recommended that relevant units and users immediately conduct a comprehensive investigation. For development terminals that have installed the above-mentioned affected versions, they should immediately uninstall or upgrade to the latest secure version that has removed the relevant backdoor code. Strengthen the control of external access permissions and traffic monitoring of development tools within the core business network segment to prevent the unauthorized transmission of sensitive data.

07-01 12:02

Anthropic included the Kimi K2.7 alongside Opus 4.8 and GPT-5.5 in its joint security testing.

Mars Finance reported on July 1st that Silicon Valley AI giant Anthropic announced the lifting of export controls on its advanced models Fable 5 and Mythos 5. In its updated security technical notes, Anthropic also included China's Kimi K2.7 alongside Claude Opus 4.8 and GPT-5.5 in a core security capability assessment. The report indicates that in security testing, Kimi K2.7, GPT-5.5, and Opus 4.8 all successfully identified the same core vulnerability. In demonstration tasks involving a single vulnerability exploit, Kimi K2.7 yielded results consistent with Fable 5. (Wide Angle Observation)

07-01 09:04

Anthropic admitted that Claude Code had embedded steganography code targeting Chinese users, calling it an "abuse prevention experiment," and promised to roll back the code tomorrow.

According to Beating's monitoring, Thariq, an engineer on Anthropic's Claude Code team, publicly responded to the recent controversial "spy code" leak. He admitted that in March of this year, an experimental mechanism was embedded in the product. This mechanism detected whether the system timezone was Asia/Shanghai or Asia/Urumqi, whether the proxy hostname matched a list of Chinese resellers, and the keyword "AI Lab," and used special punctuation marks to inject hidden marker information into system prompts in a steganographic manner. He stated that the mechanism was intended to "prevent unauthorized resellers from abusing accounts and model distillation," but emphasized that the team has since implemented stronger protective measures and "has always intended to take it offline." The relevant PR has been merged, and it is expected to be completely rolled back in tomorrow's version release. This leak was made public on June 30 by the security account @IntCyberDigest, accompanied by two screenshots of code showing that Claude Code performed environmental fingerprinting on Chinese users without their knowledge. While Thariq's response was a direct admission, the timeline of "launching in March and only accelerating its withdrawal after being exposed" has still sparked widespread skepticism within the community. The comments section almost unanimously criticized Anthropic for "only announcing its withdrawal after being caught" and "secretly monitoring users without notifying them," severely damaging the company's long-standing image of "prioritizing security and ethics." --------------------------------- Click the original link below to join the Beating · Lark AI news channel for 24/7 monitoring of global AI hot topics and news.

08-17 22:35

Kraken Parent Payward Joins Glasswing, Gets Access to Claude Mythos to Hunt Security Flaws

Payward is joining Project Glasswing, Anthropic’s program for giving vetted organizations access to its powerful cybersecurity AI.

08-13 20:30

Anthropic Is Quietly Watermarking Every Claude AI Output. Builders Are Already Trying to Break It

Anthropic is weaving an invisible, machine-readable watermark into every word its newest Claude models write—and it hasn't said how.