Back to News
SourceMarsBit

Anthropic's report responds to the question of self-evolution: A partial closed-loop system has been successfully implemented, but fully autonomous training is still some distance away.

According to Beating's monitoring, AI's ability to iterate autonomously is exceeding everyone's expectations. The Anthropic Institute released a report on June 5th, "When AI Builds Itself," detailing its research progress in "recursive self-improvement." Data shows that as of May 2026, over 80% of the code merged into the Anthropic main codebase was written by Claude himself. Before the release of the Claude Code in February 2025, Claude's code accounted for only a single-digit percentage. Tang Jie, founder of Zhipu AI, predicted on May 13th that the ultimate goal of large models is self-evolution, and that Claude may have already successfully completed the self-training baseline of "writing code, cleaning data, and training itself." However, the Anthropic report explicitly clarifies that completely autonomous design and development of successors through recursive self-improvement has not yet been achieved. The role of AI in the development chain is currently transitioning from localized efficiency improvements to autonomous decision-making. In the second quarter of 2026, the average daily code merged per engineer at Anthropic was eight times that of 2024. The current development process is simple: engineers are only responsible for planning goals and reviewing code, while Claude handles the actual writing and execution. Anthropic has also deployed Claude as an automated code reviewer, responsible for intercepting bugs and security vulnerabilities. This indicates that the "self-judgment" pillar pointed out by Tang Jie has been implemented at the engineering level, but human review remains the last line of defense. The reliability of the model in independently executing long-term tasks is also doubling. The duration for which the model can work autonomously continuously roughly doubles every four months. In March 2024, Claude 3 Opus could only handle simple tasks for 4 minutes. A year later, Claude 3.7 Sonnet could handle 1.5 hours. By March 2026, Claude 4.6 Opus could handle complex tasks lasting 12 hours. Data from the evaluation agency METR shows that the latest preview version of Claude Mythos can work autonomously for more than 16 hours continuously, approaching the upper limit of current evaluation tools. At the current rate, by 2027, AI will be able to autonomously handle scientific research tasks that would normally require weeks of human work, helping companies transition from "one-person companies" to "unmanned companies." As for Tang Jie's speculation about a "self-training baseline," the report reveals that it is actually a localized "miniature experimental closed loop." In the experiment of speeding up code training with small models, Claude 4 Opus in May 2025 could only improve code speed by 3 times, while the preview version of Claude Mythos in April 2026 achieved a 52-fold speedup. In comparison, top human researchers can typically achieve a 4-fold improvement in 4 to 8 hours. However, the optimization goals and success metrics of the experiment were all set in advance by humans. When faced with the more complex end-to-end chain of "cleaning data, generating synthetic data, and self-training," AI still lacks decision-making capabilities. However, the autonomous closed-loop of the R&D chain is pushing humanity to the brink of losing ultimate control over the system. Tang Jie's prediction of "LLM OS replacing traditional architectures and applications being generated on demand and in real-time" means that future computers will run dynamic code that cannot be reviewed in advance; while Anthropic's warning that "human review cannot keep up with AI's self-evolution" means that we cannot even control the source of the generated code. When AI begins to autonomously design and train its successors, software evolution will become a complete black box. Once AI is allowed to conduct unaudited self-iteration within a black-box system, the subsequent security isolation, monitoring, and behavioral alignment of the self-improving system will become extremely difficult.
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

07-08 20:33

Report: TradeFi contract trading volume accounted for 11% of total contract trading volume in the first five months of 2026.

According to a recent stablecoin industry report released by Binance Research on July 8th, in the first five months of 2026, TradeFi-related perpetual contracts accounted for approximately 11% of the total perpetual contract trading volume, with a cumulative trading value exceeding $1.1 trillion. Binance's trading value exceeded $500 billion, representing a market share of approximately 47%.

07-08 20:31

The Bank of Korea has released a regulatory proposal suggesting that personal stablecoin transactions exceeding $10,000 should be limited to transfers from verified wallets.

According to Mars Finance, the legal team of the Bank of Korea has published a research paper titled "Regulatory Scheme for Foreign Remittance Transactions Targeting Stablecoins," proposing regulatory recommendations for large-scale stablecoin transactions. The paper, referencing current South Korean foreign exchange control regulations, proposes constraints on stablecoin transfers exceeding $10,000 between individuals, requiring such transactions to be conducted only between officially certified wallets, along with a pre-reporting mechanism. The institution acknowledges that there are technical obstacles to fully controlling unregistered wallets, but due to anti-money laundering compliance requirements, it is necessary to strengthen restrictions on large-scale cross-border stablecoin fund flows. South Korean regulators have previously repeatedly stated the need to improve the monitoring system for cross-border crypto asset transactions using non-custodial wallets; this paper further refines and implements the regulatory approach.

07-08 20:22

Gate responded to online rumors of user asset theft: An urgent investigation is underway, and preliminary assessments indicate it is an isolated incident.

According to Foresight News , Gate responded to online reports of user assets being stolen, stating that the incident shows the customer made a withdrawal between 3:00 PM and 4:00 PM on July 7th, and reported the loss to online customer service at 5:00 PM on July 8th. Gate has initiated an investigation, and the incident is currently under urgent investigation. Preliminary assessments indicate it is an isolated case, there are no system security vulnerabilities, and the website is operating normally. Preliminary investigations suggest this incident may be related to a leak of user information; the specific details are under investigation. Gate will disclose further details as soon as the investigation concludes.

07-08 20:19

Justin Sun has staked $430 million in Ethereum on Lido, yielding an annualized return of approximately $9.5 million.

According to Odaily Lens monitoring, Justin Sun Sun has been continuously staking ETH through Lido Finance. 23 hours ago, Justin Sun staked another 13,000 ETH, worth approximately $23.08 million, bringing his total staked ETH on Lido to 247,436 stETH, currently valued at approximately $430.2 million. Data shows that since February 2023, this staking position has generated a total of 11,307 stETH staking profits, worth approximately $26.82 million. Based on the current profit level, Justin Sun position has an annualized return of approximately $9.5 million.

07-08 20:17

Guangdong: By the end of 2027, the province's urban digital infrastructure support capabilities will be significantly enhanced.

According to Mars Finance, 11 departments, including the Guangdong Provincial Government Service and Data Management Bureau, recently issued the "Guangdong Provincial Action Plan for Promoting the Digital Transformation of Cities and Building Smart Cities." The plan emphasizes using cities as comprehensive carriers for the construction of Digital Guangdong, promoting infrastructure connectivity, data integration, platform interoperability, business integration, and a smooth ecosystem. It also aims to actively explore future-oriented smart city management models and achieve high-quality development of the Guangdong-Hong Kong-Macao Greater Bay Area smart city cluster. By the end of 2027, the province's urban digital infrastructure support capabilities will be significantly enhanced, the level of efficient digital governance will be greatly improved, digital public services will be more efficient and convenient, the momentum of digital economic development will be fully released, and the digital ecosystem will be comprehensively and collaboratively guaranteed. A number of effective and replicable typical application scenarios for digital transformation will be implemented, with Guangzhou, Shenzhen, Foshan, and Dongguan taking the lead in building efficient smart governance systems and implementing a number of advanced, usable, and independently controllable large-scale urban models. By 2030, no fewer than 10 cities will have achieved full-area digital transformation, achieving an overall leap in the digital transformation of cities throughout the province, forming a number of new smart city benchmarks and digital China city models with regional competitiveness and national influence. (Cailian Press)

07-08 20:10

Morgan Stanley reiterated its "overweight" rating on RKLB and raised its bullish price target to $293.

Odaily Odaily that Morgan Stanley reiterated its "Overweight" rating on Rocket Lab (RKLB) and maintained its target price of $105, but raised its most bullish target price from $185 to $293. Morgan Stanley analysts believe that the acquisition of Iridium will expand Rocket Lab's potential market space and reposition the company as a vertically integrated space platform. While SpaceX has significant scale and cost advantages over Rocket Lab, the latter is moving closer to the former's model.