返回 7*24 快讯
来源MarsBit

Anthropic's report responds to the question of self-evolution: A partial closed-loop system has been successfully implemented, but fully autonomous training is still some distance away.

According to Beating's monitoring, AI's ability to iterate autonomously is exceeding everyone's expectations. The Anthropic Institute released a report on June 5th, "When AI Builds Itself," detailing its research progress in "recursive self-improvement." Data shows that as of May 2026, over 80% of the code merged into the Anthropic main codebase was written by Claude himself. Before the release of the Claude Code in February 2025, Claude's code accounted for only a single-digit percentage. Tang Jie, founder of Zhipu AI, predicted on May 13th that the ultimate goal of large models is self-evolution, and that Claude may have already successfully completed the self-training baseline of "writing code, cleaning data, and training itself." However, the Anthropic report explicitly clarifies that completely autonomous design and development of successors through recursive self-improvement has not yet been achieved. The role of AI in the development chain is currently transitioning from localized efficiency improvements to autonomous decision-making. In the second quarter of 2026, the average daily code merged per engineer at Anthropic was eight times that of 2024. The current development process is simple: engineers are only responsible for planning goals and reviewing code, while Claude handles the actual writing and execution. Anthropic has also deployed Claude as an automated code reviewer, responsible for intercepting bugs and security vulnerabilities. This indicates that the "self-judgment" pillar pointed out by Tang Jie has been implemented at the engineering level, but human review remains the last line of defense. The reliability of the model in independently executing long-term tasks is also doubling. The duration for which the model can work autonomously continuously roughly doubles every four months. In March 2024, Claude 3 Opus could only handle simple tasks for 4 minutes. A year later, Claude 3.7 Sonnet could handle 1.5 hours. By March 2026, Claude 4.6 Opus could handle complex tasks lasting 12 hours. Data from the evaluation agency METR shows that the latest preview version of Claude Mythos can work autonomously for more than 16 hours continuously, approaching the upper limit of current evaluation tools. At the current rate, by 2027, AI will be able to autonomously handle scientific research tasks that would normally require weeks of human work, helping companies transition from "one-person companies" to "unmanned companies." As for Tang Jie's speculation about a "self-training baseline," the report reveals that it is actually a localized "miniature experimental closed loop." In the experiment of speeding up code training with small models, Claude 4 Opus in May 2025 could only improve code speed by 3 times, while the preview version of Claude Mythos in April 2026 achieved a 52-fold speedup. In comparison, top human researchers can typically achieve a 4-fold improvement in 4 to 8 hours. However, the optimization goals and success metrics of the experiment were all set in advance by humans. When faced with the more complex end-to-end chain of "cleaning data, generating synthetic data, and self-training," AI still lacks decision-making capabilities. However, the autonomous closed-loop of the R&D chain is pushing humanity to the brink of losing ultimate control over the system. Tang Jie's prediction of "LLM OS replacing traditional architectures and applications being generated on demand and in real-time" means that future computers will run dynamic code that cannot be reviewed in advance; while Anthropic's warning that "human review cannot keep up with AI's self-evolution" means that we cannot even control the source of the generated code. When AI begins to autonomously design and train its successors, software evolution will become a complete black box. Once AI is allowed to conduct unaudited self-iteration within a black-box system, the subsequent security isolation, monitoring, and behavioral alignment of the self-improving system will become extremely difficult.
免责声明:以上内容仅为作者观点,不代表 711BTC 的任何立场,不构成与 711BTC 相关的任何投资建议。

相关推荐

-139undefined前

CMC Labs 孵化面向东南亚市场的预测交易平台 Fuyo Markets

ChainCatcher 消息,据官方消息,由 CMC Labs 孵化的专为东南亚市场打造的全新预测交易平台 Fuyo Markets 已推出。 Fuyo Markets 提供一系列独具特色的市场类别,包括 1 分钟加密货币预测(业内速度最快、Fuyo Markets 独家提供的体验)、紧跟现实世界热点事件的“热门市场”(Trending Markets)、聚焦东南亚本地股票与企业的区域性市场,以及围绕熊市行情打造的“丧”系市场(Sad Markets)。 该平台支持 10 种东南亚语言,集成 Binance Connect 法币支付功能;Fuyo Markets 致力于为该地区的新一代交易者提供简单易用、门槛亲民且符合本土文化偏好的预测交易体验。

41undefined前

特朗普或迎新纪录:30年期美债融资成本恐创2001年来最高

Odaily星球日报讯 美财政部周四将发行 250 亿美元 30 年期国债,预期将创出 2001 年以来最高融资成本。在截至 9 月底的当前财年中,美国政府的国债利息支出就已经达到 1.17 万亿美元,同比增加 15%。根据美国国会的数据,在特朗普第二任期的第一年里,美国债务实际增加了 2.25 万亿美元;到今年 7 月时,这个数字进一步上升到 3.16 万亿美元。

1undefined前

富途老虎期权内幕交易案范围大幅缩小至45人控制的47个账户,共获利1.55亿美元

Odaily星球日报讯 经过一个多月的券商数据调取和逐户交易分析,作为原告的美国期权做市商海纳国际和城堡证券已将涉嫌富途老虎期权内幕交易案的范围大幅缩小至 45 人控制的 47 个账户。经过逐户比较交易利润、回报率、合约数量、到期时间、券商、所在地及建仓时间等指标,涉嫌内幕交易的整体获利增加至 1.55 亿美元。 目前 45 人的具体名单尚未公开,但绝大多数位于美国境外,其中多数居住在中国内地和中国香港。其中有一人控制了三个账户。个别人获利高达数千万美元,获利最少的也有数十万美元。(财新网)

5undefined前重要

日本股市散户加速使用杠杆参与AI行情,融资交易金额创2016年以来最高水平

BlockBeats 消息,8 月 13 日,日本股市散户正加速使用杠杆参与 AI 行情,截至 7 月日本股市个人投资者杠杆借贷交易金额达到 123 万亿日元(约 1.09 万亿元人民币),较年初增长一倍,创 2016 年统计以来最高水平。6 月日本个人投资者杠杆借贷交易规模已创历史新高,7 月继续维持高位。同时杠杆借贷交易占个人投资者整体交易金额的比例升至 83%,同样刷新纪录。AI 概念股成为推动日本散户杠杆借贷交易激增的主要力量,其中存储概念股铠侠控股的融资买入余额截至 8 月 7 日达到 1323 万股,成为热门标的之一。

5undefined前重要

Polymarket上“8月底前达成伊朗-阿曼霍尔木兹海峡协议”概率降至36%,24小时下跌21%

PPP 预测市场工具监测显示,Polymarket 上“8 月底前达成伊朗-阿曼霍尔木兹海峡协议”概率降至 36%,24 小时下跌 21%;9 月底前达成伊朗-阿曼霍尔木兹海峡协议概率降至 66%,24 小时下跌 14%。 该事件规则为:如果阿曼和伊朗在指定日期(美国东部时间晚上 11:59)前宣布就霍尔木兹海峡的交通问题达成外交协议,将判定为“是”,外交协议是指阿曼和伊朗之间就相关行动、政策、义务或承诺达成一致的正式协议、条约、交易或实质上类似的的外交文书,可信来源为阿曼和伊朗政府的官方信息以及可信报道共识。 伊朗最高国家安全委员会此前表示,若伊朗与阿曼就霍尔木兹海峡过境通道达成协议,该协议将与海峡封锁问题分开处理。只要美国不改变其行为、不接受伊朗的条件,霍尔木兹海峡就不会重新开放。美国必须结束战争,支付被冻结的伊朗资金,整个地区包括黎巴嫩和加沙的战争都必须结束,此外还需满足通过中间人传达给美方的其他条件。 加入 PPP 信号推送社群,快人一步,掌握先机。

7undefined前重要

宏观政策预期持续变化,Gate 机构持续升级专业交易基础设施

ChainCatcher 消息,美国 7 月 CPI 环比上涨 0.1%、同比上涨 3.4%,核心 CPI 同比上涨 2.5%,整体符合市场预期。随着市场继续评估美联储后续政策路径,宏观变化对资产配置和交易策略的影响持续增强,也进一步提升了机构对流动性管理与交易执行效率的关注。 在此背景下,Gate 机构持续完善专业交易基础设施。据平台发布的 7 月透明度报告,Gate CrossEx 新增一家主流交易所及 23 个交易对,上线 RPI Orders,多个交易所费率最高下调 50%,并新增行情、资金费率、批量撤单等 API 及多项 WebSocket 功能;通过优化并发下单与成交回报延迟,系统性能提升 50%,同时上线 Colo 服务进一步降低交易延迟。 此外,SuperLink 持续优化 Fireblocks Gas 管理与结算流程,进一步提升机构跨平台资产管理与交易协同能力。未来,Gate 机构将持续围绕交易执行、流动性与跨平台协同等核心能力推进基础设施升级,为专业投资者提供更高效、稳定的机构级交易服务。