返回 7*24 快讯
Mythos 5 allows general PhDs to catch up with top experts, but it still doesn't make them autonomous scientists.
According to Beating's monitoring, Anthropic disclosed in its Claude Fable 5 and Claude Mythos 5 system cards that Mythos 5 demonstrated strong expert assistance capabilities in biosafety assessments. In a plant pathology red team exercise, six PhDs in biology were paired with large-scale model experts to design end-to-end bioresistance protocols against hypothetical engineered agricultural pathogens using Mythos 5. Three teams included plant pathologists, and the other three teams consisted of PhDs in general microbiology. The results showed that within 16 hours, two of the three general PhD teams outperformed all three expert teams in both scientific quality and feasibility. Expert reviewers estimated that without AI tools, completing these strategies and implementation protocols would typically take 40 to 95 working days, averaging approximately 72.5 working days. Anthropic argues that this is one of the strongest single pieces of evidence that Mythos 5 is approaching the CB-2 risk threshold, demonstrating that the model can provide general researchers with domain knowledge support approaching that of world-class experts in some tasks. However, this does not mean that Mythos 5 can autonomously complete cutting-edge research. Anthropic also points out that the model still relies on human experts to select ideas, has a weak open-ended conceptualization ability, and tends to recombine existing literature into complex solutions, but rarely proposes truly novel approaches; it also tends to continue along flawed frameworks provided by users, and may continue to implement solutions even if flaws are discovered. This assessment also resonates with the CUSP scientific prediction benchmark. CUSP covers 4760 scientific events and evaluates the model's feasibility assessment, mechanism identification, solution generation, and time prediction for research progress. The results showed that GPT-5.4 achieved 81.9% accuracy in identifying four-choice mechanisms, and Claude S4.5 achieved 72.4%. However, in the binary classification task of judging whether scientific progress will actually be realized, the accuracy of each model was only 45.3% to 51.9%, approaching random guessing. In other words, current large models are already very good at completing partial scientific research steps, but they are still unreliable in judging which scientific paths will actually succeed.
免责声明:以上内容仅为作者观点,不代表 711BTC 的任何立场,不构成与 711BTC 相关的任何投资建议。
相关推荐
标普500指数期货上涨 0.1%,纳指期货下跌 0.04%
ChainCatcher 消息,据 Gate 行情数据显示,标普500指数期货上涨 0.1%,纳指期货下跌 0.04%,道指上涨 0.13%。
香港恒生指数收盘下跌 43.66 点,科技指数小幅上涨
ChainCatcher 消息,据 Gate 行情数据显示,香港恒生指数 8 月 13 日收盘下跌 43.66 点,跌幅 0.17%,报 25,396.51 点;恒生科技指数上涨 15.95 点,涨幅 0.33%,报 4,792.39 点;国企指数下跌 19.78 点,跌幅 0.23%,报 8,426.49 点;红筹指数下跌 66.9 点,跌幅 1.61%,报 4,089.2 点。
现货黄金日内跌超1%,现报4364.24美元
火星财经消息,8 月 13 日,据 行情数据显示现货黄金日内跌超 1 %,现报 4364.24 美元/盎司。
AI算力租赁商CoreWeave、Nebius美股盘前双双回落
火星财经消息 8月13日,AI算力租赁商CoreWeave、Nebius美股盘前双双回落,分别跌超2%和4%,上一交易日分别收涨19%和34%。(科股宝播报)
港股收盘,恒指收跌0.17%,科指收涨0.33%
BlockBeats 消息,8 月 13 日,港股收盘,恒指收跌 0.17%,科指收涨 0.33%,腾讯控股 (00700.HK) 绩后收跌 4.46%。联想集团涨 20.179%,智谱涨 9.023%,MINIMAX-W 涨 5.988%。
思科美股盘前跌超 5%
火星财经消息,据 Gate 行情数据显示,思科 (CSCO.O) 美股盘前跌超 5%。