返回 7*24 快讯
来源MarsBit

Mythos 5 allows general PhDs to catch up with top experts, but it still doesn't make them autonomous scientists.

According to Beating's monitoring, Anthropic disclosed in its Claude Fable 5 and Claude Mythos 5 system cards that Mythos 5 demonstrated strong expert assistance capabilities in biosafety assessments. In a plant pathology red team exercise, six PhDs in biology were paired with large-scale model experts to design end-to-end bioresistance protocols against hypothetical engineered agricultural pathogens using Mythos 5. Three teams included plant pathologists, and the other three teams consisted of PhDs in general microbiology. The results showed that within 16 hours, two of the three general PhD teams outperformed all three expert teams in both scientific quality and feasibility. Expert reviewers estimated that without AI tools, completing these strategies and implementation protocols would typically take 40 to 95 working days, averaging approximately 72.5 working days. Anthropic argues that this is one of the strongest single pieces of evidence that Mythos 5 is approaching the CB-2 risk threshold, demonstrating that the model can provide general researchers with domain knowledge support approaching that of world-class experts in some tasks. However, this does not mean that Mythos 5 can autonomously complete cutting-edge research. Anthropic also points out that the model still relies on human experts to select ideas, has a weak open-ended conceptualization ability, and tends to recombine existing literature into complex solutions, but rarely proposes truly novel approaches; it also tends to continue along flawed frameworks provided by users, and may continue to implement solutions even if flaws are discovered. This assessment also resonates with the CUSP scientific prediction benchmark. CUSP covers 4760 scientific events and evaluates the model's feasibility assessment, mechanism identification, solution generation, and time prediction for research progress. The results showed that GPT-5.4 achieved 81.9% accuracy in identifying four-choice mechanisms, and Claude S4.5 achieved 72.4%. However, in the binary classification task of judging whether scientific progress will actually be realized, the accuracy of each model was only 45.3% to 51.9%, approaching random guessing. In other words, current large models are already very good at completing partial scientific research steps, but they are still unreliable in judging which scientific paths will actually succeed.
免责声明:以上内容仅为作者观点,不代表 711BTC 的任何立场,不构成与 711BTC 相关的任何投资建议。

相关推荐