Cryptocurrency prices are highly volatile. All content is for reference only and is not investment advice. Trading involves risk of total loss.

Full disclaimer
Back to News
SourceDecrypt

Anthropic Ships Claude Fable 5.1, More Than Doubling Its Predecessor on Key Benchmark

Fable 5.1 and its restricted sibling Mythos 5.1 arrive three months after export controls forced Anthropic to pull Fable 5 offline for 18 days
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

06-26 09:41

The benchmark scores of the Sakana Fugu and Fable 5 have been questioned, with differences in testing scaffolding potentially causing a 10-20 point discrepancy.

According to Beating, Fugu Ultra, a multi-agent collaborative system developed by the Japanese AI startup Sakana AI, claims to have outperformed Anthropic's flagship model Fable 5 in multiple benchmark tests, including scientific reasoning and programming. However, the benchmark results have been widely questioned by the community. Critics point out that comparing self-tested data under non-uniform testing environments is not objective. Benchmark scores are highly dependent on the scaffold/harness used, with different scaffolds causing score discrepancies of 10 to 20 points. This means the so-called "outperformance" is largely a product of systems engineering optimization, rather than a generational leap in the underlying model's capabilities. Independent evaluation data shows that the scaffolds used by agents built around the large model have a significant impact on the final score. For example, using the same Claude Opus 4.5 model, changing only three different open-source scaffolds resulted in a 50.2% to 55.4% fluctuation in the fix rate in the SWE-bench Pro benchmark test. Analysis by third-party testing organization Scale AI further confirms that operational strategies such as prompt word templates, attempt limits, context retention management, and tool integration can lead to a 10 to 20-point performance discrepancy in the weights of the same set of models. Since the data released by Sakana AI and Anthropic are based on their respective closed-source vendor scaffolds tuned specifically for their own systems, and were not tested in a standardized, independent third-party environment (such as Scale SEAL), the data cannot truly reflect the underlying capabilities of the two models.

07-02 05:59Important

Anthropic's main model, Claude Sonnet 5, officially launched on the API platform | _2024111120230_

ChainCatcher reports that Anthropic's latest Sonnet series masterpiece, Claude Sonnet 5, is now officially available to developers via its API platform. This model achieves groundbreaking advancements in code generation and complex agent tasks. Leveraging a massive 1MB context window and innovative adaptive thinking capabilities, Claude Sonnet 5 can handle large and complex projects more stably and efficiently. Furthermore, to meet the needs of different business scenarios, Claude Sonnet 5 is the first in the industry to launch a dual-channel access model: "Official Direct Connection" and "Self-Selected Service Provider," comprehensively helping developers flexibly balance ultimate performance and operational costs. Starting today, users can log in to the console to apply for an API key for quick access and experience. For detailed technical guides and documentation, please refer to the official website.

07-01 11:18Important

Anthropic's main model, Claude Sonnet 5, officially launched | _2024111120230_ | API Platform

Odaily Odaily reported on July 1st that Anthropic's latest Sonnet series masterpiece, Claude Sonnet 5, has been officially released to developers via its API platform. This model has achieved groundbreaking progress in code generation and complex agent tasks. Leveraging a massive 1MB context window and innovative adaptive thinking capabilities, Claude Sonnet 5 can handle large and complex projects more stably and efficiently. Meanwhile, to meet the needs of different business scenarios, [2024111120231] has taken the lead in the industry by launching a dual-channel access mode of "official direct connection" and "self-selected service provider," comprehensively helping developers flexibly balance ultimate performance and operating costs. Starting today, users can log in to the [2024111120232] console to apply for an API key and quickly access the platform.

07-01 09:04

Anthropic admitted that Claude Code had embedded steganography code targeting Chinese users, calling it an "abuse prevention experiment," and promised to roll back the code tomorrow.

According to Beating's monitoring, Thariq, an engineer on Anthropic's Claude Code team, publicly responded to the recent controversial "spy code" leak. He admitted that in March of this year, an experimental mechanism was embedded in the product. This mechanism detected whether the system timezone was Asia/Shanghai or Asia/Urumqi, whether the proxy hostname matched a list of Chinese resellers, and the keyword "AI Lab," and used special punctuation marks to inject hidden marker information into system prompts in a steganographic manner. He stated that the mechanism was intended to "prevent unauthorized resellers from abusing accounts and model distillation," but emphasized that the team has since implemented stronger protective measures and "has always intended to take it offline." The relevant PR has been merged, and it is expected to be completely rolled back in tomorrow's version release. This leak was made public on June 30 by the security account @IntCyberDigest, accompanied by two screenshots of code showing that Claude Code performed environmental fingerprinting on Chinese users without their knowledge. While Thariq's response was a direct admission, the timeline of "launching in March and only accelerating its withdrawal after being exposed" has still sparked widespread skepticism within the community. The comments section almost unanimously criticized Anthropic for "only announcing its withdrawal after being caught" and "secretly monitoring users without notifying them," severely damaging the company's long-standing image of "prioritizing security and ethics." --------------------------------- Click the original link below to join the Beating · Lark AI news channel for 24/7 monitoring of global AI hot topics and news.

07-01 04:01Important

Polymarket's probability that "Claude Fable 5 will resume service for US customers by July 1st" has risen to 97%, a 71% increase in the last 24 hours.

PPP prediction market monitoring tools show that the probability of "Claude Fable 5 being restored for US clients by July 1st" on Polymarket has risen to 97%, a 71% increase in 24 hours. Due to export controls imposed by the US government on national security grounds, Anthropic suspended global access to the Claude Fable 5 model a few days after its release on June 9th. Latest news indicates that Anthropic has been notified that the US Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. Access will be restored soon. Odaily Seer continues to monitor the prediction market, seeing changes before pricing.

07-01 00:38

Anthropic's top models, Fable5 and Mythos5, will be released tomorrow.

According to Beating's monitoring, Anthropic issued a notice stating that the U.S. Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5, and access will resume starting tomorrow. Updates will be shared soon.