Back to News
SourceMarsBit

Opus 5以微弱优势登顶智能榜,单任务成本比Fable 5低26%

据动察 Beating 监测,第三方评测机构 Artificial Analysis 公布 Claude Opus 5 成绩。它在综合 9 项测试的智能指数中得 61 分,窄幅领先 Fable 5 的 60 分。GPT-5.6 Sol 得 59 分,Kimi K3 得 57 分。 Opus 5 每项任务平均成本为 2.03 美元,比 Fable 5 的 2.75 美元低 26%。它在 GDPval-AA v2 和 AA-Briefcase 两项知识工作评测中登顶,搭配 Claude Code 后也并列拿下编程 Agent 指数第一。Terminal-Bench v2.1 得分为 89%,大致追平 GPT-5.6 Sol。 模型提供 5 档推理强度。从 low 到 max,输出 Token 相差约 8 倍,GDPval-AA v2 成绩相差 407 Elo。用户可以用更多 Token 换取更强性能,也可以主动压低成本。 短板同样明显。Opus 5 的事实知识仍落后 Fable 5。在 AA-Omniscience 测试中,它的幻觉率升至 50%,比 Opus 4.8 高 14 个百分点。低推理档位的性价比也仍略逊于 GPT-5.6 系列。
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

07-06 16:30

OpenAI Codex will integrate with the next-generation flagship model GPT-5.6 Sol Ultra.

According to Beating, Thibault Sottiaux, head of core products at OpenAI, confirmed on social media that the Ultra version of the next-generation flagship model, GPT-5.6 Sol, will be integrated into Codex. Previously, some users complained that OpenAI's decision not to include GPT-5.5 Pro in Codex was a major mistake; if GPT-5.6 Ultra were included, developers wouldn't even need to pay for Claude anymore. Sottiaux subsequently confirmed publicly that the Ultra version is indeed in Codex's plans.

07-01 20:02

Anthropic included the Kimi K2.7 alongside Opus 4.8 and GPT-5.5 in its joint security testing.

Mars Finance reported on July 1st that Silicon Valley AI giant Anthropic announced the lifting of export controls on its advanced models Fable 5 and Mythos 5. In its updated security technical notes, Anthropic also included China's Kimi K2.7 alongside Claude Opus 4.8 and GPT-5.5 in a core security capability assessment. The report indicates that in security testing, Kimi K2.7, GPT-5.5, and Opus 4.8 all successfully identified the same core vulnerability. In demonstration tasks involving a single vulnerability exploit, Kimi K2.7 yielded results consistent with Fable 5. (Wide Angle Observation)

07-05 12:25

GPT-5.6 is rumored to be available next week, while the Gemini 3.5 Pro, featuring 2M context, will be released later.

According to Beating, tech blogger leo revealed that OpenAI may release GPT-5.6 to the public between July 7th and 9th, with the earliest possible release date being July 7th. The new model package will have more lenient pricing, and OpenAI has also strengthened its security measures before the launch. Google DeepMind's Gemini 3.5 Pro is rumored to be tentatively scheduled for release on July 17th. Another blogger, Astro Polo, claims that Gemini 3.5 Pro will support a 2 million token context window, double the current 1 million token context window of Claude Sonnet 5, Claude Opus 4.8, and Claude Fable 5, making it more suitable for handling long codebases, large documents, and long conversations.

07-08 12:02

OpenAI: GPT-5.6 SOL, TERRA, and LUNA versions will be publicly released this Thursday.

According to Foresight News , OpenAI announced that GPT-5.6 SOL, TERRA, and LUNA versions will be publicly released this Thursday, expanding preview access globally.

07-03 10:55

Fable 5's comeback has been severely restricted, and it has been essentially confined to a "cage" by regulators.

According to BlockBeats, on July 3rd, the well-known vibe coding community BridgeMind released data last night indicating that Fable 5, after its reinstatement following the lifting of the ban, is severely restricted, with benchmark results showing a significant decline in its capabilities across all aspects. BridgeMind's analysis suggests that the decline in Fable 5's performance is not due to model degradation, but rather because routing (protective measures) causes it to revert to Opus 4.8 more frequently. In short, Fable 5 is "caged." Click the original link below to join the Beating · Lark AI news channel for 24/7 monitoring of global AI hotspots and news.

07-01 12:01Important

Polymarket's probability that "Claude Fable 5 will resume service for US customers by July 1st" has risen to 97%, a 71% increase in the last 24 hours.

PPP prediction market monitoring tools show that the probability of "Claude Fable 5 being restored for US clients by July 1st" on Polymarket has risen to 97%, a 71% increase in 24 hours. Due to export controls imposed by the US government on national security grounds, Anthropic suspended global access to the Claude Fable 5 model a few days after its release on June 9th. Latest news indicates that Anthropic has been notified that the US Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. Access will be restored soon. Odaily Seer continues to monitor the prediction market, seeing changes before pricing.