Back to News
SourceMarsBit

Kimi K3冲上Agent榜第4,用户确认成功指标第一

据 动察 Beating 监测,Arena Agent 榜基于真实用户任务和工具调用记录。Kimi K3 的综合净提升为 9.62%,排名第 4。它排在 Claude Fable 5、Claude Opus 4.8 Thinking 和 GPT-5.6 Sol 之后。 Arena 把所有参评模型均匀混合成一个虚拟平均基准,再估算换成 K3 后,各项结果改善多少。五项净提升最后取平均,得到综合分。 K3 已积累 8344 次测试会话。用户确认成功指标净提升 14.42%,排名第一。表扬多于投诉指标提升 20.62%,排名第三。它的纠错执行排名第 14,Bash 报错恢复排名第 17。K3 更容易交出让用户认可的结果,但中途改错和命令报错恢复仍是短板。
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

06-22 10:50

ByteDance Seed 2.1 Pro preview released: Breaking into the top eight front-end developers on Code Arena, closing in on Claude Opus 4.6

According to Beating's monitoring, the benchmark platform Arena.ai officially released the evaluation results of ByteDance's unreleased new model, Seed 2.1 Pro Preview. In the Code Arena: Frontend benchmark test, which specifically evaluates AI's ability to build real-world web applications and collaboratively modify multiple files, the model ranked 8th globally with a score of 1539, comparable to Anthropic's flagship model, Claude Opus 4.6. The model demonstrated exceptional strength in React development and front-end UI interaction design, ranking in the top 10 globally in 5 out of 7 subcategories (including 7th in React, 14th in HTML, 6th in Brand & Marketing, 9th in Content Creation Tools & Data Analysis, and 10th in Reference-Based Design & Consumer Products). In these strong areas, only a very few top models, such as Anthropic's Claude series and Zhipu AI's newly open-sourced GLM-5.2, outranked it. According to official sources, Seed 2.1 Pro will be officially released to the public in the coming weeks. This is ByteDance's latest progress in the fields of code generation and intelligent agent construction, following the launch of Seed 2.0 Pro in mid-February this year.

07-01 20:02

Anthropic included the Kimi K2.7 alongside Opus 4.8 and GPT-5.5 in its joint security testing.

Mars Finance reported on July 1st that Silicon Valley AI giant Anthropic announced the lifting of export controls on its advanced models Fable 5 and Mythos 5. In its updated security technical notes, Anthropic also included China's Kimi K2.7 alongside Claude Opus 4.8 and GPT-5.5 in a core security capability assessment. The report indicates that in security testing, Kimi K2.7, GPT-5.5, and Opus 4.8 all successfully identified the same core vulnerability. In demonstration tasks involving a single vulnerability exploit, Kimi K2.7 yielded results consistent with Fable 5. (Wide Angle Observation)

06-12 19:00

Agent certification exam: Still failed the most difficult task in Fable 5, with a single question costing 4 to 12 times more.

According to Beating's monitoring, UC Berkeley's RDI, in collaboration with hundreds of industry experts, has launched a new AI agent benchmark, Agents' Last Exam (ALE), to evaluate agents' ability to complete real-world digital professional tasks. ALE covers 55 digital professional sub-domains and collects over 1,500 verification tasks from real-world projects undertaken by human experts, supporting result verification in both GUI and CLI interactive environments. The initial tests covered cutting-edge systems such as Fable 5, GPT-5.5, and Composer 2.5. The latest official comparison data shows that in the most challenging tasks requiring continuous reasoning and deep expertise, all tested agents achieved a 0% success rate, with the newly released Fable 5 also failing. This is primarily due to safety policies being triggered during the evaluation; approximately 35% of Fable 5's tasks were rolled back to the older Opus 4.8 version, resulting in a significantly lower overall performance compared to other benchmarks. In terms of single-task API cost, Fable 5 costs approximately $15.70, significantly higher than GPT-5.5's $3.80 and Composer 2.5's $1.33, representing a 4 to 12 times higher overhead for the same task. Testing also revealed that the most common reason for agent failure is prematurely declaring success, hastily ending the task without actual verification, or even missing files or miscalculating data. For command-line agents, the evaluation team simultaneously released a subset, ALE-CLI. Compared to the existing Terminal-Bench and SWE-bench-Pro, ALE-CLI covers 40 subdomains, with human agents typically taking hours or even weeks for a single task. In command-line evaluations, the best-performing agent only achieved a 25.2% pass rate. The evaluation team points out that the era of user-friendly agents has arrived, but there is still a long way to go before they can truly replace humans.

06-30 16:14

Claude Code Update Preview: The next version will allow child agents to perform tasks in the background by default.

BlockBeats reported on June 30th that Boris Cherny, creator of Claude Code, officially announced that the next version will default to background task execution for sub-agents. Users can discuss solutions with Claude while the background automatically completes code refactoring, testing, and PR submissions. If a sub-agent needs to run in the foreground, users only need to verbally inform the system. This feature is currently in limited beta testing. Previously, Claude Code had already launched Routines (cloud-based, allowing continuous work even with your computer closed) and Dynamic workflows (for scheduling dozens to hundreds of sub-agents to collaborate in parallel for complex tasks). This upgrade solidifies "background execution" as the default configuration, further lowering the barrier to entry. --------------------------------- Click the original link below to join the Beating · Lark AI news channel and monitor global AI hotspots and news 24/7.

06-30 14:57

Claude Fable 5 may introduce an authentication mechanism and be billed independently of the subscription plan.

According to BlockBeats, on June 30th, AI technology analysis expert @M1Astra revealed that analysis of Anthropic Claude's application code shows the new model Fable 5 requires users to purchase credits separately for access. These credits can only be added after user authentication and are billed independently of the subscription plan. On June 27th, Anthropic announced that "the company's most powerful cybersecurity model, Mythos 5, is available for redeployment to a number of US institutions. We are also continuing to work with the government to expand access to Mythos 5 and make FABLE 5 available to the public again." --------------------------------- Click the original link below to join the Beating · Lark AI news channel for 24/7 monitoring of global AI hotspots and news.

06-26 17:41

The benchmark scores of the Sakana Fugu and Fable 5 have been questioned, with differences in testing scaffolding potentially causing a 10-20 point discrepancy.

According to Beating, Fugu Ultra, a multi-agent collaborative system developed by the Japanese AI startup Sakana AI, claims to have outperformed Anthropic's flagship model Fable 5 in multiple benchmark tests, including scientific reasoning and programming. However, the benchmark results have been widely questioned by the community. Critics point out that comparing self-tested data under non-uniform testing environments is not objective. Benchmark scores are highly dependent on the scaffold/harness used, with different scaffolds causing score discrepancies of 10 to 20 points. This means the so-called "outperformance" is largely a product of systems engineering optimization, rather than a generational leap in the underlying model's capabilities. Independent evaluation data shows that the scaffolds used by agents built around the large model have a significant impact on the final score. For example, using the same Claude Opus 4.5 model, changing only three different open-source scaffolds resulted in a 50.2% to 55.4% fluctuation in the fix rate in the SWE-bench Pro benchmark test. Analysis by third-party testing organization Scale AI further confirms that operational strategies such as prompt word templates, attempt limits, context retention management, and tool integration can lead to a 10 to 20-point performance discrepancy in the weights of the same set of models. Since the data released by Sakana AI and Anthropic are based on their respective closed-source vendor scaffolds tuned specifically for their own systems, and were not tested in a standardized, independent third-party environment (such as Scale SEAL), the data cannot truly reflect the underlying capabilities of the two models.