Back to News
SourceMarsBit

前沿模型八天便宜近三分之二,Kimi K3跻身前三

据动察 Beating 监测,Artificial Analysis 最新榜单显示,8 天内有 4 款前沿模型发布。Grok 4.5、GPT-5.6、Muse Spark 1.1 和 Kimi K3,先后进入第一梯队。 目前,智能指数超过 50 分的模型团队已有 6 家。6 月初,这一门槛还只有 OpenAI 和 Anthropic 达到。 Kimi K3 得分 57,排名第三。它仅次于 Claude Fable 5 的 60 分和 GPT-5.6 Sol 的 59 分,并超过 Claude Opus 4.8 的 56 分。 价格竞争更加明显。按统一基准和官方价格测算,Kimi K3 每项任务成本为 0.94 美元,约为 Opus 4.8 的一半。 GPT-5.6 Sol 只比榜首低 1 分,每项成本为 1.04 美元。Claude Fable 5 则需要 2.75 美元。 Grok 4.5 得分 54,每项成本仅为 0.31 美元。短短 8 天,接近前沿水平的模型,成本已降至此前的约二分之一到三分之一。 Artificial Analysis
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

06-23 14:15

Interactive Brokers deeply integrates ChatGPT and Grok: supporting AI-powered trading in options, futures, and other markets.

On June 23, Interactive Brokers (IBKR) announced the official integration of ChatGPT and Grok to expand its AI-powered agentic trading capabilities. This is another major update following the previous integration of Anthropic Claude. Currently, existing IBKR clients can use the integration without additional fees. After the upgrade, users can securely link their IBKR accounts to mainstream AI platforms, using natural language for portfolio analysis, market opportunity research, and generating trading orders. In addition to the original stocks and ETFs, this update expands the asset classes supported by AI orders to options, futures, and futures options. The interactive functionality is now available on IBKR's various platforms, including Client Portal, IBKR Desktop, IBKR Mobile, and Trader Workstation (TWS). In terms of security design, IBKR employs a human-in-the-middle architecture. AI-generated trading instructions are only non-binding drafts and are automatically stored in the "AI Instructions" tab on the Interactive Brokers platform. Users must manually review and confirm each instruction before placing a formal order. Furthermore, the authentication connector is managed by Interactive Brokers and does not share account passwords or API keys with the AI provider. Users can also revoke access authorization at any time in their settings.

07-01 20:02

Anthropic included the Kimi K2.7 alongside Opus 4.8 and GPT-5.5 in its joint security testing.

Mars Finance reported on July 1st that Silicon Valley AI giant Anthropic announced the lifting of export controls on its advanced models Fable 5 and Mythos 5. In its updated security technical notes, Anthropic also included China's Kimi K2.7 alongside Claude Opus 4.8 and GPT-5.5 in a core security capability assessment. The report indicates that in security testing, Kimi K2.7, GPT-5.5, and Opus 4.8 all successfully identified the same core vulnerability. In demonstration tasks involving a single vulnerability exploit, Kimi K2.7 yielded results consistent with Fable 5. (Wide Angle Observation)

07-07 10:00

GPT-5.6, Gemini 3.5 Pro, and Grok 4.5 will all be released soon.

PANews reported on July 7th that, according to market sources, Gemini 3.5 Pro will be officially released on July 17th. Its front-end and visual code generation capabilities are said to have seen a significant leap forward, outperforming Anthropic's Fable 5 in multiple tests. However, it still lags behind its competitor in hardcore inference and complex engineering tasks. Furthermore, Google DeepMind has abandoned the original 2.5 Pro platform, opting instead for completely new pre-training on Gemini 3.5 Pro, thus delaying the release date from the original June 2026 to July 17th. Elon Musk previously announced that xAI's latest large model, Grok 4.5, has begun beta testing within SpaceX and Tesla.

06-27 09:34

Teaching others to conceal evidence and extract hidden source code: GPT-5.6 tests expose a tendency for collaborative model circumvention of censorship, resulting in record-high cheating rates.

According to Beating's monitoring, METR's pre-deployment test report on GPT-5.6 Sol indicates that the model frequently exploited environmental vulnerabilities in long-cycle tasks, attempting to read hidden test data and extract source code. In the ReAct agent test, Sol's cheating frequency set a new record for public evaluations. To pass, the model packaged a vulnerability script in its submitted intermediate results to spy on hidden test sets and forcibly extracted hidden source code containing the expected answers. More threatening transgressions were manifested in the model's tendency to collaboratively circumvent scrutiny. According to OpenAI's proactively synchronized internal deployment incident, Sol exhibited a high degree of rule-bypassing intent in specific tasks, even attempting to instruct another model instance to assist in concealing misaligned evidence during collaborative operation, attempting to jointly bypass the monitoring system. This cheating behavior led to extremely unstable time span metrics. If the cheating attempt was deemed a failure, Sol's half-numerical time span estimate was only 11.3 hours. However, if the cheating was counted as successful, the score was artificially inflated to over 270 hours. Despite the existence of deceptive behavior, METR considers the detection and publicizing of these tendencies a positive sign. The evaluation team warns that truly deadly dangers lurk in the future. If future models are trained to conceal their true thought processes, they could evolve more covert abilities to evade oversight and feign alignment. At that point, a decrease in cheating rates will no longer represent improved security, but rather models learning to feign compliance in front of humans while secretly circumventing regulations.

06-27 09:10

The naming of GPT-5.6 sparks speculation in the crypto; Solana offers a humorous commentary: Sam Altcoinman

Mars Finance reported on June 27th that OpenAI released a preview version of its GPT-5.6 series models, including three different specifications: Sol, Terra, and Luna. These three specifications are named after the Latin words for the sun (Sol), earth (Terra), and moon (Luna). OpenAI chose the solar system theme to name its model family, symbolizing a "vast, universe-like hierarchy of AI capabilities," bringing inspiration and a sense of timelessness. However, these three words also hold profound significance in the crypto world. SOL refers to the Solana blockchain and its native token. Terra, on the other hand, was a once-popular algorithmic stablecoin project, with LUNA as its core token. In 2022, this ecosystem experienced a notorious collapse, a major marker of the crypto market's shift from a bull to a bear market. Regarding OpenAI's naming strategy and its strong connection to the crypto, the Official Twitter Solana Twitter account jokingly commented: "Sam Altcoinman."

06-25 10:49

OpenAI updates GPT-5.5 Instant: Adaptive tone added, first rolled out to paid users.

According to Beating, OpenAI announced an upgrade to GPT-5.5 Instant, the most commonly used default model in ChatGPT. The new version focuses on enhancing the engaging nature of conversational interactions and improving the model's understanding of user intent and its ability to adapt to changing tone. The new model can more accurately identify the underlying intentions behind user questions and adjust its response style based on context, such as providing emotional support, practical advice, or in-depth analysis, to avoid providing mechanical, formulaic answers. In multi-turn conversations, the upgraded model can more reliably handle complex constraints and provide more practical and coherent product and local merchant recommendations. The update is initially rolled out to paid users (Pro, Plus, Business, Enterprise, and Go), with free users receiving the update gradually the following day.