美国商务部测评Kimi K3称美国领先,但报告显示测试并非完全对等
Related
Coinbase has cut AI spending by nearly half by defaulting to open weight models such as GLM and Kimi.
On June 29th, PANews reported that Coinbase CEO Brian Armstrong shared the company's experience in optimizing AI spending. He pointed out that while token usage has grown exponentially, AI spending has been reduced by nearly half through better default settings, routing, and caching strategies, rather than relying on usage caps and alert mechanisms. Regarding default settings, Coinbase is using open-source weighting models (such as Zhipu's GLM 5.2 and Moon's Dark Side's Kimi 2.7) as the default option through its LLM gateway, while encouraging engineers to choose the correct model for specific tasks. Since 91% of the company's employees never reach the usage cap, the team opted to switch to cheaper default settings rather than lowering the cap.
Coinbase has cut AI spending by nearly half and is trying to make open weight models such as GLM 5.2 and Kimi 2.7 the default option.
According to Mars Finance, on June 27th, Coinbase CEO Brian Armstrong stated that the key to maintaining stable AI spending while token usage grows exponentially lies not in setting usage friction or spending alerts, but in better default models, routing, and caching mechanisms. Coinbase is experimenting with defaulting to open-weight models such as GLM 5.2 and Kimi 2.7 through its LLM gateway, while still encouraging engineers to choose the appropriate model based on the task. He stated that 91% of employees never reach their usage limits, so instead of lowering limits and increasing alerts, the company has shifted to lower-cost default models. Regarding model routing, Coinbase preprocesses prompts in custom processes and routes tasks to the most suitable model based on cache hit rate and model pricing. For example, a cutting-edge model might be needed during the planning phase, but overuse during the execution phase could be excessive. He believes that in the future, models should not be chosen by humans; AI can automatically complete this task. Armstrong also stated that cache misses are the easiest way to drive up costs. Coinbase's requests are cache-aware to reuse hot caches as much as possible. For example, after correctly implementing caching, LibreChat's cache hit rate has improved from 5% to 60%. Furthermore, Coinbase requires engineers to keep contexts concise, including opening new sessions when switching tasks, narrowing file context scope, and disconnecting unused tools. The goal is not to suppress AI usage, but to build infrastructure that can support exponential growth. Through these practices, Coinbase has reduced its AI spending by nearly half, while token usage continues to grow.
Alibaba Cloud fully opens up its Bailian platform, with Zhipu, Minimax, Kimi, and other products among the first to be listed.
Odaily Odaily reports that at the 2026 Alibaba Cloud Summit, Alibaba Cloud announced the full opening of its Bailian platform, partnering with companies such as Moonlit Dark Side, Minimax, Zhipu, Jieyue Xingchen, Aishi Technology, and Shengshu Technology. Models including GLM-5.1, MiniMax M2.7, Kimi K2.6, Pixverse-v6-it2v, Kling-v3-omni-video-generation, Vidu Q3-Pro, Tripo-H3.1, and mimo-v2.5-pro are now available on Bailian and are also sold through the Qianwen Cloud website. (Shanghai Securities News)
Zhipu GLM-5.2's vulnerability discovery capabilities are comparable to Mythos.
Odaily Odaily reports that Zhipu, a Chinese large-scale model company, recently released the open weighted model GLM-5.2. Some researchers claim that GLM-5.2 is comparable to Anthropic's Mythos model in certain vulnerability detection and cybersecurity scenarios. Data from cybersecurity company Semgrep shows that in some benchmark tests, GLM-5.2 outperformed Anthropic's Claude Opus 4.8 model released in May. Researchers point out that with further instructions, Opus 4.8 and GLM-5.2 can rival Mythos in vulnerability discovery capabilities. As an open weighted model, GLM-5.2 can be downloaded and run on readily available hardware by anyone.
Zhipu GLM-5.2 tops DeepSWE's open-source rankings: solving 44% of complex development tasks and outperforming mainstream closed-source models.
According to Beating, Zhipu AI's open-source model GLM-5.2 has officially entered the DeepSWE long-term software engineering benchmark. In the maximum thinking power mode, it achieved a 44% success rate for complex development tasks, ranking first among open-source models. This is 13 percentage points higher than the previously listed Kimi K2.7 Code. GLM-5.2's average cost per task is $3.92, slightly higher than Kimi K2.7 Code's $2.82, but its success rate surpasses the performance of several mainstream closed-source models under specific thinking configurations, including Claude Sonnet 4.6 [high] (30%), Gemini 3.5 Flash [medium] (37%), and Claude Opus 4.8 [low] (41%). The DeepSWE benchmark, designed by Datacurve, specifically tests the ability of AI agents to solve long tasks. The test includes 113 real-world programming problems covering 5 languages. Unlike traditional tests that modify only a single piece of code, DeepSWE requires AI to collaboratively modify multiple files, fixing an average of over 600 lines of code. The evaluation runs in isolated containers with strict limits on CPU and memory resources.
Tang Jie, founder of Zhipu: After the open-source release of GLM-5.2, the performance gap with OpenAI and Anthropic will be gradually narrowed.
According to Odaily Odaily, Tang Jie, founder of Zhipu AI, stated in an article on the X platform that since its official open-source release, GLM-5.2 has achieved leading results in several authoritative international evaluations and competition rankings. In the Artificial Analysis Intelligence Index comprehensive evaluation, GLM-5.2 scored 51 points, placing it in the same level range as Anthropic's Claude Opus 4.8. In the Code Arena front-end code generation adversarial test, it ranked 2nd globally with an Elo score of 1595, and in the DesignArena design and code fusion scenario, it scored 1360 points, ranking 1st. Overall, Zhipu GLM-5.2 continues to rank among the top in the world in various real-world scenarios such as front-end development, design generation, and software engineering, gradually narrowing the performance gap with cutting-edge models such as OpenAI and Anthropic, and will continue to push the upper limit of model capabilities. Previously, in response to Musk's statement that China's large-scale models might reach Anthropic's Fable level in the first quarter of next year, Tang Jie, founder of Zhipu AI, responded that "it will not take that long."