Back to News
SourceMarsBit

Cohere开源Command A+:218B参数MoE大模型,主打企业级Agent与数据主权

据动察 Beating 监测,Cohere 正式开源 2180 亿参数的稀疏混合专家模型 Command A+,采用 Apache 2.0 协议支持企业级 Agent 构建与私有化部署。新模型专为关键基础设施的数据主权需求设计,支持企业在无厂商锁定的情况下进行物理隔离部署。 新模型总参数量为 2180 亿(218B),但单次推理仅激活 250 亿(25B)参数,实现了计算效率与生成性能的平衡。优化后的推理效率使其仅需两张 NVIDIA H100 或单张 B200 GPU 即可运行,Hugging Face 也同步上线了 W4A4 等低精度量化版本,进一步降低了部署门槛。 Command A+ 原生支持多模态图文输入,提供 12.8 万(128K)Token 的输入上下文窗口以及 6.4 万 Token 的输出长度。它针对复杂的逻辑推理、自主工具调用、数据库查询等 Agentic 工作流以及长文档处理进行了深度优化,并支持包含所有欧盟官方语言在内的 48 种语言。
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

07-06 15:11

Tencent's Hunyuan 3.0 official version open source: Switching to Apache 2.0 removes overseas restrictions, halving the illusion rate.

According to Beating's monitoring, Tencent officially released the final version of its 295B parameter Hybrid Expert Model (MoE) 3.0. The most significant change is the switch to the more permissive Apache 2.0 open-source license, removing the previous regional restrictions that prohibited its use in the EU, UK, and South Korea. When the preview version was released in April, many overseas or multinational teams were deterred from using it due to strict regional restrictions and the 100 million monthly active user limit. The final version removes all compliance obstacles and has undergone significant optimization to address deployment challenges: it incorporates a fast and slow thinking mechanism and adds a 3.8B parameter multi-token prediction (MTP) layer for parallel generation, effectively reducing inference latency. After fine-grained data cleaning and training constraints, the final version reduced the illusion rate of large models from 12.5% to 5.4%, and the error rate in multi-round interactive testing from 17.4% to 7.9%. To address the persistent problem of tool calls easily going astray during agent development, Hunyuan 3.0 has also made targeted improvements. Whether in mainstream scaffolding tools like Cline or CodeBuddy, the accuracy fluctuation of cross-framework tool calls is kept below 4%, resulting in more stable output. The officially released FP8 quantized version also lowers the GPU memory barrier for local deployment and fine-tuning.

07-02 16:54

Enabling large models to "divide reading and writing tasks": NVIDIA's TwoTower architecture connects two 30B models in parallel, achieving a 2.4x speedup without loss.

According to Beating's monitoring, NVIDIA's open-source discrete text diffusion architecture, Nemotron-Labs-TwoTower, aims to solve the generation speed bottleneck of large models that "only generate one word at a time." Previous text diffusion models, in pursuit of parallel output, forced a single network to simultaneously handle unidirectional context understanding and bidirectional parallel error correction, leading to a significant decline in the model's cognitive capabilities. TwoTower employs a decoupled dual-tower design: on one hand, the pre-trained autoregressive large model is completely frozen as a "read-only context tower" to retain complete reasoning and common-sense capabilities; on the other hand, a separate "denoising tower" is trained, reading contextual information at the layer level through cross-attention. The tower uses a "confidence demasking" mechanism, prioritizing writing high-confidence words when predicting a block, and then gradually filling in the remaining blanks, achieving parallel writing from easy to difficult. On a 30B hybrid architecture (Mamba-Transformer MoE) model, this design retained 98.7% of the quality and improved actual generation speed by 2.42 times with only 1/12 of the data volume (2.1T tokens) pre-trained on the baseline model, without increasing unnecessary GPU memory caching overhead. Due to the need for dual towers to reside in memory, the model's static GPU memory usage increased, and there was still a slight accuracy degradation in extremely complex code and mathematical reasoning.

07-02 11:41

Nvidia provides financing support for GPU procurement and takes a percentage of its cloud computing revenue.

Nvidia reportedly provides financing support for GPU procurement and takes a percentage of its cloud computing revenue. (Cailian Press)

06-29 14:29

Micron (MU) is more important than Nvidia! Citrini analysts say AI model inference performance relies more on memory than GPU.

According to Odaily Odaily, in response to the "Micron vs. Nvidia" debate discussed in the community, Jukan, an analyst at Citrini Research, posted on the X platform, stating, "While Micron is not Nvidia, its future importance may surpass Nvidia's. Inference is now directly linked to revenue, but performance improvements are not solely achieved by adding Nvidia GPUs. GPUs often remain idle with low utilization due to memory bottlenecks during inference. For inference, increasing memory is more valuable. The ROI of inference ultimately depends on memory, not GPUs. Therefore, why are people still focused on acquiring Micron through NVIDIA's framework? A more comprehensive approach is needed. Inference is memory."

06-23 21:09

NVIDIA releases the BIONEMO intelligent agent toolkit, designed to help intelligent agents accelerate scientific discovery.

According to a Bloomberg report on June 23, NVIDIA released the BIONEMO Intelligent Agent Toolkit, a tool designed to help intelligent agents accelerate scientific discovery.

06-18 20:51

HIVE BUZZ HPC signs $220 million Canadian sovereign AI GPU contract with Bell and Cohere.

PANews reported on June 18 that BUZZ High Performance Computing (BUZZ HPC), a subsidiary of Canadian Bitcoin mining company HIVE, has entered into a three-year sovereign AI infrastructure partnership with Bell AI Fabric and Cohere Inc., signing a GPU cloud contract worth approximately $220 million.