Cohere开源Command A+:218B参数MoE大模型,主打企业级Agent与数据主权
Related
Tencent's Hunyuan 3.0 official version open source: Switching to Apache 2.0 removes overseas restrictions, halving the illusion rate.
According to Beating's monitoring, Tencent officially released the final version of its 295B parameter Hybrid Expert Model (MoE) 3.0. The most significant change is the switch to the more permissive Apache 2.0 open-source license, removing the previous regional restrictions that prohibited its use in the EU, UK, and South Korea. When the preview version was released in April, many overseas or multinational teams were deterred from using it due to strict regional restrictions and the 100 million monthly active user limit. The final version removes all compliance obstacles and has undergone significant optimization to address deployment challenges: it incorporates a fast and slow thinking mechanism and adds a 3.8B parameter multi-token prediction (MTP) layer for parallel generation, effectively reducing inference latency. After fine-grained data cleaning and training constraints, the final version reduced the illusion rate of large models from 12.5% to 5.4%, and the error rate in multi-round interactive testing from 17.4% to 7.9%. To address the persistent problem of tool calls easily going astray during agent development, Hunyuan 3.0 has also made targeted improvements. Whether in mainstream scaffolding tools like Cline or CodeBuddy, the accuracy fluctuation of cross-framework tool calls is kept below 4%, resulting in more stable output. The officially released FP8 quantized version also lowers the GPU memory barrier for local deployment and fine-tuning.
Enabling large models to "divide reading and writing tasks": NVIDIA's TwoTower architecture connects two 30B models in parallel, achieving a 2.4x speedup without loss.
According to Beating's monitoring, NVIDIA's open-source discrete text diffusion architecture, Nemotron-Labs-TwoTower, aims to solve the generation speed bottleneck of large models that "only generate one word at a time." Previous text diffusion models, in pursuit of parallel output, forced a single network to simultaneously handle unidirectional context understanding and bidirectional parallel error correction, leading to a significant decline in the model's cognitive capabilities. TwoTower employs a decoupled dual-tower design: on one hand, the pre-trained autoregressive large model is completely frozen as a "read-only context tower" to retain complete reasoning and common-sense capabilities; on the other hand, a separate "denoising tower" is trained, reading contextual information at the layer level through cross-attention. The tower uses a "confidence demasking" mechanism, prioritizing writing high-confidence words when predicting a block, and then gradually filling in the remaining blanks, achieving parallel writing from easy to difficult. On a 30B hybrid architecture (Mamba-Transformer MoE) model, this design retained 98.7% of the quality and improved actual generation speed by 2.42 times with only 1/12 of the data volume (2.1T tokens) pre-trained on the baseline model, without increasing unnecessary GPU memory caching overhead. Due to the need for dual towers to reside in memory, the model's static GPU memory usage increased, and there was still a slight accuracy degradation in extremely complex code and mathematical reasoning.
Nvidia provides financing support for GPU procurement and takes a percentage of its cloud computing revenue.
Nvidia reportedly provides financing support for GPU procurement and takes a percentage of its cloud computing revenue. (Cailian Press)
Micron (MU) is more important than Nvidia! Citrini analysts say AI model inference performance relies more on memory than GPU.
According to Odaily Odaily, in response to the "Micron vs. Nvidia" debate discussed in the community, Jukan, an analyst at Citrini Research, posted on the X platform, stating, "While Micron is not Nvidia, its future importance may surpass Nvidia's. Inference is now directly linked to revenue, but performance improvements are not solely achieved by adding Nvidia GPUs. GPUs often remain idle with low utilization due to memory bottlenecks during inference. For inference, increasing memory is more valuable. The ROI of inference ultimately depends on memory, not GPUs. Therefore, why are people still focused on acquiring Micron through NVIDIA's framework? A more comprehensive approach is needed. Inference is memory."
NVIDIA releases the BIONEMO intelligent agent toolkit, designed to help intelligent agents accelerate scientific discovery.
According to a Bloomberg report on June 23, NVIDIA released the BIONEMO Intelligent Agent Toolkit, a tool designed to help intelligent agents accelerate scientific discovery.
HIVE BUZZ HPC signs $220 million Canadian sovereign AI GPU contract with Bell and Cohere.
PANews reported on June 18 that BUZZ High Performance Computing (BUZZ HPC), a subsidiary of Canadian Bitcoin mining company HIVE, has entered into a three-year sovereign AI infrastructure partnership with Bell AI Fabric and Cohere Inc., signing a GPU cloud contract worth approximately $220 million.