Some large-scale AI models in China are 90% cheaper than those in the US; Chinese AI's high cost-effectiveness is capturing the US market.
Related
Meituan open-sourced its trillion-parameter large-scale model LongCat-2.0, and simultaneously released the inference code for domestically developed Chinese card processors.
According to Beating's monitoring, Meituan has officially open-sourced its trillion-parameter large-scale model, LongCat-2.0, with a total of 1.6T parameters and an average activation of approximately 48B, designed specifically for real-world agentic coding tasks. Architecturally, it innovatively introduces LongCat sparse attention and N-gram embedding. The former reduces fragmented memory access through flow-aware indexing and hierarchical indexing, accelerating training and inference with millions of contexts; the latter, while achieving nearly 97% sparsity in MoE, invests 135B parameters into the embedding layer, balancing parameter gains and structural stability. Post-training employs multi-teacher online distillation, categorizing experts into Agent, Inference, and Interaction types, seamlessly integrating them on a domestic computing power cluster through the MOPD architecture. As the industry's first trillion-parameter model to complete inference on a 50,000-card domestic computing power cluster, LongCat-2.0 validates the mature capability of domestic chips to handle complex large-scale model tasks. To address the multiple limitations of domestically produced Chinese chips in terms of memory, bandwidth, and interconnects, Meituan has made breakthroughs in three areas: model, chip adaptation, and deployment. At the model level, ScMoE leverages the core control capabilities of domestically produced chips to achieve physical core-level parallelism for Dense and MoE branches, combined with KV-cache partitioning to alleviate the pressure on ultra-long context memory. At the chip adaptation level, Super Kernel reduces operator startup overhead, and Weight Prefetch hides I/O latency, maximizing hardware utilization under constrained conditions. At the deployment level, PD separation is adopted to balance TTFT and TPOT, along with asynchronous Expert-Parallel load balancing to solve load unevenness under high EP (efficiency level). This open-source release simultaneously provides multiple precision versions, including BF16, FP8, and INT8, and fully opens up inference results optimized for domestic computing power, aiming to enable existing domestically produced cards and even older cards to smoothly deploy trillion-model inference services. --------------------------------- Click the original link below to join the Beating · Lark AI news channel and monitor global AI hot topics and news 24/7.
Sun Guoliang of Muxi Technology: Orders for some products are already booked until next year or even later; the MXC600 chip series has already been shipped in large quantities.
Mars Finance reported on July 8th that during the "Vibrant China Research Tour" in Shanghai, Sun Guoliang, Chief Product Officer and Senior Vice President of Muxi Technology, revealed that orders for some of the company's products are already booked into next year and beyond. The MXC600 chip series is the company's main product this year and is already shipping in large quantities. Regarding the contradiction in the market between "hardware shortages" and "concerns about computing power surplus," he believes that currently, general-purpose, easy-to-use, stable, and reliable computing power is scarce, even extremely difficult to obtain. The surplus computing power refers to computing power applicable to some non-general-purpose architectures, capable of running only previous-generation models or certain niche models, which constitutes a certain degree of overcapacity. (Science and Technology Innovation Board Daily)
Sources say China may restrict overseas users from accessing its most advanced AI models.
According to Foresight News , citing three sources familiar with the matter, Chinese authorities have held multiple meetings with top technology companies to discuss potentially restricting overseas users' access to China's most advanced artificial intelligence models, including those yet to be released.
Opinion: Open source models account for only 10% of enterprise large-scale model spending, but mature production environments will be dominated by open source models.
According to Beating's monitoring, while public opinion often touts that open-source large models are dominating everything, enterprise spending data presents the opposite picture. Jesse Zhang, co-founder and CEO of Decagon, an enterprise-level AI customer service platform, points out that the share of open-source models in total enterprise spending has now dropped to 11%. This decline stems from the fact that most enterprises' AI applications are still in the early, undefined exploratory stage, thus defaulting to reliance on closed-source models. However, he emphasizes that once application scenarios mature, open-source models will take over production environments with their advantages of extremely low latency and deep fine-tuning. In Decagon's own production environment, 90% of calls have already switched to open-source weighted models. The core driver of this transformation is interaction speed and customization capabilities, not cost savings. In customer service scenarios, a single conversation that takes 8 seconds to finish will completely destroy the product experience. Since leading closed-source labs do not allow fine-tuning of flagship models, and small closed-source models cannot be deeply customized, small-sized open-source models, through fine-tuning for specific tasks, have become the only option to support high-frequency real-time interactions. The future of enterprise AI will see a division of labor: leading closed-source labs will continue to dominate the exploration and discovery of new fields, while open-source weighted models will increasingly take over the actual production of mature businesses. Because model fine-tuning requires extremely high levels of data and talent, the migration from closed-source to open-source will be a slow process lasting several years, during which both will experience sustained growth.
Serenity: Funds in China's primary market are flowing into physical AI and world models, with funding for cutting-edge models concentrating on leading companies.
According to Mars Finance, on July 3rd, Serenity published an article stating that, based on AI investment in China's primary market, institutional funds are flowing towards embodied intelligence, physical AI, and world models. The article cites approximately $23.56 billion in funding for large-scale models/LLMs, $15.74 billion for AI infrastructure and technology layers, $13.36 billion for embodied intelligence/physical AI, $8.79 billion for AIGC applications, and $3.82 billion for autonomous driving and other Top-20 clusters. However, this figure is not entirely comparable to the aforementioned categories. Early-stage, purely basic model funding has essentially closed, with more funds flowing to established leading companies and world model companies. The article believes a similar trend may emerge in the US market, with funds concentrating on leading companies like Anthropic and OpenAI, and that "world models have become the biggest consensus in early-stage investment." Several months ago, Serenity believed that 4D AI/world modeling would be the most noteworthy area to watch, mentioning that AEVA might offer exposure to this sector. However, there are currently no clear pure-play targets in the market, and we may need to wait for the next batch of IPOs in this field. Serenity stated that AIGC applications are the most mature area of AI technology commercialization, but there are currently no clear winners. Funds continue to flow into AI infrastructure and the semiconductor supply chain, while a large amount of capital is rotating into physical AI, embodied intelligence, humanoid robots, and world modeling, with cutting-edge modeling continuing to concentrate on leading companies.
MiniMax plans to launch a large-scale model with 2.7 trillion parameters.
Mars Finance reported on July 8th that Rare Earth Technology plans to launch a new generation of large-scale models with 2.7 trillion parameters. (Science and Technology Innovation Board Daily)