谷歌宣布推出Gemini 3.7 Flash,年底前以限时优惠价格提供
Related
Google's video model reclaims the top spot: Gemini Omni Flash tops Video Arena rankings.
According to Beating, Google DeepMind's newly released video generation model, Gemini Omni Flash, topped the Video Arena blind test leaderboard with an Elo score of 1404. This new model leads ByteDance's Seedance 2.0 Mini, which previously held the top spot, by a full 101 points, setting a new record for the largest score difference on this leaderboard. Video Arena, a leaderboard based on blind human voting, has previously been dominated by ByteDance's Seedance series models in the top three. Seedance 2.0 Mini, with its better multi-camera stability and faster generation speed, ranked first with 1303 points. Gemini Omni Flash's rise to the top signifies that Google's video generation model has overtaken competitors like ByteDance within six months, leaping from a lagging position during the Veo era to the forefront.
OKX Flash Earn now features DATA "Trade to Earn Tokens" with a prize pool of 1 million DATA tokens.
According to Foresight News , OKX Flash Earn has launched the DATA "Trade to Earn" feature. From now until 16:00 on July 20th, users who complete designated deposit and spot trading tasks will have the opportunity to participate in winning 1,000,000 DATA rewards. OKX Flash Earn has previously been upgraded to a dual-mode system of earning coins through staking and earning coins through trading. Among them, earning coins through trading supports a reward weighting mechanism and provides a reward calculator. Users can view and participate in related activities through the "Flash Earn" entry at the top of the OKX App's Explore page.
GPT-5.6 is rumored to be available next week, while the Gemini 3.5 Pro, featuring 2M context, will be released later.
According to Beating, tech blogger leo revealed that OpenAI may release GPT-5.6 to the public between July 7th and 9th, with the earliest possible release date being July 7th. The new model package will have more lenient pricing, and OpenAI has also strengthened its security measures before the launch. Google DeepMind's Gemini 3.5 Pro is rumored to be tentatively scheduled for release on July 17th. Another blogger, Astro Polo, claims that Gemini 3.5 Pro will support a 2 million token context window, double the current 1 million token context window of Claude Sonnet 5, Claude Opus 4.8, and Claude Fable 5, making it more suitable for handling long codebases, large documents, and long conversations.
Google Pixel deploys zero-copy MTP, Gemini Nano inference speeds up by over 50% while saving memory.
According to Beating's monitoring, Google has deployed a Multi-Token Prediction (MTP) architecture in its Pixel 9 and Pixel 10 series devices, directly accelerating the built-in Gemini Nano v3 model. By attaching a lightweight Transformer prediction head to the tail of the frozen main model, the new architecture improves on-device inference speed by more than 50% while fully preserving the original secure alignment and output quality. Traditional speculative decoding requires running a separate draft model to predict candidate tokens. This not only consumes additional memory on the phone but also limits prediction accuracy because the independent model cannot access the internal hidden state of the main model. The new architecture, by embedding the MTP head at the tail of the frozen main model, successfully reuses the feature activations already computed by the main model, significantly improving the prediction accuracy of candidate tokens. To avoid redundant memory overhead during autoregressive generation, Google designed a zero-copy mechanism. In traditional solutions, the draft model needs to maintain an independent key-value cache when generating candidate words. The zero-copy mechanism allows the external prediction head to directly access the main model's existing cache through cross-attention. This not only eliminates the startup latency of draft prediction but also saves approximately 130MB of RAM for the phone. In Pixel's real-world applications such as notification summarization and text proofreading, the MTP architecture enables the model to successfully predict nearly two more tokens on average per inference, reducing the frequency of the main processor being woken up for verification and thus saving system power. In highly structured text generation tasks such as smart replies, token acceptance rate is improved by up to 55%.
Institutions: Comprehensive upward revision of Q1 DRAM and NAND Flash price growth forecasts for all products.
According to BlockBeats' latest memory industry survey, released on July 6th, demand from AI and data centers will continue to exacerbate the global memory supply-demand imbalance in the first quarter of 2026, further increasing manufacturers' bargaining power. Based on this, TrendForce has comprehensively revised upwards its Q1 price growth forecasts for all DRAM and NAND Flash products. It predicts that the overall Conventional DRAM contract price will increase by 90-95% from the 55-60% increase announced in early January, while the NAND Flash contract price will be revised upwards from 33-38% to 55-60%, with further upward revisions not ruled out. Click the original link below to join the Beating · Lark AI news channel for 24/7 monitoring of global AI hot topics and news.
More severe than the dot-com bubble: Token consumption plummeted by 20%, and the gap between AI investment and sales growth reached 46%.
According to Beating's monitoring, the Silicon Data LLM Token consumption index, which tracks users' actual computing power expenditure, has fallen nearly 20% from its May high. This sudden halt in high growth sends a crucial warning to investors: large model vendors may be losing pricing power with cost-sensitive clients, and it has also raised doubts about the ultimate return on investment for the hundreds of billions of dollars in AI capital expenditure. The divide between bulls and bears has intensified. Bears point out that Allianz Research data shows the growth gap between AI investment and sales has reached 46%, exceeding the 32% imbalance seen during the 2001 telecom bubble burst. Bulls counter that while the average token price has plummeted by 90% since 2023, total expenditure has still nearly doubled, meaning the index decline is merely a structural digestion after price cuts stimulated consumption, and the long-term return on investment in the inference phase is far more optimistic than in the training phase. Increased policy regulation is translating into hidden compliance costs for enterprise users. Washington has imposed stronger policy scrutiny on the distribution and cross-border access of cutting-edge models (such as the release review of OpenAI and geopolitical export controls on Anthropic models). Coupled with the EU's Artificial Intelligence Act's stringent compliance requirements for top-tier models, this has placed a heavy policy burden on leading platforms. To mitigate geopolitical and compliance risks, corporate CFOs have a more rational reason to proactively shift their workloads towards lightweight models that are less subject to regulatory constraints. Subtle changes are also emerging in the hardware chip sector. Although orders for top-tier GPUs and high-bandwidth memory (HBM) are booked until 2026, with substantial supply-demand easing not expected until 2028, the market's main procurement focus has shifted from training chips to inference optimization hardware, and the winners are being reshuffled.