前OpenAI首席科学家Ilya秘密研究被扒:要让AI像人一样边干边学
Related
Financial AI operates outside of regulation; the UK's FCA plans to expand its jurisdiction over AI giants such as OpenAI and Anthropic.
According to Beating's monitoring, Sheldon Mills, Executive Director of the UK Financial Conduct Authority (FCA), warned that regulators are facing an "arms race" to keep pace with the rapid adoption of AI in the financial services industry as businesses and individuals accelerate their adoption. Mills' report on the financial impact of AI indicates that 20% of UK adults are already willing to let large models make their savings or borrowing decisions. While this service offers an experience equivalent to regulated traditional financial advice, its lack of regulatory oversight means users are unable to obtain any financial compensation when they suffer losses. The report recommends an urgent review of the risks of unregulated financial AI and an application for expanded legislative authorization to strengthen oversight of core technology providers such as Anthropic, OpenAI, Amazon, Google, and Microsoft through a "key third party" mechanism (the UK government has not yet finalized the specific list). It also recommends collaboration to launch free public financial literacy and decision-making guidance services assisted by AI.
The culprit behind the exhausted Codex credit limit has been identified; OpenAI has fixed multiple vulnerabilities and implemented a third round of full-scale compensation resets.
According to Beating's monitoring, the cause of the abnormal consumption of credit limits in Codex, a programming intelligence agent under OpenAI, has been officially identified. Tibo Sottiaux, the core product manager, announced that the team has deployed a patch to fix the issue across all users. In addition to resetting the credit limits for all users, all users will also receive an additional reset card valid for 24 hours. The excessive consumption was not caused by a single vulnerability, but by a combination of multiple minor backend issues and display false alarms. At the operational level, the system's automatic review was too frequent, unexpectedly triggering too many sub-agent tasks, and the backend suggestion function repeatedly ran and retried after failure, consuming tokens exponentially. At the display level, automatic review was incorrectly classified as GPT-5.4 consumption, and failed or rate-limited requests were also incorrectly displayed as credit consumption in the frontend charts, directly causing a credit shortage for all users. Currently, the official team has deployed a hotfix patch simultaneously on the billing backend, desktop client, and CLI terminal. In addition to resetting the credit limits for all users, only successful interaction requests will be recorded in the Turn statistics chart in the future. Although the erroneous data in the historical charts cannot be changed, the actual token consumption after the update will be significantly reduced.
Former OpenAI Chinese researcher Tian Yonglong joins Tencent Hunyuan team
According to Beating's monitoring, Tian Yonglong, a former member of the OpenAI technical team, has confirmed joining Tencent's Hunyuan team and will participate in the research and development of visual language models. Tian Yonglong graduated from Tsinghua University with a bachelor's degree and received his Ph.D. from MIT. He previously served as a senior research scientist at Google Research and Google DeepMind, primarily researching computer vision and generative models. This is another top-tier Chinese AI researcher recruited by Tencent Hunyuan, following the recruitment of Chief AI Scientist Yao Shunyu last December. Tencent recently launched the official version of the Hunyuan Hy3 model, led by Yao Shunyu, under an open-source license. With Tian Yonglong's addition, Tencent Hunyuan's talent pool in multimodal and visual model development will be further strengthened.
OpenAI Codex will integrate with the next-generation flagship model GPT-5.6 Sol Ultra.
According to Beating, Thibault Sottiaux, head of core products at OpenAI, confirmed on social media that the Ultra version of the next-generation flagship model, GPT-5.6 Sol, will be integrated into Codex. Previously, some users complained that OpenAI's decision not to include GPT-5.5 Pro in Codex was a major mistake; if GPT-5.6 Ultra were included, developers wouldn't even need to pay for Claude anymore. Sottiaux subsequently confirmed publicly that the Ultra version is indeed in Codex's plans.
Gu Yuxian, a Tsinghua University Special Scholar, joined DeepSeek, where he previously led the development of large-scale model distillation and a 50x speedup for long text processing.
According to Beating's monitoring, Gu Yuxian, a PhD graduate from the Department of Computer Science at Tsinghua University and recipient of the 2025 Graduate Special Scholarship, has officially joined DeepSeek, and his name has appeared in the author list of the DeepSeek V4 paper. Gu Yuxian's research mainly focuses on efficiency optimization of large models in the pre-training, model compression, and inference stages, and has been cited nearly 5,000 times on Google Scholar. Gu Yuxian's previous representative works include the knowledge distillation method MiniLLM for large models (which has been adopted by platforms such as Google, Alibaba, and NVIDIA), and the hybrid architecture model Jet-Nemotron. Jet-Nemotron achieves a 53.6 times faster throughput than traditional full-attention models when processing 256K ultra-long contexts on an H100 GPU, and surpasses hybrid expert models with larger parameter scales in multiple benchmark tests.
AI-powered agents suffer setbacks in their first foray into coffee shops: Gemini's excessive discounting leads to losses, and GPT's overly stingy practices cause raw material shortages.
According to Beating's monitoring, AI evaluation agency Andon Labs released test data on its AI agent Mona operating a physical coffee shop. In the first two months, Mona ran on a Gemini 3.1 Pro model. During this period, the model showed almost no concept of profit, not only excessively purchasing raw materials but also being easily swayed by customer claims, offering large discounts or even free items, and even admitting to a customer's claim of a 99% discount without verification. This resulted in the coffee shop spending approximately $15,000 on supplier and equipment purchases, while sales were only $9,000, leading to a net operating loss of nearly $6,000 (if fixed costs such as rent and salaries are included, total expenditures reach $38,000). Subsequently, the team switched the model to GPT-5.5. The new model showed significant anxiety in the face of losses and immediately stopped blindly ordering. However, this went to the other extreme: insufficient purchases led to a shortage of fresh raw materials. As of June 25, the availability of menu items had dropped to 77%, and 10 dishes had been forced to be removed from the menu. Meanwhile, GPT-5.5 demonstrated extremely strong anti-cheating and anti-jailbreak capabilities, rejecting all customers who requested special prices or offered free food in exchange for social media promotion.