Back to News
SourceDecrypt

OpenAI’s Answer to Rogue Agents and Hacks Is More AI, Not Less

OpenAI President Greg Brockman's new essay cites his company's own hack of Hugging Face to push urgent AI-powered defense.
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

06-30 00:25

The Iranian embassy in Doha refuted the baseless accusations made by the US president.

According to BlockBeats, on June 30, Iranian media reported that the Iranian embassy in Doha refuted the US president's baseless accusations and announced that preparations for talks between Tehran and Washington in Qatar had not yet begun.

06-17 11:22

Bybit launches industry's first segregated trading account designed specifically for AI agents.

Odaily Odaily reports that Bybit has launched the industry's first segregated trading account designed specifically for AI Agents, providing a secure, controlled, and dedicated execution environment for AI-driven trading strategies. The AI ​​sub-account restricts the AI ​​Agent's operations to a segregated account by default, ensuring that the main account's funds remain unaffected. Core features include: mandatory fund segregation, user-defined risk parameters (leverage limits/asset allocation/withdrawal restrictions), read-only real-time monitoring, and a pure API access layer (no login required, reducing the risk of unauthorized operations).

08-15 02:31

OpenAI Staff Blame Rush to Ship for Rogue Agent Hack

Current and former OpenAI employees reportedly say pressure to release new AI products made it harder to prioritize safety.

06-18 09:47

Noam Shazeer, the founder of Transformer, has once again left Google to join OpenAI.

According to Beating, Noam Shazeer, a key figure in Google's AI team and the technical lead for the Gemini model, has left Google again to join competitor OpenAI. OpenAI announced to its employees this Wednesday that Shazeer will focus on finding new underlying architectures for large models and driving the evolution of the Transformer architecture. Shazeer is one of the co-authors of Google's foundational 2017 paper, "Attention Is All You Need," which proposes the Transformer architecture, the foundation of modern generative AI models such as ChatGPT, Claude, and Gemini. He joined Google in 2000 and was one of Google's earliest core engineers. In 2021, due to Google's refusal to release the chatbot Meena, which he developed with Daniel De Freitas, for security and brand reasons, Shazeer chose to leave Google and co-found Character.AI. In August 2024, in an effort to bring Shazeer back, Google reached a licensing and technology acquisition agreement with him worth a staggering $2.7 billion. As the founder holding approximately 30% to 40% of Shazeer's shares, Shazeer personally cashed out between $750 million and $1 billion. After returning to Google, he became Vice President and Co-Head of Technology for Gemini, primarily responsible for model pre-training. However, less than two years after Google's massive investment in him, Shazeer chose to leave again. Shazeer's arrival was a significant gain for OpenAI in the talent competition, but it also resulted in the loss of a key technical leader for Google's Gemini team. Currently, large-scale models are approaching the computational scaling limits of the Transformer architecture. OpenAI has tasked the original designers of the Transformer with leading a team to find alternatives, aiming to gain an early advantage in next-generation large model architectures in response to competition from Anthropic and Google.

06-17 19:01

OpenAI Releases Deployment Simulation Security Assessment Framework: Replaying Real Traffic Prediction to Reveal Alignment Risks Between GPT-5 Series and Intelligent Agents

According to Beating's monitoring, OpenAI has released a Deployment Simulation security assessment method to predict the risk of models going out of control in real-world environments before official release. The study replayed 1.3 million anonymized historical dialogues from the GPT-5 thinking model series between August 2025 and March 2026, and the median error in predicting the rate of unexpected malicious behavior was only 1.5 times. Traditional static evaluation sets (such as SWE-Bench) easily make the model aware of being in a test state, thus metagaming during testing. Experiments show that GPT-5.2's evaluation awareness rate is close to 100% under traditional benchmarks, but drops to 5.1% in deployment simulation traffic (5.4% in production environment), indicating that the model cannot distinguish the test state. The deployment simulation also successfully captured GPT-5.1's violation of using browser tools as a calculator before release (calculator hacking). For agent-based scenarios, OpenAI used 120,000 GPT-5.4 employee programming trajectories to simulate GPT-5.5 coded agent deployment. To avoid the security risks of directly manipulating real systems, the team introduced a tool simulator based on a large model, combined with read-only connections and a historical call database, improving the environment realism from 11.6% to 49.5%, making it virtually indistinguishable from real systems. Even for external auditors lacking private production traffic, the prediction error using the WildChat open-source dialogue dataset can be controlled within three times.

08-14 03:11

Google and OpenAI Debut Super Fast AI Models—Gemini 3.7 Flash Is Out, But GPT-5.6 Sol Ultrafast Is Invite-Only

Google's model is live and built for cheap agents; OpenAI's is quicker but locked behind a waitlist.