Cryptocurrency prices are highly volatile. All content is for reference only and is not investment advice. Trading involves risk of total loss.

Full disclaimer
Back to News
SourceDecrypt

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
Disclaimer: The views above are the author's only and do not represent 711BTC. Nothing here constitutes investment advice.

Related

07-02 16:51

David Sacks strongly supports Palantir CEO's criticism of AI labs: True enterprise AI security lies in controlling one's own data, models, and computing power.

According to Beating, David Sacks, co-chair of the U.S. President's Council of Advisors on Science and Technology, published an article supporting Palantir CEO Alex Karp's sharp criticism of cutting-edge AI labs, stating that the mainstream media's portrayal of his interview as a "disastrous outburst" precisely demonstrates that Karp hit the nail on the head. Sacks points out that true "AI security" in a corporate environment is not abstract "alignment research" or government-led certification systems, but rather the ability to control one's own data, model weights, and computing power—preventing cutting-edge labs from "absorbing" a company's proprietary knowledge and turning it into the next product. He quotes Karp as saying, "They want to own their means of production, not hand them over to others." Sacks cites the conflict between Figma and Anthropic as a prime example: three days before the release of Claude Design, Anthropic's Chief Product Officer was still a member of Figma's board of directors, and Figma's founder stated that Anthropic "hadn't always been honest with them"; subsequently, Figma's stock price plummeted while Anthropic's valuation soared. He further listed products such as Claude Science, Claude Security, Claude Legal, and Claude Code, pointing out that Anthropic consistently targets vertical sectors originally served by companies that relied on its models, following a consistent pattern: "Observe where value is created first, then jump in." Sacks believes that the perception of open-source models as "dangerous" is not true for companies—retaining choice at the model layer and deciding who can use their core strengths is the real bottom line for corporate security. Previously, Palantir partnered with NVIDIA to deploy Nemotron's open AI models in sovereign environments, serving the US government and critical infrastructure customers, helping organizations train and deploy AI locally while maintaining complete control over data and intellectual property. Palanitir CEO Alex Karp recently gave a scathing interview on CNBC's "Squawk Box," criticizing leading AI model companies as "completely wrong" in their approach to selling AI. Karp emphasized that companies are currently dissatisfied with "cutting-edge labs" like OpenAI and Anthropic, believing they only pursue token maximization, wasting companies' time and money while handing over proprietary value and intellectual property. Karp stated that companies are "angry" and will strive to own their own AI production resources rather than relying on third parties. On June 29th, Palantir partnered with Nvidia to deploy Nvidia Nemotron open AI models in a sovereign environment, primarily serving the US government and critical infrastructure.

09-07 21:16

OpenAI Chief Scientist Warns AI Labs May Need to Slow Down

Jakub Pachocki is calling for mandatory safety standards as OpenAI finds it harder to monitor advanced AI models’ reasoning.

08-27 09:46

Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Finds

Coordinators pressed agents with little budget left into experiments they called "permadeath," METR's investigation found.

07-07 11:04

ING: Nvidia's profit margins are threatened by customers developing their own chips.

According to Mars Finance, on July 7th, Jan Frederik Slijkerman of ING wrote in a report that Nvidia's ability to maintain profit margins is uncertain as tech giants develop their own chips. He pointed out that major customers such as Microsoft, Alphabet, and Amazon are developing their own custom chips to help control AI infrastructure costs (capital expenditure efficiency). He stated that, therefore, Nvidia's pricing power may face more intense competition than in recent years, making it more difficult for it to maintain its currently extremely high profit margins in the long term, despite the company's expansion into new business lines.

07-06 11:39

Myanmar's AI-driven telecom fraud industry exposed: Starlink becomes key infrastructure, encrypted payments and OpenAI/Google models incorporated into toolchains.

A leaked investigative report from a Myanmar scam Odaily park reveals that global telecom fraud is rapidly evolving towards an "AI industrialization + cross-border encrypted payment" system. These fraud networks use cryptocurrencies to transfer funds and employ automated tools based on large models for multilingual script generation, identity spoofing, and emotional manipulation. The investigation shows that these systems heavily utilize OpenAI's ChatGPT and Google's Gemini to support "large-scale social media fraud," while funds are rapidly laundered and transferred through on-chain payments and cross-border channels, forming a two-tiered structure of "AI customer acquisition + encrypted settlement," enabling the fraud industry to achieve high automation and transnational expansion capabilities. Furthermore, Elon Musk's Starlink has become the leading network service provider in the Myanmar scam industrial park, with US ISPs handling nearly one-fifth of the park's traffic. In response to the allegations, OpenAI stated that fraudsters using ChatGPT behave in a manner highly similar to ordinary users, making identification difficult. However, they have been using behavioral pattern recognition and risk control systems to ban approximately 100,000 suspicious accounts monthly. Google stated that its AI models have security safeguards in place and emphasized its commitment to "responsible AI development" to limit the use of tools for fraudulent and other illegal purposes. (Red Star News)

07-05 13:27

Myanmar's telecom fraud AI industrialization exposed: Starlink becomes key infrastructure, encrypted payments and OpenAI/Google models are incorporated into the toolchain.

According to a report by Red Star News, citing Mars Finance, an investigative report leaked from a scam industrial park in Myanmar reveals that global telecom fraud is rapidly evolving towards an "AI industrialization + cross-border encrypted payment" system. These fraud networks use cryptocurrencies to transfer funds and employ automated tools based on large models for multilingual script generation, identity spoofing, and emotional manipulation. The investigation shows that these systems heavily utilize OpenAI's ChatGPT and Google's Gemini to support "large-scale social fraud," while funds are rapidly laundered and transferred through on-chain payments and cross-border channels, forming a two-tiered structure of "AI customer acquisition + encrypted settlement," enabling the fraud industry to achieve a high degree of automation and transnational expansion. Furthermore, Elon Musk's Starlink has become the leading network service provider in the Myanmar scam industrial park, with US ISPs handling nearly one-fifth of the park's traffic. In response to the allegations, OpenAI stated that the behavior of fraudsters using ChatGPT is highly similar to that of ordinary users, making identification difficult, but they have already blocked approximately 100,000 suspicious accounts monthly through behavioral pattern recognition and risk control systems. Google stated that its AI models have safety barriers in place and emphasized its commitment to "responsible AI development" to limit the tools from being used for illegal purposes such as fraud.