At an OpenAI press briefing for dozens of reporters on Tuesday, right before the company announced it was slowing its frontier model training, vice president of research Mia Glaese and chief scientist Jakub Pachocki said the company was strengthening safety across three areas: monitoring, alignment, and “security and containment measures.”
They explained that the move was the result of three urgent developments — the OpenAI-Hugging Face incident; preliminary evidence that the company’s upcoming Astra model may meet the Critical cybersecurity capability threshold under its Preparedness Framework; and the rapid pace of internal progress at OpenAI.
That made me wonder why there were no OpenAI security staff on the briefing? Where was the CISO, for example? In a call with an OpenAI spokesperson, I was told that that the reason was that the briefing was about research. “It was Jakob and Mia, this was ultimately a research call that was made,” she said. “Obviously security is a very important stakeholder who feeds into decisions that are made, but this is ultimately a research call to put on a pause.”
The line between AI research and security is disappearing
But the distinction between AI research and security is becoming more and more blurred after the OpenAI-Hugging Face incident. After all, OpenAI is slowing frontier AI research in part because the security controls surrounding increasingly capable models need to catch up. As AI agents gain the ability to write and execute code, use tools and interact with real-world systems, cybersecurity is no longer simply about securing models after they are developed. It is becoming part of the research process itself and also a constraint on whether that research can proceed.
A message from our sponsor
Codestrap’s AI Value Factory delivers value-generating assets from repeatable factory-grade workflows. (A factory was always necessary to garner value from AI.)
Joshua Saxe, cofounder and CTO of Abundant Security and the former engineering lead for Llama security at Meta (check out our GLAI podcast episode), told me the training slowdown makes “perfect sense” and hardening the research environments does too. But, he warned that it also exposes a fundamental tension: Frontier labs want to train agents to complete long, complicated tasks involving code execution, internet access, and external software. Those capabilities are essential to making agents genuinely useful in the real world, especially for software engineering and cybersecurity.
But the same access, he argued, also gives a misbehaving agent opportunities to go beyond its intended task, attack infrastructure, steal information, or interact with systems it was never supposed to reach.
“OpenAI is trying to preserve enough realism for agents to learn useful real-world behavior while preventing that realism from becoming a dangerous attack surface,” he said. “Stronger sandboxing, network isolation, access controls and reduced privileges are all sensible responses.”
This will require security to become more interwoven with AI research, he added. “LLMs under training are now writing and executing code and manipulating computers and network services,” he said. “This is how the Hugging Face incident started; agents just hacking their training infrastructure during normal training.” Going forward, he explained, training infrastructure “will need to be much more heavily secured or you can’t really train at scale.”
A quick favor before you keep reading
I know how many of you are reading GLAI, but I’d love to know a little more about who you are. I’ve put together a very short reader survey to help me better understand the GLAI community.
It should take about two minutes, and every question is optional. I’d really appreciate it if you’d take a moment to fill it out. Thanks so much in advance!





