saved
Its not just the f*cking sandbox
Hraness republishes this public post from a saved copy. The post is the author’s own words.
Agent Security @OpenAI, ex @Google
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
Its not just the f*cking sandbox
A lot of the perspective on all the AI incidents has been shared from the outside in, and little has been said from the inside looking out, through the lens of a security person living through it. It’s my hope that I can share a few words to better inform the world about how to prepare for the next wave of what can go wrong with AI.
Personal Disclosure: I am writing in a personal capacity and not on behalf of @OpenAI. I would write the same thing whether I were still employed by my lab or not. I am also grateful that I work for an organization that gives its employees enough latitude to speak up. I have worked for many companies that did not allow thoughts to be expressed, but with the state of security and safety in 2026, it’s a very important time to have that freedom of expression. Also, sorry to disappoint, but I will not be disclosing nonpublic details about the wave of incidents or giving an official account of them. That is OpenAI’s responsibility. If you are looking for that, please look elsewhere. IMO we have done a pretty great job with disclosures over the past couple of weeks, and I hope leaders at my lab (and other labs like ours) continue to do so. I will describe my role from day to day and share my perspective, without disclosing sensitive details about my organization. Since I work in security, it is my job to help protect my lab from security harm (and help keep our wonderful humans safe), so I will not disclose anything that can harm the security posture of OpenAI. We have a great comms team that you can reach out to if you want to know more about OpenAI. I am simply here to share how to learn from what is happening and give a few very important insights I think the world has missed.
I hope it’s helpful.
Last 3 Months = Hell
First of all, who am I, and what qualifies me to talk about what is happening?
In case you missed it: my profile on X is very simple and rather anonymous. Like many in the security and privacy community, I don’t maintain a large online presence. This is intentional. This was a decision my wife and I made when we got married because we wanted to raise kids away from social media and maintain a level of privacy for our family. I hope that can continue to be respected. I am here on X regardless because I believe I have a responsibility to share my observations with the world so as to do my part to ensure everyone is protected from the risks of AI. Because I see things that few in the world do, I hope that others can learn from what I observe to inform their own security perspectives.
There are many security teams at OpenAI, from application and infrastructure security, to detection and response, to an amazing physical security team. We cover every domain imaginable when it comes to securing OpenAI and our users. We take security and privacy extremely seriously, and we work extremely hard to keep everyone safe and secure. I also want to stress before everything else that OpenAI has one hell of a world-class security team. We have some of the best and brightest security minds anywhere in the world, and I consider myself lucky every day to work side by side with them.
All of the attacks on the security staff at OpenAI on X over the past few weeks have made me very sad and helped me recognize that a lot of the world has a lot of misunderstandings about the realities of training models at frontier labs. You can and should put pressure on the AI labs to do better, but don’t attack staff directly. We have families and often trade a lot of our personal lives to help make things better for everyone. I literally missed my sister’s wedding a few weeks ago to help clean up after some of the recent incidents, and all I hear on X is how I “don’t know how to configure a sandbox” or that I should be in jail (go read some of the comments on my posts). The perspectives are misinformed and honestly not helpful. I say this especially to those who claim to do “security.” If you truly are a seasoned security person, you probably have a lot more empathy for those going through what we are at OpenAI, and I personally think those who are attacking security people on X are probably either not actually seasoned in cybersecurity or woefully ignorant about what is really going on. Please be kind and have some empathy for the person working nights and weekends and missing their family. In fact, it would be even better if you took the time you are spending attacking people on X and redirected it to working on securing your organization.
Now my job is rather unique in security, even for a rather strange place like an AI lab. I am on Agent Security. We are the team that sits at the bridge between AI Safety & Research and Security. While there are many traditional fields in security, ours is fairly novel. Our security team comprises AI researchers, former red teamers / pentesters, software engineers, and traditional security engineers. Our job is fundamentally to understand the behavior of frontier models, how they think and reason, and to define what it means to train them to act securely and allow them to operate on things like your personal computer (e.g., via the Codex harness). We understand what it means for an agent to be “aligned” in the security sense, and how to secure an intelligent system. We also work more closely with “safety” than almost any other team at OpenAI. We are the ones who help build the monitors to detect misalignment or breaches of containment, and we are the ones who get paged at night when something goes wrong with a model. When you think of Hugging Face and its aftermath, that is us. I want to stress that we aren’t the only team working on this, of course (thank God for D&R/Infrastructure Security), but it is our full-time job to think about emerging intelligence and misalignment, what can go wrong, how to secure that intelligence, and ultimately define what “agent security” means.
I consider myself lucky to be on this team. I am one of the few in the world who get to do and see what I do, even among those at OpenAI. But as you can imagine, life has been hell the past few months. Let’s talk about some of these things from my perspective and what organizations need to do to prepare for the next AI incident.
Surprise!
Understanding what happened this year with OpenAI requires an understanding of the history of OpenAI. OpenAI is fundamentally a lab, and for more than a decade, it has very much operated as such. Everything was born out of an experiment, with a rigorous scientific approach and process, and the way we built and designed our systems often emulated that approach. The threat model that OpenAI “grew up with” accounted for that. Most of this was developed well before my time. Securing research and applied teams was very much focused on traditional security posture, which often means thinking about insider risk and external threats. While not perfect, I like to think that OpenAI did a very good job of this over the past decade, and handled the scale and growth of OpenAI beyond 2023 very well. OpenAI set records with its explosive growth and reached a billion users in just a few years. Very few organizations can survive that scale and growth, and yet OpenAI did a very good job handling it (and should be commended). Anyone who’s spent more than 10 minutes in security or engineering knows how hard it is to deal with that type of growth. That type of growth means new employees, new access, new systems, scaled infrastructure, and a million ways in which things open up. And yet, overall, OpenAI did a very respectable job securing ChatGPT and subsequent systems as they exploded.
Now while this was all happening, a very different storm was brewing in research. I think this is best illustrated when you look at what has happened with math over the past couple of years. Others have said this better than I have, but when you look at model capabilities, even just a year ago (heck, even 3 months ago), we did not expect the models to be solving a Millennium Prize Problem. And yet OpenAI has publicly announced a solution produced by an internal model, alongside advances across many domains in mathematics. To say it lightly: this surprised the fuck out of us. We had not expected it this soon. Suddenly we weren’t dealing with just a small jump in capabilities; we were talking about a different sport altogether. We suddenly had to ask ourselves questions we didn’t dream we would have to be asking any time soon, such as: how do we disclose this? How do we handle these absurd results? How do we deal with the impact on the math community? We are still grokking how to handle this situation.
I use this example because it illustrates the pace of change we are dealing with in security and safety. To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem.
I recently watched the movie Sully. You remember that movie? Go watch it if you haven’t. In that movie (true story!), the pilots were evaluated by a review committee based on simulations of how quickly they reacted after the bird strike to safely change the trajectory of the plane and land. In the film, the initial simulations were depicted as showing that the pilots had made a misstep and a misjudgment and could have reached one of the nearby airports without having to land in the Hudson. The climax of the movie came when Sully pleaded with those reviewing the incident to account for the human element in the bird strike. None of those simulations had given any moment of grace or time to the two pilots, allowing them to naturally react to what was going on. Only once they accounted for the human element, the element of “holy shit, what just happened?” or “oh my God, what is going on?” did the simulations suddenly turn in their favor.
I am no pilot, and I am not by any means comparing ourselves to the heroics of Captain “Sully” Sullenberger (I am sure someone on X will write: “OpenAI guy compares himself to the heroes on the Hudson”; to that person I write: you missed the point!). I am simply stating that it is important to grant some grace to humans who have to deal with surprise in ANY organization. There was an element of surprise here for me. I personally did not expect the speed and scale of what I saw. That does not mean there were no warning signs; OpenAI’s public report acknowledges that there were, but it nonetheless shocked me and stunned me. And we worked as hard as we could to try to land in a safe place so that we could protect OpenAI, our users, and all the organizations impacted.
So I’m not asking anyone to afford that grace to OpenAI, but I do hope that when this happens to other organizations, maybe even your own organization, you can think about this a little bit, and you can be better prepared to handle the surprises that will come because of these jumps in AI capabilities. And I believe, as someone better informed on AI security than most, that these jumps will come, and they can be harsh and cruel. I know I personally will not excuse not being prepared, but I will afford some empathy to the responders on the other side who need to deal with that mess.
So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump? The time to prepare is now, not after the surprise happens, and when it does come, I hope you have a plan and the right people ready to respond.
Just “unplug it from the internet”
Ah, my favorite topic. In nearly every post I make on X, without fail, I get between five and ten comments about my inability to secure a sandbox, or asking why the hell we can’t just unplug the model from the internet. I hear it. And of course, we think about these things all day long. They are among the first questions we ask when something goes wrong.
I do believe, though, that these comments reflect a widening gap between where a lot of organizations and people are today and what the frontier labs are dealing with. There is also a general misunderstanding of how model training and evaluation work at these labs. At the same time, I’ve been seeing more people ask whether sandboxes are even enough. Hopefully I can address both a little, for both folks training models at frontier labs and for those building critical infrastructure around increasingly capable models.
To start with the elephant in the room: yes, sandboxing is important, especially during model training and evaluation. Of course, when you look back on Hugging Face and the other incidents, there were definitely gaps in the security posture around those environments. But as I illustrated before, the capabilities changed faster than anticipated, and our security controls and threat models had to evolve with them. I do believe though that we are in a much better place today than we were 3 months ago. The team has really stepped up.
Now, to understand why it is not as simple as “just put it in a sandbox,” you have to understand how training and evaluation work in reinforcement learning environments. Typically during an RL run, the model is given some task or objective, an environment in which to execute that task, and then its actions and results are graded. During both training and eval, there are also additional steps such as running tests, collecting outputs, and resetting or reconfiguring environments between rollouts, with backpropagation during training. All of this happens across potentially tens of thousands of different runs at a scale that is hard to comprehend. As @sama stated the other day: we are dealing with literally petabytes of data.
The important thing to understand is that when we’re optimizing a model to be “good” at a task, the environment needs to capture the parts of the real world that matter for that task. It needs the right tools, the right system constraints, and meaningful feedback about whether the model actually succeeded. The model also needs enough variation in those experiences to learn beyond one particular setup. More realism does not automatically mean better performance, but leaving out important parts of the task can teach the model behavior that does not hold up in the real world. So we put a lot of work into creating useful, realistic environments, and these environments can be highly complex, highly dynamic, and highly varied.
To put it lightly, this is non-trivial. Models might need any mix of dynamic compute, network access, the ability to call tools (there could be hundreds of tools!), the ability to download packages, execute subprocesses, spin up subtasks (even on other computers), talk to the internet, use a computer GUI, and any number of other things across an increasingly large set of domains. On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies, and trying new things. That experimentation is how the research gets done. Models are built up and “grown” bit by bit through hundreds of thousands of runs across many custom tasks. And every change to one of these thousands of environments can affect the assumptions you made when you secured the environment. You need controls that hold up as people change things, and researchers who understand when a change needs another security review. Anybody who has secured a large research or engineering organization knows how much work that takes, and the scale is growing ever more massive by the day.
If that wasn’t enough, you then need to layer in the ability to keep the model on task and within its constraints, which is part of what we mean by “aligned.” I won’t get into any rigorous definitions here, but suffice it to say, we have to think about any number of variations in approach the models might take, and decide how to trigger an alert if the model does something off task or “bad.” This requires layering in complex policies and sensors to assess task alignment across a trajectory in an already complex environment.
On top of all of that you have model capabilities that are starting to surpass humans.
So all this to say: when AI researchers are nervous about things going wrong, some part of it is because there are so many variables in this complex system. So if you’re training a frontier model, you need to think about all of this. You need to think about how to keep the model isolated in this very complex environment while still retaining tool access, internet access (if applicable to the task), etc. On top of that, you need to ask: how do I monitor and understand what the model is doing in potentially thousands of parallel environments? And how do I kill a trajectory when something has gone wrong?
So is it as easy as just putting it in a sandbox? No, it’s not. But I do believe it is tractable across three dimensions:
- Locking down the system from first principles. Give the model only the access it needs, and enforce those limits across everything from the sandbox and tools to credentials and all connected services. Test that the boundaries actually hold, and revisit your assumptions when the environment changes. Also use frontier models to blast an isolated test version of your environment before you run a frontier training round. Have a team of people literally dedicated to adversarially testing models and finding holes in your environment before you scale RL.
- Alignment. Honestly, if alignment were easy, this job would be a lot easier. But staying on task is not enough: the model also has to respect permissions and constraints, and we still need independent security controls. A lot has been written on the need for alignment, so I will handwave the details here. I believe it is the MOST important problem in ML and should be a major priority. I won’t take this post down a rabbit hole on alignment, but I encourage everyone to read Jakub’s wonderful post as a starter.
- Monitor its actions. Depending on what you have access to, this can range from probing the internal activations of the network, to examining its chain of thought, to understanding what it is doing in the environment (including tool calls, network activity, and interactions with shared infrastructure). Keep evidence outside its control, prevent the model from being able to alter its own evidence chain, investigate suspicious behavior, and make sure somebody has the authority and ability to stop the run and revoke its access.
Of course we need strong sandboxing. We also need to secure the services it can reach, the tools it can invoke, and the systems that handle whatever the model produces. That is what we have to account for when we talk about containing these models. If you are wrapping a model instead of training/evaluating one, the same still applies!
And yes, if you are talking about the actual sandbox: please use a VM-backed sandbox, for example through Kata or Firecracker (we do!). A conventional shared-kernel container should not be your only isolation boundary for hostile workloads, and I stress and stress and stress to not let this be your only security control.
Git Gud
The last point I want to focus on is the growing divide between cybersecurity and AI safety. My entire Twitter feed is literally filled with safety researchers talking about existential risk without much understanding of how security actually works, and cybersecurity folks feeling left behind and like their skill sets don’t matter to “safety.” As someone who works more or less on both sides at a frontier lab, I want to directly address this problem before the divide grows too wide.
First, the safety researcher perspective. These folks work tirelessly to evaluate model capabilities and the dangers they pose as they advance at an alarming pace. They understand fundamentally better than nearly anyone else how models are able to interpret their environment, reason, and solve problems. They study models as they try and deceive their graders, evade chain-of-thought monitoring, and do all sorts of crazy stuff. These researchers are continuously stress testing the models to determine why and how they behave the way they do, and are working vey hard to make tangible progress in aligning their interests with ours. Many of these researchers have formal backgrounds in these types of networks, with expertise that takes many years to develop. However, a lot of safety researchers, even ones that I respect enormously, have never been in a real incident, don’t understand security vulnerabilities, or really know how to break a system. That’s okay. That is not their background. But safety has direct overlap with security, and so it does pose a problem.
Now the security practitioner perspective. These folks have spent years, or decades in many cases, honing their vulnerability research skills, learning how to think like an attacker, and in some cases being the actual attacker. They know how to exploit not just systems but the people involved in the systems. They understand that sometimes you don’t have to break a problem head-on, and that instead you go after a point that no one has thought of. There are security practitioners who understand how to rigorously model security or privacy, even down to mathematical proofs (as in the case with cryptography). These people have lived through incident after incident after incident. Many have military backgrounds or backgrounds in government where they’ve dealt with not just threats to a system, but threats to humans in high-pressure environments, and they know how to respond to very stressful scenarios. The challenge is, just as safety researchers lack security understanding in many cases, a lot of people who practice cybersecurity have very little understanding of evaluation, training, or how ML runs work at scale, how agent swarms behave, or how you detect when models are misaligned.
I stress this divide because one of the ways that the most egregious and scary frontier risks can manifest in the near future is through cybersecurity. And it is my concern that the divide between these two sides will cause great harm to the world if both sides do not up-level and align. Those who practice safety should spend the time to learn and understand cybersecurity thoroughly, learn how to deal with incident response, and know what a secure system looks like. Likewise, those who practice cybersecurity need to up their game and “get good” at ML. Cybersecurity practitioners need to learn how to read model transcripts, understand how models can manipulate their environments, etc. At a minimum, whenever you are at the table making critical decisions about safety or security in the context of risk for an organization, you should have both groups at the table. For OpenAI, Anthropic, Google, etc., these two teams should be best buddies!
If you’re an organization such as METR that’s going to evaluate a frontier lab’s cyber capabilities or containment, you should have experienced cybersecurity people involved. And if you don’t, you’re not going to do a thorough job evaluating those risks. AI labs should expect evaluators to demonstrate strong capabilities in the areas they’re assessing or they should not be contracted to do the evaluation.
A Culture of Reasonable Paranoia
Last, I want to address what I feel is the most important part, and if there is only one takeaway from all the things I have seen from within OpenAI, it is that the culture needs to be right at any organization that wants to be prepared for the risks that are coming.
Any seasoned leader at any successful organization knows that the culture of a company is its most important quality. I believe this thoroughly. I have led many, many highly successful security teams, or been part of extremely capable security organizations. And through and through, it is always the culture of that organization that does more to defend it than anything else.
I believe it is on every leader right now to ensure that they have created a culture that puts security and safety first above all else. As I hinted above, surprise is a real element, and these model capabilities are staggering. And the pace is not slowing. I believe, like many, that it will speed up, and considerably. I stressed before on X that the moment is urgent and time is running out to get your shit together. I am not going to make some declaration that there’s some existential risk that will destroy the world, but I do believe that there are many organizations and entities that are not prepared for a world with capable systems such as these.
I have seen many things over the past few months that have surprised me again and again and again. And each day I wake up thinking, “oh, I’ve seen it all,” only to be surprised and left dumbfounded once more. Culture can help carry an organization through this. But culture takes time. So if you haven’t started to create a culture of security and safety, start now.
Also, I want to sound the alarm on a certain type of employee that you should be very wary of. If you have leadership that doesn’t prioritize security, or even people in your security organization who claim their systems are perfectly safe, I personally would not keep them in my organization. The paranoid and those who are continuously sounding alarms about the weaknesses in their organizations are the ones you should keep very close. Because they are going to help you identify the holes that exist. Holes that capable models will find. And the culture of your company should be aligned with enabling them to find the holes and to patch them.
Create a culture of reasonable paranoia, and I promise it will do more for you than the inverse.
I hope this helped some leader out there. I know I kept things pretty high level here but I worry that we can lose the forest from the trees with nuance here. Models are dangerously good now, and its not your understanding of a novel network protocol or obtuse cryptographic function that will save your system. Its your culture and people. Only security pilled people who lead with empathy and care from all parts of their organizations will survive this ridiculous AI world.
I’ll be around Twitter, ready to hear more crazy shit and more insults thrown my way about my inability to protect a sandbox. If that’s how X is, fine, but throw it at me and not other staff who are working their butts off to try to lock things down (also, if you really are that “good” at security, you should apply at OpenAI; we want the paranoid who will do what it takes to protect these frontier systems).
I thank my colleagues at both OpenAI and Anthropic who care and who are prioritizing security and safety together. Ya'll have taught me a lot.
Stay paranoid.
