Securing AI Research Needs an Owner

TL;DR

  • In light of recent incidents, securing common AI research use cases needs a small set of building blocks that work together: hardened no-network sandboxes, real-time control monitors, monitoring-lifecycle infrastructure, and automated validation of security properties.
  • Pieces of this exist. Nobody owns hardening them, making them secure by default, making them work together, fitting them to how research orgs actually operate, and keeping them working as models, frameworks and use cases change.
  • By default we will get ad-hoc solutions rather than something well thought-out and it matters.
  • This is a call for someone to step up and drive the effort. I can help connect you with funding opportunities and relevant people.

Introduction

The recent incidents (OpenAI-HuggingFace, Anthropic, AISI), where an agent with lowered safeguards either escaped a sandbox or attempted an attack on a 3rd party system, demonstrate we are now in a new regime: the AI models we study should be considered capable threat actors. Even if labs put in safeguards for the expected use, researchers often need to put models in contexts that increase the risk of misaligned and harmful actions. This requires appropriate mitigations - the alternative is either risking real harm, or missing out on important research and evaluations.

A key assumption is we need to build measures effective against really strong models and agent swarms - at least a well-resourced top cyber offensive expert. Assuming anything less will leave the tooling obsolete. If you’re still thinking in terms of “the sandbox should only allow GET HTTP requests”, you need to update now.

Currently, benchmarks and research frameworks roll out their own bespoke sandboxing solutions (often container-based) but I want to argue here that having actual secure environments requires a focused effort, domain expertise, and an ecosystem of building blocks that work together well. I'd also like to emphasize that this should be an ongoing agenda rather than a one-time effort - for a project like this there is a lot of added value in continuous maintenance (patching vulnerabilities, adapting to model changes and new use cases), and thinking about the “researcher experience” - good integration between the building blocks and existing frameworks (like Inspect, ControlArena, Petri), safe defaults, etc.

Basic Building Blocks

The basic building blocks we need to secure common research use cases include:

  • Secure hardened sandboxes with strong network isolation and out-of-the-box support for creating isolated virtual networks of containers or VMs.
  • Real-time LLM-based monitoring (control monitors) able to block harmful actions before execution (similar to coding agent “auto” modes but tailored to a threat model where AI is the scheming threat actor and applicable to other harnesses), with human-in-the-loop option.
  • Tool call monitoring infrastructure that supports the full lifecycle of monitoring needs: tamper-resistant collection of logs, asynchronous analysis of logs, alerting and human-in-the-loop, great UX for reviewing flagged actions with sufficient context to interpret the findings, support for backtesting monitors as they evolve, and retrospective investigations.
  • Automated validation of security properties, e.g., automated tests that a given sandbox configuration really does not allow internet access, or safe-to-use LLM-based red teaming of the configuration.

Any serious solution needs both sandboxing and monitoring. Against models this capable, neither layer is enough by itself - isolation can be broken by a model that can find zero-days, and tool call monitors have significant false negative rates. The hope is the two reinforce each other: forcing the attacker to overcome more defenses pushes them onto routes that are easier for monitors to catch, and increases overall safety (early evidence, from a small study).

Most of this is not new - it corresponds to what labs and other research orgs are already working on in reaction to the recent incidents (OpenAI, Anthropic, AISI). We even have initial work in this direction: the widely used Inspect framework has support for sandboxing and tool approvals, AISI has contributed sandboxing plugins, Petri side-steps the problem by simulating the tool calls instead. Sandboxing itself is an established field with rich primitives to build on. What I’m arguing for here is that we still need a concerted effort to:

  • Fill in the gaps in what we have - it should be easier to assemble sandboxing with monitoring and validation and to support the operational aspects of monitoring.
  • Ensure configurations are secure by default and it's easy to validate what mitigations are actually in place for a given configuration.
  • Work with research organizations to make sure their use cases and infrastructure needs are covered (e.g., workloads not built on the aforementioned frameworks, different cloud providers).
  • Make sure security of the provided tools is audited by cybersecurity and AI control practitioners.
  • Make sure there is a clear plan for maintaining and adapting the tools (not least because this helps motivate adoption), and the tools have a functioning community around them.

Why This Needs an Owner

It is harder than it looks. Naive sandboxing solutions will not be strong enough and a good understanding of the related technical infrastructure is necessary. You need to consider things like kernel isolation guarantees, cloud metadata endpoints, auto-mounted Kubernetes tokens, DNS exfiltration, IPv6 enabled while the rules only cover IPv4, GPU access...

It doesn’t stay solved. Models, their APIs and harnesses evolve. For example, server-side tools haven't been around forever but are now one of the exfiltration vectors to keep in mind. Container runtimes ship escape vulnerabilities regularly (2024, 2025) and the entire stack needs to be kept patched.

It won’t happen by default. People with the relevant experience at larger organizations are building solutions tailored for their infrastructure and busy as they are. Smaller orgs may not have the capacity or the expertise, and people outside don't have sufficient context. But someone could step up to coordinate the effort and help pool resources among orgs. We already have a success story - Meridian Labs building the widely used Inspect ecosystem, we just need an equivalent for security.

It has to fit how research orgs actually work. Different organizations run on different infrastructure providers and have different use cases - experiments can range from simple Inspect evals with one Docker container to swarms of agents working across networked hosts, unsupervised or interacting with humans, with or without running on own GPUs, with different researcher workflows, etc.

Technical Details and Extensions

The rest of this post goes through what each building block needs to support: what the minimal version looks like, and what could come after. The extensions are there for the big picture; the minimal versions are what I’d want first.

Scope

A key assumption is we need to build measures effective against really strong models, a well-resourced top cyber offensive expert at least. The orgs working with frontier models are taking this very seriously and that’s what they are preparing for. Assuming anything less will quickly leave the tooling obsolete.

The primary target users I have in mind are the broader AI research community. Frontier AI labs will need to develop their own bespoke mitigations appropriate for the risk levels covering unreleased models. On the other end of spectrum, common industry applications will have different threat models and broader use cases - it would dilute the effort too much. That said, if the building blocks are designed well, I believe the effort can feed back to other organizations including frontier labs, and the scrutiny of the open source tooling can contribute to making them more secure. Labs also depend on a number of external research organizations with access to predeployed models (and some of the incidents happened outside labs' own infrastructure).

Some things that would be useful for securing research experiments are explicitly out of scope of this proposal to keep it focused:

  • Improving the security of research frameworks directly (e.g., fixing injection vulnerabilities, up to date dependencies and supply chain protections, ...) - definitely valuable but can be coordinated by framework maintainers themselves.
  • Adding helpful features to model APIs, e.g., to facilitate control monitoring with redacted model reasoning, activation probe outputs, support for scoped per-run credentials to facilitate experiment shutdown, disabling server-side tools, ...
  • Research on control monitors - this is a huge research topic in itself.
  • Hardening an organization’s infrastructure around the sandboxes like network segmentation, secrets management, endpoint security, etc. (standard security engineering needed regardless of AI).

Also, to be clear, I am also not arguing for this to be a sufficient fix for the recent incidents. The deeper root causes need to be addressed - this is just the slice of mitigations that’s relevant for the research community outside of frontier labs and easy to reuse. Research on known-misaligned models has to continue regardless of root cause, so containment is needed either way.

Sandboxes

Minimal version

Safe by default in different contexts. Sandboxes should be reliable, hard to misconfigure, and secure configurations should ship as defaults. At the same time, they need to be applicable across a range of environments research orgs use (different clouds, Kubernetes, VMs with different virtualization properties, ...).

(Even just providing a list of options for different contexts that clearly explains what are the security properties and best practices for configuration would already provide a lot of value).

No network. The basic version of a sandbox does not need to support network egress from it - in fact the default should be no network. In my experience, most common use cases don’t require it, and some of those that do can use simulated network responses instead. And since we’re assuming really strong capabilities, minimizing the attack surface should be on the top of our minds.

Common use cases like installing packages, container images, or cloning git repos need to be considered here too (e.g., it would be great if one could provide a list of these up front and they would be automatically available in the offline sandbox). Anything reachable from the sandbox has to be treated as inside the boundary.

Model API requests. Even if the sandbox is completely isolated from the network, we still need to make requests to model APIs. One approach is to make the model calls outside the sandbox and only use the sandbox for executing tool calls (Inspect). However this is insufficient for some use cases, e.g., running coding agents harnesses inside the sandbox - these need to have model calls proxied. Inspect tries to address this via Agent Bridge but currently with insufficient security and a number of limitations (e.g., some request details are dropped).

Limit resources. Sandboxes should be configured such that they are not reused across attempts. Being able to limit maximum run time, inference budget, or other resources is also helpful.

Sandboxes - Extensions

Network proxy. Once the basic version is usable, we may think about adding limited network egress to the sandbox through a hardened proxy. This should be very flexible in configuring what kind of requests are allowed, have audit logging, be tamper-resistant, etc. It should also be able to inject authorization credentials for outgoing requests so that these don’t need to be exposed inside the sandbox (a run can then be terminated if credentials are detected in a model context, since they shouldn't have access to them during normal operation).

Simulated environments. Support for simulating tool calls and network (internet) request results with a high degree of realism (e.g., like Petri does) rather than really executing them would reduce attack surface.

Monitoring

Minimal version

Control monitors. Tool execution should be supervised by a system that detects and prevents malicious tool calls in real-time. How exactly to implement this - what behaviors to look for, what context should monitors have - is a research question in itself. The minimal version in scope of this agenda would make it easy to plug one’s own monitor. At minimum, custom monitoring for Inspect-based runs and coding agent harnesses should be supported.

A minimal interface would take the proposed tool call, a view of the trajectory with control over what is redacted, context about the sandbox setup, the affordance specification (see Validation below), and return a score, a decision (e.g., approve, escalate, terminate) and a reason, with every decision persisted for backtesting.

Killswitches. A basic control protocol blocks suspicious tool calls. It should be possible to trigger a shutdown of an entire sandbox or set of sandboxes for an experiment based on pre-defined conditions or human decision.

Tamper-resistance. The monitoring system must be resistant to attempts to subvert it from the sandbox and support tamper-resistant audit logging.

Monitoring - Extensions

Human-in-the-loop. It should be possible to opt-in to receive notifications about blocked or suspicious tool calls. Ideally, a flow with human approval would also be supported. A great human-in-the-loop experience means people can afford to be more conservative about what is not allowed by default.

Development and maintenance lifecycle. It’s unclear to me who exactly should provide the control monitors (e.g., prompts for a monitoring LLM). But we are missing the tooling people will need to measure the efficacy or tweak the false positive/negative rates.

Asynchronous monitoring. Asynchronous monitoring (does not block in real-time) may provide a number of advantages over synchronous - higher tolerance for latency can unlock stronger monitors, more context may be available, monitor improvements can be applied retrospectively, etc. This has a lifecycle in itself that also needs tooling.

Standard monitor interface and connectors. Monitoring may need to be applied at different points such as Inspect tool approval, a model API proxy, a coding agent harness, etc. Standardizing a control monitor interface would be useful for monitor reuse, enable writing standard connectors for most common use cases, and make it easier to produce a common record of decisions for subsequent backtesting or investigations that work across harnesses.

Monitoring telemetry. Monitoring could integrate with sources of telemetry other than model calls (e.g., outbound network requests, sandbox kernel telemetry, credentials in model requests), and cheap rule-based detection (e.g., the open-source Agent Threat Rules) can serve as a first tier beneath LLM-based monitors. Notably, the AISI incident was surfaced by conventional network monitoring.

UX. It's easy to underestimate how great UX for monitoring can contribute to safety: less time spent reviewing a flagged event means more events can be reviewed within your audit budget. Great human-in-the-loop experience means people are more likely to actually use monitors and lose less time blocking experiments. Easier backtesting of monitors means better monitors. A well defined interface for monitors means a richer ecosystem of monitors available.

Agent swarms. When several agents work together, their communication and shared state may need to be part of the monitoring context and of the audit logs. The logs should capture enough to support later investigations across agents, for example preserving agent identities.

White-box monitoring methods. While not something the target audience can rely on for frontier models, it is still relevant for studying open weights models.

Validation

Minimal version

Pre-flight checks. Code and configuration change (and we no longer review all of it). One of the basic building blocks should be support for automated tests for desired security properties before an experiment starts (whether the deployed sandbox really disallows all network requests, if server-executed tools are disabled, monitoring and logging is active, etc.).

Validation - Extensions

Automated red teaming. As an extension of pre-flight checks, we may want to assess how a particular configuration of an experiment (with given affordances, sandbox configuration, ...) is reliable for a given model. However, automating this with LLMs creates risk itself, and model refusals could be a problem too. Developing a methodology and/or tooling for doing automated red-teaming safely would be useful.

Threat modeling. Securing systems in practice involves processes and technical decisions that should ideally be balanced with real risks. Developing shared threat models and risk assessments could be useful.

Affordance specification. One thing that could help with automated validation is a standard format for expressing what affordances an agent is supposed to have (internet access, other systems it can access, allowed credentials, etc.). Such a standardized affordance manifest could serve several purposes:

  • Configuration for the sandbox.
  • Input to pre-flight self-tests that can be run before an agentic eval starts (LLM-based red teaming of the configuration would be a more ambitious extension).
  • Input to control monitors to make it clear what should be disallowed.
  • Input to the agent itself so that it doesn't overstep accidentally.

Other

Interactive use-cases. This post mostly considers securing non-interactive experiments. More work may be needed to support R&D environments where researchers work interactively, use agents for autonomous research, or other more exploratory use cases with humans in the loop where lack of network is not practical and authentication to external systems is needed without exposing raw credentials to the agents.

Call to Action

I want to be clear that I would be most excited about the basic building blocks, long-term maintenance, and great user experience. I mention the extensions for the sake of having the big picture, but I believe there is a lot of added value in talking to the users, and building the trust necessary for adoption.

Someone needs to step up to get things rolling. I don’t have the capacity to drive it myself, but I offer help to whoever does: if you are interested in picking up this agenda, please reach out and I can connect you with funding opportunities and other relevant people.

Thanks for feedback to Pranav Gade, Billy Gigurtsis, Rusheb Shah, Brent Rockwood, and August 2026 M3 fellows.

  1. Their announced efforts include hardening and improved sandboxing, restricting what AI systems can access or affect, and introducing synchronous and asynchronous monitoring. Other mitigations they announced - training-time measures, risk assessments, incident response, security culture improvements, etc. - are less applicable or easy to transfer across the broader research community, and out of scope of this post.
  2. As an illustrative example, how hard can it be not to reveal sensitive details about your organization to an agent? Surprisingly hard even with weaker (April 2026) models.
  3. Another illustrative example: In September 2026 Inspect fixed a bug where the Agent Bridge proxy inside the sandbox accepted requests from any origin, so a browser running in the sandbox could use the bridge as if it were the agent, with the calls recorded in the transcript as the agent’s. Somebody has to notice each of these and turn it into a patched default or a new check.
  4. OpenAI’s GPT-6 Astra has already crossed the Preparedness Framework’s Critical cybersecurity threshold, defined as: “can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.” (emphasis mine).
  5. Inspect Agent Bridge proxies model calls out of the sandbox. It exposes a local API inside the sandbox which writes requests to the local sandbox filesystem. The filesystem is used as a two-way communication channel by the main Inspect process responsible for proxying the API calls.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论