Rather than allow or ban open weights, rethinking access
I had just completed Frontier AI Governance course by BlueDot, and one of the newfound interest is about open weight models. I believe that open weights are important for the benefit of public and academia research, and it helps us to improve further both AI capabilities and safety. I also believe that there's a likelihood for high impact misuse risks (e.g. hacking and other offensive cyber capabilities). There is also the fact that open weights has risk of irreversible nature as safeguards can be removed and once the models are out it's practically impossible to retract them. As low as the risk might be, disregarding it blindly in favor of open weights seems irresponsible. Here, I am writing my thoughts on it.
This writing is made with what I took away from the course, and a rather quick research.
TLDR; I think that providing public access to models is needed, but not entirely in the form of open weights. These are early thoughts.
The wider, high level discussion landscape seems to still be mostly using a binary lens - either it's open weight or closed models. Perhaps because it's simpler to imagine and understand, either it's on or off. But there's a advocates reaching for a separate view, looking it from an access viewpoint. I think this makes more sense. It's similar to how Identity and Access Management (IAM) goal is to give the right users get the right access to the right resources at the right time. It's also in a way similar to applying the principle of "data minimization" - use the least amount of data needed to perform the processing.
Thus, I think it's more productive and future-oriented to switch from the framing "allow open weights vs prohibit open weights" into "how to provide different model resources to different groups"?
To address the first part of the question providing resources - the idea of structured access has already been proposed for some time. There's ideas to provide access to third parties from Research APIs, and thinking of the access through different deployment means from local to cloud hosted deployment settings. A way I'd like to think of it right now is a two way axis, one axis is "openness" and the other is "deployment".
Openness. At one end, it's a fully open-sourced AI where all the code, data, weights, license are open and fully shared, free-to-use. At the other end, it's a very limited model that can only be used by certain groups for certain purposes (e.g. Anthropic's Mythos or OpenAI's Cyber Models). In the middle, we have different levels of access like weights, activations, logits, sampling and output tokens.
Deployment. It's easier to think of it through by who deploys and runs it. Is it the user or a third party (e.g. OpenAI, Google, Anthropic). It's also about whether the user can deploy the model on their own (local machine or their chosen cloud compute). There's also a slight nuance, suppose that a frontier lab releases Model-X for any larger organizations like OpenRouter or HuggingFace to run, but the weights itself is not released. Something like this would be in the middle level between user and third party.
There's different risks that's associated with each level of openness-deployment. At one end of the spectrum, a fully open source model is highly risky because just about anyone (assuming the necessary compute infrastructure) is able to recreate the model. It's arguably perhaps easier to abliterate an open weight model, making it do bad things, while it's harder to jailbreak the same model within a closed setting (assuming well trained and safeguarded model). At the other end, a fully closed model use has very limited uses and people, making it easy to track usage (assuming a good and trusted party).
I think thinking of the problem in this way help to structure the different needs and circumstances, so we could design the right access and controls to different use cases.
Now the second part of the question about different groups, there's different use cases. For example, categorizes uses as chat, sampling, inspecting, fine tuning, and modifying.
These are non-complete sample use cases that I could think of to help illustrate:
- [A] As a researcher, I would like access to models' activations and perform (probes, SAEs, interventions, etc).
- [B] As a researcher, I would like to evaluate models performance through chat and sampling.
- [C] As a researcher, I would like modify (combine different models, change architecture, add other modality encoders, etc) models.
- [D] As a deployer, I would like to use local hosted models and add additional controls using probes and perform interventions.
- [E] As an end-user, I would like to host my own models locally and chat with it, or use it with my existing harness (Hermes, OpenClaw, etc)
Given these use cases, there are still empty rooms.
[A] As a researcher, I would like access to models' activations and perform (probes, SAEs, interventions, etc).
Suppose the main requirement is access to activations, then having a research API like NNSight perhaps is a good way to start. We do not need to share the open weights model, and settle with less openness (just the activation results instead of the weights), and thus reducing the risk surface. There's probably more logistics to think about, e.g. the activation data transfer between cloud and local might add bottlenecks and significant slow down research process. Perhaps cloud providers could provide a specialized environment for these? There's also the question of who would be responsible to maintain such environment. I wonder if there's a way we can also provide such environment in a container that can be run locally, which from quick reading it seems infeasible right now.
[B] As a researcher, I would like to evaluate models performance through chat and sampling.
[E] As an end-user, I would like to host my own models locally and chat with it, or use it with my existing harness (Hermes, OpenClaw, etc)
The first is attainable through current status quo, by utilizing the Chat / Completions APIs most providers provide. For the second one, there's some examples like Gemini Nano within Android SDK and Apple AFM within iOS SDK. It'd be great if there's a standardized library to run models that are in a certain format to disallow weights access, or controls around the system itself. Though it seemed like these are fairly weak right now (i.e. Gemini Nano being uploaded to HuggingFace).
As a deployer, I would like to use local hosted models and add additional controls using probes and perform interventions.
As a researcher, I would like modify (combine different models, change architecture, add other modality encoders, etc) models.
This is where it gets muddier as the open weights need to be shared. I think that licensing as an administrative control prior to sharing the models is needed, and some way to watermark or ID a copy of these models (although arguably they aren't the most effective). Perhaps some hardware-software integrated control where certain model is only runnable on approved licensed hardware. Or no more online download, people need to get the weights copies physically through a sharing facility where they can be verified. I think there's still open questions on what'd be the best way to share it, with probably the risk goal can't be a total avoidance, but reduce it to minimal levels as possible.
Overall, as I'm writing this, I shifted my perspective from a simple statement of "allow open weights" toward more differentiated access control. We need to still allow the use cases open weights currently provide, just maybe - not in the exact form of open weights.
- Milles, Hoffman, Gelles. The Use of Open Models in Research. Center for Security and Emerging Technology.
- I am with a familiar position with Greenblatt's writing e.g. "open-weight models can reduce these larger risks by accelerating AI safety research (which somewhat differentially benefits from open-weight models) and by increasing societal awareness of AI."
- Few recent incident or examples includes: OpenAI-HuggingFace Incident (METR Report), Rouge AI Agents Hacking (Transluce Report), Anthropic Threat Intelligence September 2026 Report
- "Once open weight models are released, these options are lost permanently: safeguards can be removed, and copies can be downloaded, redistributed, and run on private systems beyond monitoring. For models with dangerous capabilities – including highly cyber-capable models – open weight release therefore creates a persistent and irreversible risk of misuse." - https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are-leading-open-weight-models-on-cyber
- For example, the open letter endorsed by NVIDIA, "allowing vs prohibiting open weights" https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf.
- "Structured access is an emerging paradigm for the safe deployment of artificial intelligence (AI). Instead of openly disseminating AI systems, developers facilitate controlled, arm’s length interactions with their AI systems". https://arxiv.org/pdf/2201.05159
- Bucknall, Trager. "Structured Access for Third Party Research on Frontier AI Models Investigating Researchers Model Access Requirements"
- Kembery, Bucknall, Simpson. "Position Paper: Model Access should be a Key Concern in AI Governance"
- "Open Source AI" definition by OSI