Open-Weight AI Models Are Catching Up — Safety Tools Are Not
Open-weight AI systems are closing the capability gap with closed frontier models far faster than evaluation and safety tooling can keep pace. Because releases are effectively irreversible, that mismatch carries consequences closed models do not.

The gap between open-weight AI models, whose underlying parameters are published for anyone to download and run, and closed frontier systems kept behind an API has narrowed sharply over the past year. Several open releases now perform competitively with commercial frontier models on standard benchmarks, a development widely celebrated for democratising access to powerful AI tools. Less discussed is a parallel and more troubling trend: the safety evaluation and red-teaming infrastructure built to assess these models before release has not kept pace with how quickly their capabilities are advancing.
Why open releases are different from closed ones
A closed model’s behaviour can be adjusted after release — providers can patch prompts, restrict access, or roll back a version entirely if serious problems emerge. An open-weight model, once its parameters are published, offers no such recourse. Anyone can download, fine-tune, or strip away safety features embedded during training, and there is no mechanism to retrieve copies already distributed across the internet. That irreversibility means mistakes made at the release-decision stage cannot be corrected later, placing far greater weight on pre-release evaluation than is required for equivalent closed systems.
Where evaluation tooling is falling behind
Standard capability benchmarks measure how well a model performs specific tasks, but they say relatively little about harder questions: how easily safety training can be removed through fine-tuning, whether a model could meaningfully assist with tasks such as cyberattacks or biological-weapon design, or how its behaviour changes when adapted for a specific misuse case. Building reliable tools to answer those questions is slower and more resource-intensive than benchmarking raw capability, and the organisations best positioned to fund that work often have less commercial incentive to do so than they do to ship the next model release.
The current landscape of open-weight risk assessment
- Independent research groups have shown that safety fine-tuning in several open models can be substantially undone with a modest amount of additional training, at relatively low cost.
- Benchmark suites used to justify release decisions rarely include dedicated tests for dual-use capability, such as assistance with chemical or biological weapon design.
- Some developers now publish model cards disclosing known limitations, but disclosure practices vary widely and are not independently audited in most cases.
- Third-party red-teaming before release remains inconsistent, with some labs commissioning extensive external testing and others relying primarily on internal review.
- Regulatory proposals to require pre-release safety testing for open-weight models specifically have faced pushback from developers who argue that openness itself is a safety benefit through broader scrutiny.
Openness is genuinely valuable for research and for reducing concentration of power in a handful of companies. But openness and irreversibility are the same property viewed from two directions, and policy has to take the second one seriously too.
The argument in favour of open releases
Proponents of open-weight development make a case that carries real weight: broad access accelerates independent safety research, prevents capability from concentrating in a small number of companies, and allows academic and public-interest researchers to study model behaviour directly rather than through a restricted API. Several documented vulnerabilities in widely used AI systems were first identified by outside researchers working with open models, an outcome that would have been far harder to achieve if access had remained closed. The debate is not between safety and openness as opposites, but about how much pre-release scrutiny is proportionate given the release’s irreversibility.
What better practice would look like
Safety researchers increasingly argue for a staged approach: capability thresholds beyond which open release requires third-party evaluation specifically targeting misuse potential and fine-tuning resistance, rather than only general capability benchmarks. That would not resolve every disagreement about how much risk is acceptable, but it would at least align the intensity of scrutiny with the size of the decision being made, treating an irreversible global release with a different standard of care than an internal capability upgrade behind a closed system.
Where the policy conversation is heading
Momentum is building, unevenly, toward frameworks that would apply heightened scrutiny specifically at the point of open release rather than treating open and closed models identically. How that plays out will depend heavily on the broader federal AI regulatory push currently taking shape, since open-weight-specific rules are more likely to emerge as part of a wider framework than as a standalone measure. For now, the gap between what open models can do and what the field can reliably say about their risks continues to widen with each release cycle.
Was this helpful?

Mara Ellison
Technology Editor, Lonic
Mara has covered enterprise software for eleven years and spent two of them embedded with deployment teams shipping agent systems into production support desks.
- Artificial intelligence
- Enterprise software
- Automation
Read our editorial standards or send a correction.



