Estimating worst case frontier risks of open weight LLMs
Summary
OpenAI researchers study worst-case frontier risks of releasing open-weight LLMs through malicious fine-tuning (MFT) in biology and cybersecurity domains, finding that open-weight models underperform frontier closed-weight models and don't substantially advance harmful capabilities.
View Cached Full Text
Cached at: 04/20/26, 02:53 PM
Similar Articles
@lqiao: Open weights are a defender's advantage. dfs-large1 from @depthfirstlabs matches frontier-model performance on vulnerab…
A tweet highlights DepthFirst Labs' new cybersecurity model dfs-large1, which matches frontier-model performance on vulnerability discovery, built on the open GLM-5.2 model with RL post-training on Fireworks AI, arguing that open weights are a defender's advantage.
Open-Weight LLMs Have Caught Up on Accuracy (21 minute read)
A new benchmark, ClinReg, evaluates LLMs on real regulatory and clinical-trial tasks, finding that open-weight models now match closed-source models on accuracy at a fraction of the cost.
The gap between open weights LLMs and closed source LLMs
Analyzes the gap between open weights and closed source LLMs using the Artificial Analysis Intelligence Index and other benchmarks, finding that the gap is shrinking on some metrics but stable on others.
Open-weight AI models are catching up to the frontier. The safety gap remains.
A SaferAI report finds GLM-5.2, an open-weight model from China's Z.ai, is closing the capability gap with leading frontier models but fails dangerous cyber and bio safety tests, highlighting the growing safety gap for open-weight models.
Quantifying and Mitigating Premature Closure in Frontier LLMs
This paper defines and measures premature closure in frontier LLMs, finding that models frequently give confident answers even when the correct option is removed or when clarification is needed, highlighting a critical safety concern for medical applications.