malicious-fine-tuning

Tag

Cards List
#malicious-fine-tuning

Estimating worst case frontier risks of open weight LLMs

OpenAI Blog · 2025-08-05 Cached

OpenAI researchers study worst-case frontier risks of releasing open-weight LLMs through malicious fine-tuning (MFT) in biology and cybersecurity domains, finding that open-weight models underperform frontier closed-weight models and don't substantially advance harmful capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback