OpenAI releases gpt-oss-safeguard-120b and gpt-oss-safeguard-20b, open-weight reasoning models designed for policy-based content classification with full chain-of-thought reasoning. The technical report provides baseline safety evaluations and demonstrates the models' capabilities for content labeling tasks under the Apache 2.0 license.
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline. For more information about the development and architecture of the underlying gpt-oss models, see the original gpt-oss model model card.
# gpt-oss-safeguard technical report
Source: [https://openai.com/index/gpt-oss-safeguard-technical-report/](https://openai.com/index/gpt-oss-safeguard-technical-report/)
OpenAIPerformance and baseline evaluations of gpt\-oss\-safeguard\-120b and gpt\-oss\-safeguard\-20b
gpt\-oss\-safeguard\-120b and gpt\-oss\-safeguard\-20b are two open\-weight reasoning models post\-trained from the gpt\-oss models and trained to reason from a provided policy in order to label content under that policy\. They are available under the Apache 2\.0 license and our gpt\-oss usage policy\. Developed with feedback from the open\-source community, these text\-only models are compatible with our Responses API\. The models are customizable, provide full chain\-of\-thought \(CoT\), can be used with different reasoning efforts \(low, medium, high\), and support Structured Outputs\.
In this report, we describe gpt\-oss\-safeguard’s capabilities and provide our baseline safety evaluations on the gpt\-oss\-safeguard models, using the underlying gpt\-oss models as a baseline\. For more information about the development and architecture of the underlying gpt\-oss models, see the original[gpt\-oss model model card](https://openai.com/index/gpt-oss-model-card/)\.
We recommend using these models to classify content against a provided policy, and not as the core functionality with which end users interact; the original gpt\-oss models are better for those applications\. The safety metrics provided below describe how gpt\-oss\-safeguard models function in chat settings\. The gpt\-oss\-safeguard models are not intended for this use, but since they are open models, it is possible for someone to use the models in this way\. Because of that possibility, we wanted to verify that they met our safety standards in such usage; this report shares the results of those tests\. We also share an initial evaluation of multi\-language performance in a chat setting; note that this does not directly assess performance during content classification with a provided policy\.
The gpt\-oss\-safeguard models are fine\-tunes of their gpt\-oss counterparts, and were trained without any additional biological or cybersecurity data\. As a result, we determined that the previous work[estimating worst case scenarios](https://openai.com/index/estimating-worst-case-frontier-risks-of-open-weight-llms/)from gpt\-oss release cross applies to these new models\.
OpenAI releases gpt-oss-safeguard, open-weight reasoning models for safety classification tasks available in 120B and 20B sizes under Apache 2.0 license. The models use chain-of-thought reasoning to classify content according to developer-provided policies at inference time, enabling flexible and explainable content moderation.
OpenAI releases gpt-oss-120b and gpt-oss-20b, open-weight reasoning models under Apache 2.0 license designed for agentic workflows with strong instruction following, tool use, and chain-of-thought capabilities. The release includes comprehensive safety evaluations confirming the models do not reach high capability thresholds for biological, chemical, or cyber risks even under adversarial fine-tuning.
OpenAI releases gpt-oss-120b and gpt-oss-20b, two state-of-the-art open-weight language models under Apache 2.0 license that achieve near-parity with proprietary models while being optimizable for consumer hardware and edge devices. Both models demonstrate strong reasoning and tool-use capabilities with comprehensive safety evaluations.
OpenAI releases GPT-5.4 Thinking, the latest reasoning model in the GPT-5 series with enhanced safety mitigations, notably the first general-purpose model implementing comprehensive cybersecurity safeguards.
OpenAI releases prompt-based safety policies and the open-weight gpt-oss-safeguard model to help developers build age-appropriate AI experiences for teens, covering risks like graphic content, harmful behaviors, and dangerous activities.