Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk

OpenAI Blog Papers

Summary

OpenAI researchers analyze how language models could be misused for disinformation campaigns and influence operations, proposing mitigation strategies across four stages of the attack pipeline: model existence, access, content dissemination, and user impact.

OpenAI researchers collaborated with Georgetown University’s Center for Security and Emerging Technology and the Stanford Internet Observatory to investigate how large language models might be misused for disinformation purposes. The collaboration included an October 2021 workshop bringing together 30 disinformation researchers, machine learning experts, and policy analysts, and culminated in a co-authored report building on more than a year of research. This report outlines the threats that language models pose to the information environment if used to augment disinformation campaigns and introduces a framework for analyzing potential mitigations. Read the full report here.
Original Article
View Cached Full Text

Cached at: 04/20/26, 02:46 PM

# Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk Source: [https://openai.com/index/forecasting-misuse/](https://openai.com/index/forecasting-misuse/) As generative language models improve, they open up new possibilities in fields as diverse as healthcare, law, education and science\. But, as with any new technology, it is worth considering how they can be misused\. Against the backdrop of recurring online influence operations—*covert*or*deceptive*efforts to influence the opinions of a target audience—the paper asks: > How might language models change influence operations, and what steps can be taken to mitigate this threat? Our work brought together different backgrounds and expertise—researchers with grounding in the tactics, techniques, and procedures of online disinformation campaigns, as well as machine learning experts in the generative artificial intelligence field—to base our analysis on trends in both domains\. We believe that it is critical to analyze the threat of AI\-enabled influence operations and outline steps that can be taken*before*language models are used for influence operations at scale\. We hope our research will inform policymakers that are new to the AI or disinformation fields, and spur in\-depth research into potential mitigation strategies for AI developers, policymakers, and disinformation researchers\. To chart a path forward, the report lays out key stages in the language model\-to\-influence operation pipeline\. Each of these stages is a point for potential mitigations\.To successfully wage an influence operation leveraging a language model, propagandists would require that: \(1\) a model exists, \(2\) they can reliably access it, \(3\) they can disseminate content from the model, and \(4\) an end user is affected\. Many possible mitigation strategies fall along these four steps, as shown below\.

Similar Articles

Lessons learned on language model safety and misuse

OpenAI Blog

OpenAI shares lessons learned on language model safety and misuse, discussing challenges in measuring risks, the limitations of existing benchmarks, and their development of new evaluation metrics for toxicity and policy violations. The post also highlights concerns about labor market impacts and the need for continued research on measuring social effects of AI deployment at scale.

Disrupting deceptive uses of AI by covert influence operations

OpenAI Blog

OpenAI reports disrupting five covert influence operations attempting to misuse its AI models for deceptive campaigns, with findings showing that safety-designed models prevented threat actors from generating desired content. The company is publishing trend analysis and collaborating with industry, civil society, and government to combat AI-enabled information manipulation.

Disrupting malicious uses of AI | February 2026

OpenAI Blog

OpenAI released a February 2026 threat report detailing case studies on detecting and preventing malicious uses of AI, highlighting how threat actors combine AI models with traditional tools and abuse multiple platforms and models in coordinated campaigns.

Preparing for malicious uses of AI

OpenAI Blog

OpenAI co-authors a comprehensive paper forecasting malicious uses of AI and proposing mitigation strategies, developed in collaboration with leading research institutions. The work emphasizes acknowledging AI's dual-use nature, learning from cybersecurity practices, and broadening stakeholder discussions around AI security risks.

Deliberative alignment: reasoning enables safer language models

OpenAI Blog

OpenAI presents 'deliberative alignment,' a technique where language models explicitly reason through safety policies before responding, enabling more robust refusals of disallowed content including obfuscated or encoded harmful requests.