人工智能在在线评论中的作用
摘要
本文介绍了一种实证方法,用于衡量大型语言模型供应冲击对在线评论的影响,发现未经验证的评论向更负面方向转变,并且出现活动激增,表明人工智能正在重塑平台动态。
arXiv:2609.22198v1 Announce Type: new
Abstract: The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentially harmful forms of manipulation that may undermine platform effectiveness and reshape platform dynamics. However, measuring such activity is difficult because AI-generated content is rarely directly observable. We introduce an empirical approach that leverages discrete LLM supply shocks - abrupt changes in model prices and capabilities, and contrasts verified with non-verified reviews to identify changes in platform activity associated with generative AI supply improvements. We apply this approach to more than 13 million reviews from Trustpilot, one of the leading online platforms for business reviews. A robust finding is that following LLM supply shocks, unverified reviews shift toward greater negativity: more 1-stars, fewer 5-stars, and lower ratings, with effects driven primarily by new model releases and concentrated among firms with the lowest and highest review volumes, suggesting that strategic AI use may reshape platform competition dynamics. We further find that LLM supply shocks trigger short, concentrated bursts of review activity. Together, these findings suggest that generative AI is already reshaping how reputation and competition operate on online platforms.
查看缓存全文
缓存时间: 2026/09/22 09:07
# 1Introduction
Source: [https://arxiv.org/html/2609.22198](https://arxiv.org/html/2609.22198)
###### Abstract
The rapid adoption of large language models \(LLMs\) creates new opportunities for strategic content generation on online platforms, including potentially harmful forms of manipulation that may undermine platform effectiveness and reshape platform dynamics\. However, measuring such activity is difficult because AI\-generated content is rarely directly observable\. We introduce an empirical approach that leverages discrete LLM supply shocks \- abrupt changes in model prices and capabilities, and contrasts verified with non\-verified reviews to identify changes in platform activity associated with generative AI supply improvements\. We apply this approach to more than 13 million reviews from Trustpilot, one of the leading online platforms for business reviews\. A robust finding is that following LLM supply shocks, unverified reviews shift toward greater negativity: more 1\-stars, fewer 5\-stars, and lower ratings, with effects driven primarily by new model releases and concentrated among firms with the lowest and highest review volumes, suggesting that strategic AI use may reshape platform competition dynamics\. We further find that LLM supply shocks trigger short, concentrated bursts of review activity\. Together, these findings suggest that generative AI is already reshaping how reputation and competition operate on online platforms\.
The Role of AI in Online Reviews
Valeria Lermana,∗, Oren Rigbia, and Yaniv Dovera,b,c
11footnotetext:The Hebrew University Business School, The Hebrew University of Jerusalem, Jerusalem 9190501, Israel22footnotetext:Department of Cognitive and Brain Sciences, Faculty of Humanities, The Hebrew University of Jerusalem, Jerusalem 9190501, Israel33footnotetext:The Federmann Center for the Study of Rationality, The Hebrew University of Jerusalem, Jerusalem 9190501, Israel11footnotetext:Corresponding author\. Email: valeria\.lerman@mail\.huji\.ac\.il
August 30, 2026
Keywords:Generative AI; Large Language Models; Online Reviews; Digital Platforms; User\-Generated Content
## 1Introduction
In recent years, generative artificial intelligence \(GenAI\) has diffused rapidly across a wide range of domains, transforming both the automation of digital tasks and the production of content\. Advances in large language models \(LLMs\) have enabled AI systems not only to generate coherent, contextually relevant, and human\-like text, but also, increasingly, to browse the internet, interact with digital platforms, and perform actions autonomously\. These capabilities have fundamentally altered the economics of text production by sharply reducing its marginal cost and enabling content to be generated at unprecedented scale and speed \([Brown et al\. \(2020\)](https://arxiv.org/html/2609.22198#bib.bib22);[Bommasani et al\. \(2021\)](https://arxiv.org/html/2609.22198#bib.bib23)\)\. As a result, text\-intensive activities that traditionally required substantial human effort, including drafting news and opinion articles, producing creative content, generating social media posts, and composing online reviews, can now be partially or fully automated\.
Empirical studies indeed show that AI\-generated text is often indistinguishable from human\-written content, particularly when evaluated by non\-expert readers \([Clark et al\. \(2021\)](https://arxiv.org/html/2609.22198#bib.bib24);[Kreps et al\. \(2022\)](https://arxiv.org/html/2609.22198#bib.bib25)\)\. Because content on digital platforms can significantly shape market dynamics and economic outcomes, this high degree of realism raises important questions about authenticity and its broader implications for information ecosystems\. The increasing presence of machine\-generated content blurs the boundary between human and automated expression, complicating the ability of businesses, users, and platforms to assess the origin and credibility of online information\.
One prominent narrative reflecting public concern is the so\-called “dead internet theory,” which suggests that a substantial portion of online content is already generated by automated systems rather than humans \([Baronio \(2025\)](https://arxiv.org/html/2609.22198#bib.bib65);[Down \(2025\)](https://arxiv.org/html/2609.22198#bib.bib66);[Murray \(2025\)](https://arxiv.org/html/2609.22198#bib.bib67);[Levy \(2026\)](https://arxiv.org/html/2609.22198#bib.bib68)\)\. Although this claim is sometimes raised in public debate, it lacks rigorous empirical support\. Nevertheless, it reflects broader concerns about the growing role of artificial agents in digital environments\. Even if the internet is far from “dead,” there is increasing evidence that automated content generation is becoming a growing component of online activity \([Ferrara et al\. \(2016\)](https://arxiv.org/html/2609.22198#bib.bib26);[Muzumdar et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib27);[Walter \(2025\)](https://arxiv.org/html/2609.22198#bib.bib28)\)\.
The implications of this shift are particularly salient in the context of user\-generated content \(UGC\), which plays a central role in modern digital economies\. Social media platforms, for example, are highly dependent on user contributions that influence public opinion, consumer behavior, and even political outcomes \([Allcott and Gentzkow \(2017\)](https://arxiv.org/html/2609.22198#bib.bib29)\)\. Integrating AI\-generated content into these ecosystems reduces dissemination costs, but increases the risk of low\-quality, misleading, or strategically manipulated content\. Similarly, informational platforms such as Wikipedia or question\-and\-answer forums depend on crowd\-sourced knowledge production\. Integration of AI tools may improve productivity, but could also affect content reliability and editorial norms \([Shin \(2026\)](https://arxiv.org/html/2609.22198#bib.bib30);[Park \(2024\)](https://arxiv.org/html/2609.22198#bib.bib31);[Marcellino et al\. \(2023\)](https://arxiv.org/html/2609.22198#bib.bib32)\)\.
The issue becomes even more pressing in the context of commercial content\. Online reviews, product descriptions, and marketing materials are essential inputs in consumer decision\-making processes \([Rachmiani et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib33);[Lackermair et al\. \(2013\)](https://arxiv.org/html/2609.22198#bib.bib34)\) and were repeatedly shown to affect important economic outcomes \([Pocchiari et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib48);[Alzate et al\. \(2021\)](https://arxiv.org/html/2609.22198#bib.bib49);[Huang and Pape \(2020\)](https://arxiv.org/html/2609.22198#bib.bib50)\)\. The incentives for businesses to generate such content at scale using AI without disclosing its origin are clear\. Although such practices are unethical, they offer an efficient way to shape their own and others’ online reputations, increase demand, and achieve other desirable outcomes\. If firms engage in the production of hard\-to\-identify AI\-generated online reviews, there is a considerable risk to market authenticity and transparency, and consequently the usefulness of online reviews in digital markets\. Moreover, Generative AI may alter competitive dynamics on commercial platforms in ways that further undermine the existing digital ecosystem\. These concerns warrant systematic scientific investigation\.
Here, we focus on online review platforms, which are central to the digital economy and play a key role in shaping consumer decisions and market outcomes, including demand, pricing, and firm reputation \([Qiu and Zhang \(2023\)](https://arxiv.org/html/2609.22198#bib.bib35);[Burton \(2024\)](https://arxiv.org/html/2609.22198#bib.bib36)\)\. However, their credibility may be undermined by fake and manipulated reviews, which firms may strategically use to increase ratings, improve brand image, or harm competitors \([Mayzlin et al\. \(2014\)](https://arxiv.org/html/2609.22198#bib.bib51);[Lim et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib37);[Martínez Otero \(2021\)](https://arxiv.org/html/2609.22198#bib.bib38);[He et al\. \(2022\)](https://arxiv.org/html/2609.22198#bib.bib39)\)\. These distort consumer signals and can produce inefficient market outcomes\.
Generative AI may significantly worsen these challenges by making it far cheaper and easier to produce large volumes of human\-like fake reviews\. Compared to earlier fake reviews, AI\-generated content has the potential to be more coherent, diverse, and contextually relevant, increasing both the scale of potential manipulation and the difficulty of detection \([Özaydın \(2025\)](https://arxiv.org/html/2609.22198#bib.bib40);[Gupta et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib41);[Zhao et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib42);[Meng et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib43);[Knight et al\. \(2023\)](https://arxiv.org/html/2609.22198#bib.bib44)\)\.
Some recent studies have attempted to distinguish AI\-generated online reviews from authentic human\-written reviews by identifying differences in their linguistic characteristics\. These studies find that AI\-generated reviews are often more readable and coherent but tend to be less specific, emotional, and empathetic, and tend to follow a more mechanical structure \([Zhao et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib42)\)\. However, these findings may not generalize across contexts and settings or persist over time, given the rapid and substantial advances in AI capabilities\. The findings also suggest that AI\-generated deception differs from traditional fake reviews and can evade detection methods that rely on psychological cues associated with human behavior\. Related research finds that in their case AI\-generated reviews are often higher\-rated, posted by users with limited platform history, and are more readable but less linguistically complex than authentic reviews and appear more common among lower\-traffic businesses \([Gambetti and Han \(2023b\)](https://arxiv.org/html/2609.22198#bib.bib45)\)\. Despite these studies, there is still no conclusive evidence that AI\-generated reviews can be systematically identified using textual markers\. On the contrary, other studies show that both humans and advanced language models have a hard time detecting AI\-generated reviews \([Meng et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib43);[Santos and Antonio \(2025\)](https://arxiv.org/html/2609.22198#bib.bib69);[Agrahari et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib70)\)\. Another strand of research attempts to develop more advanced detection methods, combining textual features with statistical signals such as outlier patterns in review distributions\. Although these approaches seem to somewhat improve performance, they still depend on model\-specific assumptions and imperfect ground\-truth proxies, underscoring the ongoing difficulty in identifying AI\-generated reviews at scale \([Luo et al\. \(2026\)](https://arxiv.org/html/2609.22198#bib.bib46);[Gambetti and Han \(2023a\)](https://arxiv.org/html/2609.22198#bib.bib47)\)\.
Taken together, the literature is mixed\. AI\-generated reviews may differ from human\-written ones, but evidence on their detectability and prevalence is inconsistent, in part because studies rely on different data, labels, and evaluation settings, often using synthetic samples\. This highlights the need to use novel approaches to better understand how firms use Generative AI tools in practice — whether for self\-promotion, competitive manipulation, or both—and with what consequences for market outcomes\.
To examine the real\-world use of generative AI for online review manipulation, we develop a novel empirical strategy that leverages rapidly changing exogenous variation in the cost and capabilities of AI\-generated text\. Specifically, we use data from Trustpilot, a large online platform hosting consumer reviews of firms and services for the period of 2023 \- 2024, and leverage OpenAI’s discrete API price reductions and the introduction of more affordable and more efficient model tiers as means of supply\-side shifts\. This abrupt reduction in the cost and improvement in the capabilities of generative AI provide a unique opportunity to study whether and how firms respond to changes in the economic feasibility of using generative AI for content production\. Both abrupt price reductions and new model launches provide plausibly exogenous shocks, as their precise timing is generally unknown in advance and is unlikely to be driven by the activities of Generative AI producers on Trustpilot\.111Although price reductions or model launches may occasionally be anticipated by a day or two, such limited advance notice is unlikely to threaten identification\. Bias would arise only if actors immediately changed activity before the official event and before the lower price and new capabilities are introduced, which does not seem likely in our context\.We focus on OpenAI because it was the leading provider in the LLM API market during that period[Tully et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib71);[Wang and Xu \(2024\)](https://arxiv.org/html/2609.22198#bib.bib72)\. In what follows, we use the term ‘LLM supply shocks’ to jointly refer to these API price reductions and the introduction of more affordable or efficient model tiers\.
Our approach leverages the exogeneity of unanticipated discrete price reductions and new language model introductions along with the verified\-reviews feature of the Trustpilot platform\. In particular, we exploit the distinction between verified and unverified reviews on the platform, in the spirit of[Mayzlin et al\. \(2014\)](https://arxiv.org/html/2609.22198#bib.bib51)and[Luca and Zervas \(2016\)](https://arxiv.org/html/2609.22198#bib.bib52)\. Verified reviews are linked to confirmed transactions or experiences and are therefore more likely to reflect genuine consumer activity\. Verified reviews form part of the substantial efforts that Trustpilot reports undertaking to detect and remove fake reviews \([Trustpilot \(2026a\)](https://arxiv.org/html/2609.22198#bib.bib53)\)\.222The platform also reports having flagged roughly 6% of submitted unverified reviews as fake in recent years, indicating an active commitment to identifying and removing fraudulent content, including potentially AI\-generated reviews\. Thus, one implication is that any AI\-generated activity detected in our analysis reflects activity that remains observable despite the platform’s filtering efforts\. We discuss the implications of this moderation process below\.In contrast, unverified reviews face fewer credibility constraints and are more susceptible to strategic manipulation, including the potential use of AI\-generated content\. Under this assumption, unverified reviews are more likely to respond to changes in the cost and capabilities of generative AI\.
By exploiting variation along two dimensions — immediate time \(shortly before vs\. after LLM supply shocks\) and review type \(verified vs\. unverified\) — we implement a difference\-in\-differences framework that compares changes over time in review characteristics across groups\. We focus on short time windows around the LLM supply shocks to isolate the effect of each of these events and reduce contamination from longer\-term effects\. This design allows us to isolate the differential impact of changes in the cost and capabilities of generative AI on reviews that are more likely to incorporate generative AI\. Our specifications include time and company fixed effects, so identification comes from within\-company changes around each LLM supply shock, and because verified reviews serve as a control group, the estimates absorb broader common shocks that affect both verified and unverified reviews, increasing confidence that any remaining effect reflects AI\-driven changes in unverified review activity\.
Our empirical analysis yields two central sets of findings\. First, we document systematic changes in the distribution of ratings following exogenous changes in the cost and capabilities of generative AI, driven by OpenAI price reductions and new model releases\. Across specifications, unverified reviews experience a statistically significant decline in average ratings relative to verified reviews in the seven days following the LLM supply shock\. This shift is driven by an increase in the share of one\-star reviews and a corresponding decrease in the share of five\-star reviews\. Although the estimated effects are modest in percentage\-point terms, they represent a lower bound on GenAI activity in online reviews\. Our approach captures only marginal responses to changes in model prices and capabilities, and we do not observe reviews that were flagged and removed by the platform\. We therefore cannot estimate the absolute prevalence of AI\-generated activity across the platform\. Given these limitations, the observed effects are nevertheless economically meaningful at Trustpilot’s scale and imply substantial shifts in the overall distribution of ratings\. Furthermore, we find that these effects are driven by the release of newer, more capable language models rather than by price reductions, and are strongest for reviews of the largest and smallest firms, as proxied by platform\-level activity\.
These results suggest that lower\-cost and more efficient generative AI is associated with an increase in negative review activity rather than positive self\-promotion\. A plausible interpretation is that firms use AI\-generated content strategically to target competitors rather than primarily to promote themselves\. The greater capabilities and efficiency of newer models may enable more sophisticated competitive strategies, while negative reviews may also carry a lower risk of detection than attempts to artificially inflate a firm’s own ratings, particularly given the platform’s intensive efforts to identify fraudulent activity\.
In addition, although data limitations prevent us from detecting an increase in aggregate review volume following price reductions, we do find a significant short\-term surge of review activity after the introduction of newer, more capable models\. This pattern suggests that AI tools are used not merely to generate reviews, but as part of a specific, concentrated review\-generation tactic\.
Taken together, these findings contribute to the growing literature on the economic and behavioral implications of generative AI by providing causal evidence of how improvements in AI capabilities and reductions in AI usage costs shape online information environments\. More specifically, our results provide some of the first empirical evidence that AI\-generated reviews are actively used in practice\. They further suggest that, at least on Trustpilot, such reviews are deployed primarily as a competitive strategy to harm rival firms rather than as a tool for enhancing firms’ own reputations\. This interpretation should, however, be tempered by the possibility that the platform’s aggressive filtering policies more effectively remove self\-promotional fake reviews, rendering them less observable even if they are present\. From a practical perspective, the findings raise important concerns for digital platforms and regulators, suggesting that cheaper and more accessible AI tools may intensify challenges related to review authenticity, platform trust, platform dynamics, and consumer welfare\.
The remainder of the paper is organized as follows: Section[2](https://arxiv.org/html/2609.22198#S2)describes the data sources, data preparation procedures, and descriptive statistics\. Section[3](https://arxiv.org/html/2609.22198#S3)presents the empirical methodology\. Section[4](https://arxiv.org/html/2609.22198#S4)reports the effects of LLM supply shocks on review ratings and review\-volume categories\. Section[5](https://arxiv.org/html/2609.22198#S5)presents robustness checks and additional heterogeneity analyses\. Finally, Section[6](https://arxiv.org/html/2609.22198#S6)concludes\.
## 2Data and Descriptive Statistics
This study combines data from two primary sources\. First, we collect reviews from the Trustpilot platform, which provides reviews of different services, both online and offline\. Second, we construct a dataset of OpenAI supply shock dates\. By combining these datasets, we are able to examine how changes in the cost and capabilities of generative AI affect online reviews\.
### 2\.1Data Sources
#### 2\.1\.1Trustpilot Reviews Data
The primary dataset used in our analysis consists of 13,818,281 reviews collected from the Trustpilot platform, which hosts public reviews of services\.333The data were obtained through Bright Data, a third\-party data provider, and are not publicly accessible\.From this dataset we utilize several key variables for our analysis: review creation date, review rating \(between 1 and 5\), review verification status, and company identifier\. The specific ways in which these variables are used in the empirical analysis are described in Section[3](https://arxiv.org/html/2609.22198#S3)\. We perform the analysis using reviews written between January 1, 2023, and December 31, 2024 \(inclusive\)\. It is also important to note that the Trustpilot data has already undergone platform\-level moderation and cleaning\.444According to Trustpilot’s transparency reports, the removal of fake reviews remains relatively stable year\-over\-year at approximately 6%, with the majority of such reviews \(about 82%\) identified through Trustpilot’s automated detection technology \([Trustpilot \(2024\)](https://arxiv.org/html/2609.22198#bib.bib1)\)\. Trustpilot attributes this capability to continuous investments in advanced detection systems, including machine learning and AI\-based models that leverage the platform’s growing volume of review data\. These systems analyze hundreds of data points and identify suspicious behavioral patterns and anomalies \(e\.g\., unusual reviewing activity or coordinated behavior\) to detect guideline violations\.
#### 2\.1\.2OpenAI LLM Supply Shock Dates
To construct the dataset of LLM supply shock events, we manually collected the dates on which OpenAI announced price reductions and/or released new models\. We focus on OpenAI\-related events because it was the dominant actor in the generative AI market during the 2023–2024 period, both in terms of technological leadership and widespread adoption \([Elad \(2024\)](https://arxiv.org/html/2609.22198#bib.bib16);[Bailyn \(2026\)](https://arxiv.org/html/2609.22198#bib.bib17);[Bernzweig \(2025\)](https://arxiv.org/html/2609.22198#bib.bib18);[Pahwa \(2026\)](https://arxiv.org/html/2609.22198#bib.bib19);[Fried \(2024\)](https://arxiv.org/html/2609.22198#bib.bib20)\)\. The release of ChatGPT and subsequent model iterations \(e\.g\., GPT\-4 and GPT\-4o\) drove rapid diffusion of AI technologies across industries, making OpenAI the primary source of market\-wide pricing shocks \([Raman et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib54);[Zhang and Shao \(2024\)](https://arxiv.org/html/2609.22198#bib.bib55)\)\. As a result, price changes by OpenAI are likely to capture the most economically meaningful variation in AI costs during our sample period, whereas competing providers had more limited adoption or entered the market later\. Since we could not locate any single centralized source that lists all historical supply shocks, we compiled this information from multiple publicly available sources that we gathered after a comprehensive search\. The primary sources include official announcements and updates published on OpenAI’s website, as well as reports from artificial intelligence–focused technology blogs \([OpenAI Developer Community \(2023\)](https://arxiv.org/html/2609.22198#bib.bib2);[Schwartz \(2023\)](https://arxiv.org/html/2609.22198#bib.bib3);[OpenAI \(2023\)](https://arxiv.org/html/2609.22198#bib.bib4);[OpenAI \(2024c\)](https://arxiv.org/html/2609.22198#bib.bib5);[OpenAI \(2024b\)](https://arxiv.org/html/2609.22198#bib.bib6);[OpenAI \(2024a\)](https://arxiv.org/html/2609.22198#bib.bib7);[Mudaliar \(2024\)](https://arxiv.org/html/2609.22198#bib.bib8);[H \(2024\)](https://arxiv.org/html/2609.22198#bib.bib9);[OpenAI Developers \(2024\)](https://arxiv.org/html/2609.22198#bib.bib10);[Shariss \(2024\)](https://arxiv.org/html/2609.22198#bib.bib11);[Clark \(2023\)](https://arxiv.org/html/2609.22198#bib.bib13)\)\. To ensure comprehensive coverage, we conducted multiple systematic searches across OpenAI’s official communications \(including blog posts, release notes, and developer updates\) as well as leading AI\-focused technology outlets\. We cross\-validated information across sources and compared overlapping reports to minimize omissions and ensure consistency in the documented pricing changes\.
Our working assumption is that these sources report LLM supply shocks simultaneously or shortly after they are announced, as such changes are typically communicated publicly and quickly disseminated within the AI community\. Using these sources, we identified the relevant timings of LLM supply shocks, and constructed a timeline of events that we use as exogenous shocks in our empirical analysis\. Table[1](https://arxiv.org/html/2609.22198#S2.T1)summarizes the dates of LLM supply shocks included in the analysis, along with the corresponding posted input and output prices for the affected models, while Figure[1](https://arxiv.org/html/2609.22198#S2.F1)illustrates the evolution of these prices over time after normalizing each model’s initial observed price to 100%\.
Table 1:OpenAI Model Pricing Reductions and ReleasesDateModelInput price\(1000 tokens, $\)Output price\(1000 tokens, $\)Model typeEvent type2023\-03\-01ChatGPT 3\.5 Turbo 4K0\.0020\.002Language modelPrice reduction2023\-06\-13ChatGPT 3\.5 Turbo 4K0\.00150\.002Language modelPrice reductionChatGPT 3\.5 Turbo 16K0\.0030\.004Language modelPrice reductiontext\-embedding\-ada\-002a0\.0001N/AbEmbedding modelPrice reduction2023\-11\-06ChatGPT 3\.5 Turbo 16K0\.0010\.002Language modelPrice reductionChatGPT 3\.5 Turbo 4K FT0\.0030\.006Language modelPrice reductionGPT\-4\-Turbo 128K0\.010\.03Language modelNew model release2024\-01\-24ChatGPT 3\.5 Turbo 4K0\.00050\.0015Language modelPrice reductiontext\-embedding\-3\-smalla0\.00002N/AbEmbedding modelNew model release2024\-05\-13GPT\-4o0\.0050\.015Language modelNew model release2024\-07\-18GPT\-4o mini0\.000150\.0006Language modelNew model release2024\-08\-09GPT\-4o0\.00250\.01Language modelPrice reduction2024\-09\-12o1\-mini0\.0030\.012Language modelNew model release2024\-10\-30gpt\-realtimeN/AbN/AbRealtime modelPrice reduction2024\-12\-18GPT\-4o mini realtime \(text\)0\.00060\.0024Realtime modelNew model release
- •Notes:The table summarizes key OpenAI API pricing reductions and model introductions over 2023–2024\. For each event, we report the date, model, input and output token prices, model type, and whether the event corresponds to a price reduction or a new model release\. These events form the basis for the LLM supply shocks used in the empirical analysis\.
- aEmbedding models process input text into vector representations and therefore do not generate text output\. As a result, they have an input price but no corresponding output price\.
- bFor this event, the price reduction was implemented through prompt caching rather than a change in the posted base token price\. Cached text inputs were discounted by 50%, and cached audio inputs were discounted by 80%\.
Figure 1:OpenAI Model Pricing Reductions and Releases, 2023–2024 \(normalized to initial price = 100%\)
Notes:This figure is based on the pricing data reported in Table[1](https://arxiv.org/html/2609.22198#S2.T1)and on the raw \(non\-normalized\) price series presented in Figure[S1](https://arxiv.org/html/2609.22198#Sx1.F1)\. Prices are normalized to 100% at the initial observed price for each model\. The horizontal axis displays calendar dates over the period 2023–2024, while the vertical axis reports normalized prices \(in percentage terms\)\. The figure consists of two panels: the upper panel shows input prices and the lower panel shows output prices\. Colors correspond to different models, with similar color schemes used for models belonging to the same family\. The values reported in parentheses in the legend denote benchmark performance scores for each model\. Specifically, they correspond to Massive Multitask Language Understanding \(MMLU\) scores for language models and Massive Text Embedding Benchmark \(MTEB\) scores for embedding models\. These benchmark scores are included only to provide an approximate indication of each model’s capabilities\. Vertical dashed lines indicate the dates of the LLM supply shocks used in the empirical analysis and listed in Table[1](https://arxiv.org/html/2609.22198#S2.T1)\. The red dashed vertical line denotes the LLM supply shock associated with thegpt\-realtimemodel; since detailed pricing information is unavailable, this event is indicated without corresponding price series\. Solid circular markers denote the pricing observations included in the analysis, namely prices corresponding to price reductions or to the release of new, more efficient models\. For some models, such as GPT\-3\.5 Turbo FT and text\-embedding\-ada\-002, the initial release price is shown in the series but is not marked with a solid point because it does not correspond to a price reduction or a lower\-priced model release and was therefore not included in the analysis\. In addition, the release of text\-embedding\-ada\-002 predates the sample period and is not displayed, as the figure is restricted to dates from January 1, 2023 onward\. Embedding models do not report output prices, as they generate vector representations rather than textual outputs\. Finally, slight horizontal jitter is applied to improve visual separation between overlapping series\.
In order to conduct the empirical analysis, we use the Trustpilot platform review data to construct the online reviews activity around the OpenAI supply shocks\. The detailed procedures used to prepare the dataset are described in the following section\.
### 2\.2Data Preparation
#### 2\.2\.1Data Filtration and Aggregation
For the purpose of the analyses, we keep only reviews with valid creation dates and valid ratings \(i\.e\., that are between 1 and 5\)\. This process results in the exclusion of 270 reviews from the 2023–2024 sample which includes 13\.8M reviews\. In addition, our analysis includes only reviews written within the relevant seven days before and seven days after each event’s time windows which amounts to about 2\.8M reviews\. For better interpretability, we excluded the event day itself\. We aggregated review\-level data at the company–date\-verification level\. Specifically, observations are grouped by company identifier, review verification status \(verified or non\-verified\), and the date on which the review was written\. As a result, each observation in the aggregated dataset corresponds to the set of all reviews posted for a given company on a specific date with a given verification status\.
For each company–date–verification status group, we compute several statistics that we will use as the outcome variables that characterize review provision activity\. For these statistics, we assume that a company began its activity on the platform when the first of its reviews were recorded in the data\. The outcome variables are: the total number of reviews, the average review rating, the count and proportion of reviews in each rating category \(from 1 to 5\), the average number of words per review, the count of daily review surges \(batches of reviews per company, verification status, and day; see Section[4\.1\.2](https://arxiv.org/html/2609.22198#S4.SS1.SSS2)for a detailed description of these measures\), and a measure of review text homogeneity which captures the degree of the similarity among reviews within a company\-date\-verification group\.
### 2\.3Descriptive Statistics
The summary statistics for the main variables in our Trustpilot dataset are reported in Table[2](https://arxiv.org/html/2609.22198#S2.T2)\. The table presents descriptive statistics for both the full dataset of reviews collected between January 1, 2023, and December 31, 2024, and the event\-window sample used in our main analyses\. The full dataset contains 13,818,281 observations corresponding to 136,750 companies, whereas the event\-window sample, which includes reviews posted during the seven days before and the seven days after each LLM supply shock \(excluding the event day itself\), contains 2,797,199 observations from 72,298 companies\.
Table 2:Summary Statistics for the Full Sample and the Event\-Window Analysis SampleStatisticFull sampleEvent\-window sampleNumber of observations13,818,2812,797,199Number of unique companies136,75072,298Number of unique review dates731140Mean review rating4\.184\.19Standard deviation of review rating1\.471\.47Average number of reviews across all companies per day18,903\.2619,979\.99Average number of reviews per company101\.0538\.69Average review length \(words\)33\.2432\.92Unverified reviews \(%\)40\.2040\.17Verified reviews \(%\)59\.8059\.83Rating: 1 \(%\)15\.1014\.97Rating: 2 \(%\)2\.342\.29Rating: 3 \(%\)3\.433\.44Rating: 4 \(%\)7\.587\.77Rating: 5 \(%\)71\.5471\.53
The distribution of review ratings is strongly skewed toward positive evaluations\. In the full sample, the mean review rating is 4\.18 \(SD = 1\.47\), with 5\-star and 1\-star reviews accounting for 71\.54% and 15\.10% of all reviews, respectively\. Ratings of 2, 3, and 4 stars are considerably less common, representing 2\.34%, 3\.43%, and 7\.58% of the sample\. This distribution is consistent with the well\-documented polarization of online review ratings reported in the literature \([Hu et al\. \(2009\)](https://arxiv.org/html/2609.22198#bib.bib56);[Chevalier and Mayzlin \(2006\)](https://arxiv.org/html/2609.22198#bib.bib57);[Tanase et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib58)\)\. Importantly, the event\-window sample used in our analyses exhibits nearly identical characteristics, with very similar rating distributions, average rating, and review length\. This suggests that the observations included in the event windows are highly representative of the overall population of reviews during the study period\.
Table[2](https://arxiv.org/html/2609.22198#S2.T2)also shows that verified reviews account for 59\.80% of the full dataset, while unverified reviews account for the remaining 40\.20%\. The corresponding proportions in the event\-window sample are virtually identical \(59\.83% and 40\.17%, respectively; see Section[3](https://arxiv.org/html/2609.22198#S3)for a detailed description of the verification process\)\. This similarity is advantageous for our empirical design, as it indicates that the event\-window sample closely resembles the full dataset while allowing verified reviews to serve as a meaningful control group\. Finally, reviews contain an average of 33\.24 words in the full sample and 32\.92 words in the event\-window sample,555The calculation of the average review length is based only on non\-empty reviews\.suggesting that reviews on the platform tend to be relatively concise regardless of the sample considered\.
Figure[2](https://arxiv.org/html/2609.22198#S2.F2)illustrates the evolution of review activity over time with a daily resolution during the 2023–2024 period\. The gray line represents the daily number of reviews and the black line shows a smoothed trend line using a LOESS procedure\. The figure reveals a moderate and relatively stable increase in review activity over the sample period, with no apparent anomalous shifts in platform\-wide review volume\.
Figure 2:Daily Number of Reviews Over Time \(2023–2024\)
Notes:The grey line shows the raw daily number of reviews, while the black line represents the smoothed trend\. The figure is based onN=13,818,281N=13\{,\}818\{,\}281observations\.
## 3Methodology
This section outlines the empirical strategy used to estimate the effect of OpenAI’s supply shocks on the characteristics of online reviews\. Leveraging temporal variation in AI usage costs and model quality and efficiency, we employ a difference\-in\-differences design that compares changes in outcome variables for unverified reviews relative to verified reviews before versus after each LLM supply shock\.
### 3\.1Baseline Difference\-in\-Differences Analysis
To estimate the effect of LLM supply shocks on online reviews, we implement a difference\-in\-differences \(DiD\) regression with company and time fixed effects\. The objective is to measure how review properties change following changes in the cost and capabilities of OpenAI models\. Importantly, the baseline specification treats all LLM supply shocks as equivalent events, regardless of their magnitude or any concurrent changes in model quality or performance\. In this approach, we compare changes over time \(before\-versus\-after an event\) between verified and unverified reviews within a company\. In additional analyses presented in Section[4](https://arxiv.org/html/2609.22198#S4), we relax this assumption and examine heterogeneity across different types of events, distinguishing between abrupt price reductions, new lower\-priced more\-efficient model releases, and specifications that combine both categories of shocks\.
Our identification strategy relies on the assumption that unverified reviews are much more likely to be generated or assisted by AI tools, while verified reviews are linked to specific consumers with a confirmed genuine experience with the business who opt to review\. Verified reviews are therefore considerably less likely to use AI\-generated content created using API \([Trustpilot \(2026b\)](https://arxiv.org/html/2609.22198#bib.bib12)\)\. Under this assumption, verified reviews serve as a control group, while unverified reviews constitute the treatment group that is likely to respond to changes in AI usage costs\.
For each event, we construct an event window that includes the seven days prior to the LLM supply shock and the seven days following the supply shock\. The event date itself is excluded from the analysis\. We use a relatively narrow seven\-day window to limit the influence of potentially confounding events occurring around the LLM supply shocks\. Extending the window to longer periods, such as 10 or 14 days, would increase the likelihood that other contemporaneous events unrelated to the focal supply shock affect review activity, making it more difficult to attribute observed changes to the shock itself\. In addition, longer windows would result in overlap between the event windows of some LLM supply shocks in our sample, potentially confounding the estimated effects of distinct shocks\. The seven\-day window therefore provides a balance between allowing sufficient time to capture changes following each shock and maintaining a sufficiently narrow window to isolate its effect\. The analysis uses the aggregated outcome and predictor variables for the before and after periods\. Specifically, we estimate the following difference\-in\-differences specification:
Ycvt=αc\+γw\(t\)\+β1Treatmentv\+β2Aftert\+β3\(Aftert×Treatmentv\)\+εcvtY\_\{cvt\}=\\alpha\_\{c\}\+\\gamma\_\{w\(t\)\}\+\\beta\_\{1\}Treatment\_\{v\}\+\\beta\_\{2\}After\_\{t\}\+\\beta\_\{3\}\(After\_\{t\}\\times Treatment\_\{v\}\)\+\\varepsilon\_\{cvt\}\(1\)
whereYcvtY\_\{cvt\}denotes a review \(aggregated\) outcome for companycc, verification statusvv, and datett\(which could be either before or after the event\)\. The outcome variables include: \(1\) the number of reviews, \(2\) the average rating of the reviews, \(3\) the number of reviews in each rating category \(1–5\), \(4\) the share of reviews in each rating category \(1\-5\), \(5\) the average number of words per review, \(6\) a measure of review text homogeneity and \(7\) three indicator variables for review\-volume groups, equal to one when the number of reviews falls into the ranges 0–19, 20–99, and 100 or more, respectively\. The variableAftertAfter\_\{t\}is a binary indicator equal to one for observations occurring after the LLM supply shocks, within the event window, and zero otherwise\. The variableTreatmentvTreatment\_\{v\}is a binary indicator equal to one for unverified reviews and zero for verified reviews\.
The specification includes company fixed effects \(αc\\alpha\_\{c\}\) to control for time\-invariant differences across companies, and calendar week fixed effects \(γw\(t\)\\gamma\_\{w\(t\)\}\) to account for common time shocks affecting all companies\. The coefficientβ1\\beta\_\{1\}captures pre\-treatment differences between unverified and verified reviews, whileβ2\\beta\_\{2\}reflects the average change in outcomes after the LLM supply shock for the reference group \(verified reviews\)\. The main parameter of interest,β3\\beta\_\{3\}, measures the differential change over time \(after vs\. before\) in review characteristics for unverified reviews relative to verified reviews, following the OpenAI supply shocks\. The error termεcvt\\varepsilon\_\{cvt\}captures unobserved determinants of the outcome\. Standard errors are clustered at the company level, consistent with the assumption that shocks affecting a given company may be persistent, while remaining independent across companies\.
### 3\.2Event\-study specification
In addition to the aggregated before–after baseline specification, we estimate an event\-study specification that allows the treatment effect to vary across individual days surrounding the LLM supply shocks\. Whereas the baseline specification compares review activity over the entire seven\-day periods before and after each event, the event\-study specification estimates separate effects for each day within the event window\.
Specifically, we use daily outcome variables for the fourteen days surrounding each event: seven days before and seven days after the event\. This higher\-frequency specification allows us to examine the dynamics of review activity within the event window, including whether the response emerges immediately following the event or evolves gradually over subsequent days\. It also allows us to assess the identifying parallel\-trends assumption by examining differential dynamics between verified and unverified observations during the pre\-event period\.
Each observation is indexed by companycc, verification statusvv, and calendar datett\. LetTreatmentv\\mathrm\{Treatment\}\_\{v\}be an indicator equal to one for unverified observations and zero for verified observations\. We defineDtkD\_\{t\}^\{k\}as an indicator equal to one when calendar datettiskkdays relative to the event date\. The event\-study window is given by:
k∈\{−7,−6,…,−1,1,…,7\},k\\in\\\{\-7,\-6,\\ldots,\-1,1,\\ldots,7\\\},
where the event day,k=0k=0, is excluded\.
Rather than selecting a single pre\-event day as the reference period, we normalize the coefficients such that their average over the seven pre\-event days equals zero\. We impose this normalization separately on the event\-time coefficients,λk\\lambda\_\{k\}, and the treatment\-by\-event\-time interaction coefficients,δk\\delta\_\{k\}:
17∑k=−7−1λk=0,17∑k=−7−1δk=0\.\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\lambda\_\{k\}=0,\\qquad\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\delta\_\{k\}=0\.
Under this normalization, the event\-study specification is:
Ycvt=\\displaystyle Y\_\{cvt\}=\{\}αc\+γw\(t\)\+θTreatmentv\\displaystyle\\alpha\_\{c\}\+\\gamma\_\{w\(t\)\}\+\\theta\\,\\mathrm\{Treatment\}\_\{v\}\(2\)\+∑k=−7−2λk\(Dtk−Dt−1\)\+∑k=17λkDtk\\displaystyle\+\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\+\\sum\_\{k=1\}^\{7\}\\lambda\_\{k\}D\_\{t\}^\{k\}\+∑k=−7−2δkTreatmentv\(Dtk−Dt−1\)\+∑k=17δkTreatmentvDtk\+εcvt\.\\displaystyle\+\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\mathrm\{Treatment\}\_\{v\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\+\\sum\_\{k=1\}^\{7\}\\delta\_\{k\}\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{k\}\+\\varepsilon\_\{cvt\}\.Here,YcvtY\_\{cvt\}denotes the outcome for companycc, verification statusvv, and datett;αc\\alpha\_\{c\}denotes company fixed effects; andγw\(t\)\\gamma\_\{w\(t\)\}denotes calendar\-week fixed effects, wherew\(t\)w\(t\)is the calendar week containing datett\. Standard errors are clustered at the company level\.
Because the pre\-event coefficients are normalized to have mean zero,θ\\thetacaptures the average difference between unverified and verified observations over the seven pre\-event days\. The coefficientsλk\\lambda\_\{k\}describe the event\-time dynamics for verified observations, the reference group, relative to their average pre\-event event\-time effect\. The interaction coefficientsδk\\delta\_\{k\}, which are the primary coefficients of interest, capture how the difference between unverified and verified observations on relative daykkdeviates from the average difference between the two groups during the seven\-day pre\-event period\.
The coefficients for relative day−1\-1,λ−1\\lambda\_\{\-1\}andδ−1\\delta\_\{\-1\}, are not estimated directly\. Under the normalization above, they are recovered as
λ^−1=−∑k=−7−2λ^k,δ^−1=−∑k=−7−2δ^k\.\\hat\{\\lambda\}\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\hat\{\\lambda\}\_\{k\},\\qquad\\hat\{\\delta\}\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\hat\{\\delta\}\_\{k\}\.
Consequently, the estimated pre\-event coefficients satisfy
17∑k=−7−1λ^k=0,17∑k=−7−1δ^k=0\.\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\hat\{\\lambda\}\_\{k\}=0,\\qquad\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\hat\{\\delta\}\_\{k\}=0\.
To assess the parallel\-trends assumption, we test whether the differential dynamics between unverified and verified observations are jointly zero during the pre\-event period\. Given the normalization, this is equivalent to testing the six freely estimated pre\-event interaction coefficients:
H0:δ−7=δ−6=⋯=δ−2=0\.H\_\{0\}:\\delta\_\{\-7\}=\\delta\_\{\-6\}=\\cdots=\\delta\_\{\-2\}=0\.
If these six coefficients are jointly equal to zero, the normalization implies thatδ−1=0\\delta\_\{\-1\}=0as well\. The joint test therefore evaluates whether there is evidence of systematic differential dynamics between unverified and verified observations prior to the event\.
The derivation of the average pre\-event normalization and the corresponding reparameterization of the regression specification are provided in Section[S1\.3](https://arxiv.org/html/2609.22198#S1.SS3)\.
### 3\.3Text Homogeneity Measure
In addition to standard review\-level outcomes, we construct a measure of textual homogeneity that captures the degree of similarity across reviews within a given group\. This measure is intended to proxy for the extent to which review content is standardized or exhibits similar linguistic patterns\. In other words, it allows us to test whether increased use of generative AI may alter the general textual properties of online reviews\. We expect AI\-generated reviews to exhibit higher levels of textual homogeneity, following prior research suggesting that AI\-generated texts tend to display more uniform linguistic structures, more repetitive lexical patterns, and lower stylistic variation than human\-written texts \([Kujur \(2025\)](https://arxiv.org/html/2609.22198#bib.bib59);[Culda et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib60);[Muñoz\-Ortiz et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib61)\)\.
For each event, we group reviews by company, verification status, and event\-time period, distinguishing between reviews posted during the seven days preceding the event and those posted during the seven days following the event\. We retain groups containing at least 10 reviews and exclude reviews shorter than 5 characters to remove trivial or non\-informative text entries\. For groups containing more than 100 reviews, we randomly sample 100 reviews to maintain computational tractability\.666Only 3\.34% of the groups contain more than 100 reviews and therefore require sampling; all remaining eligible groups are used in full\.
We convert each review text into a vector representation using the pre\-trained all\-MiniLM\-L6\-v2 Sentence Transformer model\. The resulting 384\-dimensional embeddings map reviews into a semantic vector space in which reviews with more similar content are located closer to one another\. Within each group, we compute a centroid embedding, defined as the average of all review embeddings in that group\. For each review, we then calculate its cosine similarity to the group centroid\. Higher cosine similarity indicates that a review is more similar to the typical review in its group and, consequently, that the group’s review content is more homogeneous\.
Finally, for each group, we compute the mean and standard deviation of the cosine similarity scores\. The mean captures the overall level of textual homogeneity, whereas the standard deviation captures the dispersion in textual similarity within the group\.
## 4Results
### 4\.1Baseline Difference\-in\-Differences Analysis
#### 4\.1\.1The Effect on Rating Distribution
Table[3](https://arxiv.org/html/2609.22198#S4.T3)shows the effect of OpenAI supply shocks on the average rating and on the proportions of 1\-star and 5\-star ratings\. We focus on these three metrics because they provide the clearest evidence of changes in review patterns\. We also estimate the same specifications for the 2\-, 3\-, and 4\-star ratings, and also using other standard review metrics, including raw review counts, the logarithm of review counts, review length, and measures of review text homogeneity; however, the estimates of the effects are not statistically significant across specifications\. For completeness, the corresponding results are reported in Tables[S1](https://arxiv.org/html/2609.22198#S1.T1)and[S2](https://arxiv.org/html/2609.22198#S1.T2)\. The coefficient of theTreatmentvariable is large and statistically significant across all specifications, indicating that, prior to the LLM supply shocks, unverified reviews differ from verified reviews in that they exhibit lower average ratings, which is driven by a higher proportion of 1\-star ratings and a lower proportion of 5\-star ratings\. Several mechanisms may explain the greater negativity of unverified reviews, including dissatisfied consumers’ stronger desire for anonymity, positive selection by businesses in encouraging satisfied customers to leave verified reviews, or other unobserved factors\. Our empirical setup is designed to account for these baseline differences\. The coefficient of theAftervariable is small and not statistically significant across all outcomes, indicating that verified reviews — the reference group — were not significantly affected by the LLM supply shock events\. In other words, the behavior of verified reviewers remains broadly stable before and after the shock\. Turning to the effect of interest, the coefficient of the interactionTreatment × Afterfor the average rating is statistically significant\. This result indicates that, following OpenAI supply shocks, and given all the controls in this setup, while the average rating did not significantly change for verified reviews, it declined for unverified reviews relative to verified reviews\. In other words, we find that a reduction in production costs or an improvement in the capabilities of generative AI models increases the average negativity of ratings directed at businesses on the Trustpilot platform\.
Table 3:Difference\-in\-Differences Results: Seven\-Day Event Window\(1\)\(2\)\(3\)Avg\. ratingProp\. 1\-starProp\. 5\-starTreatment\-1\.0403\*\*\*0\.2648\*\*\*\-0\.2268\*\*\*\(0\.022\)\(0\.005\)\(0\.006\)After0\.0045\-0\.00130\.0006\(0\.005\)\(0\.001\)\(0\.002\)Treatment×\\timesAfter\-0\.0171\*\*\*0\.0034\*\*\*\-0\.0046\*\*\*\(0\.005\)\(0\.001\)\(0\.001\)Company FEYesYesYesWeek FEYesYesYesMean \(After = 0\)3\.7630\.2640\.628R\-squared0\.5750\.5550\.499Observations748,513748,513748,513
\(4\)\(5\)\(6\)Low\-volume \[0–20\)Moderate\-volume \[20–100\)High\-volume 100\+Treatment0\.001329\*\*\*\-0\.001155\*\*\*\-0\.000174\*\*\*\(0\.00014\)\(0\.00013\)\(0\.00004\)After0\.000234\*\*\*\-0\.000222\*\*\*\-0\.000013\(0\.00003\)\(0\.00003\)\(0\.00001\)Treatment×\\timesAfter\-0\.000040\*0\.0000370\.000002\(0\.00002\)\(0\.00002\)\(0\.00001\)Company FEYesYesYesWeek FEYesYesYesMean \(After = 0\)0\.99870\.00120\.0001R\-squared0\.2680\.2440\.216Observations16,116,18616,116,18616,116,186
- •Notes:This table reports Difference\-in\-Differences estimates for review ratings and review\-volume categories within the seven\-day event window\. Columns \(1\)–\(3\) report results for the average review rating and the proportions of 1\-star and 5\-star reviews, respectively\. Columns \(4\)–\(6\) report results for indicators equal to one if a company–day–verification\-status observation belongs to the low\-volume \(0–19 reviews\), moderate\-volume \(20–99 reviews\), or high\-volume \(100 or more reviews\) category, respectively, and zero otherwise\. Company–day–verification\-status observations with zero reviews are retained in the sample and classified in the low\-volume category\. The coefficients in Columns \(4\)–\(6\) can therefore be interpreted as changes in the probability that an observation belongs to the corresponding review\-volume category\. The coefficient on the interaction term \(Treatment×\\timesAfter\) captures the differential change following the LLM supply shock for unverified reviews relative to verified reviews\. For the rating outcomes, the interaction coefficient is negative and statistically significant for average review rating and the proportion of 5\-star reviews, and positive and statistically significant for the proportion of 1\-star reviews\. For the review\-volume categories, the interaction coefficient is negative and marginally significant for the low\-volume category, while no statistically significant differential effects are observed for the moderate\- and high\-volume categories\. All regressions include company and week fixed effects\. Standard errors are clustered at the company level and reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
This pattern is consistent with the effect on the distribution of ratings\. Specifically, the proportion of one\-star reviews increases significantly following LLM supply shocks for unverified reviews relative to verified reviews, with an estimated coefficient of 0\.0034 \(p\-value < 0\.01\), corresponding to an increase of 0\.34 percentage points\. Conversely, the proportion of five\-star reviews decreases significantly, with a coefficient of \-0\.0046 \(p\-value < 0\.01\), implying a decline of 0\.46 percentage points\. It is important to note that although the magnitude of these effects may appear small in percentage point terms, they are economically meaningful given the scale of the platform\. Even modest changes in proportions translate into a substantial number of reviews when aggregated over millions of observations and for over 72,000 companies and businesses\. As such, these shifts can meaningfully influence the overall distribution of ratings, potentially affecting consumer perceptions, firm reputation, and competitive dynamics on the platform\.
Two important observations are required to interpret these findings\. First, the platform enacts a relatively strict policy of filtering what it believes to be fake reviews \(around 6% are filtered each year\)\. A successful platform\-filtering process would presumably leave a negligible amount of AI\-generated reviews\. Therefore, our results capture the residual AI\-generated activity that remains after Trustpilot’s aggressive filtering process has already taken place\. Second, from an economic perspective, AI\-generated online reviews mostly serve two broad strategic purposes: they can be used either to artificially enhance the reputation of one’s own or affiliated businesses, or to strategically damage competing businesses by generating unfavorable evaluations\. Our results suggest that some AI\-generated reviews manage to bypass the fake review detection mechanisms of the platform, although we cannot rule out that what we see is a result of some interaction between the platform and the AI\-generated reviews activity \(e\.g\., that changes in AI supply affects the filtering process, in some way\)\. Assuming these patterns are indeed AI\-generated reviews that evaded platform detection, it seems then that the availability of cheaper or more efficient generative AI tools is used more for competitive strategic purposes rather than self\-enhancing\. One potential explanation is that artificially inflating one’s own ratings exposes a business to a higher risk of detection and punishment by platform moderation systems, whereas posting negative reviews targeting competitors can achieve a similar effect, by widening the rating gap, while carrying a lower risk of being detected\. It may also suggest that AI’s enhanced capabilities make it easier to generate credible negative reviews than positive ones\. This strategy may be especially lucrative, given prior evidence that negative reviews exert a stronger influence on economic outcomes\. This behavior, if it exists on other platforms, could amplify the competitive dynamics on digital platforms, harm the quality of consumer\-generated information on platforms, and increase monitoring costs as firms may increasingly rely on automated content generation to influence perceived product quality and reputation in online review environments\.
#### 4\.1\.2The Effect on Review\-Volume Groups
As stated in the previous section, the review counts did not exhibit statistically significant effects\. Moreover, as shown below in the event\-study analysis, the temporal evolution of review counts does not appear to satisfy the parallel\-trends assumption\. This motivates the question of whether AI\-generated reviews are generated in a manner not detectable using a naive counts measure\. We therefore investigate an alternative form of review\-volume activity: abnormally large numbers of reviews posted during the same day, for a given company and verification status, which may be an indication of coordinated or automated review generation\.
To study this phenomenon, we encode three dummy variables corresponding to different levels of within\-day review volume based on the number of reviews received on that date:low review volume\(0–19 same\-company reviews\),moderate review volume\(20–99 same\-company reviews\), andhigh review volume\(100 or more same\-company reviews\)\. These thresholds were chosen to capture increasingly high levels of review activity while preserving a sufficient number of observations within each group\. We tested several alternative cutoffs, and the results remain qualitatively unchanged\.
Table 4:Descriptive Statistics of Review\-Volume CategoriesStatisticLow volume \[0–20\)Moderate volume \[20–100\)High volume 100\+Number of company–day–verification observations16,095,83418,6251,727Number of unique companies72,2911,103177Share of company–day–verification observations \(%\)99\.870\.120\.01Total reviews1,716,188685,022395,989Share of reviews \(%\)61\.3524\.4914\.16Mean reviews per company–day–verification observation0\.1136\.78229\.29Median reviews per company–day–verification observation0\.0031\.00159\.00Standard deviation of reviews0\.7617\.39250\.6125th percentile reviews0\.0024\.00121\.5075th percentile reviews0\.0044\.00243\.50Verified reviews \(%\)47\.3876\.5784\.87Unverified reviews \(%\)52\.6223\.4315\.13
- Notes:This table reports descriptive statistics for the three review\-volume categories used in the analysis\. The unit of observation is a company–day–verification\-status observation within the event\-window panel\. Low\-volume observations correspond to 0–19 reviews, moderate\-volume observations correspond to 20–99 reviews, and high\-volume observations correspond to 100 or more reviews\. Company–day–verification\-status observations with zero reviews are retained in the sample and classified in the low\-volume category\. Verified and unverified review percentages are calculated based on the number of reviews within each review\-volume category\. The sample is restricted to review dates in 2023–2024\.
Table[4](https://arxiv.org/html/2609.22198#S4.T4)reports descriptive statistics for the three review\-volume categories\. The descriptive statistics are based on the panel used in the analysis and therefore include company–day–verification\-status observations with zero reviews, which are classified in the low\-volume category\. Consequently, low\-volume observations account for 99\.87% of all company–day–verification\-status observations but only 61\.35% of total reviews, reflecting that most observations contain few or no reviews\. In contrast, moderate\- and high\-volume observations are relatively rare, representing only 0\.12% and 0\.01% of all observations, respectively, yet together account for nearly 39% of all reviews in the sample\. High\-volume observations are particularly concentrated: although they comprise only 0\.01% of company–day–verification\-status observations, they account for 14\.16% of all reviews\. The table also shows that verified reviews are disproportionately represented among moderate\- and high\-volume observations, whereas unverified reviews are relatively more prevalent among low\-volume observations\.
The estimation results for these three review\-volume outcome variables are reported in the second panel of Table[3](https://arxiv.org/html/2609.22198#S4.T3)\.777We also estimated nonlinear specifications \(logit and binomial models\) for these outcomes\. However, due to the inclusion of a large number of fixed effects \(e\.g\., company and time effects\), these models face substantial econometric and computational challenges\. In particular, fixed effects estimators in nonlinear panel models such as logit and probit are known to suffer from the incidental parameters problem, which can lead to severe bias and unreliable inference \([Cruz\-Gonzalez et al\. \(2017\)](https://arxiv.org/html/2609.22198#bib.bib14);[Hahn and Kuersteiner \(2011\)](https://arxiv.org/html/2609.22198#bib.bib15)\)\. As a result, we adopt a linear fixed effects specification, which remains computationally tractable and provides a reasonable approximation to average partial effects \(as discussed in\([Wooldridge, 2010](https://arxiv.org/html/2609.22198#bib.bib21), p\. 563\)\)\.
As in the previous specifications, the coefficient of theTreatmentindicator shows that the distribution of review volume differs systematically between unverified and verified reviews in the pre\-treatment period\. In particular, low\-volume days \(0–19 reviews\) are more likely to appear for unverified reviews, while moderate\- and high\-volume days are less likely\. This suggests a lower prevalence of high\-volume review activity among unverified reviews, which could reflect systematic differences between the two types of reviews or, alternatively, stronger platform filtering of unverified reviews\.
The estimated coefficients of the interaction termTreatment × Afterfor the moderate\- and high\-volume categories are not statistically significant\. On the other hand, we do find evidence of a marginally significantnegativeeffect for the low\-volume category \(0–19 reviews\) at the 10% significance level\. This effect can be interpreted as the equivalent of anincreasein the complementary review\-volume category, namely days with 20 or more same\-day reviews \(i\.e\., moderate\- and high\-volume days\), following LLM supply shocks\. Since batches of more than 20 unverified reviews in a single day for a single company appear to be mostly unlikely under normal conditions, we interpret this as evidence of AI\-generated review activity that succeeds in bypassing platform filters\. In the next section, we present additional evidence at the daily level suggesting that AI\-generated activity may manifest itself through unusually high review\-volume days\.
### 4\.2Event Study Specification
As described in Section[3](https://arxiv.org/html/2609.22198#S3), we further estimate a difference\-in\-differences specification with company and calendar\-week fixed effects, but instead of using a single indicator for the entire post\-event window, the regression includes separate indicators for each day before and after the LLM supply shock date\. Rather than using a single pre\-event day as the omitted reference category, the specification normalizes all event\-time coefficients relative to the average of the seven pre\-treatment days\. Accordingly, the estimated coefficients capture deviations from the average pre\-treatment period\. The analysis considers the same six outcome variables examined in the previous subsection\.
Figure[3](https://arxiv.org/html/2609.22198#S4.F3)presents the estimated interaction\-term coefficients of the event\-study specification for the six outcomes\. The left column includes figures for which the dependent variable is the following rating outcomes: average review rating, the proportion of one\-star reviews, and the proportion of five\-star reviews\. The right column includes the figures for which the dependent variable is the surge review\-volume categories: low, moderate, and high review\-volume\.
Figure 3:Event\-Study Estimates of the Effect of LLM Supply Shocks on Rating Outcomes and Review\-Volume Categories
Notes:This figure reports event\-study estimates of the effect of LLM supply shocks on rating outcomes and review\-volume categories\. The left column reports rating outcomes: Panel \(a\) shows the average review rating, Panel \(b\) the proportion of one\-star reviews, and Panel \(c\) the proportion of five\-star reviews\. The dependent variables in the right column are indicators equal to one if a company–day–verification\-status observation belongs to the low\-volume \(0–19 reviews\) category in Panel \(a\), the moderate\-volume \(20–99 reviews\) category in Panel \(b\), or the high\-volume \(100 or more reviews\) category in Panel \(c\), and zero otherwise\. The coefficients can therefore be interpreted as changes in the probability that an observation belongs to the corresponding review\-volume category\. The sample consists of aggregated observations at the company–day–verification\-status level within a symmetric 7\-day window before and after each LLM supply shock \(see Section[2\.2](https://arxiv.org/html/2609.22198#S2.SS2)for details on data construction\)\. The treatment group comprises unverified reviews, while verified reviews serve as the control group\. The coefficients are obtained from the event\-study specification in Equation \([2](https://arxiv.org/html/2609.22198#S3.E2)\), where each outcome is regressed on interactions between an indicator for unverified reviews and a set of event\-time dummies\. The event\-time coefficients are normalized such that their average over the seven pre\-treatment days \(−7\-7through−1\-1\) equals zero\. To implement this normalization, the coefficient for day−1\-1is omitted from the regression and subsequently reconstructed as the negative sum of the estimated coefficients for the other six pre\-treatment days\. Each coefficient represents the estimated differential change in the outcome for unverified relative to verified reviews at that event time, expressed relative to the average differential during the pre\-treatment period\. All regressions include company fixed effects and calendar week fixed effects\. Standard errors are clustered at the company level, and error bars represent 95% confidence intervals\. The vertical dashed line indicates the timing of the LLM supply shock \(day 0\)\. The p\-value reported below each panel corresponds to a joint test of the null hypothesis that the six freely estimated pre\-treatment coefficients are jointly equal to zero\. The rating\-outcome panels are based onN=748,513N=748\{,\}513observations\. The review\-volume panels are based onN=16,116,186N=16\{,\}116\{,\}186observations\. This number exceeds that for the rating outcomes because company–day–verification\-status observations with zero reviews are coded as zeros in the review\-volume category indicators and therefore included in the review\-volume analyses, whereas they are treated as missing in the rating and proportion analyses\.
The review rating outcome panels in the left column reveal several noteworthy patterns\. The estimated daily coefficients in the pre\-treatment period are generally small and statistically indistinguishable from zero\. Moreover, the formal pre\-trend tests fail to reject the null hypothesis of no differential pre\-treatment trends for all three outcome variables, providing support for the parallel trends assumption \(average rating: pre p\-value = 0\.756; one\-star proportion: pre p\-value = 0\.572; five\-star proportion: pre p\-value = 0\.517\)\. We do see a statistically significant pattern for the post\-event dynamics\. In the days following the LLM supply shocks, the average review rating initially declines before gradually recovering, with the effect being most pronounced approximately three to five days after the LLM supply shocks\. This pattern appears to be driven by an increase in the relative share of one\-star reviews during the seven days following LLM supply shocks, accompanied by a corresponding decrease in the share of five\-star reviews\. We do not detect meaningful changes in the other rating categories \(2, 3, and 4\), suggesting that any AI\-generated activity may occur primarily within the one\-star or five\-star categories\. These dynamic responses are consistent with the aggregated seven\-day results and, taken together, suggest that after price reductions and new model launches, there seems to be a shift of AI\-generated reviews usage from positive self\-promotion to negative competitive efforts\. At the same time, the results reveal the localized temporal structure of the suspected AI\-generated activity, indicating that the activity is done in a "concentrated wave" rather than diffusely\. We do note that this pattern could be either due to producers’ strategic timing or to some sort of interaction between review generation and the platform’s filtering mechanisms\. Given the nature of the research setup and the data, we are unable to discern between the two scenarios\. Notably, these effects are averaged across a highly diverse set of categories and over 72,000 companies and businesses\.
The review\-volume panels in the right column present the corresponding event\-study estimates for the three review\-volume surge categories at the daily level, i\.e\., a cluster of same\-company\-same\-day reviews\. Specifically, we focus on three categories of per\-day review volume: low\-volume days \(0–19 reviews per day\), moderate\-volume days \(20–99 reviews per day\), and high\-volume days \(100 or more reviews per day\), for the same company\.
First, we find that the pre\-treatment coefficients satisfy the parallel trends condition for the moderate\-volume category \(pre p\-value = 0\.175\)\. For this category, we also observe statistically significant post\-treatment dynamics \(post p\-value = 0\.022\)\. Similarly to the review rating and distribution, here too the effect peaks approximately three days after the event in the form of a temporal "wave," and is broadly consistent with the aggregate\-analysis results\.
For the other two categories, the pre\-trend tests are not statistically significant at the conventional 5% level, but remain relatively close to the threshold and therefore make it difficult to fully rule out some degree of pre\-treatment dynamics \(low\-volume days: pre p\-value = 0\.057; high\-volume days: pre p\-value = 0\.060\)\. For low\-volume days, we detect statistically significant post\-treatment dynamics \(post p\-value = 0\.009\), while for high\-volume days the post\-treatment effects are not statistically significant \(post p\-value = 0\.570\)\.
Overall, the findings regarding the same\-day\-same\-company review surges suggest that at least part of the AI\-generated review activity that escapes platform filtering may occur in the form of moderate\-volume review days and, to a lesser extent, low\-volume review days\. More broadly, the results indicate that suspected AI\-generated reviews in some cases tend to emerge in temporally concentrated bursts\.
### 4\.3Review\-Volume Effects Among Extreme Ratings
To better understand which reviews drive the observed changes in review volume, we focus on the most extreme ratings: 1\-star and 5\-star reviews\. These categories account for the majority of reviews in our sample and are particularly relevant because our previous analyses show that the most pronounced rating effects are concentrated at the extremes of the rating distribution\. We therefore examine the combined volume of 1\-star and 5\-star reviews\.
Because review\-count distributions differ across rating groups, we adjust the volume thresholds using the same quantile cutoffs as in the baseline specification\. For the combined 1\-star and 5\-star distribution, this yields three categories: low volume \(0–17 reviews\), moderate volume \(18–84 reviews\), and high volume \(85 or more reviews\)\.
Table[S4](https://arxiv.org/html/2609.22198#S1.T4)reports the baseline difference\-in\-differences estimates for these categories\. TheTreatment×\\timesAftercoefficients are not statistically significant in any of the three specifications:−0\.000035\-0\.000035for low volume,0\.0000360\.000036for moderate volume, and−0\.000001\-0\.000001for high volume\. Thus, the baseline specification provides no evidence of a statistically significant differential post\-shock change between unverified and verified reviews in any of the three combined 1\-star and 5\-star volume categories\.
The event\-study results in Figure[S4](https://arxiv.org/html/2609.22198#S1.F4), however, reveal short\-run dynamics that are not captured by the baseline specification\. The joint post\-treatment test rejects the null of no post\-treatment effects for the low\-volume category \(p\-value = 0\.021\) and provides weaker evidence for the moderate\-volume category \(p\-value = 0\.096\), while the high\-volume category is not significant \(p\-value = 0\.178\)\. The parallel\-trends assumption is supported for the low\- and moderate\-volume categories, with pre\-trend p\-values of0\.1340\.134and0\.6000\.600, respectively\. By contrast, the high\-volume category exhibits a significant pre\-trend \(p\-value = 0\.013\), making its post\-treatment dynamics less easily interpretable\.
The difference between the baseline and event\-study findings reflects the temporal aggregation imposed by the baseline specification\. The baseline model estimates a singleTreatment×\\timesAftercoefficient across the seven\-day post\-treatment period, whereas the event study allows the treatment effect to vary by day\. Short\-lived effects can therefore be obscured when aggregated across the full post\-treatment window, particularly when effects differ in magnitude or direction across days\.
This pattern is visible in Figure[S4](https://arxiv.org/html/2609.22198#S1.F4)\. For the low\-volume category, the treatment effect is close to zero during the first two post\-treatment days, falls sharply on day 3, and then gradually returns toward zero\. The moderate\-volume category exhibits the opposite pattern: the effect rises sharply on day 3 and subsequently declines toward zero\. Thus, the third post\-treatment day is characterized by a simultaneous decline in low\-volume extreme\-rating activity and an increase in moderate\-volume extreme\-rating activity\.
These dynamics are consistent with the main event\-study results in Figure[3](https://arxiv.org/html/2609.22198#S4.F3)\. While the main analysis documents the broader response in review activity following LLM supply shocks, the results here suggest that an important part of this response is concentrated among 1\-star and 5\-star reviews\. The category\-specific analysis therefore suggests that the overall review\-volume response is partly driven by a short\-run increase in the concentration of 1\-star and 5\-star reviews\.
One possible interpretation is that access to lower\-cost and more capable LLMs facilitates short\-lived bursts of strategically valuable review generation\. The movement from low\- to moderate\-volume 1\-star and 5\-star activity around day 3 is consistent with a temporary increase in the concentration of extreme reviews\. Five\-star reviews could potentially be used to improve a firm’s own reputation, whereas 1\-star reviews could be directed toward competitors\. Because these ratings lie at the extremes of the distribution, they may be particularly effective at influencing ratings and consumer perceptions\. This interpretation is suggestive rather than causal evidence of firms’ intentions, as the data identify changes in review activity but not who generated individual reviews or for what purpose\.
### 4\.4Heterogeneity Across AI\-Related Events \- Which Type is Driving the Effects?
We are interested in identifying which types of events drive the effects we observe: price reductions, new model releases, or events that combine both\. To examine this type of heterogeneity in the effects of AI\-related shocks, we estimate both the baseline Difference\-in\-Differences specification and the event\-study specification separately for three categories of LLM supply shocks: \(i\) price reduction events only, \(ii\) new model release events only, and \(iii\) events that involve both price reductions and new model releases\. Table[1](https://arxiv.org/html/2609.22198#S2.T1)presents the timeline of OpenAI pricing reductions and new model releases used to construct these event categories, including the corresponding event dates and shock types\.
We begin by examining the baseline Difference\-in\-Differences estimates\. The results are reported in Tables[S5](https://arxiv.org/html/2609.22198#S1.T5),[S6](https://arxiv.org/html/2609.22198#S1.T6), and[S7](https://arxiv.org/html/2609.22198#S1.T7)\. The estimates reveal substantial heterogeneity across AI\-related event types\. For price reduction events and events that combine price reductions with new model releases, the interaction coefficients are generally small and statistically insignificant across all outcomes\. In contrast, the interaction coefficients associated with new model release events are statistically significant for most outcomes and closely resemble the patterns observed in the main specification that pools all event types together\. Specifically, new model releases are associated with a significant decline in the average review rating and in the proportion of 5\-star reviews, alongside a significant increase in the proportion of 1\-star reviews\. At the same time, the probability of observing low\-volume review days \(0–19 reviews\) declines significantly, while the probability of observing moderate\-volume review days \(20–99 reviews\) increases significantly\. These findings suggest that the main effects documented in the baseline Difference\-in\-Differences analysis are largely concentrated around new model release dates rather than price\-reduction events\. To investigate the dynamics underlying these effects, we next estimate event\-study specifications separately for each event category\.
The results of the event\-study analysis separated into the three categories are shown in Figures[4](https://arxiv.org/html/2609.22198#S4.F4),[5](https://arxiv.org/html/2609.22198#S4.F5), and[6](https://arxiv.org/html/2609.22198#S4.F6)\.
Figure 4:Event\-Study Estimates for Price Reduction Events Only
Notes:This figure follows the same design and specification as Figure[3](https://arxiv.org/html/2609.22198#S4.F3), but restricts the analysis to LLM supply shocks consisting exclusively of price reductions \(see Table[1](https://arxiv.org/html/2609.22198#S2.T1)\)\. The rating\-outcome panels are based onN=286,490N=286,490observations, and the review\-volume panels are based onN=6,010,256N=6,010,256observations\.
Figure 5:Event\-Study Estimates for New Model Release Events Only
Notes:This figure follows the same design and specification as Figure[3](https://arxiv.org/html/2609.22198#S4.F3), but restricts the analysis to LLM supply shocks consisting exclusively of new model releases \(see Table[1](https://arxiv.org/html/2609.22198#S2.T1)\)\. The rating\-outcome panels are based onN=308,237N=308,237observations, and the review\-volume panels are based onN=7,200,354N=7,200,354observations\.
Figure 6:Event\-Study Estimates for Events Involving Both Price Reductions and New Model Releases
Notes:This figure follows the same design and specification as Figure[3](https://arxiv.org/html/2609.22198#S4.F3), but restricts the analysis to LLM supply shocks involving both a price reduction and a new model release \(see Table[1](https://arxiv.org/html/2609.22198#S2.T1)\)\. The rating\-outcome panels are based onN=130,283N=130,283observations, and the review\-volume panels are based onN=2,905,576N=2,905,576observations\.
The results are consistent and further reveal substantial heterogeneity across the types of AI shocks\. For price reduction events, the evidence is generally weak and does not support robust post\-event effects\. Specifically, in terms of rating\-related outcomes, namelyAverage Review Rating,Proportion of Rating 1, andProportion of Rating 5, the post\-period p\-values are all statistically insignificant, indicating no meaningful post\-event changes in the firm rating distributions \(see Table[S8](https://arxiv.org/html/2609.22198#S1.T8)\)\. Notably, the pre\-period p\-values for these outcomes show statistical insignificance, suggesting that the parallel trends assumption is satisfied but that the treatment effects themselves are weak or absent\. In terms of review\-volume surge variables, the results are similarly weak, and we do not observe notable effects for any of the review\-volume outcomes\. Figure[6](https://arxiv.org/html/2609.22198#S4.F6)shows a similar pattern for events that combine price reductions and new model releases: we do not observe notable effects for any of the outcomes examined\.
In contrast, we observe robust effects for the new model release events\. For all three rating\-related outcomes, the parallel trends assumption holds and the post\-period effects are significant\. These findings indicate that new model releases are likely driving the effects observed in firms’ rating distributions\. They suggest that producers of AI\-generated reviews respond primarily to changes in model capabilities and efficiency rather than to reductions in the monetary costs of using the models\. Interestingly, we do not see effects for the two events which include both price reductions and new model releases\. Although the very small number of events prohibits any credible conclusion, we hypothesize that the null\-result models \(GPT\-4\-Turbo\-128K and text\-embedding\-3\-small, see Table[1](https://arxiv.org/html/2609.22198#S2.T1)\) differ from the others in ways that potentially make them less likely to affect review production: text\-embedding\-3\-small is not a generative model and therefore cannot produce review text, while GPT\-4 Turbo was a relatively costly, resource\-intensive model compared with the newer, faster, and cheaper GPT\-4o\-based models\. Therefore, we loosely hypothesize that the absence of detectable effects for these releases may be consistent with limited relevance to large\-scale review generation, whereas the remaining models more directly reduced the cost or increased the accessibility and efficiency of producing such content\.
The dynamic pattern is similar to what we observed across all events in Figure[3](https://arxiv.org/html/2609.22198#S4.F3)\. Both theAverage Review Ratingand theProportion of Rating 1initially decline shortly after the LLM supply shock and then subsequently increase\. In contrast, theProportion of Rating 5exhibits the opposite pattern, increasing immediately following the LLM supply shock before gradually declining\.
The review\-volume outcomes are also consistent with this interpretation\. Both the low\-volume and moderate\-volume outcomes satisfy the parallel trends assumption and exhibit highly significant post\-period effects \(p\-value < 0\.001 for both variables; see Table[S9](https://arxiv.org/html/2609.22198#S1.T9)\)\. Interestingly, the dynamic patterns displayed in Figure[5](https://arxiv.org/html/2609.22198#S4.F5)differ across these two outcomes\. For the low\-volume category, the estimated effect declines immediately after the new model release and then gradually increases\. In contrast, the moderate\-volume category exhibits a sharp increase immediately after the release, followed by a gradual decline\.
Taken together, these results suggest that improvements in GenAI capabilities and efficiency associated with the release of new models may increase firms’ incentives or ability to generate AI\-assisted reviews, whereas reductions in model prices alone appear to have little impact on review\-generation behavior\.
## 5Robustness and Further Heterogeneity Checks
### 5\.1Are the Effects Driven by a Specific Event?
To further assess the robustness of our findings and examine whether they are driven by a unique LLM supply shock on a specific date, we re\-estimate the event\-study specification separately for each event date\. Tables[S8](https://arxiv.org/html/2609.22198#S1.T8)and[S9](https://arxiv.org/html/2609.22198#S1.T9)present the event\-level diagnostics for the rating\-related and review\-count outcomes, respectively\.
Several patterns emerge from the event\-level analysis\. First, the effects are not concentrated on a single event but appear across several event dates\. In particular, significant post\-period effects combined with non\-significant pre\-trends are observed for the price reduction on 2023\-03\-01 and across several subsequent events, including multiple new model releases\. For the 2023\-03\-01 price reduction, significant post\-event dynamics are observed for average ratings and the share of 5\-star ratings, with no evidence of differential pre\-trends for these outcomes\. Significant post\-period effects with non\-significant pre\-trends also arise for several outcomes around the new model release events on 2024\-05\-13, 2024\-07\-18, and 2024\-09\-12\.
Interestingly, the effects within the review\-volume categories appear weaker when the analysis is conducted separately by LLM supply shock date than in the pooled specification\. A likely explanation is that moderate\- and high\-volume review days are relatively rare by construction, reducing the number of observations available for each individual event and consequently lowering statistical power\. Evidence of higher\-volume review activity is present for some event dates, but the effects are generally sparse and not consistently observed across LLM supply shocks\.
Overall, event\-level analysis supports our interpretation of the main findings by showing that the phenomenon occurs across multiple points in time and across multiple events, suggesting that they reflect a broader pattern related to the usage of AI on the platform rather than stemming from a singular event\.
### 5\.2Heterogeneity of the Effect by Company Market Size
Because firms of different sizes may differ in both their incentives and capacity to adopt AI for review generation, we next examine heterogeneity by firm size to determine whether the observed effects are concentrated among smaller or larger firms\. Specifically, we examine whether the estimated treatment effects vary by company size, proxied by the volume of company\-related activity on the platform\. For simplicity, we assume that the total number of reviews \(both verified and non\-verified\) received during 2023–2024 on the platform can be used as a proxy for the relative rank of the company’s market size on and outside the platform\.
It is important to note that given that our model uses company fixed effects, small firms with only a handful of reviews do not contribute to the diff\-in\-diff analysis, due to insufficient observations and lack of variance\. This affects how we partition firm market size categories\. We therefore restrict the sample to companies that received at least 50 reviews during the sample period\. These companies are then divided into four equally sized quartiles based on their total review volume and we estimate our baseline model \(Equation[1](https://arxiv.org/html/2609.22198#S3.E1)\) separately for each quartile\. Table[5](https://arxiv.org/html/2609.22198#S5.T5)reports the review\-volume ranges associated with each quartile, while Tables[S10](https://arxiv.org/html/2609.22198#S1.T10)–[S13](https://arxiv.org/html/2609.22198#S1.T13)present the corresponding regression results\.
Table 5:Company Size Quantiles Based on Total Review VolumeQuantileReview\-volume rangeNumber of companiesQ1 \(Lowest\)50–805,056Q280–1525,056Q3152–4315,056Q4 \(Highest\)431–240,7455,056The results reported in Tables[S10](https://arxiv.org/html/2609.22198#S1.T10)–[S13](https://arxiv.org/html/2609.22198#S1.T13)suggest that the relationship between the treatment and review outcomes varies across company\-market\-size groups\. The largest effects are observed among companies within both the highest and lowest review count quartiles\. In both quartiles, the coefficient of theTreatment×\\timesAfterinteraction is negative and statistically significant for both the average review rating and the share of 5\-star reviews\. Even though the effect on 1\-star reviews is weak in this case, given that we look at the share of review stars the effect we are observing in this case is consistent with the general effect we see in other specifications of movement between positive and negative reviews\. This suggests that AI\-related activity is concentrated among very small firms and firms with exceptionally high levels of activity, but not among firms with moderate levels of activity\. Although identifying the precise mechanism is beyond the scope of this study, one possible explanation is that very small and very large firms may represent more attractive targets in terms of bang\-for\-the\-buck adversarial AI usage\. The potential harm caused by a negative review depends both on the number of existing reviews and on the firm’s market size\. In the perspective of the aggressor AI\-using firms: if their competitor is a very small firm or business, even a small number of AI\-generated negative reviews can meaningfully damage reputation on the platform, and in general\. If the competitor is a large, high\-demand firm, pushing even a small fraction of customers to switch firms may be of value to the aggressor firm\.
Finally, in terms of review\-volume surge activity, as in the previous heterogeneity analyses, the relative rarity of moderate\- and high\-volume review days limits statistical power, which may explain why in this case we find no meaningful effects\.
## 6Discussion, Limitations and Future Research
At least in the context examined here, this study provides causal evidence that advances in generative artificial intelligence are beginning to reshape online information ecosystems by changing the economics of textual content production\. Using OpenAI’s API pricing reductions and new model releases as exogenous shocks, we find systematic changes in unverified reviews on Trustpilot following improvements in AI capabilities and accessibility\.
First and foremost, a central contribution of this study is to provide evidence that AI is being used systematically within a marketplace, likely in pursuit of economic objectives\. It is important to note that because our approach is designed to detect changes in review activity in response to LLM supply shocks, it captures only a lower bound of the underlying AI\-related activity\. Furthermore, our findings suggest that the rapid advances of large language models, effectively providing "intelligence as a service" for content creation, are reshaping firms’ incentives to strategically manipulate online information, with potential consequences for the economics of digital platforms\. Among other implications, these findings suggest that the availability of AI tools may be altering the dynamics of competition in the relevant markets\.
Our findings contribute to the growing literature on AI\-generated reviews and AI\-generated content in a notable way\. Most existing studies have focused on distinguishing AI\-generated reviews or content from authentic human reviews by identifying linguistic characteristics or developing automated detection methods \([Zhao et al\. \(2025\)](https://arxiv.org/html/2609.22198#bib.bib42);[Fariello \(2024\)](https://arxiv.org/html/2609.22198#bib.bib62);[Guo et al\. \(2024\)](https://arxiv.org/html/2609.22198#bib.bib63);[Chaka \(2024\)](https://arxiv.org/html/2609.22198#bib.bib64)\)\. Although these studies have improved our understanding of how AI\-generated reviews or content differ from human\-written content, they provide limited evidence on whether firms actually change their behavior as generative AI becomes more accessible, and in what manners\. By leveraging exogenous reductions in the cost and improvements in the capabilities of state\-of\-the\-art language models, our study instead examines the behavioral consequences of generative AI adoption in a real\-world marketplace\. This perspective complements existing research by shifting attention from identifying AI\-generated reviews and content to understanding how advances in generative AI influence firms’ strategic behavior\.
A robust finding across our analyses is that the observed effects are mainly associated with negative reviews activity\. Analogous to other marketing\-oriented uses of AI, firms might reasonably be expected to use generative AI to enhance their own reputations by producing favorable reviews that appear authentic, particularly because doing so is relatively inexpensive and requires little effort compared with, e\.g\., relying on human labor or experts\. Instead, our findings are more consistent with generative AI being used as a competitive tool that enables firms to target rivals more easily\. This emerging development may have important implications for the competitive structure of markets\. Although our empirical design does not allow us to identify the motivations of reviewers, several mechanisms may explain this pattern\. Positive AI\-generated self\-promotion may expose the perpetrator to greater risk if detected by the platform, because the identity of the benefiting firm is readily apparent\. This risk is further amplified if platforms can detect AI\-generated self\-promotion more effectively than negative attacks on competitors\. Under such conditions, improvements in AI’s ability to produce credible, authentic\-looking negative content may make such manipulation a relatively less risky and potentially more effective strategy\. This possibility becomes even more salient given that damaging a rival’s reputation may yield greater returns per AI\-generated review, as prior research suggests that negative reviews exert a stronger influence than positive ones \([Chevalier and Mayzlin \(2006\)](https://arxiv.org/html/2609.22198#bib.bib57)\)\. Finally, fabricating a negative review that contains detailed, well\-articulated criticisms of a business may be more difficult than generating a generic positive review, making such reviews more dependent on the enhanced capabilities of advanced AI models\. Taken together, these considerations suggest that generative AI may not only increase the volume of manipulated content but also shift the strategic direction of manipulation itself\.
Our results further suggest that advances in intelligence as a service, the release of new, more capable models, matter more than API price reductions alone\. Although our empirical strategy is motivated partly by leveraging declines in the effective cost of AI\-generated text, the strongest effects coincide with model releases, indicating that firms may be more responsive to improvements in quality, reasoning, contextual understanding, and production efficiency than to lower monetary costs\. In simpler terms, this indicates that firms value more capable AI models\. Consequently, as a small number of providers increasingly supply “intelligence as a service,” AI capabilities may become more concentrated in the hands of a few firms, raising concerns about market power and dependence on these providers\.
Another contribution concerns the temporal organization of suspected AI\-generated review activity\. Our approach focuses on short\-term responses and does not capture persistent longer\-run effects\. The review\-volume and event\-study results suggest that at least some AI\-related activity occurs in concentrated bursts around AI supply shocks rather than as a continuous increase in review production\. Also, the rise in review surges following LLM supply shocks is consistent with coordinated waves of activity and indicates that users of AI\-generated reviews adapt quickly to changes in model availability\. This temporal concentration may also provide a useful signal for detecting similar activity in future work\.
Our heterogeneity analyses further suggest that these strategic incentives vary across firm activity levels\. The estimated effects are strongest among firms with very small and very large review volumes, consistent with the possibility that the expected returns to manipulation depend on market position\. Although this interpretation remains speculative, both groups may represent especially attractive targets for AI\-generated attacks: a small number of strategically generated reviews can materially affect firms with limited review histories, while highly visible firms may be targeted because even modest reputational changes can influence a large customer base\. This finding is consistent with the general interpretation of our findings that generative AI is being deployed as a strategic tool\.
A major alternative explanation is that LLM supply shocks induce changes in the platform’s filtering process rather than in AI\-generated review activity\. While platform responses are likely slower than those of firms or contractors, the findings themselves also make a simple platform\-wide filtering explanation less plausible\. Effects concentrated among the lowest\- and highest\-activity firms and short\-lived same\-company review bursts are more consistent with targeted AI\-generated review activity than with a uniform moderation change\. Although we cannot rule out filtering changes that interact with firm characteristics or AI capabilities, such an explanation would require a more specific moderation response\.
This study has several limitations\. First, our analysis focuses on a single online review platform\. Although Trustpilot is one of the largest review platforms which covers a very wide range of categories and companies, and provides institutional features that are well suited to our identification strategy, future research should examine whether similar patterns emerge on other platforms with different moderation policies, verification mechanisms, and user populations\. Second, our empirical strategy provides indirect evidence regarding AI\-assisted review generation changes rather than direct identification of thetotalamount of AI\-generated reviews\. Third, our approach is better suited to identifying short\-term effects of LLM supply shocks on online review dynamics than longer\-term effects, limiting our ability to assess their persistence and broader long\-run consequences, which we recommend future research examine\. Fourth, while exploiting exogenous variation generated by LLM supply shocks offers important advantages for causal inference, our design does not allow us to identify the actors responsible for generating the new reviews\. For example, whether they are produced by the interested firms themselves or by hired specialized contractors with the ability to produce AI\-generated content, or whether the findings we observe are a result of some sort of interaction with the platform’s filtering mechanisms\. Fifth, our analysis relies on the assumption that verified reviews constitute an appropriate control group because they are substantially less susceptible to strategic manipulation than unverified reviews\. Although we are relatively confident that verified reviews are very hard to manipulate at scale using AI models, as we have seen in the estimations, in some cases they may not be a perfect control for the unverified reviews\. Interestingly, they appear to perform better as control variables for normalized measures of review activity, such as rating shares, possibly because normalization helps account for underlying fluctuations and differences between the overall volumes of verified and unverified reviews\.
An additional limitation is that the Trustpilot data include only reviews that remain after the platform’s moderation and filtering procedures have been applied\. According to Trustpilot’s transparency reports, reviews identified as fake are removed before becoming part of the data analyzed in this study\. Consequently, our estimates capture only the AI\-related review activity that remains observable after platform moderation\. This implies that the observed AI\-related effects represent a lower bound of the actual AI\-related activity and that the behavioral changes we document persist despite the presence of sophisticated review\-filtering mechanisms\.
Finally, our findings suggest several directions for future research\. Applying similar identification strategies across review platforms, social media, and digital marketplaces could provide an approach that does not rely on text\-based markers or classification methods, helping clarify the generalizability of these patterns and the extent to which they depend on institutional settings and moderation policies\. Examining the developments of other AI providers, such as Google, Anthropic, and Meta, would help determine whether the observed responses reflect broader changes in the generative AI ecosystem, and how they differ across AI tools\. Similar mechanisms may extend beyond online reviews to broader forms of user\-generated content that have become critical components of digital platforms and the modern economy\. These include social media posts, online forums, question\-and\-answer communities, and recommendation systems, where AI\-generated content may similarly influence information flows and competitive dynamics\. Finally, future work should examine how platforms adapt to increasingly capable generative AI and how advances in content generation, strategic firm behavior, and moderation technologies jointly shape the credibility of digital information ecosystems\.
## References
- S\. Agrahari, S\. Kumar, and R\. S\. SanasamCan you really trust that review? protofewroberta and detectairev: a prototypical few\-shot method and multi\-domain benchmark for detecting ai\-generated reviews\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,pp\. 2118–2140\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Allcott and Gentzkow \(2017\)H\. Allcott and M\. GentzkowSocial media and fake news in the 2016 election\.Journal of economic perspectives31\(2\),pp\. 211–236\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p4.1)\.
- Alzateet al\.\(2021\)M\. Alzate, M\. Arce\-Urriza, and J\. CebolladaOnline reviews and product sales: the role of review visibility\.Journal of Theoretical and Applied Electronic Commerce Research16\(4\),pp\. 638–669\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p5.1)\.
- Bailyn \(2026\)E\. BailynTop generative ai chatbots by market share – april 2026\(Website\)Note:First Page Sage\. Accessed: 2026\-03\-29External Links:[Link](https://firstpagesage.com/reports/top-generative-ai-chatbots/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Baronio \(2025\)J\. BaronioIs the internet dead?\(Website\)ABC News Australia\.Note:Behind the News \(BTN\)External Links:[Link](https://www.abc.net.au/btn/high/is-the-internet-dead/104897518)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Bernzweig \(2025\)M\. BernzweigChatGPT dominance data \+ statistics\(Website\)Note:Software Oasis\. Accessed: 2026\-03\-30External Links:[Link](https://softwareoasis.com/chatgpt-dominance-2/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Bommasaniet al\.\(2021\)R\. Bommasani, D\. A\. Hudson, E\. Adeli, R\. Altman, S\. Arora, S\. von Arx, M\. S\. Bernstein, J\. Bohg, A\. Bosselut, E\. Brunskill,et al\.On the opportunities and risks of foundation models\.arXiv preprint arXiv:2108\.07258\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p1.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p1.1)\.
- Burton \(2024\)J\. BurtonOnline reviews can make or break your business: pay attention to them\(Website\)Note:Forbes Technology CouncilExternal Links:[Link](https://www.forbes.com/councils/forbestechcouncil/2024/11/12/online-reviews-can-make-or-break-your-business-pay-attention-to-them/)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Chaka \(2024\)C\. ChakaReviewing the performance of ai detection tools in differentiating between ai\-generated and human\-written texts: a literature and integrative hybrid review\.Journal of Applied Learning & Teaching7\(1\),pp\. 115–126\.Cited by:[§6](https://arxiv.org/html/2609.22198#S6.p3.1)\.
- Chevalier and Mayzlin \(2006\)J\. A\. Chevalier and D\. MayzlinThe effect of word of mouth on sales: online book reviews\.Journal of marketing research43\(3\),pp\. 345–354\.Cited by:[§2\.3](https://arxiv.org/html/2609.22198#S2.SS3.p2.1),[§6](https://arxiv.org/html/2609.22198#S6.p4.1)\.
- Clarket al\.\(2021\)E\. Clark, T\. August, S\. Serrano, N\. Haduong, S\. Gururangan, and N\. A\. SmithAll that’s ‘human’is not gold: evaluating human evaluation of generated text\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 7282–7296\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p2.1)\.
- Clark \(2023\)M\. ClarkOpenAI announces an api for chatgpt and its whisper speech\-to\-text tech\(Website\)External Links:[Link](https://www.theverge.com/2023/3/1/23620783/chatgpt-api-openai-pricing-whisper)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Cruz\-Gonzalezet al\.\(2017\)M\. Cruz\-Gonzalez, I\. Fernández\-Val, and M\. WeidnerBias corrections for probit and logit models with two\-way fixed effects\.The Stata Journal17\(3\),pp\. 517–545\.Cited by:[footnote 7](https://arxiv.org/html/2609.22198#footnote7)\.
- Culdaet al\.\(2025\)L\. C\. Culda, R\. A\. Nerişanu, M\. P\. Cristescu, D\. A\. Mara, A\. Bâra, and S\. OpreaComparative linguistic analysis framework of human\-written vs\. machine\-generated text\.Connection Science37\(1\),pp\. 2507183\.Cited by:[§3\.3](https://arxiv.org/html/2609.22198#S3.SS3.p1.1)\.
- Down \(2025\)A\. DownFrom shrimp jesus to erotic tractors: how viral ai slop took over the internet\(Website\)The Guardian\.External Links:[Link](https://www.theguardian.com/technology/2025/dec/27/from-shrimp-jesus-to-erotic-tractors-how-viral-ai-slop-took-over-the-internet)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Elad \(2024\)B\. EladOpenAI statistics 2024: revenue, growth, users and facts\(Website\)Note:Accessed: 2026\-03\-29External Links:[Link](https://www.enterpriseappstoday.com/stats/openai-statistics.html)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Fariello \(2024\)S\. FarielloDistinguishing human from machine: a review of advances and challenges in ai\-generated text detection\.Cited by:[§6](https://arxiv.org/html/2609.22198#S6.p3.1)\.
- Ferraraet al\.\(2016\)E\. Ferrara, O\. Varol, C\. Davis, F\. Menczer, and A\. FlamminiThe rise of social bots\.Communications of the ACM59\(7\),pp\. 96–104\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Fried \(2024\)I\. FriedOpenAI says chatgpt usage has doubled since last year\(Website\)Note:Axios\. Accessed: 2026\-03\-30External Links:[Link](https://www.axios.com/2024/08/29/openai-chatgpt-200-million-weekly-active-users)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Gambetti and Han \(2023a\)A\. Gambetti and Q\. HanCombat ai with ai: counteract machine\-generated fake restaurant reviews on social media\.arXiv preprint arXiv:2302\.07731\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Gambetti and Han \(2023b\)A\. Gambetti and Q\. HanDissecting ai\-generated fake reviews: detection and analysis of gpt\-based restaurant reviews on social media\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Guoet al\.\(2024\)X\. Guo, S\. Zhang, Y\. He, T\. Zhang, W\. Feng, H\. Huang, and C\. MaDetective: detecting ai\-generated text via multi\-level contrastive learning\.Advances in Neural Information Processing Systems37,pp\. 88320–88347\.Cited by:[§6](https://arxiv.org/html/2609.22198#S6.p3.1)\.
- Guptaet al\.\(2024\)R\. Gupta, V\. Jindal, and I\. KashyapRecent state\-of\-the\-art of fake review detection: a comprehensive review\.The Knowledge Engineering Review39,pp\. e8\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p7.1)\.
- H \(2024\)A\. HOpenAI o1 api pricing explained: everything you need to know\(Website\)Note:Medium \(Towards AGI\)\. Accessed: 2026\-03\-12External Links:[Link](https://medium.com/towards-agi/openai-o1-api-pricing-explained-everything-you-need-to-know-cbab89e5200d)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Hahn and Kuersteiner \(2011\)J\. Hahn and G\. KuersteinerBias reduction for dynamic nonlinear panel models with fixed effects\.Econometric Theory27\(6\),pp\. 1152–1191\.Cited by:[footnote 7](https://arxiv.org/html/2609.22198#footnote7)\.
- Heet al\.\(2022\)S\. He, B\. Hollenbeck, and D\. ProserpioThe market for fake reviews\.Marketing Science41\(5\),pp\. 896–921\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Huet al\.\(2009\)N\. Hu, P\. A\. Pavlou, and J\. J\. ZhangWhy do online product reviews have a j\-shaped distribution? overcoming biases in online word\-of\-mouth communication\.Communications of the ACM52\(10\),pp\. 144–147\.Cited by:[§2\.3](https://arxiv.org/html/2609.22198#S2.SS3.p2.1)\.
- Huang and Pape \(2020\)M\. Huang and A\. PapeThe impact of online consumer reviews on online sales: the case\-based decision theory approach\.Journal of Consumer Policy43\(3\),pp\. 463–490\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p5.1)\.
- Knightet al\.\(2023\)S\. Knight, Y\. Bart, and M\. YangGenerative ai and the perceived quality of user\-generated content: evidence from online reviews\.Northeastern U\. D’Amore\-McKim School of Business Research Paper\(4621982\)\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p7.1)\.
- Krepset al\.\(2022\)S\. Kreps, R\. M\. McCain, and M\. BrundageAll the news that’s fit to fabricate: ai\-generated text as a tool of media misinformation\.Journal of experimental political science9\(1\),pp\. 104–117\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p2.1)\.
- Kujur \(2025\)A\. KujurA comparative analysis of ai\-generated and human\-written text: linguistic patterns, detection accuracy, and implications for modern communication\.Detection Accuracy, and Implications for Modern Communication \(November 29, 2025\)\.Cited by:[§3\.3](https://arxiv.org/html/2609.22198#S3.SS3.p1.1)\.
- Lackermairet al\.\(2013\)G\. Lackermair, D\. Kailer, and K\. KanmazImportance of online product reviews from a consumer’s perspective\.Advances in economics and business1\(1\),pp\. 1–5\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p5.1)\.
- Levy \(2026\)S\. LevyAI slop melodramas are taking over x—and their creators are cashing in\(Website\)WIRED\.External Links:[Link](https://www.wired.com/story/ai-slop-melodramas-are-taking-over-x-and-their-creators-are-cashing-in/)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Limet al\.\(2025\)W\. M\. Lim, R\. Agarwal, A\. Mishra, and A\. MehrotraThe rise of fake reviews: toward a marketing\-oriented framework for understanding fake reviews\.Australasian Marketing Journal33\(2\),pp\. 178–198\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Luca and Zervas \(2016\)M\. Luca and G\. ZervasFake it till you make it: reputation, competition, and yelp review fraud\.Management science62\(12\),pp\. 3412–3427\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p11.1)\.
- Luoet al\.\(2026\)J\. Luo, G\. Nan, and D\. LiAI\-generated fake review detection\.Decision Support Systems,pp\. 114628\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Marcellinoet al\.\(2023\)W\. Marcellino, N\. Beauchamp\-Mustafaga, A\. Kerrigan, L\. N\. Chao, and J\. SmithThe rise of generative ai and the coming era of social media manipulation 3\.0: next\-generation chinese astroturfing and coping with ubiquitous ai\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p4.1)\.
- Martínez Otero \(2021\)J\. M\. Martínez OteroFake reviews on online platforms: perspectives from the us, uk and eu legislations\.SN Social Sciences1\(7\),pp\. 181\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Mayzlinet al\.\(2014\)D\. Mayzlin, Y\. Dover, and J\. ChevalierPromotional reviews: an empirical investigation of online review manipulation\.American Economic Review104\(8\),pp\. 2421–2455\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p11.1),[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Menget al\.\(2025\)W\. Meng, J\. Harvey, J\. Goulding, C\. J\. Carter, E\. Lukinova, A\. Smith, P\. Frobisher, M\. Forrest, and G\. Nica\-AvramLarge language models as’ hidden persuaders’: fake product reviews are indistinguishable to humans and machines\.arXiv preprint arXiv:2506\.13313\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p7.1),[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Mudaliar \(2024\)A\. MudaliarOpenAI launches structured outputs json api and reduces gpt prices\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://www.spiceworks.com/tech/artificial-intelligence/news/openai-launches-structured-outputs-json-api-reduces-gpt-prices/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Muñoz\-Ortizet al\.\(2024\)A\. Muñoz\-Ortiz, C\. Gómez\-Rodríguez, and D\. VilaresContrasting linguistic patterns in human and llm\-generated news text\.Artificial Intelligence Review57\(10\),pp\. 265\.Cited by:[§3\.3](https://arxiv.org/html/2609.22198#S3.SS3.p1.1)\.
- Murray \(2025\)C\. MurrayOhanian and altman warn of ‘Dead Internet Theory’—what is it and how is ai making it happen?\(Website\)Forbes\.External Links:[Link](https://www.forbes.com/sites/conormurray/2025/10/13/ohanian-and-altman-warn-of-dead-internet-theory-what-is-it-and-how-is-ai-making-it-happen/)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Muzumdaret al\.\(2025\)P\. Muzumdar, S\. Cheemalapati, S\. R\. RamiReddy, K\. Singh, G\. Kurian, and A\. MuleyThe dead internet theory: a survey on artificial interactions and the future of social media\.arXiv preprint arXiv:2502\.00007\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- OpenAI Developer Community \(2023\)OpenAI Developer CommunityGPT\-4 and gpt\-3\.5 turbo api cost comparison and understanding\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://community.openai.com/t/gpt4-and-gpt-3-5-turb-api-cost-comparison-and-understanding/106192)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- OpenAI Developers \(2024\)OpenAI DevelopersAnnouncement of gpt\-4 turbo api price reduction\(Website\)Note:Post on X \(Twitter\)\. Accessed: 2026\-03\-12External Links:[Link](https://x.com/OpenAIDevs/status/1851668229938159853)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- OpenAI \(2023\)OpenAINew models and developer products announced at devday\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://openai.com/index/new-models-and-developer-products-announced-at-devday/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- OpenAI \(2024a\)OpenAIGPT\-4o mini: advancing cost\-efficient intelligence\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- OpenAI \(2024b\)OpenAIHello gpt\-4o\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://openai.com/index/hello-gpt-4o/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- OpenAI \(2024c\)OpenAINew embedding models and api updates\(Website\)Note:Accessed: 2026\-03\-12External Links:[Link](https://openai.com/index/new-embedding-models-and-api-updates/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Özaydın \(2025\)H\. ÖzaydınFake reviews and ratings undermining consumer trust\.pp\.\.External Links:ISBN 978\-625\-5958\-72\-3,[Document](https://dx.doi.org/10.58830/ozgur.pub710.c3028)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p7.1)\.
- Pahwa \(2026\)A\. Pahwa100\+ openai statistics 2026: valuation, revenue & market share\(Website\)Note:Feedough\. Accessed: 2026\-03\-30External Links:[Link](https://www.feedough.com/openai-statistics/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Park \(2024\)H\. J\. ParkThe rise of generative artificial intelligence and the threat of fake news and disinformation online: perspectives from sexual medicine\.Investigative and Clinical Urology65\(3\),pp\. 199\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p4.1)\.
- Pocchiariet al\.\(2025\)M\. Pocchiari, D\. Proserpio, and Y\. DoverOnline reviews: a literature review and roadmap for future research\.International journal of research in marketing42\(2\),pp\. 275–297\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p5.1)\.
- Qiu and Zhang \(2023\)K\. Qiu and L\. ZhangHow online reviews affect purchase intention: a meta\-analysis across contextual and cultural factors\. data and information management, 8 \(2\), 100058\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p6.1)\.
- Rachmianiet al\.\(2024\)R\. Rachmiani, N\. K\. Oktadinna, and T\. R\. FauzanThe impact of online reviews and ratings on consumer purchasing decisions on e\-commerce platforms\.International Journal of Management Science and Information Technology4\(2\),pp\. 504–515\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p5.1)\.
- Ramanet al\.\(2024\)R\. Raman, S\. Mandal, P\. Das, T\. Kaur, J\. Sanjanasri, and P\. NedungadiExploring university students’ adoption of chatgpt using the diffusion of innovation theory and sentiment analysis with gender dimension\.Human Behavior and Emerging Technologies2024\(1\),pp\. 3085910\.Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Santos and Antonio \(2025\)A\. M\. Santos and N\. AntonioImproving trust in online reviews: a machine learning approach to detecting artificial intelligence\-generated reviews\.Information Technology & Tourism27\(3\),pp\. 739–766\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p8.1)\.
- Schwartz \(2023\)E\. H\. SchwartzOpenAI upgrades gpt\-4 and gpt\-3\.5 turbo models, reduces api prices\(Website\)Note:Voicebot\.ai\. Accessed: 2026\-03\-12External Links:[Link](https://voicebot.ai/2023/06/13/openai-upgrades-gpt-4-and-gpt-3-5-turbo-models-reduces-api-prices/)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Shariss \(2024\)J\. SharissRealtime api updates: webrtc, cheaper prices, 4o\-mini, and more\(Website\)Note:OpenAI Developer Community\. Accessed: 2026\-03\-12External Links:[Link](https://community.openai.com/t/realtime-api-updates-webrtc-cheaper-prices-4o-mini-and-more/1059962)Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Shin \(2026\)J\. ShinAI in the age of fake \(imagined\) content\.Available at SSRN 6353778\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p4.1)\.
- Tanaseet al\.\(2024\)I\. Tanase, L\. N\. Barbu, and E\. F\. GrejdanOnline reviews in romania: motivations, perceptions, and the impact of the j\-shaped distribution on consumer behavior\.Ovidius University Annals, Economic Sciences Series24\(2\),pp\. 450–454\.Cited by:[§2\.3](https://arxiv.org/html/2609.22198#S2.SS3.p2.1)\.
- Trustpilot \(2024\)TrustpilotTransparency report\(Website\)Note:[https://corporate\.trustpilot\.com/press/news/transparency\-report](https://corporate.trustpilot.com/press/news/transparency-report)Accessed: 2026\-03\-12External Links:[Link](https://corporate.trustpilot.com/press/news/transparency-report)Cited by:[footnote 4](https://arxiv.org/html/2609.22198#footnote4)\.
- Trustpilot \(2026a\)TrustpilotTrust and transparency\(Website\)Note:[https://corporate\.trustpilot\.com/trust](https://corporate.trustpilot.com/trust)Accessed: 2026\-04\-21Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p11.1)\.
- Trustpilot \(2026b\)TrustpilotWhy are some reviews marked "verified"?\(Website\)Note:Accessed: 2026\-03\-17External Links:[Link](https://help.trustpilot.com/s/article/Why-are-some-reviews-marked-Verified?language=en_US)Cited by:[§3\.1](https://arxiv.org/html/2609.22198#S3.SS1.p2.1)\.
- Tullyet al\.\(2024\)T\. Tully, J\. Redfern, and D\. Xiao2024: the state of generative ai in the enterprise\.Technical reportMenlo Ventures\.Note:Accessed: 2026\-08\-30External Links:[Link](https://menlovc.com/2024-the-state-of-generative-ai-in-the-enterprise/)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p10.1)\.
- Walter \(2025\)Y\. WalterArtificial influencers and the dead internet theory\.AI & SOCIETY40\(1\),pp\. 239–240\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p3.1)\.
- Wang and Xu \(2024\)S\. Wang and S\. Xu16 changes to the way enterprises are building and buying generative ai\.Andreessen Horowitz\.Note:Accessed: 2026\-08\-30External Links:[Link](https://a16z.com/generative-ai-enterprise-2024/)Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p10.1)\.
- Wooldridge \(2010\)J\. M\. WooldridgeEconometric analysis of cross section and panel data\.MIT press\.Cited by:[footnote 7](https://arxiv.org/html/2609.22198#footnote7)\.
- Zhang and Shao \(2024\)H\. Zhang and H\. ShaoExploring the latest applications of openai and chatgpt: an in\-depth survey\.\.Computer Modeling in Engineering & Sciences \(CMES\)138\(3\)\.Cited by:[§2\.1\.2](https://arxiv.org/html/2609.22198#S2.SS1.SSS2.p1.1)\.
- Zhaoet al\.\(2025\)Y\. Zhao, S\. Tang, H\. Zhang, and L\. LyuAI vs\. human: a large\-scale analysis of ai\-generated fake reviews, human\-generated fake reviews and authentic reviews\.Journal of Retailing and Consumer Services87,pp\. 104400\.Cited by:[§1](https://arxiv.org/html/2609.22198#S1.p7.1),[§1](https://arxiv.org/html/2609.22198#S1.p8.1),[§6](https://arxiv.org/html/2609.22198#S6.p3.1)\.
## Supplementary Materials
Figure S1:OpenAI Model Pricing Reductions and Releases, 2023–2024 \(raw prices\)
Notes:This figure presents the raw \(non\-normalized\) pricing series underlying Figure[1](https://arxiv.org/html/2609.22198#S2.F1)\. Prices are reported in U\.S\. dollars per 1,000 tokens\. The upper panel shows input prices and the lower panel shows output prices\. The horizontal axis displays calendar dates over the 2023–2024 period\. Colors correspond to different models, with similar color schemes used for models belonging to the same family\. Vertical dashed lines indicate the dates of the LLM supply shocks used in the empirical analysis and listed in Table[1](https://arxiv.org/html/2609.22198#S2.T1)\. Solid circular markers denote observed pricing points included in the analysis, corresponding either to price reductions or to the release of lower\-priced models\. Embedding models do not report output prices because they generate vector representations rather than textual outputs\. Figure[1](https://arxiv.org/html/2609.22198#S2.F1)presents the same data normalized to each model’s initial observed price \(=100%\)\.
## S1Event\-study normalization and derivation
This section provides the derivation of the event\-study specification used in the main analysis\. The specification normalizes the average of the seven pre\-event coefficients to zero rather than selecting a single pre\-event day as the omitted reference period\. The normalization is imposed separately on the event\-time coefficients,λk\\lambda\_\{k\}, and the treatment\-by\-event\-time interaction coefficients,δk\\delta\_\{k\}\.
Each observation is indexed by companycc, verification statusvv, and calendar datett\. Define
Treatmentv=𝟏\{v=unverified\},\\mathrm\{Treatment\}\_\{v\}=\\mathbf\{1\}\\\{v=\\text\{unverified\}\\\},
such thatTreatmentv=1\\mathrm\{Treatment\}\_\{v\}=1for unverified observations andTreatmentv=0\\mathrm\{Treatment\}\_\{v\}=0for verified observations\.
Let
Dtk=𝟏\{datetiskdays relative to the event date\}\.D\_\{t\}^\{k\}=\\mathbf\{1\}\\\{\\text\{date \}t\\text\{ is \}k\\text\{ days relative to the event date\}\\\}\.
The event\-time indicator is indexed by calendar datettand relative timekk, wherekkdenotes the number of days relative to the event date\. For a given calendar datett, the verified and unverified observations therefore have the same event\-time indicator\. Differences between the two groups are captured byTreatmentv\\mathrm\{Treatment\}\_\{v\}and its interactions with the event\-time indicators\.
The event day,k=0k=0, is excluded from the specification\. The event\-study window consists of the seven days preceding and seven days following the event:
k∈\{−7,−6,…,−1,1,…,7\}\.k\\in\\\{\-7,\-6,\\ldots,\-1,1,\\ldots,7\\\}\.
### S1\.1Normalization of the pre\-event event\-time coefficients
Rather than setting one particular pre\-event coefficient equal to zero, we normalize the event\-time coefficients such that their average over the seven pre\-event days equals zero:
17∑k=−7−1λk=0\.\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\lambda\_\{k\}=0\.
Equivalently,
∑k=−7−1λk=0\.\\sum\_\{k=\-7\}^\{\-1\}\\lambda\_\{k\}=0\.
Solving this restriction for the coefficient on relative day−1\-1gives
λ−1=−∑k=−7−2λk\.\\lambda\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}\.
Without imposing the normalization, the pre\-event event\-time component can be written as
∑k=−7−1λkDtk\.\\sum\_\{k=\-7\}^\{\-1\}\\lambda\_\{k\}D\_\{t\}^\{k\}\.
Separating the coefficient for relative day−1\-1yields
∑k=−7−2λkDtk\+λ−1Dt−1\.\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}D\_\{t\}^\{k\}\+\\lambda\_\{\-1\}D\_\{t\}^\{\-1\}\.
Substituting
λ−1=−∑k=−7−2λk\\lambda\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}
gives
∑k=−7−2λkDtk−\(∑k=−7−2λk\)Dt−1\.\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}D\_\{t\}^\{k\}\-\\left\(\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}\\right\)D\_\{t\}^\{\-1\}\.
Collecting the terms associated with eachλk\\lambda\_\{k\}, the expression becomes
∑k=−7−2λk\(Dtk−Dt−1\)\.\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\.
Thus, usingDtk−Dt−1D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}as the regressors fork=−7,…,−2k=\-7,\\ldots,\-2is algebraically equivalent to estimating all seven pre\-event event\-time coefficients subject to the restriction that their average equals zero\.
### S1\.2Normalization of the pre\-event interaction coefficients
We impose the analogous normalization on the treatment\-by\-event\-time interaction coefficients:
17∑k=−7−1δk=0\.\\frac\{1\}\{7\}\\sum\_\{k=\-7\}^\{\-1\}\\delta\_\{k\}=0\.
Equivalently,
∑k=−7−1δk=0\.\\sum\_\{k=\-7\}^\{\-1\}\\delta\_\{k\}=0\.
Solving for the interaction coefficient on relative day−1\-1gives
δ−1=−∑k=−7−2δk\.\\delta\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\.
Without imposing the normalization, the pre\-event interaction component can be written as
∑k=−7−1δk\(TreatmentvDtk\)\.\\sum\_\{k=\-7\}^\{\-1\}\\delta\_\{k\}\\left\(\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{k\}\\right\)\.
Separating the interaction coefficient for relative day−1\-1yields
∑k=−7−2δk\(TreatmentvDtk\)\+δ−1\(TreatmentvDt−1\)\.\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\left\(\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{k\}\\right\)\+\\delta\_\{\-1\}\\left\(\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{\-1\}\\right\)\.
Substituting
δ−1=−∑k=−7−2δk\\delta\_\{\-1\}=\-\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}
gives
∑k=−7−2δk\(TreatmentvDtk\)−\(∑k=−7−2δk\)\(TreatmentvDt−1\)\.\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\left\(\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{k\}\\right\)\-\\left\(\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\right\)\\left\(\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{\-1\}\\right\)\.
Collecting the terms associated with eachδk\\delta\_\{k\}, this expression becomes
∑k=−7−2δkTreatmentv\(Dtk−Dt−1\)\.\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\mathrm\{Treatment\}\_\{v\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\.
Thus, usingTreatmentv\(Dtk−Dt−1\)\\mathrm\{Treatment\}\_\{v\}\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\)as the interaction regressors fork=−7,…,−2k=\-7,\\ldots,\-2is algebraically equivalent to estimating all seven pre\-event interaction coefficients subject to the restriction that their average equals zero\.
### S1\.3Resulting regression specification
Combining these reparameterizations with the post\-event coefficients gives the event\-study regression
Ycvt=\\displaystyle Y\_\{cvt\}=\{\}αc\+γw\(t\)\+θTreatmentv\\displaystyle\\alpha\_\{c\}\+\\gamma\_\{w\(t\)\}\+\\theta\\,\\mathrm\{Treatment\}\_\{v\}\+∑k=−7−2λk\(Dtk−Dt−1\)\+∑k=17λkDtk\\displaystyle\+\\sum\_\{k=\-7\}^\{\-2\}\\lambda\_\{k\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\+\\sum\_\{k=1\}^\{7\}\\lambda\_\{k\}D\_\{t\}^\{k\}\+∑k=−7−2δkTreatmentv\(Dtk−Dt−1\)\+∑k=17δkTreatmentvDtk\+εcvt\.\\displaystyle\+\\sum\_\{k=\-7\}^\{\-2\}\\delta\_\{k\}\\mathrm\{Treatment\}\_\{v\}\\left\(D\_\{t\}^\{k\}\-D\_\{t\}^\{\-1\}\\right\)\+\\sum\_\{k=1\}^\{7\}\\delta\_\{k\}\\mathrm\{Treatment\}\_\{v\}D\_\{t\}^\{k\}\+\\varepsilon\_\{cvt\}\.
Table S1:Difference\-in\-Differences Results: Review Count, Length, and Text Homogeneity\(1\)\(2\)\(3\)\(4\)\(5\)CountLog countLengthLog lengthHomogeneityTreatment\-0\.066\*\*\*0\.012\*\*\*34\.841\*\*\*0\.670\*\*\*0\.0363\*\*\*\(0\.013\)\(0\.001\)\(0\.591\)\(0\.010\)\(0\.0015\)After\-0\.023\*\*\*\-0\.008\*\*\*0\.269\-0\.007\*0\.0006\(0\.003\)\(0\.000\)\(0\.213\)\(0\.004\)\(0\.0006\)Treatment×\\timesAfter\-0\.004\-0\.001\*\*\*0\.2270\.003\-0\.0009\(0\.004\)\(0\.000\)\(0\.215\)\(0\.004\)\(0\.0006\)Company FEYesYesYesYesYesWeek FEYesYesYesYesYesMean \(After = 0\)0\.1760\.05347\.9023\.4210\.604R\-squared0\.1780\.3550\.2970\.3920\.344Observations16,116,18616,116,186748,513748,5131,209,426
Notes:This table reports Difference\-in\-Differences estimates for additional standard review metrics, including review count, the logarithm of review count, review length, the logarithm of review length, and review text homogeneity\. Across most specifications, the coefficient on the interaction term \(Treatment×\\timesAfter\) is small and not statistically significant, indicating no robust differential effect of the LLM supply shocks on these outcomes\. An exception is the specification using the logarithm of review counts, where the interaction coefficient is statistically significant in the baseline model\. However, this result is not robust: in the corresponding event\-study specification, the parallel trends assumption is violated, and the effect does not persist \(see Table[S3](https://arxiv.org/html/2609.22198#S1.T3)\)\.
Standard errors in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S2:Difference\-in\-Differences Results: Log Review Counts by Rating Category \(1–5\)\(1\)\(2\)\(3\)\(4\)\(5\)Log count 1Log count 2Log count 3Log count 4Log count 5Treatment0\.01636\*\*\*\-0\.00037\*\*\*\-0\.00181\*\*\*\-0\.00339\*\*\*\-0\.00576\*\*\*\(0\.00036\)\(0\.00013\)\(0\.00018\)\(0\.00028\)\(0\.00094\)After\-0\.00217\*\*\*\-0\.00033\*\*\*\-0\.00033\*\*\*\-0\.00079\*\*\*\-0\.00582\*\*\*\(0\.00008\)\(0\.00004\)\(0\.00005\)\(0\.00007\)\(0\.00017\)Treatment×\\timesAfter\-0\.00035\*\*\*0\.00008\*\*0\.000030\.00003\-0\.00038\*\*\(0\.00009\)\(0\.00004\)\(0\.00004\)\(0\.00006\)\(0\.00015\)Company FEYesYesYesYesYesWeek FEYesYesYesYesYesMean \(After = 0\)0\.0140\.0020\.0030\.0060\.038R\-squared0\.2390\.1670\.2070\.2320\.332Observations16,116,18616,116,18616,116,18616,116,18616,116,186
- •Notes:This table reports Difference\-in\-Differences estimates for the logarithm of review counts across rating categories \(1 to 5 stars\)\. The interaction term \(Treatment×\\timesAfter\) is statistically significant for 1\-star, 2\-star, and 5\-star reviews in the baseline specification, while no significant effects are observed for 3\-star and 4\-star reviews\. However, these results are not robust: in the corresponding event\-study analysis, the parallel trends assumption is violated for these outcomes, and the estimated effects do not persist \(see Table[S3](https://arxiv.org/html/2609.22198#S1.T3)\)\. Standard errors are reported in parentheses\. Statistical significance levels are denoted as follows: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S3:Pre\-Trend and Post\-Treatment TestsOutcomePre\-trend p\-valuePost\-treatment p\-valueAverage post coefficientCount0\.007\*\*\*0\.114\-0\.004255Log count<0\.001<0\.001\*\*\*<0\.001<0\.001\*\*\*\-0\.000721Length0\.6400\.089\*0\.255823Log length0\.5990\.1220\.003607Homogeneity0\.2700\.285\-0\.000647Log count 1<0\.001<0\.001\*\*\*<0\.001<0\.001\*\*\*\-0\.000349Log count 20\.022\*\*0\.2730\.000081Log count 30\.5900\.9420\.000025Log count 40\.089\*0\.030\*\*0\.000031Log count 5<0\.001<0\.001\*\*\*<0\.001<0\.001\*\*\*\-0\.000377
Notes:This table reports joint F\-tests of pre\-treatment and post\-treatment interaction coefficients from the event\-study specifications\. The pre\-treatment test examines the null hypothesis that the freely estimated pre\-treatment interaction coefficients are jointly equal to zero\. Treatment\-by\-event\-time coefficients are normalized such that their average over the seven pre\-treatment days \(k=−7,…,−1k=\-7,\\ldots,\-1\) equals zero; the coefficient atk=−1k=\-1is recovered from this restriction\. The post\-treatment test examines the null hypothesis that all post\-treatment interaction coefficients \(k=1,…,7k=1,\\ldots,7\) are jointly equal to zero\. The average post coefficient is the arithmetic mean of the seven post\-treatment interaction coefficients,17∑k=17δk\\frac\{1\}\{7\}\\sum\_\{k=1\}^\{7\}\\delta\_\{k\}\. All regressions include company and week fixed effects, and standard errors are clustered at the company level\. Statistical significance levels are denoted as follows: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Figure S2:Event\-Study Coefficients for Total Count, Length, and Homogeneity
Notes:This figure plots the event\-study treatment\-by\-event\-time coefficients for review count, log review count, average review length, log average review length, and text homogeneity\. Points represent estimated treatment effects at each event time, and vertical bars show 95% confidence intervals\. The treatment effects are normalized such that the average coefficient over the seven pre\-treatment days \(\(k=\-7,…,\-1\)\) equals zero\. To impose this normalization, the pre\-treatment indicators are parameterized relative to \(k=\-1\), and the coefficient for \(k=\-1\) is recovered from the restriction that the seven pre\-treatment coefficients sum to zero\. The vertical line marks the LLM supply shock dates, and the horizontal dashed line denotes zero\.
Figure S3:Event\-Study Coefficients for Log Review Counts by Rating Category
Notes:This figure plots the event\-study treatment\-by\-event\-time coefficients for the logarithm of review counts by rating category, separately for 1\-star through 5\-star reviews\. Points represent estimated treatment effects at each event time, and vertical bars show 95% confidence intervals\. The treatment effects are normalized such that the average coefficient over the seven pre\-treatment days \(\(k=\-7,…,\-1\)\) equals zero\. To impose this normalization, the pre\-treatment indicators are parameterized relative to \(k=\-1\), and the coefficient for \(k=\-1\) is recovered from the restriction that the seven pre\-treatment coefficients sum to zero\. The vertical line marks the LLM supply shock dates, and the horizontal dashed line denotes zero\.
Table S4:Difference\-in\-Differences Results for Combined 1\-Star and 5\-Star Review\-Volume Categories Using Adjusted Thresholds\(1\)\(2\)\(3\)1\+5\-Star Rating1\+5\-Star Rating1\+5\-Star RatingAdjusted ThresholdAdjusted ThresholdAdjusted ThresholdLow\-volume: 0–17Moderate\-volume: 18–84High\-volume: 85\+Treatment0\.001243\*\*\*\-0\.001077\*\*\*\-0\.000167\*\*\*\(0\.00013\)\(0\.00012\)\(0\.00004\)After0\.000224\*\*\*\-0\.000217\*\*\*\-0\.000007\(0\.00003\)\(0\.00003\)\(0\.00001\)Treatment×\\timesAfter\-0\.0000350\.000036\-0\.000001\(0\.00002\)\(0\.00002\)\(0\.00001\)Company FEYesYesYesWeek FEYesYesYesMean \(After=0=0\)0\.99880\.00110\.0001R\-squared0\.2650\.2410\.215Observations16,116,18616,116,18616,116,186
Notes:This table reports fixed effects difference\-in\-differences regressions for indicators capturing combined 1\-star and 5\-star review\-volume categories using adjusted thresholds\. The dependent variables are indicators equal to one if a company–day–verification\-status observation belongs to the low\-volume \(0–17 combined 1\-star and 5\-star reviews\), moderate\-volume \(18–84 combined 1\-star and 5\-star reviews\), or high\-volume \(85 or more combined 1\-star and 5\-star reviews\) category, and zero otherwise\. The coefficients can therefore be interpreted as changes in the probability that an observation belongs to the corresponding review\-volume category\. Treatment equals one for unverified reviews and zero for verified reviews\. After equals one for observations after the LLM supply shock and zero for observations before the shock\. The coefficient on Treatment×\\timesAfter is the difference\-in\-differences estimate\. All specifications include company and week fixed effects, and standard errors are clustered at the company level\. Standard errors are reported in parentheses\. Statistical significance levels are denoted as follows: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.10p<0\.10\.
Figure S4:Event\-Study Estimates for Combined 1\-Star and 5\-Star Review\-Volume Categories Using Adjusted Thresholds
Notes:This figure plots event\-study treatment\-by\-event\-time coefficients for three indicators of combined 1\-star and 5\-star review volume using adjusted thresholds\. The indicators equal one if a company–day–verification\-status observation belongs to the low\-volume \(0–17 combined 1\-star and 5\-star reviews\), moderate\-volume \(18–84 combined 1\-star and 5\-star reviews\), or high\-volume \(85 or more combined 1\-star and 5\-star reviews\) category, respectively, and zero otherwise\. The coefficients can therefore be interpreted as changes in the probability that an observation belongs to the corresponding review\-volume category\. Points represent estimated treatment effects at each event time, and vertical bars show 95% confidence intervals\. The treatment effects are normalized such that the average coefficient over the seven pre\-treatment days \(k=−7,…,−1k=\-7,\\ldots,\-1\) equals zero\. To impose this normalization, the pre\-treatment indicators are parameterized relative tok=−1k=\-1, and the coefficient fork=−1k=\-1is recovered from the restriction that the seven pre\-treatment coefficients sum to zero\. The corresponding event\-time effects are parameterized using the same pre\-treatment normalization\. Treatment equals one for unverified reviews and zero for verified reviews\. All specifications include company and week fixed effects\. Standard errors are clustered at the company level\. The vertical dotted line marks the LLM supply shock dates, and the horizontal dotted line denotes zero\. The p\-value reported below each panel corresponds to the joint test of the pre\-treatment treatment\-by\-event\-time coefficients\.
Table S5:Difference\-in\-Differences Estimates: Price Reduction DatesReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)High\-volume100\+Treatment\-0\.98827\*\*\*0\.25209\*\*\*\-0\.21595\*\*\*0\.00136\*\*\*\-0\.00120\*\*\*\-0\.00017\*\*\*\(0\.02503\)\(0\.00618\)\(0\.00625\)\(0\.00015\)\(0\.00014\)\(0\.00004\)After\-0\.006880\.00200\-0\.001620\.00018\*\*\*\-0\.00017\*\*\*\-0\.00001\(0\.00755\)\(0\.00188\)\(0\.00245\)\(0\.00005\)\(0\.00005\)\(0\.00001\)Treatment×\\timesAfter\-0\.00477\-0\.00018\-0\.001570\.00005\-0\.00005\-0\.00000\(0\.00797\)\(0\.00203\)\(0\.00244\)\(0\.00004\)\(0\.00004\)\(0\.00001\)Company FEYesYesYesYesYesYesWeek FEYesYesYesYesYesYesMean \(After = 0\)3\.7840\.2590\.6330\.9990\.0010\.000R\-squared0\.5980\.5810\.5210\.2820\.2620\.220Observations286,490286,490286,4906,010,2566,010,2566,010,256
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S6:Difference\-in\-Differences Estimates: New Model Release DatesReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)High\-volume100\+Treatment\-1\.18741\*\*\*0\.30004\*\*\*\-0\.26122\*\*\*0\.00125\*\*\*\-0\.00106\*\*\*\-0\.00019\*\*\*\(0\.02618\)\(0\.00634\)\(0\.00664\)\(0\.00014\)\(0\.00013\)\(0\.00004\)After0\.00677\-0\.001930\.000490\.00030\*\*\*\-0\.00030\*\*\*\-0\.00001\(0\.00752\)\(0\.00187\)\(0\.00242\)\(0\.00004\)\(0\.00004\)\(0\.00001\)Treatment×\\timesAfter\-0\.03173\*\*\*0\.00704\*\*\*\-0\.00855\*\*\*\-0\.00011\*\*\*0\.00011\*\*\*\-0\.00000\(0\.00756\)\(0\.00193\)\(0\.00233\)\(0\.00004\)\(0\.00004\)\(0\.00001\)Company FEYesYesYesYesYesYesWeek FEYesYesYesYesYesYesMean \(After = 0\)3\.7340\.2710\.6220\.9990\.0010\.000R\-squared0\.6130\.5940\.5360\.2970\.2710\.250Observations308,237308,237308,2377,200,3547,200,3547,200,354
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S7:Difference\-in\-Differences Estimates: Both Price Reduction and New Model Release DatesReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)High\-volume100\+Treatment\-0\.82463\*\*\*0\.21122\*\*\*\-0\.17860\*\*\*0\.00146\*\*\*\-0\.00129\*\*\*\-0\.00016\*\*\*\(0\.03394\)\(0\.00834\)\(0\.00858\)\(0\.00018\)\(0\.00016\)\(0\.00005\)After0\.03083\*\*\-0\.00893\*\*0\.00903\*\*0\.00013\*\-0\.00010\-0\.00003\(0\.01397\)\(0\.00356\)\(0\.00433\)\(0\.00007\)\(0\.00007\)\(0\.00002\)Treatment×\\timesAfter\-0\.010920\.00197\-0\.00336\-0\.000060\.000040\.00001\(0\.01165\)\(0\.00296\)\(0\.00359\)\(0\.00006\)\(0\.00006\)\(0\.00002\)Company FEYesYesYesYesYesYesWeek FEYesYesYesYesYesYesMean \(After = 0\)3\.7850\.2580\.6340\.9990\.0010\.000R\-squared0\.6340\.6160\.5590\.3290\.3070\.301Observations130,283130,283130,2832,905,5762,905,5762,905,576
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S8:Event\-study diagnostics: rating outcomesAverage ratingShare 1\-starShare 5\-starDateEvent typePre p\-valPost p\-valAvg\. post coef\.Pre p\-valPost p\-valAvg\. post coef\.Pre p\-valPost p\-valAvg\. post coef\.2023\-03\-01Price reduction0\.3080\.044∗∗\-0\.0178250\.3740\.2150\.0011010\.4400\.035∗∗\-0\.0071982023\-06\-13Price reduction0\.8780\.9910\.0044750\.7920\.986\-0\.0017470\.9440\.9150\.0028242023\-11\-06Price reduction and new model release0\.3560\.953\-0\.0062530\.5350\.9710\.0010800\.3460\.9590\.0007852024\-01\-24Price reduction and new model release0\.075∗0\.357\-0\.0071440\.065∗0\.411\-0\.0004110\.1290\.222\-0\.0055612024\-05\-13New model release0\.2780\.056∗\-0\.0387760\.7060\.052∗0\.0099120\.2610\.200\-0\.0112982024\-07\-18New model release0\.2510\.029∗∗\-0\.0122690\.5490\.007∗∗∗\-0\.0003510\.3220\.137\-0\.0090162024\-08\-09Price reduction0\.6280\.440\-0\.0120230\.4890\.2770\.0017330\.7330\.513\-0\.0034152024\-09\-12New model release0\.086∗<0\.001∗∗∗\-0\.0362870\.2910\.001∗∗∗0\.0103030\.084∗<0\.001∗∗∗\-0\.0053502024\-10\-30Price reduction0\.2680\.2420\.0185710\.3920\.457\-0\.0036520\.1810\.1830\.0080472024\-12\-18New model release0\.8080\.184\-0\.0364020\.2880\.1390\.0111180\.5030\.789\-0\.006115
Notes:The pre\-period p\-value reports the joint test of the null hypothesis that the pre\-event treatment\-effect coefficients are equal to zero\. Under the average pre\-treatment normalization, this corresponds to testing for differential pre\-event dynamics between unverified and verified observations\.
The post\-period p\-value reports the joint test of the null hypothesis that the post\-event treatment\-effect coefficients are equal to zero\.
The average post coefficient is the mean of the seven estimated post\-event treatment\-effect coefficients,δ^1,…,δ^7\\hat\{\\delta\}\_\{1\},\\ldots,\\hat\{\\delta\}\_\{7\}\.
Each row is estimated separately for the indicated event date\. The event\-date\-specific regressions include company fixed effects but omit week fixed effects\. Standard errors are clustered by company\.
Bold entries indicate cases in which the pre\-period joint test is not statistically significant at the 10% level \(p\-value≥\\geq0\.10\) and the post\-period joint test is statistically significant at the 10% level \(p\-value < 0\.10\)\.
Significance levels are denoted by∗p<0\.10\{\}^\{\*\}p<0\.10,p∗∗<0\.05\{\}^\{\*\*\}p<0\.05, and∗∗∗p<0\.01\{\}^\{\*\*\*\}p<0\.01\.
Table S9:Event\-study diagnostics: review\-volume outcomesLow\-volume \[0–20\)Moderate\-volume \[20–100\)High\-volume 100\+DateEvent typePre p\-valPost p\-valAvg\. post coef\.Pre p\-valPost p\-valAvg\. post coef\.Pre p\-valPost p\-valAvg\. post coef\.2023\-03\-01Price reduction0\.009∗∗∗0\.044∗∗0\.0002390\.029∗∗0\.073∗\-0\.0002030\.3620\.648\-0\.0000362023\-06\-13Price reduction<0\.001∗∗∗<0\.001∗∗∗\-0\.0000040\.001∗∗∗0\.004∗∗∗0\.0000160\.2190\.130\-0\.0000132023\-11\-06Price reduction and new model release0\.027∗∗0\.041∗∗\-0\.0000480\.027∗∗0\.050∗∗0\.0000660\.2980\.251\-0\.0000182024\-01\-24Price reduction and new model release0\.002∗∗∗0\.085∗\-0\.0000620\.003∗∗∗0\.2060\.0000170\.2140\.1770\.0000452024\-05\-13New model release<0\.001∗∗∗0\.002∗∗∗\-0\.000132<0\.001∗∗∗0\.011∗∗0\.0001300\.013∗∗0\.6240\.0000022024\-07\-18New model release0\.002∗∗∗0\.006∗∗∗0\.0000520\.009∗∗∗0\.009∗∗∗\-0\.0000460\.1150\.504\-0\.0000052024\-08\-09Price reduction0\.088∗0\.001∗∗∗\-0\.0000240\.1970\.003∗∗∗\-0\.0000110\.2850\.3500\.0000362024\-09\-12New model release<0\.001∗∗∗0\.006∗∗∗0\.0000150\.001∗∗∗0\.021∗∗0\.0000030\.015∗∗0\.675\-0\.0000172024\-10\-30Price reduction0\.036∗∗0\.006∗∗∗0\.0000350\.066∗0\.014∗∗\-0\.0000300\.2580\.352\-0\.0000062024\-12\-18New model release0\.263<0\.001∗∗∗\-0\.0003330\.2920\.001∗∗∗0\.0003160\.9850\.6850\.000017
Notes:The pre\-period p\-value reports the joint test of the null hypothesis that the pre\-event treatment\-effect coefficients are equal to zero\. Under the average pre\-treatment normalization, this corresponds to testing for differential pre\-event dynamics between unverified and verified observations\.
The post\-period p\-value reports the joint test of the null hypothesis that the post\-event treatment\-effect coefficients are equal to zero\.
The average post coefficient is the mean of the seven estimated post\-event treatment\-effect coefficients,δ^1,…,δ^7\\hat\{\\delta\}\_\{1\},\\ldots,\\hat\{\\delta\}\_\{7\}\.
Each row is estimated separately for the indicated event date\. The event\-date\-specific regressions include company fixed effects but omit week fixed effects\. Standard errors are clustered by company\.
Bold entries indicate cases in which the pre\-period joint test is not statistically significant at the 10% level \(p\-value≥\\geq0\.10\) and the post\-period joint test is statistically significant at the 10% level \(p\-value < 0\.10\)\.
Significance levels are denoted by∗p<0\.10\{\}^\{\*\}p<0\.10,p∗∗<0\.05\{\}^\{\*\*\}p<0\.05, and∗∗∗p<0\.01\{\}^\{\*\*\*\}p<0\.01\.
Table S10:Difference\-in\-Differences Estimates by Company Size: First \(Lowest\) Review\-Volume QuantileReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)Treatment\-0\.9068\*\*\*0\.2333\*\*\*\-0\.1909\*\*\*\-0\.000139\*\*\*0\.000139\*\*\*\(0\.056\)\(0\.014\)\(0\.013\)\(0\.00005\)\(0\.00005\)After\-0\.01570\.0051\-0\.0016\-0\.0000330\.000033\(0\.020\)\(0\.005\)\(0\.006\)\(0\.00006\)\(0\.00006\)Treatment×\\timesAfter\-0\.0348\*0\.0060\-0\.0117\*\*0\.000032\-0\.000032\(0\.019\)\(0\.005\)\(0\.006\)\(0\.00008\)\(0\.00008\)Company FEYesYesYesYesYesWeek FEYesYesYesYesYesMean \(After = 0\)3\.8480\.2470\.6530\.99990\.0001R\-squared0\.5380\.5160\.4540\.0060\.006Observations57,14457,14457,144373,264373,264
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. The high\-volume 100\+ specification could not be estimated because of insufficient variation in the outcome variable\. Statistical significance levels are denoted as follows: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, and \*p<0\.1p<0\.1\.
Table S11:Difference\-in\-Differences Estimates by Company Size: Second Review\-Volume QuantileReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)Treatment\-0\.9231\*\*\*0\.2397\*\*\*\-0\.1917\*\*\*\-0\.000275\*\*\*0\.000275\*\*\*\(0\.055\)\(0\.014\)\(0\.013\)\(0\.00008\)\(0\.00008\)After0\.0181\-0\.00530\.00350\.000032\-0\.000032\(0\.016\)\(0\.004\)\(0\.005\)\(0\.00010\)\(0\.00010\)Treatment×\\timesAfter\-0\.01340\.0033\-0\.00220\.000138\-0\.000138\(0\.015\)\(0\.004\)\(0\.005\)\(0\.00012\)\(0\.00012\)Company FEYesYesYesYesYesWeek FEYesYesYesYesYesMean \(After = 0\)3\.9010\.2320\.6640\.99970\.0003R\-squared0\.5210\.5080\.4330\.0060\.006Observations79,54679,54679,546379,350379,350
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. The high\-volume 100\+ specification could not be estimated because of insufficient variation in the outcome variable\. Statistical significance levels are denoted as follows: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, and \*p<0\.1p<0\.1\.
Table S12:Difference\-in\-Differences Estimates by Company Size: Third Review\-Volume QuantileReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)High\-volume100\+Treatment\-0\.9421\*\*\*0\.2370\*\*\*\-0\.2102\*\*\*\-0\.0001760\.0001660\.000010\(0\.049\)\(0\.012\)\(0\.012\)\(0\.00017\)\(0\.00017\)\(0\.00002\)After\-0\.0022\-0\.0009\-0\.00160\.000194\-0\.000134\-0\.000061\*\*\(0\.012\)\(0\.003\)\(0\.004\)\(0\.00018\)\(0\.00018\)\(0\.00003\)Treatment×\\timesAfter0\.0049\-0\.00130\.0006\-0\.0001960\.0001550\.000041\(0\.011\)\(0\.003\)\(0\.003\)\(0\.00022\)\(0\.00022\)\(0\.00003\)Company FEYesYesYesYesYesYesWeek FEYesYesYesYesYesYesMean \(After = 0\)4\.0270\.1990\.6940\.99930\.00070\.0000R\-squared0\.4950\.4790\.4090\.0140\.0140\.005Observations119,181119,181119,181386,050386,050386,050
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.
Table S13:Difference\-in\-Differences Estimates by Company Size: Fourth \(Highest\) Review\-Volume QuantileReview rating outcomesReview\-volume categoriesAvg\. ratingProp\. 1\-starProp\. 5\-starLow\-volume\[0,20\)\[0,20\)Moderate\-volume\[20,100\)\[20,100\)High\-volume100\+Treatment\-1\.1463\*\*\*0\.2894\*\*\*\-0\.2556\*\*\*0\.055033\*\*\*\-0\.047906\*\*\*\-0\.007127\*\*\*\(0\.041\)\(0\.010\)\(0\.010\)\(0\.00548\)\(0\.00508\)\(0\.00162\)After0\.0109\-0\.00180\.00350\.009347\*\*\*\-0\.008920\*\*\*\-0\.000427\(0\.007\)\(0\.002\)\(0\.002\)\(0\.00137\)\(0\.00141\)\(0\.00043\)Treatment×\\timesAfter\-0\.0217\*\*\*0\.0033\-0\.0077\*\*\*\-0\.0008690\.000911\-0\.000042\(0\.008\)\(0\.002\)\(0\.002\)\(0\.00095\)\(0\.00096\)\(0\.00037\)Company FEYesYesYesYesYesYesWeek FEYesYesYesYesYesYesMean \(After = 0\)4\.0660\.1800\.6910\.94880\.04680\.0044R\-squared0\.4460\.4220\.3930\.2520\.2280\.217Observations198,571198,571198,571391,438391,438391,438
Notes:Estimates from the baseline Difference\-in\-Differences specification with company and week fixed effects\. Standard errors are reported in parentheses\. Statistical significance levels: \*\*\*p<0\.01p<0\.01, \*\*p<0\.05p<0\.05, \*p<0\.1p<0\.1\.相似文章
大语言模型威胁双盲评审
本文证明,仅凭标题和摘要,大语言模型就能有效解除科学论文作者匿名,从而威胁到双盲同行评审的有效性。作者认为,问题框架和研究焦点中的稳定模式构成了作者身份的潜在概念特征,这要求我们在人工智能增强的研究生态系统中重新评估匿名实践。
对AI辅助同行评议的操纵给科学界带来新风险
一项新研究表明,AI辅助的同行评审易通过廉价手段被操控——仅需对论文摘要进行表面改写,即可显著提高AI生成的评审分数,并可能使人类编辑决策产生偏差,凸显了建立防护措施的必要性。
当AI评审训练AI评审者:科学判断崩溃与缓解
本研究探讨了AI生成的评审影响未来AI评审者训练的反馈循环,导致判断多样性降低,这种现象称为'科学判断崩溃'。同时介绍了TrustReviewer,这是一个开源系统,通过精选训练和激活引导来缓解这一问题。
用户对生成式人工智能的看法:应用商店评论中信任与摩擦的跨平台自然语言处理分析
本文呈现了一项跨平台自然语言处理分析,针对六大生成式AI应用的17,012条应用商店评论,识别出广告、认证和定价等关键信任与摩擦障碍,并发现跨平台显著的情感差异。
NeurIPS 2026 AI生成的评审 [D]
关于在NeurIPS 2026中使用AI生成的评审的讨论,包括对提示注入的担忧,以及评审者在没有适当监督的情况下使用LLM却没有后果的问题。