What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

arXiv cs.CL Papers

Summary

This paper presents a cross-platform NLP analysis of 17,012 app store reviews for six major generative AI applications, identifying key trust and friction barriers like advertisements, authentication, and pricing, with significant sentiment variations across platforms.

arXiv:2609.19151v1 Announce Type: new Abstract: Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews. Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications. As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.
Original Article
View Cached Full Text

Cached at: 09/18/26, 08:47 AM

# What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews
Source: [https://arxiv.org/html/2609.19151](https://arxiv.org/html/2609.19151)
\\tnotemark

\[1\]\\tnotetext\[1\]Preprint\. This manuscript has been submitted toArray\(Elsevier\) and is currently under peer review\.

\\cormark

\[1\]\\creditConceptualization, Methodology, Software, Data Curation, Writing \- Original Draft

1\]organization=KFCIS, Florida International University, city=Miami, state=Florida, country=USA

\\credit

Validation, Supervision, Writing \- Review & Editing

2\]organization=DoA, Bangladesh University of Engineering and Technology, city=Dhaka, country=Bangladesh

Umme Nusrat Jahanujaha001@fiu\.eduShouvaggo Sharif Shammoshouvaggo\.shammo@buet\.ac\.bd\[

###### Abstract

Generative AI \(GenAI\) applications have achieved rapid consumer adoption, yet little large\-scale research examines user\-perceived quality, trust, and adoption barriers\. We present, to the best of our knowledge, one of the first cross\-application analyses of app store reviews for six major GenAI applications \(ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity\), comprising 17,012 English\-language reviews from Google Play and the Apple App Store\. We combine BERTopic topic modeling \(all\-MiniLM\-L6\-v2 embeddings\) with RoBERTa sentiment classification, and test cross\-application differences using chi\-square, Kruskal–Wallis, and multinomial logistic regression with Bonferroni correction\. Both components are validated against human coding: manual thematic coding of a stratified 300\-review sample \(inter\-coderκ\\kappa= 0\.544\) and human sentiment labels for the same sample \(RoBERTa accuracy 75\.3%, macro\-F1F\_\{1\}= 0\.73; inter\-coderκ\\kappa= 0\.89\)\. As established, sentiment\-validated findings, negativity concentrates in specific friction topics—advertising \(91% negative\), authentication \(89%\), server reliability \(83%\), and subscription pricing \(73%\)—and sentiment differs significantly across applications; Claude shows the highest negative sentiment \(47\.7%\) alongside a strongly enthusiastic core, a statistically verified polarization\. These cross\-application conclusions are robust to the applications’ unequal review counts\. As exploratory observations, which we frame cautiously because topic modeling captures abstract constructs weakly, a subset of DeepSeek reviews surfaced geopolitical and data\-privacy concerns tied to its Chinese origin, and a proposed Trust Friction Score decomposes each application’s friction into actionable sub\-dimensions \(its rank correlation with negativity,ρ\\rho= 0\.886 over six applications, is illustrative rather than inferential\)\. The study offers validated, actionable evidence on quality and trust barriers in consumer GenAI\.

###### keywords:

Generative AI\\sepApp Store Reviews\\sepTrust\\sepUsability\\sepBERTopic\\sepSentiment Analysis

\{highlights\}

Analyzed 17,012 app store reviews from six leading generative AI applications\.

Applied BERTopic topic modeling and RoBERTa sentiment analysis to identify user concerns\.

Account friction, advertisements, and subscription pricing emerged as major adoption barriers\.

Comparative analysis revealed significant sentiment variation across GenAI platforms\.

Provides one of the first cross\-platform NLP\-based studies of GenAI app reviews\.

## 1Introduction

The evolution of generative AI applications has witnessed a truly remarkable shift from experimental models to consumer\-oriented solutions at an unprecedentedly rapid pace\. Within two months after its release, OpenAI’s ChatGPT had already amassed more than 100 million active users per month, earning it a place among the most rapidly adopted consumer technologies ever created\[hu2023generative\]\. From here, its scope has only expanded to become part of an intensely competitive landscape, which includes Google Gemini, Microsoft Copilot, Anthropic Claude, DeepSeek, and Perplexity, among others, all of whom serve hundreds of millions of users today\[zhao2023survey,chang2024survey\]\.

Mobile app stores have emerged as the main distribution platform for such applications, with generative AI applications consistently making it to the most popular categories on the Google Play and Apple App Stores in the years 2025 and 2026\[statista2025genai\]\. In view of the rapid adoption of these applications, generative AI offers a unique setting for the study of user perceptions and difficulties related to intelligent systems\.

In spite of such fast adoption, the academic knowledge regarding the user experience of generative AI tools used in daily life still lacks attention\. The current literature demonstrates that user trust in AI\-based systems depends on a combination of several factors, which include accuracy of the system, its output clarity, privacy of data, and general context\[zhang2020effect,jacovi2021formalizing,siau2018building\]\.

As illustrated by usability research, design quality, onboarding processes, authentication systems, and affordances have been shown to serve as critical mediators between the level of model capabilities and user satisfaction\[amershi2019software,nielsen1994usability\]\. Nevertheless, the current empirical literature in this area is overwhelmingly dominated by laboratory experiments, interview studies with limited sample sizes, or researcher\-developed surveys that measure the attitudes of users at one point in time\[jakesch2023human,kapania2022because\]\. Despite their importance in pinpointing particular factors, these methods fall short when it comes to incorporating the wide variety and spontaneous nature of issues that arise through users’ natural interaction with generative AI tools\.

App store reviews represent a uniquely valuable and underutilized data source for addressing this gap\. In contrast to survey\-based methods, review writing is an optional task performed spontaneously when using the application, hence reflecting users’ spontaneous responses to concerns important for them, rather than predetermined answer scales prepared by the researcher\[pagano2013user,chen2014ar\]\. Review mining has already been successfully applied in the domain of software engineering research for requirement elicitation, bug detection, feature prioritization, and user satisfaction evaluation\[maalej2016automatic,guzman2014users,martin2017survey\]\. Because mobile app stores serve as the primary access point for consumer\-facing generative AI tools, the reviews they accumulate represent the voices of the broadest possible user base, including non\-expert users who would be unlikely to participate in formal academic studies or controlled experiments\. However, current efforts to use app store reviews to develop applications of generative AI have been constrained both in scope and in terms of methodology\. According to Alabduljabbar\[alabduljabbar2024\], just five GenAI applications and 11,549 reviews were studied using VADER sentiment analysis and LDA topic modeling, both of which are capable of capturing sentiment polarity but fail to capture the complexities of the discussion around AI\. In Meng et al\.’s paper\[meng2026\], their model was scaled up to around 100,000 reviews, yet only considered feature extraction, ignoring any consideration of trust and adoption models\. Importantly, neither of the two studies used recent advancements in technology, like DeepSeek, Claude, or Perplexity, which belong to entirely new classes of artificial intelligence\. Most importantly, to the best of our knowledge, few if any existing studies have framed generative AI app reviews through the combined lens of trust, usability, and responsible adoption, which is precisely the analytical perspective called for by the growing body of work on human factors in AI systems\[liao2023ai,weisz2024design\]\.

This paper addresses these gaps by presenting a large\-scale, mixed\-methods analysis of user reviews collected from six major generative AI applications: ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity\. As summarized in Table[1](https://arxiv.org/html/2609.19151#S3.T1), we collected 54,748 raw reviews from both Google Play and the Apple App Store\. After language filtering, minimum length enforcement, and deduplication, 17,012 English\-language reviews were retained for analysis, spanning the period from September 2025 to May 2026\. We employ BERTopic\[grootendorst2022bertopic\]with contextual sentence embeddings \(all\-MiniLM\-L6\-v2\) for topic modeling, complemented by transformer\-based sentiment classification using the RoBERTa\-based cardiffnlp/twitter\-roberta\-base\-sentiment\-latest model\[loureiro2022timelms\]\. These automated analyses are supplemented by chi\-square and non\-parametric statistical tests with Bonferroni correction, and validated through manual thematic coding of a stratified 300\-review sample\. The study is guided by three research questions:

- RQ1\.What are the dominant user concerns and discussion themes expressed in app store reviews of generative AI applications?
- RQ2\.How does user sentiment differ across major generative AI applications and between mobile platforms?
- RQ3\.Which trust, usability, and adoption factors are most strongly associated with negative user experiences?

The contributions of this work are threefold:

1. 1\.One of the first comparative cross\-application studiesof user\-generated reviews for generative AI applications, covering six competing products, including three \(DeepSeek, Claude, and Perplexity\) that, to the best of our knowledge, have not been examined in the app store mining literature, with 17,012 analyzed reviews drawn from both Google Play and the Apple App Store\.
2. 2\.Methodological advancementover the LDA and VADER pipeline employed in prior work\[alabduljabbar2024\], demonstrating that BERTopic with contextual embeddings and RoBERTa\-based transformer sentiment analysis achieves substantially finer thematic granularity\. We also contribute a methodological finding: BERTopic works effectively with lexically distinct themes \(e\.g\., problems with accounts, subscription\-related grievances\), but it cannot handle abstract themes like trust and privacy \(κ=0\.241\\kappa=0\.241\), which highlights the importance of manually validating themes when researching mobile app reviews\.
3. 3\.Empirical findings on trust, usability, and adoption inhibitorsin consumer GenAI, including authentication and advertising as dominant friction points, subscription pricing as an adoption barrier, an exploratory geopolitical\-trust theme unique to DeepSeek, and a polarization pattern in which Claude combines the highest negative sentiment with a strongly enthusiastic core\. Full statistics and significance tests are reported in Section[4](https://arxiv.org/html/2609.19151#S4)\.

The remainder of this paper is organized as follows\. Section[2](https://arxiv.org/html/2609.19151#S2)reviews related work on trust and user perception of generative AI, usability evaluation of AI applications, and app store review mining\. Section[3](https://arxiv.org/html/2609.19151#S3)describes the research methodology, including data collection, preprocessing, topic modeling, sentiment analysis, statistical testing, and manual thematic validation\. Section[4](https://arxiv.org/html/2609.19151#S4)presents results organized by research question\. Section[5](https://arxiv.org/html/2609.19151#S5)discusses findings and their implications for both research and practice\. Section[6](https://arxiv.org/html/2609.19151#S6)addresses threats to validity, and Section[7](https://arxiv.org/html/2609.19151#S7)concludes the paper\.

## 2Related Work

This chapter presents literature review of three areas related to our study: \(a\) literature on trust and user perception towards Generative AI \(Chapter[2\.1](https://arxiv.org/html/2609.19151#S2.SS1)\), \(b\) literature on usability evaluation of AI applications \(Chapter[2\.2](https://arxiv.org/html/2609.19151#S2.SS2)\) and finally \(c\) literature related to app\-store review mining as a research method \(Chapter[2\.3](https://arxiv.org/html/2609.19151#S2.SS3)\)\.

### 2\.1Trust and User Perception of Generative AI

The concept of trust has long been established as an essential factor influencing the adoption of new technology by potential users\. The Technology Acceptance Model proposed by Davis clearly states that perceived usefulness and perceived ease of use are two key elements that together affect behavioral intention towards adopting a new system\[davis1989tam\], while recent extensions to TAM include trust as a separate factor that gains importance in uncertain situations\. Within the domain of automation and intelligence, Hoff and Bashir\[hoff2015trust\]developed an elaborate three\-tier theory of trust by synthesizing over a decade’s worth of empirical research on trust in automation\. In particular, the authors identified dispositional trust \(the innate personality factor that predisposes one to trust others\), situational trust \(factors that arise from the situation, like task complexity and risk\), and learned trust \(trust gained through experience with the system\)\.

Due to the appearance of generative AI, the issue of trust has gained even more complexity than before\. While other automated processes followed a defined algorithm, generative AI systems generate output flexibly and creatively, making it impossible to check the results or recognize mistakes\[zhang2020effect,jacovi2021formalizing\]\. According to Siau and Wang, the development of initial trust in AI systems can be explained by representation \(how the system introduces itself\) and the understanding of technology, whereas further trust is based on the perceived performance of the system, process visibility, and its use\[siau2018building\]\. Empirical studies validate these theoretical expectations\. In a study of 48,000 individuals in 47 nations conducted globally in 2025, it was noted that while 66% of people use AI on a consistent basis, only 46% are willing to put their trust in the system\[gillespie2025trust\]\. This trend of lower levels of trust is not universal but depends on the field of application, user profile, and, more importantly, nationality\.

Geopolitical considerations related to trust have gained importance with the development of Chinese AI technologies\. The quick international growth of DeepSeek after its debut in January 2025 was met with immediate worries about security, resulting in government prohibitions in Taiwan, parts of the US government, and some European organizations\[deepseek2025security\]\. Independent security assessments revealed significant vulnerabilities in DeepSeek’s safety filtering mechanisms\[cisco2025deepseek\], and its privacy policy confirms that user data is stored on servers in China, subject to Chinese data governance laws\. These geopolitical trust concerns represent a qualitatively different category from the performance\-based trust factors traditionally studied in HCI research, yet no empirical study has examined how ordinary users articulate these concerns in their own words through naturalistic feedback channels such as app store reviews\.

A growing body of qualitative and mixed\-methods research has begun examining trust in generative AI through user experience lenses\. Studies using interview and diary methods have explored how users develop trust in AI\-mediated emotional support\[generativeconfidants2025\], while others have investigated the role of transparency and explainability in calibrating trust in large language model outputs\[liao2023ai\]\. Survey\-based studies have confirmed that trust in generative AI is multidimensional, encompassing beliefs about competence, benevolence, and reciprocity\[chen2025trustscale\]\. However, these studies typically recruit self\-selected participants who are already aware of and interested in AI, potentially missing the perspectives of casual or frustrated users who interact with generative AI applications through mainstream mobile channels\. Our study complements this literature by mining the unsolicited trust\-related discourse of a broad, self\-selected user population through app store reviews\.

### 2\.2Usability Evaluation of AI Applications

Usability, defined by the ISO 9241\-11 standard as the extent to which a system can be used to achieve specified goals with effectiveness, efficiency, and satisfaction in a specified context of use\[iso9241\], has been a cornerstone of human\-computer interaction research for decades\. Ten usability heuristics provided by Nielsen\[nielsen1994usability\]still prove very applicable, presenting a systematic approach for finding interaction problems that relate to visibility of the system status, matching system and real\-world conventions, users’ control and freedom, consistency, prevention of errors, recognition rather than recall, flexibility, aesthetic simplicity, recovery from errors, and documentation\. Nevertheless, AI\-driven systems have an interplay of elements that makes them incompatible with traditional usability guidelines\. This is because of the unpredictability of AI\-generated results, the challenge of managing users’ expectations in a non\-deterministic environment, and the necessity to gracefully degrade in cases of system failures\[amershi2019guidelines,yang2020reexamining\]\.

This problem was tackled by Amershi et al\.\[amershi2019guidelines\], who suggested 18 evidence\-based guidelines for the design of human\-AI interactions, which were evaluated through testing with 49 design professionals on 20 different AI\-enabled applications\. These guidelines cover all four stages of interaction: beginning \(communicating the capabilities of the system to the user\), while interacting \(displaying contextual information to the user\), making mistakes \(enabling quick rectification\), and learning from user actions\. Weisz et al\.\[weisz2024design\]further elaborated on these findings by applying them to generative AI systems, laying out principles for designing software where the output is inherently variable and based on users’ interactions\. The researchers pointed out that the usability problems associated with generative AI arise not from single instances of interaction but from the long\-term effects of using the software repeatedly, such as rate\-limiting messages, choosing subscription tiers, and authenticating access\.

The limitations of the mobile platform add more usability concerns\. Xu\[xu2019ai\]pointed out that AI software developed for mobile platforms encounters special difficulties, such as the restricted display area for displaying AI logic, the touch interface for user input, and variations in performance among different platforms\. Research on usability analysis of mobile applications revealed that loading time, navigational difficulty, and cross\-platform compatibility have a considerable impact on user experience and user retention\[baharuddin2013usability,coursaris2012meta\]\.

The overlap between the specific usability problems faced by AI with mobile platform limitations is especially important in the case of generative AI apps, which need to strike a balance between their complex features \(multiple conversation turns, uploading documents, image generation, voice interactions\) and the limitations posed by mobile platforms\. The first systematic usability testing of generative AI apps was carried out by Alabduljabbar\[alabduljabbar2024\], who applied the ISO 9241 standards on usability to five GenAI apps based on their reviews available in the app stores\. However, the usability scores of these apps were calculated using a compound score obtained from VADER sentiment analysis\.

### 2\.3App Store Review Mining and Analysis

The analysis of user feedback on the app stores has proved to be a mature research technique used in software engineering for obtaining actionable data\. Pagano and Maalej\[pagano2013user\]have proved that app store reviews are composed of diverse information such as bugs, requests for new features, positive comments about the product, and its use cases\. Further research efforts have resulted in systematic methodologies for carrying out such an analysis\. In particular, Guzman and Maalej\[guzman2014users\]introduced a method combining sentiment and topic analysis to automatically identify and aggregate fine\-grained opinions about certain app features, whereas Maalej et al\.\[maalej2016automatic\]created a set of classifiers that can classify app reviews into four categories of bug reports, feature requests, user experience, and rating with remarkable accuracy\. Martin et al\.\[martin2017survey\]conducted a thorough survey of the relevant scientific literature and found that review mining was the most popular field of study\.

Topic modeling has remained a key tool in the analysis of reviews from app stores, with Latent Dirichlet Allocation \(LDA\) remaining the go\-to method for over a decade\[blei2003lda\]\. In LDA, documents are modeled as mixtures of topics, and topics as mixtures of words, thus making it possible to find hidden topics within large data sets\. But LDA makes use of the bag\-of\-words model, which disregards word order and context, hence failing to produce relevant topics\[egger2022topic\]\.BERTopic\[grootendorst2022bertopic\]is a considerable methodology innovation in itself, where pre\-trained transformer embeddings combined with dimensionality reduction and density clustering result in the generation of more semantically meaningful and coherent topics\. Experimental studies comparing BERTopic and LDA show that BERTopic outperforms LDA in terms of topic coherence and diversity metrics across various short\-text datasets\[egger2022topic\]\. Recent research utilizes BERTopic for health and fitness application reviews, revealing specific user problems that are left unnoticed by LDA\-driven analysis methods\[healthfitness2025bertopic\]\.

For sentiment analysis in review mining, the Valence Aware Dictionary and Sentiment Reasoner \(VADER\)\[hutto2014vader\]has been widely used due to its simplicity and effectiveness on short social media texts\. However, VADER is a lexicon\-based approach that relies on predefined sentiment dictionaries and grammatical rules, making it insensitive to domain\-specific language, sarcasm, and complex sentiment expressions common in app reviews\[hutto2014vader\]\. Transformer\-based sentiment models, particularly those fine\-tuned on social media data such as the RoBERTa\-based cardiffnlp/twitter\-roberta\-base\-sentiment\-latest model\[loureiro2022timelms\], offer substantially improved contextual understanding by encoding the full sentence context through self\-attention mechanisms\. While no sentiment model is perfectly calibrated for the app store review domain \(a limitation we explicitly address in our methodology and threats to validity\), transformer\-based approaches represent a meaningful improvement over lexicon\-based baselines for capturing the nuanced sentiment expressions found in user reviews of complex AI applications\.

However, the use of review mining in relation to generative AI applications remains relatively recent and nascent in scope\. Alabduljabbar\[alabduljabbar2024\], for instance, carried out the pioneering systematic research effort using VADER sentiment analysis and LDA topic modeling in the context of an ISO 9241 usability assessment framework based on 11,549 reviews from five GenAI apps: ChatGPT, Bing AI, Microsoft Copilot, Gemini, and Da Vinci AI gathered during January through March 2024\. Meng et al\.\[meng2026\], on the other hand, analyzed reviews numbering close to 100,000 but only did so in terms of feature extraction, lacking any theoretical grounding with respect to trust and adoption\. The most recent use of BERTopic involved analyzing DeepSeek Chinese reviews using BERTopic in combination with sentiment analysis to map the results onto the five components of the user experience defined by Garrett\[deepseek2026chinese\]\. Nevertheless, this particular research only considered one use case and one language\.

##### Research gap\.

Despite the growing sophistication of both trust research and app store mining methods, no existing study integrates these streams to examine generative AI applications comparatively\. Specifically, three gaps remain unaddressed\. First, no cross\-application study has analyzed user reviews of newer generative AI products such as Claude, DeepSeek, and Perplexity, which embody distinct design philosophies \(safety\-focused, open\-source Chinese\-origin, and search\-augmented, respectively\) that may elicit qualitatively different user concerns\. Second, no study has applied BERTopic\-based topic modeling with transformer sentiment analysis to GenAI app reviews, meaning that the thematic landscape reported in prior work reflects the limitations of LDA and VADER rather than the actual complexity of user discourse\. Third, no study has framed GenAI app reviews through the combined lens of trust, usability, and responsible adoption, the precise intersection identified by the IST special issue on human factors in generative AI as requiring empirical investigation\. The present study addresses all three gaps\.

## 3Research Methodology

### 3\.1Research Design Overview

The current research adopts a computational approach for the analysis of user\-generated app store reviews of generative AI software using a mixed\-methods research design\. In terms of methodology, the entire research design encompasses six stages: \(i\) data gathering from the Google Play Store and Apple App Store; \(ii\) data pre\-processing through language filtering, length filtering, and deduplication; \(iii\) topic modeling using BERTopic for identification of common themes in user complaints, \(iv\) sentiment analysis using a transformer\-based classification algorithm to measure the degree of positive, neutral, and negative sentiments expressed by the reviewers, \(v\) statistical evaluation to determine if there is a statistically significant difference between different applications and platforms, and \(vi\) qualitative thematic validation on a stratified sample of reviews\.

The utilization of reviews from app stores as secondary sources is a common practice within the body of literature in software engineering\. For example, Martin et al\.\[martin2017survey\]concluded that review mining was the most actively researched aspect of app store analytics, and that it had been successfully employed in various areas of software engineering, including requirements extraction, defect identification, and version planning\. Moreover, Guzman and Maalej\[guzman2014users\]proved that the conjunction of topic modeling with sentiment analysis allows for extracting users’ opinions regarding the functionality of particular applications\. Unlike primary data collection through surveys or interviews, app store review mining captures unsolicited, contextually grounded user feedback without researcher\-imposed framing\[pagano2013user\], making it particularly well\-suited for exploratory studies that aim to discover user concerns rather than confirm predefined hypotheses\. We adopt this established methodological paradigm and extend it with modern NLP techniques \(BERTopic, transformer\-based sentiment analysis\) and a theoretical framing grounded in trust, usability, and adoption\.

### 3\.2Data Collection

Six AI Generative models were selected for this paper’s investigation: ChatGPT \(OpenAI\), Google’s Gemini, Microsoft Copilot, Claude \(Anthropic\), DeepSeek, and Perplexity\. These applications were selected based on three criteria: \(1\) they represent the most widely downloaded and actively used generative AI consumer applications as of early 2026, \(2\) they span diverse design philosophies, including general\-purpose conversational AI \(ChatGPT, Gemini\), productivity\-integrated AI \(Copilot\), safety\-focused AI \(Claude\), open\-source Chinese\-origin AI \(DeepSeek\), and search\-augmented AI \(Perplexity\), and \(3\) each is available as a mobile application on at least one major app store platform, ensuring the existence of user review data\.

Reviews were collected from two sources: Google Play and the Apple App Store\. For Google Play, we used the open\-sourcegoogle\-play\-scraperPython library\[googleplayscraper\], which provides programmatic access to publicly available review data, including review text, star rating, timestamp, and review ID\. We collected up to 8,000 reviews per application from Google Play, yielding 48,000 raw Google Play reviews across all six applications\. For the Apple App Store, we used the publicly accessible Apple RSS feed endpoint, which provides the most recent customer reviews for a given application\. Apple App Store reviews were available for three of the six applications \(ChatGPT, Claude, and Perplexity\), yielding 6,748 additional reviews\. Gemini, Microsoft Copilot, and DeepSeek did not have accessible iOS review data at the time of collection, either because the application was not available on iOS in all regions or because the RSS feed returned insufficient data\. In total, 54,748 raw reviews were collected across both platforms\.

Table[1](https://arxiv.org/html/2609.19151#S3.T1)provides a detailed overview of the dataset, including raw review counts by platform, the number of reviews retained after preprocessing, the date range of analyzed reviews, and the average star rating per application\. Data collection was carried out in May 2026\. Data collected by the researcher can be accessed by anyone at any time on the corresponding app store platform, and it does not include any personal information except for the display names\.

Table 1:Dataset overview: raw, cleaned, and analyzed review counts per application\.AppGPASClean\.Anal\.RangeAvgChatGPT8,0002,4501,8411,841May–May 20263\.86Gemini8,0000812812May–May 20263\.50Copilot8,00002,6862,686Feb–May 20263\.99Claude8,0002,0815,0375,037Mar–May 20263\.19DeepSeek8,00003,0423,042Sep–May 20263\.87Perplexity8,0002,2173,5943,594Jan–May 20263\.27TOTAL48,0006,74817,01217,012—3\.54

Note:GP = Google Play; AS = App Store; Anal\. = analyzed\. Date ranges cover analyzed English\-only reviews\.

### 3\.3Data Preprocessing

The corpus of 54,748 reviews was preprocessed through multiple steps in order to guarantee data quality\. First, reviews for both applications were merged into one dataset, where reviews had the same standardized format \(the review itself, its star rating, application name, and timestamp\)\. Second, we used language detection with thelangdetectPython package\[langdetect\], which identified and filtered only those reviews that were written in English\. Language filtering was required since there were numerous reviews in the following languages among the Google Play reviews: Hindi, Arabic, Portuguese, Spanish, etc\. Reviews not in the English language were removed to keep analysis consistent, since both BERTopic embedding and RoBERTa sentiment classifier were trained on the English language\. After applying the language filter, 52,080 reviews remained\.

Thirdly, we filtered the reviews with the minimum length criterion of 10 words per review\. The rationale behind such filtering is based on the fact that very short reviews \(such as one\-worded responses like "good" or "bad" and short phrases such as "doesn’t work"\) contain not enough text material for topic modeling due to the need for the context of BERTopic embeddings\[grootendorst2022bertopic\]\. The number of reviews decreased to 17,550\. This minimum\-length filter removed 34,530 of the 52,080 post\-language\-filter reviews \(66%\)\. Because the discarded short reviews are disproportionately brief praise, this step biases the retained corpus toward longer, more critical reviews; we quantify the direction and magnitude of this bias in Section[6](https://arxiv.org/html/2609.19151#S6)and Appendix[C](https://arxiv.org/html/2609.19151#A3)\. Fourth, we applied exact\-text duplication removal to filter out duplicate reviews from different scrapes or that were posted multiple times by the same user\. Following duplication removal and a second application of the language filter for posts incorrectly classified during the first attempt, the final corpus analyzed contained 17,012 reviews\.

Text cleansing involved very little effort\. The case, punctuation, and spellings of the reviews were maintained since the sentence transformer embedding used in BERTopic is trained on natural language and requires such features for optimal functioning\. Stemming, lemmatization, and stop\-word elimination were not employed in the preprocessing phase; rather, stop\-word elimination was incorporated into theCountVectorizerof BERTopic during the topic representation phase\.

### 3\.4Topic Modeling with BERTopic

We employed BERTopic\[grootendorst2022bertopic\]for topic modeling, a method that combines pre\-trained sentence embeddings with dimensionality reduction and density\-based clustering to discover latent topics in text corpora\. BERTopic was selected over the more commonly used Latent Dirichlet Allocation \(LDA\)\[blei2003lda\]for three reasons: \(1\) BERTopic leverages contextual embeddings that capture semantic relationships between words, producing more coherent topics than LDA’s bag\-of\-words representation; \(2\) BERTopic uses density\-based clustering \(HDBSCAN\) rather than requiring a fixed number of topics as a prior, allowing the algorithm to discover the natural topic structure in the data; and \(3\) comparative evaluations have demonstrated BERTopic’s superior performance on short\-text corpora such as app reviews and social media posts\[egger2022topic\]\.

The BERTopic pipeline consists of four sequential components, each of which we configured as follows\. For document embedding, we used theall\-MiniLM\-L6\-v2sentence transformer model\[reimers2019sentencebert\], which maps each review into a 384\-dimensional dense vector\. This model was selected for its strong balance between embedding quality and computational efficiency, and it is one of the most widely used models for BERTopic applications\. For dimensionality reduction, we applied UMAP \(Uniform Manifold Approximation and Projection\)\[mcinnes2018umap\]to reduce the 384\-dimensional embeddings to 5 dimensions, using the parametersn\_neighbors=15,n\_components=5, andmetric=’cosine’\. The 5\-dimensional representation preserves the local neighborhood structure of the high\-dimensional embeddings while making the subsequent clustering step computationally feasible\.

For clustering, we used HDBSCAN \(Hierarchical Density\-Based Spatial Clustering of Applications with Noise\)\[campello2013hdbscan\]withmin\_cluster\_size=30, meaning that a group of semantically similar reviews must contain at least 30 members to be recognized as a distinct topic\. Reviews that HDBSCAN could not assign to any cluster were labeled as outliers \(Topic−1\-1\)\. For topic representation, we used aCountVectorizerwith English stop\-word removal,min\_df=10\(words must appear in at least 10 documents\), andngram\_range=\(1, 2\)\(capturing both unigrams and bigrams\)\. BERTopic’s class\-based TF\-IDF \(c\-TF\-IDF\) procedure then extracted the most representative terms for each cluster, producing human\-interpretable topic labels\.

The initial clustering produced 63 topics\. To improve interpretability and reduce topic fragmentation, we applied BERTopic’s built\-in topic reduction mechanism withnr\_topics=25, which hierarchically merges the most similar topics until the target count is reached\. After reduction, 24 meaningful topics remained \(T00 through T23\), plus the outlier topic \(−1\-1\)\. Of the 17,012 reviews, 12,430 \(73\.1%\) were assigned to one of the 24 topics, while 4,582 reviews \(26\.9%\) remained as outliers\. Each topic was then manually assigned an interpretive label \(e\.g\., “Sign\-in / account issues,” “Subscription & pricing,” “Language & trust \(Chinese\)”\) by examining the top keywords, representative reviews, and thematic content\. The full topic listing with keywords, review counts, and corpus percentages is presented in Table[2](https://arxiv.org/html/2609.19151#S4.T2)in Section[4](https://arxiv.org/html/2609.19151#S4)\.

### 3\.5Sentiment Analysis

Sentiment analysis was performed using the RoBERTa\-basedtwitter\-roberta\-base\-sentiment\-latestmodel from Cardiff NLP\[loureiro2022timelms\]\. The model was fine\-tuned on approximately 124 million tweets for a three\-class sentiment classification \(positive, neutral, and negative\)\.

Compared with the lexicon\-based VADER approach\[hutto2014vader\]used in prior GenAI review studies\[alabduljabbar2024\], transformer\-based models capture full sentence context through self\-attention, improving robustness to negation, sarcasm, and domain\-specific expressions\. Each review was tokenized with a sequence length of 512 tokens at most and categorized according to its sentiment\. The model outputs a softmax probability distribution over the three classes; we assigned each review the class with the highest probability and retained the raw probability scores for downstream statistical analysis\. The processing was done in batches of 64 reviews\. We recognize that there is an inherent domain mismatch due to the fact that this classifier was initially trained using Twitter data and not data from app store reviews\. Reviews in the app stores can be much more formalized than posts made on Twitter\. This will be discussed further in Section[6](https://arxiv.org/html/2609.19151#S6), and some measure of mitigation has been achieved via manual verification described in Section[3\.10](https://arxiv.org/html/2609.19151#S3.SS10)\.

### 3\.6Sentiment Model Validation

Because the RoBERTa classifier was trained on Twitter data rather than app store reviews, we evaluated its outputs against gold\-standard sentiment labels using two complementary procedures\.

First, as a preliminary corpus\-wide robustness check, we treated the user\-assigned star rating as a weak sentiment proxy for all 17,012 reviews, mapping ratings of 1–2 stars tonegative, 3 stars toneutral, and 4–5 stars topositive\. We then compared the RoBERTa prediction for each review against this proxy label and computed accuracy, macro\-averagedF1F\_\{1\}, and a confusion matrix\. This proxy is deliberately conservative: star ratings and textual sentiment do not always coincide \(for example, a five\-star review may contain a specific complaint\), so it provides a lower\-bound sanity check rather than a definitive evaluation\.

Second, to obtain a human\-validated estimate, two coders independently labelled the sentiment \(negative,neutral, orpositive\) of the same stratified 300\-review sample used for thematic validation \(Section[3\.10](https://arxiv.org/html/2609.19151#S3.SS10)\), judging each review from its text alone and blind to both the star rating and the RoBERTa output\. Inter\-coder reliability was quantified with Cohen’sκ\\kappaand Krippendorff’sα\\alpha; disagreements were adjudicated by the first author to produce a gold\-standard label set\. RoBERTa performance was then reported as accuracy, macro\-F1F\_\{1\}, per\-class precision/recall/F1F\_\{1\}, and a confusion matrix against these human gold labels\. Results are presented in Section[4\.5](https://arxiv.org/html/2609.19151#S4.SS5)\.

### 3\.7Statistical Analysis

In order to determine if there are any statistical differences between the distribution of sentiment scores and star ratings across different applications and different platforms, we performed non\-parametric tests and tests on categorical variables\. The reason for opting for non\-parametric methods is that the star rating distribution and sentiment scores do not have a normal distribution\.

The chi\-square test of independence was conducted at the omnibus level to test whether there is any difference between the distribution of sentiments \(positive, neutral, negative\) among the six apps\. We also applied the Kruskal\-Wallis H test\[kruskal1952\]to evaluate whether the distribution of star ratings differs across applications, as star ratings are ordinal rather than interval\-scaled\. For platform comparisons \(Android vs\. iOS\), we used Mann\-Whitney U tests to compare the distributions of positive sentiment scores and negative sentiment scores between the two platforms\.

For all significant omnibus tests, post\-hoc pairwise comparisons were conducted with Bonferroni correction to control the family\-wise error rate\. With six applications, the 15 pairwise comparisons required a corrected significance threshold ofα=0\.05/15=0\.003\\alpha=0\.05/15=0\.003\. Effect sizes were computed for all tests: Cramer’s V for chi\-square tests \(with established thresholds of 0\.1 = small, 0\.3 = medium, 0\.5 = large\), eta\-squared \(η2\\eta^\{2\}\) for Kruskal\-Wallis \(0\.01 = small, 0\.06 = medium, 0\.14 = large\), and rank\-biserial correlationrrfor Mann\-Whitney U tests \(0\.1 = small, 0\.3 = medium, 0\.5 = large\)\[cohen1988statistical\]\. The reporting of effect sizes along with thepp\-values is crucial to understand practical importance, especially for large samples, since even very small differences can be statistically significant\.

### 3\.8Multinomial Logistic Regression

To go beyond descriptive cross\-tabs and measure the independent effect of each predictor on the review sentiment, we conduct a multinomial logistic regression analysis\. Although the chi\-square test and Kruskal\-Wallis test conducted in Section[3\.7](https://arxiv.org/html/2609.19151#S3.SS7)show that sentiment distributions vary by application, they do not allow us to separate the effects of the application name, topic affiliation, social media site, time, and review length\. A regression framework addresses this limitation by estimating the effect of each predictor while controlling for the others\.

Let the sentiment label of reviewiibeYiY\_\{i\}\. The possible sentiment classes arepositive,neutral, andnegative, withpositivetreated as the reference category\. The multinomial logistic regression models the log\-odds of each non\-reference class against the reference class:

ln⁡P​\(Yi=k\)P​\(Yi=pos\)\\displaystyle\\ln\\frac\{P\(Y\_\{i\}=k\)\}\{P\(Y\_\{i\}=\\textit\{pos\}\)\}=β0​k\+β1​k​𝐀𝐩𝐩i\+β2​k​𝐓𝐨𝐩𝐢𝐜i\\displaystyle=\\beta\_\{0k\}\+\\beta\_\{1k\}\\mathbf\{App\}\_\{i\}\+\\beta\_\{2k\}\\mathbf\{Topic\}\_\{i\}\(1\)\+β3​k​Platformi\+β4​k​Lengthi\\displaystyle\\qquad\+\\beta\_\{3k\}\\text\{Platform\}\_\{i\}\+\\beta\_\{4k\}\\text\{Length\}\_\{i\}\+β5​k​Monthi\\displaystyle\\qquad\+\\beta\_\{5k\}\\text\{Month\}\_\{i\}
where𝐀𝐩𝐩i\\mathbf\{App\}\_\{i\}is a vector of dummy variables encoding the six applications \(reference: ChatGPT\),𝐓𝐨𝐩𝐢𝐜i\\mathbf\{Topic\}\_\{i\}is a vector of dummy variables encoding the 24 BERTopic topics plus the outlier category \(reference: T00, General positive\),Platformi\\text\{Platform\}\_\{i\}is a binary indicator for iOS versus Android,Lengthi\\text\{Length\}\_\{i\}is the log\-transformed word count of the review \(to reduce right skew\), andMonthi\\text\{Month\}\_\{i\}is a set of dummy variables encoding the calendar month of the review to control for temporal confounds such as feature releases or pricing changes\.

The model was estimated using maximum likelihood via thestatsmodels\.MNLogitimplementation in Python\. We verified model convergence, assessed multicollinearity through variance inflation factors \(VIF<5<5for all predictors after excluding structurally correlated dummies\), and evaluated overall model fit using McFadden’s pseudo\-R2R^\{2\}and the likelihood ratio test against an intercept\-only null model\. Exponentiated coefficientsexp⁡\(β\)\\exp\(\\beta\)are reported as odds ratios \(OR\) with 95% Wald confidence intervals throughout\. An odds ratio greater than 1 indicates that the predictor increases the odds of the outcome category \(neutral or negative\) relative to positive sentiment\.

To assess whether specific trust\-friction topics affect sentiment differently depending on the application, we also estimated an extended model that includes interaction terms between the five trust\-friction topics \(T05, T12, T13, T14, T16\) and the application variable:

ln⁡P​\(Yi=k\)P​\(Yi=pos\)=β0​k\+𝜷k⊤​𝐗i\+∑t∈𝒯frictionγt​k​\(𝐀𝐩𝐩i×𝐓𝐨𝐩𝐢𝐜i,t\)\\ln\\frac\{P\(Y\_\{i\}=k\)\}\{P\(Y\_\{i\}=\\textit\{pos\}\)\}=\\beta\_\{0k\}\+\\boldsymbol\{\\beta\}\_\{k\}^\{\\top\}\\mathbf\{X\}\_\{i\}\+\\sum\_\{t\\in\\mathcal\{T\}\_\{\\text\{friction\}\}\}\\gamma\_\{tk\}\\,\(\\mathbf\{App\}\_\{i\}\\times\\mathbf\{Topic\}\_\{i,t\}\)\(2\)
where𝒯friction=\{T05, T12, T13, T14, T16\}\\mathcal\{T\}\_\{\\text\{friction\}\}=\\\{\\text\{T05, T12, T13, T14, T16\}\\\}and𝐗i\\mathbf\{X\}\_\{i\}collects all main\-effect predictors from Equation[1](https://arxiv.org/html/2609.19151#S3.E1)\. A significant interaction termγt​k\\gamma\_\{tk\}indicates that the sentiment impact of a friction topic varies across applications\. We compare the base and interaction models using the Akaike Information Criterion \(AIC\) and a likelihood ratio test\.

### 3\.9Polarization Analysis and Trust Friction Scoring

Beyond testing whether sentiment and ratings differ across applications \(Section[3\.7](https://arxiv.org/html/2609.19151#S3.SS7)\) and modeling which factors drive those differences \(Section[3\.8](https://arxiv.org/html/2609.19151#S3.SS8)\), we introduce two additional quantitative measures designed to capture phenomena that standard tests do not address: user\-base polarization and composite trust friction severity\.

#### 3\.9\.1Polarization Indices

The descriptive observation that several applications, particularly Claude, exhibit bimodal star\-rating distributions \(Section[4\.2](https://arxiv.org/html/2609.19151#S4.SS2)\) requires formal quantification\. We compute two complementary polarization indices for each application\.

##### Bimodality Coefficient \(BC\)\.

The bimodality coefficient\[pfister2013good\]provides a single scalar measure of whether a distribution is unimodal or bimodal:

B​Ca=γa2\+1κa\+3​\(na−1\)2\(na−2\)​\(na−3\)BC\_\{a\}=\\frac\{\\gamma\_\{a\}^\{2\}\+1\}\{\\kappa\_\{a\}\+\\dfrac\{3\(n\_\{a\}\-1\)^\{2\}\}\{\(n\_\{a\}\-2\)\(n\_\{a\}\-3\)\}\}\(3\)
whereγa\\gamma\_\{a\}is the sample skewness,κa\\kappa\_\{a\}is the sample excess kurtosis, andnan\_\{a\}is the number of reviews for applicationaa\. A value ofB​C\>0\.555BC\>0\.555indicates bimodality\[freeman2013assessing\]\. We computeB​CaBC\_\{a\}over the 1–5 star rating distribution for each application\.

##### Esteban\-Ray \(ER\) Polarization Index\.

The Esteban\-Ray index\[esteban1994measurement\]captures the joint effect of group identification and inter\-group alienation:

E​Ra​\(α\)=K​∑i=15∑j=15πa,i1\+α​πa,j​\|si−sj\|ER\_\{a\}\(\\alpha\)=K\\sum\_\{i=1\}^\{5\}\\sum\_\{j=1\}^\{5\}\\pi\_\{a,i\}^\{1\+\\alpha\}\\;\\pi\_\{a,j\}\\;\|s\_\{i\}\-s\_\{j\}\|\(4\)
whereπa,i\\pi\_\{a,i\}is the proportion of reviews for applicationaain star\-rating binii,sis\_\{i\}is the star value,α∈\[1\.0,1\.6\]\\alpha\\in\[1\.0,1\.6\]controls sensitivity to group identification, andKKnormalizes the maximum to 1\. We report results atα=1\.0\\alpha=1\.0andα=1\.6\\alpha=1\.6\. Statistical significance is assessed via 95% bootstrap confidence intervals \(10,000 resamples per application\)\.

#### 3\.9\.2Trust Friction Score \(TFS\)

The TFS is designed to summarise, in a single interpretable number per application, the degree to which its user base is exposed to trust\- and usability\-related friction\. Its construction rests on two principles\. First, a friction dimension should matter to the extent that it is both*prevalent*\(many users encounter it\) and*negatively experienced*\(those users are dissatisfied\)\. Multiplying a topic’s prevalence,na,t/Nan\_\{a,t\}/N\_\{a\}, by its negative\-sentiment rate,NegRatea,t\\text\{NegRate\}\_\{a,t\}, captures exactly this interaction: a rare\-but\-negative topic or a common\-but\-neutral topic contributes little, whereas a topic that is both common and strongly negative contributes a lot\. Second, the friction set is theory\-driven rather than data\-dredged\. The six constituent topics operationalise distinct, well\-established constructs from the trust and technology\-adoption literature: authentication \(T05\) and server reliability \(T14\) correspond to perceived ease of use and system dependability in the Technology Acceptance Model\[davis1989tam\]; subscription pricing \(T12\) and chat/usage limits \(T11\) correspond to the value and effort expectancy of UTAUT\[venkatesh2012consumer\]; geopolitical and language trust \(T13\) corresponds to provider\-level trust\[siau2018building\]; and advertising intrusiveness \(T16\) corresponds to perceived credibility and interruption cost\. Summing the six prevalence\-weighted negativity terms yields a composite that is bounded in\[0,1\]\[0,1\], additively decomposable into these interpretable sub\-dimensions, and equal to the fraction of an application’s reviews that are friction\-related negatives\. Accordingly, we define:

TFSa=∑t∈𝒯frictionna,tNa⋅NegRatea,t\\text\{TFS\}\_\{a\}=\\sum\_\{t\\in\\mathcal\{T\}\_\{\\text\{friction\}\}\}\\frac\{n\_\{a,t\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,t\}\(5\)
where𝒯friction=\{T05, T11, T12, T13, T14, T16\}\\mathcal\{T\}\_\{\\text\{friction\}\}=\\\{\\text\{T05, T11, T12, T13, T14, T16\}\\\},na,tn\_\{a,t\}is the number of reviews from applicationaain topictt,NaN\_\{a\}is the total reviews for applicationaa, andNegRatea,t\\text\{NegRate\}\_\{a,t\}is the negative sentiment proportion\. TFS ranges from 0 to 1; higher values indicate greater exposure to friction\.

We decompose TFS into five sub\-dimensional scores to operationalize the multi\-dimensional trust taxonomy:

TFSaauth\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{auth\}\}=na,T05Na⋅NegRatea,T05\\displaystyle=\\frac\{n\_\{a,\\text\{T05\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T05\}\}\(6\)TFSalimits\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{limits\}\}=na,T11Na⋅NegRatea,T11\\displaystyle=\\frac\{n\_\{a,\\text\{T11\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T11\}\}\(7\)TFSaprice\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{price\}\}=na,T12Na⋅NegRatea,T12\\displaystyle=\\frac\{n\_\{a,\\text\{T12\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T12\}\}\(8\)TFSageo\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{geo\}\}=na,T13Na⋅NegRatea,T13\\displaystyle=\\frac\{n\_\{a,\\text\{T13\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T13\}\}\(9\)TFSareliab\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{reliab\}\}=na,T14Na⋅NegRatea,T14\\displaystyle=\\frac\{n\_\{a,\\text\{T14\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T14\}\}\(10\)TFSaads\\displaystyle\\text\{TFS\}\_\{a\}^\{\\text\{ads\}\}=na,T16Na⋅NegRatea,T16\\displaystyle=\\frac\{n\_\{a,\\text\{T16\}\}\}\{N\_\{a\}\}\\cdot\\text\{NegRate\}\_\{a,\\text\{T16\}\}\(11\)
These six sub\-scores correspond to authentication, chat/usage\-limit, pricing, geopolitical, reliability, and advertising friction, respectively, and sum to the compositeTFSa\\text\{TFS\}\_\{a\}\. Bootstrap 95% CIs \(10,000 resamples\) are computed for TFS and all sub\-scores\.

### 3\.10Manual Thematic Validation

To assess the validity of the automated topic modeling results, we conducted manual thematic coding on a stratified random sample of 300 reviews\. The sample was stratified by application \(50 reviews per application\) and by sentiment class \(approximately equal representation of positive, neutral, and negative reviews within each application stratum\) to ensure that the validation covered the full range of thematic content and sentiment polarity in the corpus\.

A codebook of 12 theme codes was developed through an initial open coding pass on a separate pilot sample of 50 reviews \(not included in the final validation set\)\. The 12 codes were:trust,usability,feature,performance,pricing,privacy,positive\(general satisfaction\),comparison\(cross\-app comparisons\),account\(sign\-in and authentication issues\),language\(language\-related concerns including translation and Chinese\-language trust\),limits\(message or usage restrictions\), andcontent\_quality\(response accuracy, hallucination, or helpfulness\)\. Each review in the validation sample was assigned exactly one primary theme code based on the dominant topic of the review\.

Two independent coders performed the manual coding: the first author served as Coder 1 and a second researcher, both familiar with the app\-review domain, served as Coder 2\. The codebook was first drafted from the open\-coding pilot, then refined over two calibration rounds on small held\-out batches \(10–15 reviews each\) until the coders agreed the code definitions and boundary cases were stable; these calibration reviews were excluded from the validation set\. The two coders then independently assigned exactly one primary theme code to each of the 300 reviews using the finalised codebook, coding from the review text alone and blind to the BERTopic topic assignment to avoid anchoring\. Inter\-coder reliability was computed with Cohen’sκ\\kappaand Krippendorff’sα\\alpha\[cohen1960kappa\]\. The 112 disagreements were then adjudicated by the first author, who re\-read each contested review and selected the final code, yielding a single gold\-standard label set\. Agreement between the automated BERTopic assignments \(mapped to the 12 manual theme codes\) and this gold standard was measured with Cohen’sκ\\kappa\. This same 300\-review stratified sample was subsequently reused for the sentiment\-model validation \(Section[3\.6](https://arxiv.org/html/2609.19151#S3.SS6)\), so that both the topic and sentiment validations rest on a common, documented reference set\. The primary purpose of the thematic validation is diagnostic—to identify where the automated topic model succeeds and fails—rather than to establish a definitive human coding; the results, including inter\-coder reliability, the confusion matrix, and the topics on which BERTopic performs well \(lexically distinctive themes such as account and pricing\) versus poorly \(abstract themes such as trust and privacy\), are presented in Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)\.

### 3\.11Ethical Considerations

Although this study analyses only publicly available data, it raises several ethical considerations that we address explicitly\. First, all reviews were collected from the public review sections of Google Play and the Apple App Store, where users post with the expectation of public visibility; no private, restricted, or authentication\-gated content was accessed, and collection used the platforms’ publicly documented review interfaces\[googleplayscraper\]while respecting rate limits\. Second, we practised data minimisation: usernames were discarded during preprocessing, and only the review text, star rating, application name, platform, country, and timestamp were retained for analysis\. No personally identifying information was collected, stored, or inferred, and we made no attempt to re\-identify, profile, or contact individual reviewers\. Third, all quantitative results are reported in aggregate; the short verbatim excerpts quoted in the paper are already public, non\-sensitive, and contain no author identifiers\. Fourth, because the data are public, non\-interventional, and de\-identified, the work does not constitute human\-subjects research requiring ethics\-board approval under common institutional guidelines; nonetheless we followed privacy\-by\-design principles throughout\. Finally, reviews may contain subjective or unverified assertions about the applications and their providers; we treat this content as evidence of user*perceptions*rather than as factual claims, and the geopolitical\-trust observations in particular \(Section[4\.3](https://arxiv.org/html/2609.19151#S4.SS3)\) are reported as expressed user sentiment, not as endorsements or verified statements about any provider\.

## 4Results

This section presents the findings organized by research question\. Section[4\.1](https://arxiv.org/html/2609.19151#S4.SS1)addresses RQ1 \(dominant user concerns\), Section[4\.2](https://arxiv.org/html/2609.19151#S4.SS2)addresses RQ2 \(sentiment differences across applications\), Section[4\.3](https://arxiv.org/html/2609.19151#S4.SS3)addresses RQ3 \(trust and usability factors in negative experiences\), Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)reports the manual thematic validation results, Section[4\.5](https://arxiv.org/html/2609.19151#S4.SS5)validates the sentiment classifier, Section[4\.6](https://arxiv.org/html/2609.19151#S4.SS6)presents the multivariate sentiment modeling, and Section[4\.7](https://arxiv.org/html/2609.19151#S4.SS7)reports the trust friction scores\.

### 4\.1RQ1: Dominant User Concerns \(Topic Modeling Results\)

BERTopic identified 24 distinct topics from the corpus of 17,012 reviews\. Of these, 12,430 reviews \(73\.1%\) were assigned to one of the 24 topics, while 4,582 reviews \(26\.9%\) were classified as outliers that did not cluster strongly enough to be assigned to any topic\. Table[2](https://arxiv.org/html/2609.19151#S4.T2)presents the top 20 topics with their interpreted theme labels, top keywords extracted via c\-TF\-IDF, review counts, percentage of the total corpus, and the application that contributed the most reviews to each topic\.

The discovered topics can be organized into five thematic categories\. The first and largest category encompassesgeneral satisfaction and quality assessment\. Topic T00 \(General positive experience;n=2,926n=2\{,\}926, 17\.2% of corpus\) captured broad positive feedback using terms such as “good,” “helpful,” and “love,” and was the single largest topic in the corpus\. Topic T01 \(AI quality comparisons;n=1,908n=1\{,\}908, 11\.2%\) reflected users evaluating the overall quality of AI applications, with Claude contributing the most reviews to this topic\.

The second category coverscross\-application and within\-application comparisons\. Topic T02 \(Cross\-app comparisons;n=1,601n=1\{,\}601, 9\.4%\) captured reviews that explicitly compared generative AI applications against each other, with keywords such as “chatgpt,” “gpt,” and “better” indicating that users frequently benchmark one application against ChatGPT\. Topics T06 through T10 captured application\-specific feedback for Copilot \(n=500n=500\), Gemini \(n=476n=476\), Claude \(n=471n=471\), DeepSeek \(n=442n=442\), and Perplexity \(n=429n=429\), respectively\.

The third category addressesfeature\-specific feedback\. Topic T03 \(Image generation;n=822n=822, 4\.8%\) was dominated by ChatGPT reviews discussing image upload and generation capabilities\. T04 Topic Voice and UI Features \(n=633n=633, 3\.7%\) contained feedback relating to voice interaction styles, text characteristics, and interface designs, with Claude providing the highest number of reviews\.

The fourth category comprisesbarriers of trust and friction\. Topic T05 \(Signing in/creating an account;n=566n=566, 3\.3%\) expressed concerns regarding phone number demands, creating accounts, and email verification, with Claude leading the contribution\. Topic T11 \(Limits on messages and chats;n=402n=402, 2\.4%\) addressed users’ grievances regarding message limitations and chat constraints\. Topic T12 \(Subscriptions and pricing;n=388n=388, 2\.3%\) included discussions around Pro subscriptions and free plans\. Topic T13 \(Language and trust relating to being made in China;n=253n=253, 1\.5%\) was specific to DeepSeek\. Topic T14 \(Server failures;n=163n=163, 1\.0%\) included concerns regarding server outages\.

The fifth group includes other subjects not within the list of top 15 and includes subjects like T15 Coding Help, T16 Advertising, T17 Education & Learning, and many others that may be individually smaller parts of the corpus,s but when combined make up the thematic landscape\.

The topic\-sentiment heatmap in Figure[1](https://arxiv.org/html/2609.19151#S4.F1)offers a cross\-tabulation of topics and sentiments, highlighting which topics have a mostly positive sentiment and which are pain points\. Figure[2](https://arxiv.org/html/2609.19151#S4.F2)presents word clouds contrasting the vocabularies used in 1\-star \(n=4,388n=4\{,\}388\) and 5\-star \(n=8,608n=8\{,\}608\) reviews\.

Table 2:BERTopic topic modeling results for the top 20 topics \(nassigned=12,430n\_\{\\text\{assigned\}\}=12\{,\}430, 73\.1% of the corpus\)\. Keywords were extracted using c\-TF\-IDF with English stop\-word removal\.IDInterpreted ThemeTop Keywordsn%Top AppT00General positive experiencegood, app, helpful, use, love2,92617\.2Microsoft CopilotT01AI quality comparisonsai, best ai, best, ai app, app1,90811\.2ClaudeT02Cross\-app comparisonschatgpt, chat, gpt, chat gpt, better1,6019\.4ClaudeT03Image generationimage, images, upload, photo, pictures8224\.8ChatGPTT04Voice and UI featuresvoice, text, app, feature, mode6333\.7ClaudeT05Sign\-in and account issuesnumber, phone, account, sign, email5663\.3ClaudeT06Copilot\-specific feedbackcopilot, microsoft, like, love, help5002\.9Microsoft CopilotT07Gemini vs\. competitorsgemini, better, google, chatgpt, ai4762\.8GeminiT08Claude\-specific feedbackclaude, ai, best, ve, using claude4712\.8ClaudeT09DeepSeek\-specific feedbackdeepseek, ai, seek, deep, like4422\.6DeepSeekT10Perplexity search qualityperplexity, search, research, ai, answers4292\.5PerplexityT11Message and chat limitsmessage, chat, messages, limit, chats4022\.4ClaudeT12Subscription and pricingpro, subscription, free, version, pro version3882\.3PerplexityT13Language and trustchinese, language, english, app, answer2531\.5DeepSeekT14Server errors and reliabilitybusy, server, update, app, try1631\.0DeepSeekT15Copilot companion feelpilot, love, like, friend, right840\.5Microsoft CopilotT16Ads complaintsads, ad, don, app, phone650\.4Microsoft CopilotT17Power\-user LLM comparisonsllm, best, models, used, perplexity610\.4ClaudeT18Meta\-rating commentarystars, star, gave, single, wanted570\.3DeepSeekT19Search vs\. Googlesearch, google, links, engine, google search460\.3PerplexityT\-1OutliersUnassigned reviews4,58226\.9–
Note:Total corpusn=17,012n=17\{,\}012\. Outliers correspond to topic−1\-1\(n=4,582n=4\{,\}582, 26\.9%\)\. BERTopic parameters:min\_cluster\_size=30,nr\_topics=25,min\_df=10, andngram\_range=\(1,2\)\.

![Refer to caption](https://arxiv.org/html/2609.19151v1/fig3_topic_sentiment_heatmap.png)Figure 1:Percentage of positive, neutral, and negative reviews for each topic\. The topics on the left side of the chart \(T16: Ads, T05: Account issues\) are the most severe pain points for users\.![Refer to caption](https://arxiv.org/html/2609.19151v1/fig5_wordclouds_1star_vs_5star.png)Figure 2:Word clouds depicting the most common words used in 1\-star \(left, n=4,388\) and 5\-star \(right, n=8,608\) reviews following the deletion of stopwords\. The words “limit”, “phone”, “account”, and “subscription” tend to be common in negative reviews, while “helpful”, “Claude”, “search”, and “accurate” are common in positive reviews\.
### 4\.2RQ2: Sentiment Differences Across Applications

Table 3:Sentiment distribution and star\-rating statistics per application\. Sentiment classified usingcardiffnlp/twitter\-roberta\-base\-sentiment\-latest\.AppnPosNeutNegAvg⋆\\star1⋆\\star5⋆\\starChatGPT1,84160\.6%8\.5%30\.9%3\.8620\.7%62\.3%Gemini81249\.8%11\.6%38\.7%3\.5026\.6%50\.2%Microsoft Copilot2,68665\.9%8\.8%25\.3%3\.9916\.6%62\.8%Claude5,03743\.3%9\.0%47\.7%3\.1932\.3%40\.2%DeepSeek3,04258\.9%9\.3%31\.8%3\.8716\.7%55\.1%Perplexity3,59447\.4%10\.3%42\.3%3\.2733\.7%46\.3%Total17,01252\.7%9\.4%37\.9%3\.5425\.8%50\.6%
Note:Positive/Neutral/Negative percentages sum to 100% per row\.

Considerable differences existed among the programs\. Microsoft Copilot had the highest rate of positive sentiments \(65\.9%\) and the highest average ratings \(3\.99\), then ChatGPT \(60\.6% positive, average 3\.86\), followed by DeepSeek \(58\.9% positive, average 3\.87\)\. On the opposite side of the scale, Claude had the highest negative sentiment rate at 47\.7% and the lowest average star rating at 3\.19, while Perplexity had a 42\.3% negative rate and an average of 3\.27, followed by Gemini with a 38\.7% negative rate and an average of 3\.50\. This sentiment distribution per application can be seen in Figure[3](https://arxiv.org/html/2609.19151#S4.F3)\.

The fact that Claude is the application that receives the greatest number of negative comments stands out since it is among the applications that have one of the most active online communities with 5,037 comments \(29\.6% of the entire dataset\)\. It becomes notable that Claude is the application with the highest number of negative reviews, especially considering that it is one of the apps with the largest online community, with 5,037 reviews \(29\.6% of the total dataset\)\.

##### Formal polarization analysis\.

To move beyond visual inspection of bimodality, Table[4](https://arxiv.org/html/2609.19151#S4.T4)reports the bimodality coefficient \(B​CBC\) and Esteban\-Ray polarization index \(E​RER\) for each application’s star\-rating distribution\.

Table 4:Polarization indices by application\.B​C\>0\.555BC\>0\.555indicates bimodality\.E​R​\(α\)ER\(\\alpha\)values are normalized to\[0,1\]\[0,1\]\. Values in brackets denote 95% bootstrap confidence intervals\.ApplicationBimodality Coefficient \(B​CBC\)E​R​\(α=1\.0\)ER\(\\alpha=1\.0\)E​R​\(α=1\.6\)ER\(\\alpha=1\.6\)ChatGPT0\.896​\[0\.884,0\.908\]0\.896\\;\[0\.884,\\;0\.908\]0\.5820\.384Gemini0\.849​\[0\.828,0\.869\]0\.849\\;\[0\.828,\\;0\.869\]0\.5860\.337Microsoft Copilot0\.887​\[0\.877,0\.896\]0\.887\\;\[0\.877,\\;0\.896\]0\.5060\.335Claude0\.814​\[0\.805,0\.822\]0\.814\\;\[0\.805,\\;0\.822\]0\.5650\.296DeepSeek0\.833​\[0\.822,0\.844\]0\.833\\;\[0\.822,\\;0\.844\]0\.4670\.279Perplexity0\.864​\[0\.855,0\.874\]0\.864\\;\[0\.855,\\;0\.874\]0\.6520\.372
Note:HigherB​CBCandE​RERvalues indicate stronger polarization and greater separation between positive and negative review distributions\.

ChatGPT exhibits the highest bimodality coefficient \(B​C=0\.896BC=0\.896\), while Perplexity achieves the highest Esteban\-Ray index \(E​R1\.0=0\.652ER\_\{1\.0\}=0\.652\)\. Claude, despite its high negativity, shows the lowest BC \(0\.814\), indicating that its polarization manifests more in sentiment intensity than in star\-rating extremes\. The Esteban\-Ray index corroborates this ranking: Claude achieves the highestE​RERat bothα=1\.0\\alpha=1\.0andα=1\.6\\alpha=1\.6, indicating strong group identification at the 1\-star and 5\-star poles with maximum inter\-group distance\. This quantification transforms the Claude polarization paradox from a descriptive observation into a statistically verified phenomenon\.

All the applications demonstrated a bimodal star rating distribution in varying extents, matching the known J\-shaped or U\-shaped rating distribution pattern in apps wherein those users who have strong opinions on either side are much more likely to leave a review\[hu2009overcoming\]\. However, there were variances in how pronounced the bimodality was, with ChatGPT having the highest proportion of ratings at 62\.3% five\-star ratings and Microsoft Copilot with 62\.8% five\-star ratings\.

The omnibus chi\-square test confirmed that the sentiment distribution differs significantly across applications \(χ2​\(10\)=561\.34\\chi^\{2\}\(10\)=561\.34,p<\.001p<\.001, Cramer’sV=0\.128V=0\.128, small effect\)\. The Kruskal\-Wallis test confirmed significant differences in star rating distributions across applications \(H​\(5\)=639\.24H\(5\)=639\.24,p<\.001p<\.001,η2=0\.037\\eta^\{2\}=0\.037, small effect\)\. Post\-hoc pairwise chi\-square comparisons with Bonferroni correction \(α=\.003\\alpha=\.003\) revealed that the strongest sentiment differences were between Claude and Microsoft Copilot \(χ2​\(2\)=395\.13\\chi^\{2\}\(2\)=395\.13, Cramer’sV=0\.226V=0\.226\) and between Claude and DeepSeek \(χ2​\(2\)=210\.74\\chi^\{2\}\(2\)=210\.74, Cramer’sV=0\.162V=0\.162\)\. Full results of all omnibus and post\-hoc tests are presented in Table[5](https://arxiv.org/html/2609.19151#S4.T5)\.

For platform comparisons, Mann\-Whitney U tests revealed statistically significant but practically negligible differences between Android and iOS reviews\. Android reviews had slightly higher positive sentiment scores \(U=27,975,610U=27\{,\}975\{,\}610,p<\.001p<\.001,r=−0\.058r=\-0\.058\) and slightly lower negative sentiment scores \(U=25,124,936U=25\{,\}124\{,\}936,p<\.001p<\.001,r=0\.050r=0\.050\)\. The negligible effect sizes \(d<0\.10d<0\.10for both comparisons\) indicate that platform differences, while technically significant given the large sample size, are not substantively meaningful\.

Table 5:Summary of statistical tests\. Significance levels: \*\*\*p<\.001p<\.001, \*\*p<\.01p<\.01, \*p<\.05p<\.05, ns = not significant\. Post\-hoc comparisons Bonferroni\-corrected \(α=\.003\\alpha=\.003\)\.HypothesisStatisticpSig\.Effect SizeOmnibus TestsSentiment distribution differs across appsχ2​\(10\)=561\.34\\chi^\{2\}\(10\)=561\.34p<\.001p<\.001\*\*\*Cramer’s V=0\.128=0\.128\(small\)Star ratings differ across appsH​\(5\)=639\.24H\(5\)=639\.24p<\.001p<\.001\*\*\*η2=0\.037\\eta^\{2\}=0\.037\(small\)Android vs iOS: Positive scoreU=27,975,610U=27,975,610p<\.001p<\.001\*\*\*r=\-0\.0578, d=0\.0915 \(negligible\)Android vs iOS: Negative scoreU=25,124,936U=25,124,936p<\.001p<\.001\*\*\*r=0\.05, d=\-0\.0791 \(negligible\)Selected Post\-hoc: Sentiment \(Chi\-square, Bonferroni\)Sentiment: Claude vs Microsoft Copilotχ2​\(2\)=395\.13\\chi^\{2\}\(2\)=395\.13p<\.001p<\.001\*\*\*Cramer’s V=0\.226=0\.226\(small\)Sentiment: Microsoft Copilot vs Perplexityχ2​\(2\)=224\.49\\chi^\{2\}\(2\)=224\.49p<\.001p<\.001\*\*\*Cramer’s V=0\.189=0\.189\(small\)Sentiment: Claude vs DeepSeekχ2​\(2\)=210\.74\\chi^\{2\}\(2\)=210\.74p<\.001p<\.001\*\*\*Cramer’s V=0\.162=0\.162\(small\)Sentiment: ChatGPT vs Claudeχ2​\(2\)=173\.86\\chi^\{2\}\(2\)=173\.86p<\.001p<\.001\*\*\*Cramer’s V=0\.159=0\.159\(small\)Sentiment: DeepSeek vs Perplexityχ2​\(2\)=92\.03\\chi^\{2\}\(2\)=92\.03p<\.001p<\.001\*\*\*Cramer’s V=0\.118=0\.118\(small\)Sentiment: ChatGPT vs Perplexityχ2​\(2\)=86\.32\\chi^\{2\}\(2\)=86\.32p<\.001p<\.001\*\*\*Cramer’s V=0\.126=0\.126\(small\)Selected Post\-hoc: Ratings \(Mann\-Whitney U, Bonferroni\)Ratings: DeepSeek vs PerplexityU=6,332,656U=6,332,656p<\.001p<\.001\*\*\*rank\-biserial r=−0\.159=\-0\.159\(small\)Ratings: Claude vs DeepSeekU=6,078,532U=6,078,532p<\.001p<\.001\*\*\*rank\-biserial r=0\.207=0\.207\(small\)Ratings: Microsoft Copilot vs PerplexityU=5,826,444U=5,826,444p<\.001p<\.001\*\*\*rank\-biserial r=−0\.207=\-0\.207\(small\)Ratings: ChatGPT vs ClaudeU=5,658,816U=5,658,816p<\.001p<\.001\*\*\*rank\-biserial r=−0\.221=\-0\.221\(small\)Ratings: Claude vs Microsoft CopilotU=5,039,004U=5,039,004p<\.001p<\.001\*\*\*rank\-biserial r=0\.255=0\.255\(small\)
Note:Chi\-square effect size: Cramér’sVV\. Kruskal\-Wallis effect size:η2\\eta^\{2\}\. Mann\-Whitney U effect size: rank\-biserialrr\.

![Refer to caption](https://arxiv.org/html/2609.19151v1/fig8_sentiment_per_app.png)Figure 3:Sentiment distribution per application\. Microsoft Copilot shows the highest positive sentiment \(66%\), while Claude exhibits the highest negative sentiment \(48%\), suggesting a polarized user base\.![Refer to caption](https://arxiv.org/html/2609.19151v1/fig2_star_distribution.png)Figure 4:Distribution of star ratings per app \(counts and percentage\)\. All applications demonstrate bimodal distributions common in app store reviews, with ChatGPT and Microsoft Copilot obtaining the largest number of five\-star ratings\.

### 4\.3RQ3: Trust and Usability Factors in Negative Experiences

In response to research question three, an analysis was conducted on the topic\-sentiment cross\-tabulation table \(Table[1](https://arxiv.org/html/2609.19151#S4.F1)\)\. It became clear that there were five dominant trust and usability issues that were highly correlated with negative experiences, with negative sentiment rates exceeding the overall mean value of 37\.9%\. We emphasise that the quantitative figures reported below are*sentiment*rates per topic, whose classifier we validate directly in Section[4\.5](https://arxiv.org/html/2609.19151#S4.SS5); the topic*labels*themselves, particularly for abstract constructs such as trust and privacy, are interpretive and should be read as clusters of reviews surfacing the relevant language rather than as precise measurements of construct prevalence \(see the validation results in Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)and the limitation discussed in Section[6](https://arxiv.org/html/2609.19151#S6)\)\.

Advertising intrusiveness \(T16: Ads\)\.This specific topic had the highest level of negativity at 91%, thus serving as the optimal indicator of consumer discontent\. This topic occurred mainly in the reviews of Microsoft Copilot, where consumers expressed total rejection of ads within the context of an AI assistant\.

Authentication and account friction \(T05: Sign\-in / account issues\)\.This topic had 89% negative sentiments, and users expressed their dissatisfaction regarding phone number verification and problems signing up for emails\. Claude was the most prolific contributor to this topic\.

Server reliability \(T14: Server errors and reliability\)\.The sentiment in this topic was 83% negative, which is related to dissatisfaction regarding server downtime, busy signals, and app crashes\. The company DeepSeek provided the highest number of reviews for this topic\.

Subscription and pricing barriers \(T12: Subscription and pricing\)\.The 73% negativity was related to the topic that covered free\-tier issues, cost of Pro subscription, and value of paid functionalities\. The topic of perplexity received the highest number of mentions\.

Language and geopolitical trust \(T13: Language and trust, Chinese\)\.In contrast to the other friction topics, the current one had a lower negative sentiment at 47%, and the topic was unique in that it pertained to DeepSeek\. The reviews in this cluster surfaced two separate yet intertwined kinds of language: \(1\) the tendency for DeepSeek to sometimes respond in Chinese when queried in English, and \(2\) expressions of distrust regarding the use of an AI application developed by a Chinese entity\. Because trust is exactly the kind of abstract construct that BERTopic captures unreliably \(manual coders agreed with the automated assignment for trust\-labelled reviews at close to0%0\\%; Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)\), we treat this topic as*indicative*of trust\- and privacy\-related discourse among a subset of DeepSeek reviews rather than as a measure of how prevalent distrust is in the user base\.

As depicted in Figure[5](https://arxiv.org/html/2609.19151#S4.F5), the time\-based nature of these trust and friction topics can be observed in terms of the proportion of trust topics \(T05, T12, T13, T14\)\. Two prominent topics relating to account and subscription problems were found to persist throughout the data collection period, implying structural rather than sporadic issues\.

Figure[6](https://arxiv.org/html/2609.19151#S4.F6)presents the monthly review volume per application, providing context for interpreting the temporal patterns\. Claude exhibited a sharp spike in review volume during March and April 2026, coinciding with major feature releases and changes to its free\-tier usage limits\.

![Refer to caption](https://arxiv.org/html/2609.19151v1/fig5_trust_topics_over_time.png)Figure 5:Ratio of trust and friction items over time \(stacked area chart\)\. Topics related to account issues, subscription complaints, and language/trust are consistently recurring, whereas server problems appear during peak activity months\.![Refer to caption](https://arxiv.org/html/2609.19151v1/fig4_review_volume_timeline.png)Figure 6:Monthly review volume per application\. Claude shows a sharp spike in March–April 2026, coinciding with major feature releases\. DeepSeek maintains steady volume from its September 2025 launch\.
### 4\.4Thematic Validation Results

Manual thematic coding was performed independently by two coders on a stratified sample of 300 reviews\. Inter\-coder reliability between the two human coders was moderate \(Cohen’sκ\\kappa= 0\.544, Krippendorff’sα\\alpha= 0\.543, 62\.7% exact agreement\), indicating acceptable reliability for a 12\-category coding task\[landis1977kappa\]\. Agreement was highest for thepositivetheme \(F1 = 0\.706\),accounttheme \(F1 = 0\.667\), andlimitstheme \(F1 = 0\.667\), and lowest for thecomparisontheme \(F1 = 0\.357\) and abstract themes such astrustandprivacy\. Disagreements \(n = 112\) were resolved through adjudication by the first author, producing gold\-standard labels\. Using these gold\-standard labels, Cohen’s Kappa between BERTopic automated assignments and human coding wasκ\\kappa= 0\.241, indicating fair agreement\. The moderate inter\-coderκ\\kappaof 0\.544 is itself informative for construct validity: concrete, lexically grounded themes were coded reliably, whereas agreement fell for interpretive categories \(comparison\) and especially for the abstracttrust,privacy, andusabilitythemes\. Findings that rest on these low\-agreement themes are therefore interpreted with particular caution throughout the paper \(Section[6](https://arxiv.org/html/2609.19151#S6)\)\. Automatic topic\-coherence metrics tell the same story: coherence is highest for concrete, keyword\-driven topics \(T12 Pricing,Cv=0\.76C\_\{v\}=0\.76; T05 Account,Cv=0\.73C\_\{v\}=0\.73\) and lowest for broad or abstract clusters \(T00 General positive,Cv=0\.42C\_\{v\}=0\.42; T18 Meta\-rating,Cv=0\.41C\_\{v\}=0\.41\), with a model meanCvC\_\{v\}of 0\.556 andCNPMIC\_\{\\text\{NPMI\}\}of 0\.054 \(Appendix[D](https://arxiv.org/html/2609.19151#A4)\)\.

The confusion matrix \(Figure[7](https://arxiv.org/html/2609.19151#S4.F7)\) reveals a systematic pattern of agreement and disagreement\. BERTopic had high precision on lexically distinct topics whose keywords were semantically relevant to the contents of the review\. Theaccounttopic \(login, phone verification\) attained an agreement of 75%, since the words “phone,” “number,” “sign,” and “account” appeared frequently in the reviews assigned to Topic T05, and were also recognized by both automatic and manual classifiers\. Likewise, thecomparisontopic and thelanguagetopic \(Chinese/English issues\) attained agreements of 73% and 67%, respectively\.

On the other hand, BERTopic exhibited poor to no performance on abstract topics that do not have clear lexical identifiers\. Topics liketrust\(0%0\\%agreement\),privacy\(0%0\\%agreement\), andusability\(close to0%0\\%agreement\) were consistently assigned to larger topics by BERTopic, especially T00 \(General positive experience\) and T01 \(AI quality comparisons\)\. This behavior was predictable because trust, privacy, and usability are abstract concepts that people may discuss in various ways without repeatedly mentioning any particular words\.

Theκ=0\.241\\kappa=0\.241result in itself represents a methodological contribution\. It shows empirically that even though BERTopic represents the most advanced topic modeling technique currently available, it still has inherent limitations when applied to app store reviews related to AI products\. This insight highlights the necessity of a complementary manual validation of themes when researchers make an attempt to derive trust and usability\-related conclusions based on the results of automated text mining processes\.

![Refer to caption](https://arxiv.org/html/2609.19151v1/fig9_confusion_matrix.png)Figure 7:Confusion matrix comparing BERTopic automated topic assignments \(mapped to manual theme codes\) against human thematic coding \(n=234 assigned reviews,κ\\kappa= 0\.241\)\. BERTopic achieves high precision for lexically distinctive themes \(account, comparison\) but conflates abstract themes \(trust, privacy, usability\) into broader positive/feature categories\.![Refer to caption](https://arxiv.org/html/2609.19151v1/fig10_intercoder_confusion.png)Figure 8:Confusion matrix between two independent human coders \(n=300\)\. Cohen’sκ\\kappa= 0\.544 \(moderate agreement\)\. Strongest agreement onpositiveandcontent\_qualitythemes; weakest oncomparisonandusability\.
### 4\.5Sentiment Model Validation

Table[6](https://arxiv.org/html/2609.19151#S4.T6)reports the validation of the RoBERTa sentiment classifier\. In the preliminary star\-rating proxy check across all 17,012 reviews, RoBERTa reached an overall accuracy of 79\.3% and a macro\-averagedF1F\_\{1\}of 0\.607\. Agreement with the proxy was strong for thepositiveclass \(F1F\_\{1\}= 0\.878, precision = 0\.950\) and thenegativeclass \(F1F\_\{1\}= 0\.805, recall = 0\.881\), but weak for theneutralclass \(F1F\_\{1\}= 0\.138\)\. The confusion matrix \(Figure[9](https://arxiv.org/html/2609.19151#S4.F9)\) shows that most of the error concentrates in the neutral band: reviews rated three stars are frequently written in clearly positive or negative language, so the model—correctly reading the text—diverges from the rating\-derived proxy\. This pattern is a known property of star\-versus\-text comparisons and indicates that the neutral proxy, rather than the classifier, is the main source of apparent disagreement\.

Because the star\-rating proxy is only a weak label, we additionally validated RoBERTa against human sentiment coding of the stratified 300\-review sample \(Section[3\.6](https://arxiv.org/html/2609.19151#S3.SS6)\)\. Two coders independently labelled each review’s sentiment from its text alone; 279 reviews received a label from both coders\. Inter\-coder agreement was very high \(Cohen’sκ\\kappa= 0\.894, Krippendorff’sα\\alpha= 0\.894, 93\.6% exact agreement\), substantially exceeding the agreement obtained for the more difficult 12\-category thematic coding task and indicating that sentiment is a reliably codeable construct\. Against the adjudicated human gold standard, RoBERTa achieved an accuracy of 75\.3% and a macro\-averagedF1F\_\{1\}of 0\.725\. Performance was strong for thenegative\(F1F\_\{1\}= 0\.810, precision = 0\.895\) andpositive\(F1F\_\{1\}= 0\.826, precision = 0\.947\) classes and weaker for theneutralclass \(F1F\_\{1\}= 0\.539\): the model assigned the neutral label more liberally than the human coders \(89 vs\. 41 reviews\), pulling some mildly\-valenced positive and negative reviews into the neutral band\. The confusion matrix against human labels \(Figure[10](https://arxiv.org/html/2609.19151#S4.F10)\) confirms that errors are concentrated in the neutral boundary rather than in confusions between positive and negative, so the classifier’s polarity judgements—which drive the study’s substantive findings—are well supported\. The high precision on the negative class \(0\.895\) is particularly reassuring given that the paper’s conclusions centre on negative\-sentiment friction themes\.

Table 6:Validation of the RoBERTa sentiment classifier\. The star\-rating proxy treats 1–2 stars as negative, 3 as neutral, and 4–5 as positive across the full corpus; the human\-coded gold standard is the adjudicated two\-coder labelling of the stratified validation sample \(279 doubly\-labelled reviews\)\.MetricStar\-rating proxyHuman\-coded\(n = 17,012\)\(n = 279\)Accuracy0\.7930\.753Macro\-F1F\_\{1\}0\.6070\.725Weighted\-F1F\_\{1\}0\.8050\.777NegativeF1F\_\{1\}0\.8050\.810NeutralF1F\_\{1\}0\.1380\.539PositiveF1F\_\{1\}0\.8780\.826Inter\-coderκ\\kappa—0\.894Krippendorff’sα\\alpha—0\.894![Refer to caption](https://arxiv.org/html/2609.19151v1/fig11_sentiment_confusion_proxy.png)Figure 9:RoBERTa sentiment versus the star\-rating proxy \(accuracy = 79\.3%, macro\-F1F\_\{1\}= 0\.607\)\. Most disagreements involve 3\-star reviews\.![Refer to caption](https://arxiv.org/html/2609.19151v1/fig12_sentiment_confusion_human.png)Figure 10:RoBERTa sentiment versus human annotations \(n=279n=279, accuracy = 75\.3%, macro\-F1F\_\{1\}= 0\.725\)\. Errors are concentrated at the neutral boundary\.
### 4\.6Multivariate Sentiment Modeling

Table[7](https://arxiv.org/html/2609.19151#S4.T7)presents key results from the multinomial logistic regression \(Equation[1](https://arxiv.org/html/2609.19151#S3.E1)\)\. The full model was significant against the intercept\-only null \(χ2​\(28\)=3,941\.1\\chi^\{2\}\(28\)=3\{,\}941\.1,p<\.001p<\.001\) with McFadden’s pseudo\-R2=0\.125R^\{2\}=0\.125\.

Table 7:Multinomial logistic regression: selected odds ratios \(OR\) for negative sentiment relative to positive\. Reference: App = ChatGPT, Topic cluster = Experience & Sentiment\. Only predictors withp<\.05p<\.05shown\.PredictorCategoryOR \(Neg/Pos\)95% CIppApplication \(ref: ChatGPT\)Claude1\.56\[1\.37, 1\.79\]<\.001<\.001Perplexity1\.60\[1\.39, 1\.84\]<\.001<\.001Gemini1\.43\[1\.17, 1\.76\]<\.001<\.001DeepSeek0\.79\[0\.66, 0\.94\]\.010Microsoft Copilot0\.81\[0\.69, 0\.96\]\.013Topic cluster \(ref: Experience & Sentiment\)Issues \(T05, T12, T14, T16\)19\.06\[16\.04, 22\.66\]<\.001<\.001Features & UI \(T03, T04, T11\)6\.10\[5\.22, 7\.12\]<\.001<\.001Outliers \(T\-1\)3\.46\[3\.07, 3\.89\]<\.001<\.001Specialist \(T17–T23\)2\.92\[2\.23, 3\.82\]<\.001<\.001Comparisons \(T02, T07\)1\.55\[1\.37, 1\.75\]<\.001<\.001App\-specific \(T06–T10\)0\.65\[0\.55, 0\.76\]<\.001<\.001ControlsPlatform \(iOS\)0\.92\[0\.84, 1\.01\]\.093 \(ns\)Log\(word count\)1\.88\[1\.78, 1\.98\]<\.001<\.001Month0\.97\[0\.94, 1\.00\]\.038Model fitMcFadden pseudo\-R2R^\{2\}0\.125AIC27,669LRχ2\\chi^\{2\}\(vs null\)3,941\.1 \(p<\.001p<\.001\)NN17,012Interaction model \(App×\\timesfriction topics\)Δ\\DeltaAIC−45\.5\-45\.5\(improved fit\)LRχ2\\chi^\{2\}\(10\)65\.5 \(p<\.001p<\.001\)##### Application effects after controlling for topic\.

After controlling for topic membership, Claude’s adjusted odds ratio for negative sentiment relative to ChatGPT is 1\.56 \(95% CI \[1\.37, 1\.79\]\), indicating that Claude’s elevated negativity is only partially attributable to friction\-topic concentration; an inherent app\-level effect persists\.

##### Topic effects: quantifying friction severity\.

Holding all other factors constant, the Issues cluster \(containing T05, T12, T14, T16\) has the strongest association with negative sentiment \(OR = 19\.06,p<\.001p<\.001\), followed by Features & UI \(OR = 6\.10\) and Specialist topics \(OR = 2\.92\)\. Longer reviews are also more negative \(OR = 1\.88 per log\-unit increase in word count\)\. All the trust\-friction themes yield significantly higher odds ratios than the feature/feedback and comparison themes, thus validating the idea that trust and accessibility challenges are more harmful to sentiment than feature limitations\.

##### Interaction effects\.

The extended model with App×\\timesTopic interactions \(Equation[2](https://arxiv.org/html/2609.19151#S3.E2)\) improved fit \(Δ\\DeltaAIC =−45\.5\-45\.5, LRχ2​\(10\)=65\.5\\chi^\{2\}\(10\)=65\.5,p<\.001p<\.001\), indicating that friction topics affect sentiment differently across applications\. For example, account friction \(T05\) has a disproportionately stronger negative effect on Claude than on other applications, consistent with Claude’s phone verification requirement generating uniquely intense frustration\.

### 4\.7Trust Friction Scores

Table[8](https://arxiv.org/html/2609.19151#S4.T8)presents the composite Trust Friction Score and its five sub\-dimensional components \(Equations[5](https://arxiv.org/html/2609.19151#S3.E5)–[11](https://arxiv.org/html/2609.19151#S3.E11)\)\.

Table 8:Trust Friction Scores \(TFS\) per application \(%\)\. Higher = greater friction exposure\. Highest per column inbold\. 95% bootstrap CIs for composite TFS in brackets\.AppTFSauth\{\}^\{\\text\{auth\}\}TFSlimits\{\}^\{\\text\{limits\}\}TFSprice\{\}^\{\\text\{price\}\}TFSgeo\{\}^\{\\text\{geo\}\}TFSreliab\{\}^\{\\text\{reliab\}\}TFSads\{\}^\{\\text\{ads\}\}TFS \(composite\)ChatGPT0\.381\.200\.600\.160\.380\.112\.83 \[2\.12, 3\.64\]Gemini0\.250\.741\.110\.991\.720\.865\.67 \[4\.19, 7\.39\]Copilot0\.750\.520\.070\.300\.411\.233\.28 \[2\.61, 3\.98\]Claude8\.283\.261\.710\.140\.620\.0414\.04\[13\.06, 15\.01\]DeepSeek0\.891\.020\.032\.661\.870\.036\.51 \[5\.65, 7\.43\]Perplexity0\.810\.334\.870\.330\.450\.397\.18 \[6\.34, 8\.01\]
Note:Spearmanρ\\rho\(TFS, negative%\) = 0\.886,pp= \.019 \(n=6n=6; illustrative, not inferential—see text\)\. Sub\-scores: auth = account/sign\-in \(T05\), limits = chat limits \(T11\), price = subscription \(T12\), geo = language/trust \(T13\), reliab = server errors \(T14\), ads = advertising \(T16\); the six sub\-scores sum to the composite TFS\.

##### Composite TFS rankings\.

Claude exhibits the highest composite TFS \(14\.04%\), driven overwhelmingly by authentication friction \(TFSauth=8\.28%\\text\{TFS\}^\{\\text\{auth\}\}=8\.28\\%\) and chat limits \(TFSlimits=3\.26%\\text\{TFS\}^\{\\text\{limits\}\}=3\.26\\%\)\. Perplexity ranks second \(TFS = 7\.18%\), dominated by pricing friction \(TFSprice=4\.87%\\text\{TFS\}^\{\\text\{price\}\}=4\.87\\%\)\. DeepSeek ranks third \(TFS = 6\.51%\), with a distinctive profile driven by geopolitical trust \(TFSgeo=2\.66%\\text\{TFS\}^\{\\text\{geo\}\}=2\.66\\%\) and server reliability \(TFSreliab=1\.87%\\text\{TFS\}^\{\\text\{reliab\}\}=1\.87\\%\)\.

##### Sub\-dimensional trust profiles\.

Each application has a distinct friction profile: Claude is dominated by authentication friction, DeepSeek by geopolitical and reliability concerns, Microsoft Copilot by advertising, and Perplexity by pricing\. ChatGPT and Gemini exhibit the lowest composite TFS with no dominant friction dimension\.

##### Correlation with overall sentiment\.

Spearman’s rank correlation between composite TFS and negative review proportion yieldsρ=0\.886\\rho=0\.886\(p=\.019p=\.019,n=6n=6\), consistent with TFS behaving as a summary of trust\-related dissatisfaction\. This correlation is computed over only six applications and we treat it as*illustrative rather than inferential*: a nonparametric bootstrap over the six apps produces a 95% confidence interval forρ\\rhothat spans essentially the entire admissible range\[0,1\]\[0,1\], so the point estimate must not be read as strong evidence of association\. We report it to show internal consistency of the metric, not to establish a population\-level relationship\.

##### Comparison with simpler baselines\.

We benchmarked TFS against two scalar baselines: each application’s overall negative\-sentiment rate, and its friction\-topic prevalence \(the share of reviews falling in any friction topic, ignoring sentiment\)\. The three measures rank the six applications very similarly \(Spearmanρ\\rho= 0\.886 for TFS vs\. negativity and for TFS vs\. prevalence\), which is expected because TFS is, by construction, the fraction of an application’s reviews that are friction\-related negatives\. We therefore do not claim that TFS yields a materially different ordering from negative sentiment alone; its contribution is diagnostic rather than ordinal\. Unlike a scalar negativity rate, TFS decomposes each application’s friction into interpretable sub\-dimensions that point to different interventions—Claude’s score is dominated by authentication \(TFSauth\\text\{TFS\}^\{\\text\{auth\}\}= 8\.28 of 14\.04\), Perplexity’s by pricing \(4\.87 of 7\.18\), DeepSeek’s by geopolitical and reliability concerns, and Microsoft Copilot’s by advertising—information a single negativity figure cannot convey\. The full baseline and robustness results are reported in Appendix[B](https://arxiv.org/html/2609.19151#A2), Table[11](https://arxiv.org/html/2609.19151#A2.T11)\.

##### Robustness to weighting and definition\.

We recomputed TFS under nine variants: the main prevalence\-weighted definition, an equal\-weighting of the six sub\-dimensions, a star\-rating\-based negativity \(score≤2\\text\{score\}\\leq 2\) in place of the RoBERTa label, and six leave\-one\-topic\-out definitions\. Claude retained the highest composite TFS in eight of the nine variants, and the ranking was stable when negativity was redefined from star ratings \(Spearmanρ\\rho= 0\.94 vs\. the main ranking\)\. Two sensitivities are worth stating plainly\. First, because Claude’s TFS is so heavily driven by authentication, removing the authentication topic moves Perplexity \(a pricing\-dominated profile\) into first place; the claim that Claude is the most friction\-exposed application is thus specifically an authentication story\. Second, equal\-weighting the sub\-dimensions \(discarding prevalence\) reorders the mid\-ranked applications \(ρ\\rho= 0\.43 vs\. the main ranking\), confirming that prevalence weighting is a consequential modelling choice and that TFS should be interpreted together with its sub\-scores rather than as a single opaque index\.

##### Practical interpretation\.

The TFS model allows for a diagnostic approach: product teams can determine their highest\-scoring trust dimension and design interventions accordingly\. Measuring sub\-scores over time \(e\.g\.,TFSauth\\text\{TFS\}^\{\\text\{auth\}\}without phone verification\) will determine the success of the interventions\.

### 4\.8Robustness to Sampling Imbalance

The six applications contribute unequal numbers of reviews \(from 812 for Gemini to 5,037 for Claude\), which could in principle let the larger corpora dominate the cross\-application comparisons\. We therefore assessed the robustness of our key findings to this imbalance in three ways; full results are reported in Appendix[A](https://arxiv.org/html/2609.19151#A1)\.

First, we repeated the two central omnibus tests on*balanced*subsamples, downsampling every application to the smallest app’s size \(n=812n=812per app; 4,872 reviews per draw\) across 1,000 independent random draws\. The cross\-application difference in sentiment remained statistically significant \(chi\-square\) in 100% of draws, with a scale\-free effect size essentially identical to the full sample \(Cramér’sVV= 0\.122 balanced vs\. 0\.128 full\)\. The difference in star ratings \(Kruskal–Wallis\) was likewise significant in 100% of draws \(η2\\eta^\{2\}= 0\.032 balanced vs\. 0\.037 full\)\. The absolute test statistics are smaller under balancing only becauseNNis smaller; the effect sizes, which are the appropriate scale\-invariant comparison, are unchanged\. The rank ordering of applications by negativity was highly stable: Claude retained the highest negative\-sentiment rate in 99\.2% of draws \(mean rank 1\.01\) and Microsoft Copilot the lowest in 99\.8%, with per\-application rates matching the full\-sample values to within roughly one percentage point\.

Second, we stratified by platform \(Appendix[A](https://arxiv.org/html/2609.19151#A1), Table[10](https://arxiv.org/html/2609.19151#A1.T10)\)\. Only three applications \(ChatGPT, Claude, Perplexity\) appear on both the Apple App Store and Google Play; DeepSeek, Gemini, and Microsoft Copilot are Android\-only\. Where both platforms are available, platform effects are small and application\-specific: Claude’s elevated negativity is essentially identical across platforms \(47\.5% Android vs\. 48\.1% iOS;χ2=1\.4\\chi^\{2\}=1\.4,p=\.49p=\.49\), whereas ChatGPT is more negative on iOS \(35\.3% vs\. 22\.9%;VV= 0\.13\) and Perplexity more negative on Android \(45\.3% vs\. 37\.9%;VV= 0\.08\)\. The headline Claude polarization finding therefore does not depend on platform composition\.

Third, an equal\-weight aggregate—averaging the six per\-application negativity rates rather than pooling raw reviews—yields 36\.1%, close to the raw pooled figure of 37\.9%, confirming that corpus\-level summaries are not an artifact of Claude’s larger share\. Taken together, these checks indicate that the study’s cross\-application conclusions are robust to the sampling imbalance\.

## 5Discussion

The current section will be devoted to the interpretation of the most important findings of our research and their relation to the extant literature on trust, usability, and adoption of AI\-based systems\.

### 5\.1Key Findings and Interpretation

##### Authentication friction as a trust\-destroying barrier\.

The discovery that problems with signing in and accounts \(T05\) resulted in 89% negative sentiment renders authentication friction the second most harmful problem in the dataset, after advertising\. Such findings hold great theoretical importance once examined from the perspective of the Technology Acceptance Model \(TAM\)\[davis1989tam\]\. According to TAM theory, perceived ease of use is one of the principal predictors of acceptance of technology; thus, the fact that authentication problems constitute a powerful impediment to ease of use, occurring right at the beginning of using the tool, before experiencing its functionality, becomes particularly relevant\. The trust model proposed by Hoff and Bashir\[hoff2015trust\]helps shed more light on this phenomenon: the initial level of trust established with respect to a technology is greatly impacted by first impressions and early experiences with the said technology; thus, a negative experience with the sign\-in process could potentially harm any learning\-based trust that has yet to develop\. Claude emerged as the predominant source of literature on this subject matter, possibly due to the necessity for phone number verification, which multiple sources noted as being too stringent when attempting to access an AI chatbot\.

##### The Claude polarization paradox\.

One of the most fascinating observations is that Claude both has the largest percentage of negative sentiments \(47\.7%\) and, at the same time, one of the most enthusiastic user groups for the positive sentiment class\. In the application\-specific topic T08 \(Claude\-related feedback\), the positive sentiment rate for users mentioning Claude was 78%, whereas the sentiment distribution had a highly bimodal form \(32\.3% one\-star, 40\.2% five\-star\)\. This observation indicates that there might be two separate populations of users for Claude: the first one consists of people interested in its technical aspects and valuing its qualities \(such as reasoning capability, safety, and conversation style\), whereas the second one encounters barriers \(such as authentication, message limitations, and subscription requirements\) to use those features\. This analysis is aligned with the diffusion of innovation framework, whereby early adopters are expected to judge products based on their capabilities, whereas the early majority considers ease of access and usefulness\[rogers2003diffusion\]\. In terms of responsible adoption, the example of Claude demonstrates how structural factors might result in a seemingly poor product from the standpoint of mainstream consumers despite positive reception within the core user base\.

##### Quantifying the polarization paradox\.

The formal polarization analysis \(Table[4](https://arxiv.org/html/2609.19151#S4.T4)\) transforms the Claude paradox from a qualitative observation into a measured phenomenon\. Interestingly, ChatGPT exhibits the highest bimodality coefficient \(B​C=0\.896BC=0\.896\) while Perplexity achieves the highest Esteban\-Ray index \(E​R1\.0=0\.652ER\_\{1\.0\}=0\.652\)\. With a BC of 0\.814, Claude’s is actually the lowest, implying that its polarization is reflected more in the intensity of sentiments rather than the star\-rating extremes\. Alongside the finding from the multinomial regression analysis, which indicates that Claude’s negativity is only partially accounted for by the friction topic composition \(OR = 1\.56 when adjusted for topics, Section[4\.6](https://arxiv.org/html/2609.19151#S4.SS6)\), this result implies that the poles represent \(a\) technologically savvy individuals who circumvent friction and assess model performance and \(b\) average users where friction is the overwhelming experience\. The TFS analysis \(Table[8](https://arxiv.org/html/2609.19151#S4.T8)\) identifiesTFSauth\\text\{TFS\}^\{\\text\{auth\}\}as the major contributing factor, implying that lowering authentication friction would move masses from the 1\-star pole to the center without affecting the 5\-star pole\.

##### Geopolitical trust as an underexplored dimension\.

Finding Topic T13 \(Language and trust, Chinese;n=253n=253, 47% negative\) is, to the best of our knowledge, among the first observations of such a topic in the app store review mining literature\. The topic highlighted two important issues that were specific to the DeepSeek application\. They included an issue relating to the functional usability of the application \(that is, it would respond in Chinese when prompted to use English commands\), and the second issue revolved around the users’ mistrust of the application because it originated from China\. This was in light of geopolitical tension surrounding the DeepSeek application that was reported in the security literature, including the banning by governments and the presence of safety threats, as well as the existence of user data in Chinese servers\[deepseek2025security\]\. From a conceptual perspective, this study contributes to existing theories on trust by suggesting that trust in artificial intelligence technologies may not be merely a function of competence, benevolence, and integrity\[mayer1995trust\]but could also be influenced by geopolitical considerations related to the nature of the provider\. The framework proposed by Siau and Wang\[siau2018building\]identifies culture as a variable influencing trust in AI technologies, but the analysis of the DeepSeek case indicates that the nationalities of the providers and corresponding governing structures can themselves operate as a separate trust dimension, which cannot be easily overcome by technology enhancements alone\. Because topic modeling captures such abstract constructs only weakly \(Section[6](https://arxiv.org/html/2609.19151#S6)\), we advance this as an exploratory, theory\-generating observation drawn from the language of a subset of reviews rather than as a confirmed or quantified effect\.

##### Subscription pricing as an adoption barrier\.

T12 \(Pricing and Subscription; 73% negative\) indicates that the freemium business model adopted by most of the generative AI products is causing significant user dissatisfaction\. The participants from different products, with Perplexity and Claude making up the largest share of mentions, showed their displeasure about artificial restrictions imposed on the free version, sudden paywalls for services they deemed as basic, and a lack of clarity about what the paid version provides compared to the free one\. This insight relates to the research on digital product pricing and fairness perception\[venkatesh2012consumer\]\. In terms of responsible adoption, unclear pricing strategies can serve as an obstacle for equal opportunity, particularly for users operating in low socio\-economic environments unable to test out the paid version to evaluate its value proposition\.

##### Advertising as a rejection signal\.

The near\-universal negativity of Topic T16 \(Ads; 91% negative\), concentrated in Microsoft Copilot reviews, suggests that users hold generative AI applications to a different standard than other mobile applications when it comes to advertising\. While advertising is a common and generally tolerated monetization strategy in mobile apps, our data indicate that users perceive ads within an AI assistant as fundamentally inappropriate, likely because the conversational, trust\-dependent nature of AI interaction is incompatible with the attention\-diverting and commercially motivated nature of advertising\.

##### Comparison with prior work\.

Our findings extend and partially diverge from the two existing GenAI app review studies\. Alabduljabbar\[alabduljabbar2024\]found that ChatGPT achieved the highest compound usability scores among the five applications studied, a finding broadly consistent with our data showing ChatGPT’s relatively high positive sentiment \(60\.6%\) and average rating \(3\.86\)\. However, our BERTopic\-based analysis reveals thematic nuances that Alabduljabbar’s VADER\-plus\-LDA pipeline could not detect, including the distinct authentication friction, pricing, and geopolitical trust themes that emerge only with contextual embeddings\. Meng et al\.\[meng2026\]extracted feature\-related topics from their dataset of 100,000 reviews; however, since Meng et al\.’s analysis did not incorporate a framework related to trust or adoption, no distinction was made between friction topics such as signing\-in problems and subscription obstacles, as opposed to feature\-related topics\. Our research shows how analyzing reviews using a trust and usability perspective generates entirely new insights\.

### 5\.2Implications for Research

The paper makes a contribution to the literature based on five counts\. Firstly, the analysis shows that the review mining approach that is very common in software engineering when researching traditional apps\[martin2017survey\]could also be applied to generative AI, providing a new avenue for investigating questions related to trust, usability, and acceptance of technology beyond surveys and experiments\. Secondly, the naturalistic quality of the reviews allows the researcher to explore issues of trust \(geopolitical\), authentication, and pricing,g which may not come up in controlled experiments where the participants get access to the system for free\.

Second, the use of BERTopic in conjunction with transformer\-based sentiment classification constitutes an innovation from the methodology used in previous research on GenAI reviews, which utilized LDA followed by VADER\[alabduljabbar2024\]\. The contextual embeddings generated by BERTopic provide greater coherence in topics, while the sentiment classifier based on RoBERTa allows for better capture of contextual sentiment \(considering negation or partial praise\)\. Future researchers conducting app store reviews would do well to employ this pipeline as their default option, considering the relatively low computational cost compared to LDA\.

Thirdly, the result of the Cohen’s Kappa test \(κ=0\.241\\kappa=0\.241\) is a methodological result on its own\. The results clearly indicate that there are inherent limitations to the use of automated topic modeling approaches when applied to more abstract concepts like trust, privacy, and usability\. Thus, the study suggests a need for the use of manual thematic coding alongside automated topic models for research in which theory is drawn upon\[braun2006thematic\]\. It means that future work in the area cannot view automated topic modeling as a tool for the replacement of manual thematic coding, but as one of the discovery tools instead\.

Fourth, we propose the Trust Friction Score \(TFS\), an index that represents a multi\-dimensional measure of trust that emerges out of our investigation into the topic\. While negativity is correlated with both the prevalence of negative topics and their negativity strength, TFS takes into account both of these aspects simultaneously in a unified measure\. The breakdown to its constituent dimensions \(Equations[5](https://arxiv.org/html/2609.19151#S3.E5)–[11](https://arxiv.org/html/2609.19151#S3.E11)\) offers a generalizable solution to measuring trust friction in ASR research\.

Fifth, our finding that trust in generative AI applications operates along multiple independent dimensions, including performance trust, authentication trust, pricing trust, and geopolitical trust, suggests that existing unidimensional or even bidimensional trust scales may be insufficient for capturing the full range of user concerns in this domain\[mayer1995trust,hoff2015trust\]\. We encourage researchers to develop and validate multi\-dimensional trust measurement instruments specifically calibrated for consumer\-facing generative AI applications\.

### 5\.3Implications for Practice

Our findings yield five actionable recommendations for developers and product teams building consumer\-facing generative AI applications\.

First,streamline authentication flows as the highest\-priority trust intervention\. The 89% negativity rate for sign\-in and account issues indicates that authentication friction is the single most impactful quick win available to developers\. Requiring phone number verification for an AI chatbot appears disproportionate to users and should be reconsidered in favor of lighter authentication mechanisms\. Amershi et al\.’s\[amershi2019guidelines\]guidelines for human\-AI interaction recommend making clear what the system can do at the outset of interaction; this principle should be extended to ensuring that users can reach the system without unnecessary barriers\.

Second,implement transparent and well\-articulated pricing policies\. Considering the negative attitude towards subscription and pricing \(73%\), it is apparent that users do not trust the developers due to the lack of transparency associated with the free tier\. It is important that developers offer comparisons between the free and premium versions when users are prompted to upgrade their accounts, and that any form of restriction placed on the essential functions does not feel like punishment to the user\[weisz2024design\]\.

Third,focus on language and localization quality when designing applications targeting global customers\. The mistrust that exists among users of the DeepSeek application in Chinese is partly because of a failure in localization, where the application does not respond in the user’s language\. However, it is also partly a matter of mistrust\.

The fourth suggestion is tonot advertise through conversational AI platforms\. The fact that Microsoft’s Copilot received an approval rating of only 9% for its advertisements clearly shows user disapproval\. The very essence of communication via AI cannot be reconciled with advertisements, as users have certain expectations in such contexts\[fogg2003prominence\]\.

Fifth,communicate the constraints of the AI system truthfully and in advance\. The fact that we observed multiple constraints discussed by users in our data, such as constraints related to the message size limit \(T11\) or constraints related to server dependability \(T14\), suggests that users are disappointed not only by the constraints themselves but also by the absence of transparency regarding why they are there\. In this context, it is helpful for users to know what the reason behind the constraint is\[amershi2019guidelines\]\.

## 6Threats to Validity

We organize threats to validity into four categories following established guidelines for empirical software engineering research\[wohlin2012experimentation\]\.

##### Internal validity\.

Two methodological choices may affect the internal validity of our findings\. First, the sentiment classifier \(cardiffnlp/twitter\-roberta\-base\-sentiment\-latest\) was trained on Twitter data rather than app store reviews\[loureiro2022timelms\]\. App store reviews tend to be longer, more structured, and less colloquial than tweets, which may introduce classification errors at the margins, particularly for reviews with mixed or nuanced sentiment\. Although this was partially addressed via thematic validation manually, a sentiment analysis model trained specifically for app store data could increase classification accuracy\. Additionally, several BERTopic parameters, includingmin\_cluster\_size=30,n\_components=5, andnr\_topics=25,were determined experimentally rather than being optimized\. With another combination of parameters, there might be a different number of topics formed with varying boundaries, resulting in a completely different thematic map from what we discovered\. This was partially addressed using topic validation manually and selecting parameters to maximize interpretability and coherence of topics\.

##### External validity\.

The generalizability of our research results is limited in several ways\. First, we have examined only six AI generators; other popular platforms like Grok, Pi, Poe, and Character\.ai might provoke distinct user apprehensions\. Secondly, our analysis was conducted based on reviews in English, which does not include viewpoints from leading consumer markets like India, Brazil, Japan, and countries with Chinese speakers\. Since one of our principal results touches upon geopolitical trust, this language limitation can lead us to underestimate the importance and strength of cross\-cultural trust interactions\. Furthermore, our dataset is unbalanced across applications: Claude provided 5,037 reviews \(29\.6% of our entire dataset\), whereas Gemini offered only 812 reviews \(4\.8%\)\. To ensure this imbalance does not drive our conclusions, we conducted balanced\-subsample, platform\-stratified, and equal\-weight sensitivity analyses \(Section[4\.8](https://arxiv.org/html/2609.19151#S4.SS8), Appendix[A](https://arxiv.org/html/2609.19151#A1)\); the cross\-application findings held in essentially all balanced draws, so the imbalance affects the precision of individual per\-app estimates more than the substantive conclusions\. Finally, only three out of six applications had reviews in the Apple App Store, so the platform\-stratified comparisons are limited to those applications\.

##### Construct validity\.

The Cohen’s Kappa ofκ=0\.241\\kappa=0\.241between automated and manual topic assignments indicates only fair agreement\[landis1977kappa\], reflecting BERTopic’s inability to reliably capture abstract themes such as trust, privacy, and usability\. This means that the topic\-level findings for these abstract constructs should be interpreted as conservative lower bounds rather than precise estimates of their prevalence in the corpus\. Manual thematic validation was performed by two independent coders on a stratified sample of 300 reviews\. Inter\-coder reliability reached moderate agreement \(Cohen’sκ\\kappa= 0\.544, Krippendorff’sα\\alpha= 0\.543, 62\.7% exact agreement\), which is acceptable for a 12\-category qualitative coding task\. Disagreements \(n = 112\) were resolved through adjudication by the first author to produce gold\-standard labels\. While the primary purpose of the validation was diagnostic \(identifying where BERTopic succeeds and fails\) rather than establishing a gold\-standard human coding, a multi\-coder design with formal inter\-coder reliability assessment would strengthen the construct validity of the thematic analysis\[mcdonald2019reliability\]\.

##### Interpreting abstract\-construct topics\.

A direct consequence of the preceding point is that our trust\- and privacy\-related findings must be read as exploratory\. BERTopic reliably recovers lexically distinctive topics \(e\.g\., account or pricing complaints, where shared keywords make the cluster coherent\), but abstract constructs such as trust, privacy, and usability are expressed in heterogeneous language and were largely absorbed into broad positive/feature clusters by the model, yielding near\-zero agreement with human coders for these themes \(Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)\)\. We therefore do not interpret the size of a trust\- or privacy\-labelled topic \(notably T13\) as a measurement of how many users distrust an application\. Where we report quantitative figures for these topics, they are*sentiment*rates computed over the reviews in the topic—and the sentiment classifier producing them is validated against human labels \(accuracy 75\.3%, macro\-F1F\_\{1\}= 0\.725; Section[4\.5](https://arxiv.org/html/2609.19151#S4.SS5)\)—rather than prevalence estimates for the abstract construct itself\. Claims about trust, privacy, and geopolitical concern throughout the paper are consequently framed as indicative patterns that warrant targeted follow\-up \(e\.g\., survey or interview studies\) rather than as confirmatory prevalence measures\.

##### Reliability\.

Two data\-related decisions affect the reproducibility and completeness of our findings\. First, the 10\-word minimum length filter removed 34,530 of the 52,080 post\-language\-filter reviews \(66%\), discarding feedback that may be terse but meaningful \(e\.g\., “crashes constantly” or “love it”\)\. We examined the direction of this bias by comparing removed and retained reviews \(Appendix[C](https://arxiv.org/html/2609.19151#A3), Table[12](https://arxiv.org/html/2609.19151#A3.T12)\)\. The removed reviews are markedly more positive than the retained ones \(mean rating 4\.41 vs\. 3\.54 stars; 76\.6% five\-star vs\. 50\.8%\), and they over\-represent applications whose users tend to leave brief praise \(ChatGPT and Gemini\), while Claude—whose users write longer, more critical reviews—is comparatively under\-represented among the discarded set\. The practical consequence is that our corpus is skewed toward longer, more critical reviews, so the absolute negative\-sentiment percentages we report are best read as upper bounds relative to the full reviewer population; the cross\-application*comparisons*, however, are computed within this same filtered frame and are not undermined by the filter\. Although such filtering is essential for effective topic modeling \(short reviews lack the context BERTopic embeddings require\), a lower threshold would retain more terse feedback at the cost of noisier topics\. Secondly, our data was gathered from a single scraping run conducted in May 2026 and therefore reflects a one\-time snapshot of a roughly eight\-month period \(September 2025 to May 2026\) rather than a longitudinal study\. Consumer opinions can change over time depending on application improvements or changes to the pricing model, and a cross\-sectional approach cannot capture that information\. Moreover, the constraints of app store scraping APIs can limit the availability of historical data on consumer opinions for any particular application\.

## 7Conclusion

This research aimed to investigate the true feelings, perceptions, and challenges faced by users of AI\-based applications by analyzing 17,012 app store reviews for six popular applications, namely ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity\. BERTopic topic modeling, RoBERTa\-based sentiment analysis, and statistical analysis were used to identify the prominent user issues, sentiment variance between different applications, and the particular trust and usability\-related factors most predictive of dissatisfaction\.

Our study identified 24 topics, with five being found to be important trust and usability barriers\. The biggest usability challenge identified by our study is related to the friction involved with authentication procedures \(89% negative\)\. This suggests that the sign\-in procedure is a key trust\-busting factor\. Advertisements within conversational AI have also been rejected with great intensity \(91% negative\) and were specific to Microsoft Copilot\. Pricing issues have been flagged as an adoption issue \(73% negative\)\.

Server problems \(83% negative\) were seen to be associated with the teething problems that were occurring during the scaling up of AI infrastructure, specifically DeepSeek\. Notably, a subset of DeepSeek reviews surfaced an exploratory geopolitical\-trust theme \(Topic T13\), with users raising concerns over the app’s Chinese origin, data privacy, and occasional Chinese\-language replies; consistent with topic modeling’s limited reliability for abstract constructs, we report this as an indicative pattern rather than a prevalence estimate\. Additionally, we found that there was a polarization problem in relation to the application Claude, whereby the application recorded the highest rate of negative feedback \(47\.7%\) yet still had an enthusiastic group of positive users\.

This paper’s contributions can be enumerated into three categories\. Firstly, this paper offers, to the best of our knowledge, one of the first comparative cross\-application studies of generative AI app reviews, with the inclusion of six apps, including three \(Claude, DeepSeek, Perplexity\) which, to our knowledge, have not been analyzed in prior app store mining studies\. Secondly, this paper shows an advancement in methodology where the use of BERTopic with contextual embeddings and RoBERTa sentiment analysis provides much higher thematic granularity compared to the pipeline involving the use of LDA and VADER in previous literature, while simultaneously providing empirical insights into the limitations of automated topic modeling with respect to abstract concepts \(κ=0\.241\\kappa=0\.241\)\. Thirdly, this paper presents new empirical insights into issues of trust, usability, and adoption challenges that will guide the research on human factors in generative AI and the practical development of applications\.

There are multiple avenues for future research emerging from this paper\. For one thing, conducting a longitudinal study would be helpful in identifying whether the friction barriers that have been mentioned here are structural problems or are temporary pains that eventually fade out as the applications evolve\. For another thing, expanding our analysis to other languages besides English, such as Hindi, Spanish, Portuguese, and Chinese, would allow us to consider the points of view of large portions of users that have not been considered in the current study and would enable us to analyze the role of trust across cultures better\. Fourth, increasing the sample size by including other software applications like Grok, Pi, Poe, and Character\.ai, apart from other non\-mobile platforms like web interfaces, IDEs, and desktop applications, will increase the generalizability of the findings that we have discovered thus far\. Lastly, creating an AI tool\-specific sentiment classifier through fine\-tuning of existing models will help overcome the limitations associated with tweets that we faced when conducting our study\.

## Appendix ASampling\-Imbalance Sensitivity Analyses

This appendix reports the sensitivity analyses summarised in Section[4\.8](https://arxiv.org/html/2609.19151#S4.SS8)\. Table[9](https://arxiv.org/html/2609.19151#A1.T9)compares each key cross\-application test on the full corpus against balanced subsamples in which every application is downsampled to the smallest app’s size \(n=812n=812per application, 4,872 reviews per draw\) over 1,000 random draws\. Because the balancedNNis smaller than the full corpus, the raw test statistics are necessarily lower; the scale\-free effect sizes \(Cramér’sVV,η2\\eta^\{2\}\) are the appropriate basis for comparison and are essentially unchanged\. Table[10](https://arxiv.org/html/2609.19151#A1.T10)reports the platform\-stratified \(iOS vs\. Android\) negativity analysis\.

Table 9:Balanced\-subsample sensitivity \(1,000 draws,n=812n=812/app\)\. Balanced values are mean \[2\.5th, 97\.5th percentile\] across draws\. Effect sizes are essentially unchanged from the full sample, and the ordering of applications by negativity is highly stable\.QuantityFull sampleBalanced subsamplesSentiment×\\timesapp:χ2\\chi^\{2\}561\.3145\.8 \[110\.5, 185\.2\]Cramér’sVV0\.1280\.122 \[0\.107, 0\.138\]significant \(p<\.05p<\.05\)yes100% of drawsRatings×\\timesapp:HH639\.2159\.7 \[120\.0, 205\.0\]η2\\eta^\{2\}0\.0370\.032 \[0\.024, 0\.041\]significant \(p<\.05p<\.05\)yes100% of drawsClaude highest negativityyes99\.2% of drawsCopilot lowest negativityyes99\.8% of drawsPer\-app negative\-sentiment rate \(%\), balanced meanClaude47\.747\.7Perplexity42\.342\.3Gemini38\.738\.7DeepSeek31\.831\.8ChatGPT30\.930\.9Microsoft Copilot25\.325\.3Table 10:Platform\-stratified negative\-sentiment rates\. Only ChatGPT, Claude, and Perplexity appear on both stores; DeepSeek, Gemini, and Microsoft Copilot are Android\-only\. Within\-app tests are chi\-square on the sentiment×\\timesplatform table\.ScopeAndroidiOSSentiment×\\timesPlatformAll apps37\.0 \(12,917\)40\.8 \(4,095\)—ChatGPT22\.9 \(659\)35\.3 \(1,182\)χ2=30\.5\\chi^\{2\}=30\.5,p<\.001p<\.001,V=0\.13V=0\.13Claude47\.5 \(3,576\)48\.1 \(1,461\)χ2=1\.4\\chi^\{2\}=1\.4,p=\.49p=\.49\(ns\)Perplexity45\.3 \(2,142\)37\.9 \(1,452\)χ2=21\.7\\chi^\{2\}=21\.7,p<\.001p<\.001,V=0\.08V=0\.08An equal\-weight aggregate \(mean of the six per\-application negativity rates\) is 36\.1%, close to the raw pooled 37\.9%, further indicating that corpus\-level summaries are not driven by the larger corpora\.

## Appendix BTrust Friction Score: Baselines and Robustness

Table[11](https://arxiv.org/html/2609.19151#A2.T11)accompanies the TFS validation in Section[4\.7](https://arxiv.org/html/2609.19151#S4.SS7)\. It reports how the composite TFS ranking of the six applications behaves under alternative weightings and definitions, alongside its agreement with the two scalar baselines \(negative\-sentiment rate and friction\-topic prevalence\)\. Agreement is measured by Spearman rank correlation against the main prevalence\-weighted TFS and against the negativity baseline; the “top app” column records which application receives the highest score under each variant\.

Table 11:Robustness of the Trust Friction Score \(TFS\)\.TFS definitionTop app𝝆\\boldsymbol\{\\rho\}vs\. TFS𝝆\\boldsymbol\{\\rho\}vs\. Neg\.Main \(prevalence\-weighted\)Claude1\.000\.89Equal\-weight sub\-dimensionsClaude0\.430\.60Star\-rating negativity \(≤2\\leq 2\)Claude0\.940\.94Drop authentication \(T05\)Perplexity0\.940\.83Drop chat limits \(T11\)Claude1\.000\.89Drop pricing \(T12\)Claude0\.660\.49Drop geopolitical \(T13\)Claude0\.940\.94Drop reliability \(T14\)Claude1\.000\.89Drop advertising \(T16\)Claude0\.940\.94
Baseline rank agreement:ρ\(TFS,Neg\.\)=0\.89\\rho\(\\mathrm\{TFS\},\\mathrm\{Neg\.\}\)=0\.89,ρ\(TFS,Prev\.\)=0\.89\\rho\(\\mathrm\{TFS\},\\mathrm\{Prev\.\}\)=0\.89,ρ\(Neg\.,Prev\.\)=0\.83\\rho\(\\mathrm\{Neg\.\},\\mathrm\{Prev\.\}\)=0\.83\.

## Appendix CMinimum\-Length Filter: Removed vs\. Retained Reviews

Table[12](https://arxiv.org/html/2609.19151#A3.T12)supports the bias analysis in Section[6](https://arxiv.org/html/2609.19151#S6)\. It compares the 34,530 reviews removed by the 10\-word minimum\-length filter against the 17,550 retained \(before deduplication\), across star rating and application\. Removed reviews are substantially more positive and are concentrated in applications whose users tend to leave short praise\.

Table 12:Effect of the 10\-word minimum review\-length filter\.RemovedRetained\(<10\)\(≥\\geq10\)Number of reviews34,53017,550Mean star rating4\.413\.545\-star \(%\)76\.650\.81\-star \(%\)9\.725\.9Application share \(%\)ChatGPT22\.511\.2Gemini18\.85\.0Perplexity17\.921\.1Microsoft Copilot14\.515\.6Claude13\.729\.3DeepSeek12\.517\.8
Note:Results are based on the post\-language\-filter corpus \(n=52,080n=52\{,\}080\)\. Short reviews \(<10 words\) are substantially more positive, with higher mean ratings and a larger proportion of 5\-star reviews\. Rating distributions differ significantly between groups \(Mann–WhitneyUU,p<\.001p<\.001\), indicating that the retained corpus is biased toward longer, more critical reviews\.

## Appendix DTopic Coherence

To complement the manual thematic validation \(Section[4\.4](https://arxiv.org/html/2609.19151#S4.SS4)\), we computed two standard automatic topic\-coherence metrics over the top\-10 words of each of the 24 non\-outlier topics, using the analysed corpus as reference: theCvC\_\{v\}measure and normalised pointwise mutual information \(CNPMIC\_\{\\text\{NPMI\}\}\)\. The model attained a meanCvC\_\{v\}of 0\.556 \(median 0\.546\) and a meanCNPMIC\_\{\\text\{NPMI\}\}of 0\.054 \(median 0\.045\), values typical of BERTopic on noisy short\-text review corpora\. Consistent with the manual validation, the most coherent topics are lexically distinctive \(T12 Subscription/Pricing,Cv=0\.76C\_\{v\}=0\.76; T05 Sign\-in/Account,Cv=0\.73C\_\{v\}=0\.73\), whereas the least coherent are broad, heterogeneous clusters \(T00 General positive,Cv=0\.42C\_\{v\}=0\.42; T18 Meta\-rating,Cv=0\.41C\_\{v\}=0\.41\)\. This pattern corroborates the finding that BERTopic recovers concrete, keyword\-driven topics well but abstract or catch\-all clusters less reliably\.

## References

Similar Articles

Gen AI Website Traffic Share

Reddit r/singularity

An analysis of traffic share for generative AI websites, highlighting which platforms are gaining or losing visitors.

Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles

arXiv cs.CL

This scoping review analyzes 81 articles (2022-2025) examining the use of generative AI for creating and evaluating user personas, identifying strengths in reproducibility but critical issues including lack of evaluation in 45% of studies, over-reliance on GPT models (86%), and risks of circularity where the same model generates and evaluates personas.

The Role of AI in Online Reviews

arXiv cs.CL

This paper introduces an empirical approach to measure the impact of large language model supply shocks on online reviews, finding that unverified reviews shift toward greater negativity and activity bursts occur, suggesting AI is reshaping platform dynamics.