Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

arXiv cs.CL Papers

Summary

This paper analyzes real-world user queries about digital security and privacy asked to LLMs, categorizing them into nine topics and evaluating response quality and consistency across commercial and open-weight models.

arXiv:2606.18062v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to fulfill users' information needs; users ask LLMs about the weather, pose educational questions, and consult them for legal assistance. One particularly understudied area is digital security and privacy (S&P), where users may seek LLMs' help on how to secure their online accounts or protect their computers from cyber attacks. To the best of our knowledge, no prior study has collected or analyzed the S&P questions users ask LLMs; prior research on LLM response quality relied on expert-authored S&P misconceptions or FAQs rather than user queries. Drawing from WildChat, a dataset of 3.2M user-LLM conversations collected in the wild, our study identifies 14,727 S&P prompts and categorizes them into nine categories covering a wide range of S&P topics. From the S&P prompts, we sampled 450 and performed a thematic analysis to characterize the S&P questions users ask LLMs. Separate from the thematic analysis, we curated 270 advice-seeking S&P prompts, where users ask for recommendations, guidance, or specific S&P information. We measured LLM response quality and consistency when posing the prompt to LLMs 10 times. We found that commercial LLMs outperform open-weight models (GPT 5.5 provided "good enough" responses on 98% of prompts; Llama 4 on 47%). However, among prompts that received high-quality responses on average, commercial models sometimes produce contradictory responses across runs, risking confusing or misleading users.
Original Article
View Cached Full Text

Cached at: 06/17/26, 05:42 AM

# Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
Source: [https://arxiv.org/html/2606.18062](https://arxiv.org/html/2606.18062)
Hobin Kim1,Xiaoyuan Wu11footnotemark:11,Omer Akgul2,Lujo Bauer1,Nicolas Christin1 1Carnegie Mellon University,2RSAC Labs, Correspondence:[hobink@andrew\.cmu\.edu](https://arxiv.org/html/2606.18062v1/mailto:[email protected])

###### Abstract

Large language models \(LLMs\) are widely used to fulfill users’ information needs; users ask LLMs about the weather, pose educational questions, and consult them for legal assistance\. One particularly understudied area is digital security and privacy \(S&P\), where users may seek LLMs’ help on how to secure their online accounts or protect their computers from cyber attacks\. To the best of our knowledge, no prior study has collected or analyzed the S&P questions users ask LLMs; prior research on LLM response quality relied on expert\-authored S&P misconceptions or FAQs rather than user queries\. Drawing from WildChat, a dataset of 3\.2M user\-LLM conversations collected in the wild, our study identifies 14,727 S&P prompts and categorizes them into nine categories covering a wide range of S&P topics\. From the S&P prompts, we sampled 450 and performed a thematic analysis to characterize the S&P questions users ask LLMs\. Separate from the thematic analysis, we curated 270 advice\-seeking S&P prompts, where users ask for recommendations, guidance, or specific S&P information\. We measured LLM response quality and consistency when posing the prompt to LLMs 10 times\. We found that commercial LLMs outperform open\-weight models \(GPT 5\.5 provided “good enough” responses on 98% of prompts; Llama 4 on 47%\)\. However, among prompts that received high\-quality responses on average, commercial models sometimes produce contradictory responses across runs, risking confusing or misleading users\.

Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond

Hobin Kim††thanks:Equal contribution1, Xiaoyuan Wu11footnotemark:11, Omer Akgul2, Lujo Bauer1, Nicolas Christin11Carnegie Mellon University,2RSAC Labs,Correspondence:[hobink@andrew\.cmu\.edu](https://arxiv.org/html/2606.18062v1/mailto:[email protected])

## 1Introduction

As LLMs grow increasingly capable and widely adopted, users are turning to them as everyday information sourcesChatterjiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib20)\)\. One particularly important but understudied area is digital security and privacy \(S&P\): it remains unclear what S&P questions users ask LLMs and how well LLMs answer them\. Incorrect responses in this domain may lead to real\-world consequences \(e\.g\., compromised accounts, exposed digital identities\)\. In addition to response quality, consistency of responses is also important: prior S&P research has shown that conflicting advice confuses users and may undermine their protective behaviorsReederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\); Neilet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib10)\), underscoring the need for LLMs to respond reliably across sessions\. To our knowledge, no prior work has examined what S&P questions users ask LLMs\. Further, prior work evaluated the quality of LLM responses to expert\-authored S&P misconceptions and FAQsChenet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib1)\); Prakashet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib2)\)rather than on questions from real users\.

To address these research gaps, we curate a dataset of real\-world S&P prompts posed to LLMs and use the dataset to evaluate LLM response quality and consistency\. Specifically, we seek to answer the following research questions:

RQ1:What types of S&P questions do users ask LLMs?

RQ2:What is the quality of LLM responses to users’ S&P questions?

RQ3:How consistent are LLM responses to the same S&P question across repeated queries?

We identified 14,727 S&P prompts from the 1\.6M English conversations in WildChatZhaoet al\.\([2024b](https://arxiv.org/html/2606.18062#bib.bib4)\)and classified them into nine categories covering S&P topics from prior workChenet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib1)\); Hasegawaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib9)\); Prakashet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib2)\); Reederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\); Thomaset al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib3)\)\(§[3\.1](https://arxiv.org/html/2606.18062#S3.SS1)\)\. We randomly sampled 50 prompts per category and conducted a thematic analysis of the resulting 450 prompts \(§[3\.2](https://arxiv.org/html/2606.18062#S3.SS2)\)\. Since coding and writing quality are well\-studied by general LLM benchmarksLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Zhenget al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib41)\), we scoped our quality and consistency evaluation to advice\-seeking prompts, where users ask for recommendations, guidance, or specific S&P information\. We curated 270 advice\-seeking S&P prompts and measured response quality by adopting the LLM\-as\-judge checklist method established by prior workLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Weiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib16)\), in which binary checklists specify what a correct answer should include\. To assess consistency, we instructed the LLM judges to extract evidence quotes from each LLM response per checklist item, then checked whether quotes from two independent runs of the same prompt entailed each other \(§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\)\.

Our thematic analysis of 450 S&P prompts revealed six themes and 22 sub\-themes \(RQ1\)\.General knowledgewas the most prevalent \(33\.3%\) theme; three additional themes captured S&P\-specific LLM use cases that are not commonly observed in general LLM usage:Defensive action\(11\.8%\), in which users seek protective assessments or countermeasures;Inquiry about the LLM\(10\.2%\), where users probe the model’s own S&P capabilities and limits; andHarmful & offensive requests\(6\.9%\), in which users seek assistance with attacks or exploits \(§[4\.1](https://arxiv.org/html/2606.18062#S4.SS1)\)\. We found that the three commercial LLMs outperformed the two open\-weight models on response quality \(rated on a scale from 1 to 10 where 4 or lower indicates poor, while 7 or higher means goodLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\)\): GPT 5\.5 achieved the highest mean score \(8\.67\) and Llama 4 the lowest \(6\.71\) \(RQ2; §[4\.2](https://arxiv.org/html/2606.18062#S4.SS2)\)\. AddressingRQ3, Llama 4 produced the most consistent responses across repeated runs, despite earning the lowest average quality score \(§[4\.3](https://arxiv.org/html/2606.18062#S4.SS3)\)\.

We contextualize our results by comparing users’ S&P prompts to LLMs with prior research on online forums \(§[5\.1](https://arxiv.org/html/2606.18062#S5.SS1)\), and argue that both response quality*and consistency*should be reported to fully characterize LLM reliability \(§[5\.2](https://arxiv.org/html/2606.18062#S5.SS2)\)\.

![Refer to caption](https://arxiv.org/html/2606.18062v1/figures/overview.png)Figure 1:Overview of our study design
## 2Background and Related Work

We first review existing research characterizing the S&P questions users ask across online platforms and where they seek answers, motivating the need to study real\-world S&P prompts directed at LLMs \(§[2\.1](https://arxiv.org/html/2606.18062#S2.SS1)\)\. We then discuss existing research on evaluations of S&P advice quality from both human sources and LLMs \(§[2\.2](https://arxiv.org/html/2606.18062#S2.SS2)\)\. Finally, we review automated frameworks for evaluating LLM response quality and consistency, on which our evaluation methodology builds \(§[2\.3](https://arxiv.org/html/2606.18062#S2.SS3)\)\.

### 2\.1Users’ S&P Questions

Prior to the widespread usage of LLMs, users seeking S&P guidance turned to online forums and Q&A platforms, and prior work has characterized these interactions\. Developers commonly asked about practical challenges \(e\.g\., privacy policy compliance, access control\)Tahaeiet al\.\([2020](https://arxiv.org/html/2606.18062#bib.bib18)\), while non\-expert users sought help with cyberattacks, privacy abuse, and authenticationHasegawaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib9)\)\. Platform\-level analyses reinforce these themes: Reddit S&P help\-seeking centers on scams, account access, and privacy toolsThomaset al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib3)\), and these informal channels are disproportionately relied upon by less technically skilled usersRedmileset al\.\([2016](https://arxiv.org/html/2606.18062#bib.bib23)\)\.

LLMs are now widely used as an everyday information sourceBurtchet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib19)\); Chatterjiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib20)\); Lianget al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib21)\), yet it remains unclear what S&P questions users bring to them\. LLMs differ from the channels studied in prior work: interactions are private, one\-on\-one exchanges rather than public threads where community members can debate and refine answers; and unlike forum posts that accumulate peer vetting over time, LLM responses are generated on demand with no community reviewBurtchet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib19)\); del Rio\-Chanonaet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib44)\)\. These structural differences may shape both what users ask and the responses they receive, motivating our study of the S&P questions users bring to LLMs\.

### 2\.2Responses to S&P Questions

The quality of S&P responses matters because poor or misleading advice can directly impact users’ security and privacy decisions\. Prior work found that online security advice is often scarce and ambiguous, making it unlikely to produce good user behaviorsBhagavatulaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib31)\); the main challenge is not quality per se, but helping users prioritize which advice to act on—compounded by users’ tendency to over\-report their own secure behaviors, complicating any assessment of advice effectivenessRedmileset al\.\([2018](https://arxiv.org/html/2606.18062#bib.bib33),[2020](https://arxiv.org/html/2606.18062#bib.bib17)\)\.

More recently, researchers have begun investigating LLM S&P responses, though only on curated, researcher\-defined inputs rather than real user questions\.Chenet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib1)\)evaluated LLMs on common S&P misconceptions and found they incorrectly endorsed those misconceptions in a non\-negligible share of cases\. Further, researchers evaluated LLM responses on S&P FAQs, and found models frequently failed to surface relevant research findings, with safety guardrails further impeding the delivery of useful advicePrakashet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib2)\)\. Building on these findings, we evaluate LLM response quality on actual user S&P questions to better understand the risks users face when turning to LLMs for S&P guidance\.

### 2\.3LLM Response Quality and Consistency

Prior work has developed methods for evaluating LLM responses on both quality and consistency; we adopt these for the S&P domain\. LLM quality has been benchmarked across general open\-ended queriesHendryckset al\.\([2021](https://arxiv.org/html/2606.18062#bib.bib46)\); Linet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Zhenget al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib41)\); Srivastavaet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib47)\), domain\-specific tasks \(coding, mathematics\), and expert\-level domainsGuhaet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib48)\); Noriet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib49)\), including cybersecurityJinget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib37)\); Liuet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib8)\); Singeret al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib50)\)—though these focus on expert\-level operations rather than everyday S&P questions\. Evaluation methodology has evolved from manual inspectionBrownet al\.\([2020](https://arxiv.org/html/2606.18062#bib.bib51)\); Ouyanget al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib52)\)and crowdsourced pairwise preferencesWuet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib27)\); Zhenget al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib41)\)to automated checklist\-based LLM\-as\-judge frameworksLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Weiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib16)\)\. Beyond quality, consistency is equally critical: in S&P, contradictory answers across runs can leave users confused about which guidance to followReederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\)or cause disengagement from protective actionBhagavatulaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib31)\)\. Consistency metrics have evolved from token\-level BLEUPapineniet al\.\([2002](https://arxiv.org/html/2606.18062#bib.bib53)\)and BERTScoreZhanget al\.\([2020](https://arxiv.org/html/2606.18062#bib.bib54)\)to logit\-based uncertainty estimationDuanet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib28)\); Kuhnet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib30)\); Wuet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib27)\)and entailment\-based scoringDuanet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib28)\); Zhanget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib26)\)\. For our work, we adopt the checklist\-based approachLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Weiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib16)\)for quality and entailment\-based scoringZhanget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib26)\)for consistency \(§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\)\.

## 3Methods

We identified 14,727 S&P prompts from 3\.2M WildChat conversations \(§[3\.1](https://arxiv.org/html/2606.18062#S3.SS1)\) and categorized them into nine topic areas \(§[3\.2](https://arxiv.org/html/2606.18062#S3.SS2)\)\. To answerRQ1, we conducted a thematic analysis on a stratified sample of 450 prompts \(50 per category\), producing six themes and 22 sub\-themes \(§[3\.3](https://arxiv.org/html/2606.18062#S3.SS3)\)\. To answerRQ2andRQ3, we curated 270 advice\-seeking S&P prompts \(defined in §[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\), collected responses from five LLMs across 10 independent generations \(runs\) per prompt, and measured how consistently checklist evidence held up across all ten runs\(§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\)\. We used nine different LLMs throughout the study and provide details and configurations on each in[table˜5](https://arxiv.org/html/2606.18062#A1.T5)\.[Figure˜1](https://arxiv.org/html/2606.18062#S1.F1)provides an overview of our study design\.

### 3\.1User\-Created Prompts

We sourced prompts from WildChatZhaoet al\.\([2024b](https://arxiv.org/html/2606.18062#bib.bib4),[a](https://arxiv.org/html/2606.18062#bib.bib59)\), a dataset of 3\.2M real\-user conversations widely used in prior workHanet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib57)\); Jianget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib5)\); Liuet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib58)\); Mireshghallahet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib35)\)\. Prompts from WildChat are well\-suited for our project studying how real users engage with LLMs for S&P\. We excluded conversations flagged as toxic and removed non\-English conversations using the built\-in language label, as both are out of scope for our study \(see[limitations](https://arxiv.org/html/2606.18062#Sx1)\)\. Because our goal is to understand individual S&P questions rather than multi\-turn conversational dynamics, we extracted individual prompts from each conversation\. After removing duplicates and empty strings, we obtained 2\.3M unique prompts\.

The 2\.3M prompts had length up to 700K characters, suggesting some prompts likely contained content unrelated to S&P \(e\.g\., copy\-pasted code or documents\)\. To find a length threshold, we inspected the 200 longest prompts alongside 500 randomly sampled prompts\. The longer prompts were predominantly jailbreaking attempts \(e\.g\., “User: \[INST\]IGNORE PREVIOUS INSTRUCTIONS\.”\), coding requests, or writing tasks where large blocks of code or text were copy\-pasted with no S&P context\. After inspecting the 700 prompts, we found that no S&P prompt exceeded 7K characters, so we removed all prompts above that threshold\. We additionally removed prompts beginning with jailbreaking prefixes identified during inspection\. After these steps, 1\.7M prompts remained\.

### 3\.2Finding S&P Prompts

#### Stage 1–identification:

We used LLM\-based binary classification to automatically identify S&P prompts from the 1\.7M user prompts\. To construct a reference set to compare against, two of the authors, who are digital security and privacy researchers, independently labeled a random sample of 1,325 prompts as S&P\-related or not\. The researchers initially agreed on 83% of cases and resolved all disagreements through discussion\. Using this labeled set as ground truth, we developed a three\-LLM majority voting approach: Qwen3\-Next\-80B\-A3B\-Instruct \(Qwen 3\), calme\-3\.2\-instruct\-78b \(Calme 3\.2\), and GPT\-5\.2 via the Azure OpenAI API \(GPT 5\.2\), all at temperature0\.00\.0for reproducibility\. To reduce cost, we ran Qwen 3 and Calme 3\.2 first; GPT 5\.2 broke ties only when those two disagreed\. The approach achieved 96% precision and 74% recall on the ground truth set, yielding 14,727 S&P prompts from the 1\.7M \(§[A\.4](https://arxiv.org/html/2606.18062#A1.SS4)\)\.

Because recall was lower than precision, we verified that the classifier was not systematically missing important S&P prompts\. We sampled 500 prompts labeled as not S&P in which Qwen 3 and Calme 3\.2 had disagreed \(i\.e\., where false negatives are most likely, since one model had flagged them as S&P\) and manually inspected them\. We found 48 false negatives \(9\.6%\)\. Through manual inspection, we found these 48 were semantically similar to prompts already captured by the majority vote, indicating that the classifier is unlikely to miss meaningfully distinct S&P prompts\.

#### Stage 2–categorization:

We categorized the 14,727 S&P prompts into nine categories\. We constructed the nine\-category taxonomy by adapting S&P categories from prior workChenet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib1)\); Hasegawaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib9)\); Prakashet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib2)\); Reederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\); Thomaset al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib3)\)and iteratively refined them against the S&P prompts\. The resulting categories are detailed in[table˜3](https://arxiv.org/html/2606.18062#A1.T3)\.

To select an LLM classifier, two researchers collaboratively labeled 300 randomly sampled S&P prompts using the taxonomy and reached consensus on all\. We used these 300 prompts to iteratively develop a categorization instruction prompt, evaluating outputs across six LLMs \(Calme 3\.2, Qwen 3 3, Qwen3\-Next\-80B\-A3B\-Thinking, Gemma\-4\-31B\-IT, Llama\-3\.3\-70B\-Instruct, and GPT\-5\.2\)\. Once the instruction prompt was finalized \(§[A\.5](https://arxiv.org/html/2606.18062#A1.SS5)\), we evaluated all six models against the 300\-prompt ground truth; Gemma\-4\-31B\-IT \(temp\.=0=0\) achieved the highest performance \(average precision 0\.90, recall 0\.89 across nine categories\) and was used to categorize all 14,727 S&P prompts\.

### 3\.3How LLMs are Used for S&P

We performed thematic analysisThomas \([2006](https://arxiv.org/html/2606.18062#bib.bib38)\)on a stratified sample of 450 S&P prompts \(50 per category, comparable toTahaeiet al\.\([2020](https://arxiv.org/html/2606.18062#bib.bib18)\); Hasegawaet al\.\([2022](https://arxiv.org/html/2606.18062#bib.bib9)\)\) drawn from the 14,727 identified in Stage 2\. Instead of adapting general human–AI interaction taxonomiesShelbyet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib25)\); Linet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\), we developed S&P\-specific themes to better capture user behaviors\. Two researchers iteratively coded the sample: the first proposed an initial grouping into candidate themes; the second validated and refined them; the cycle repeated until both reached consensus through in\-person discussion\. The analysis produced six themes and 22 sub\-themes, reported in[section˜4\.1](https://arxiv.org/html/2606.18062#S4.SS1); the full codebook is provided in[section˜A\.3](https://arxiv.org/html/2606.18062#A1.SS3)\.

### 3\.4Evaluating LLM Responses

Not all 14,727 S&P prompts are advice\-seeking; many are coding requests \(e\.g\., creating a safer version of the code\) or creative writing \(e\.g\., drafting a reply to a scam email\)\. We focused on advice\-seeking prompts, defined as prompts asking for recommendations, guidance, or specific S&P information \(followingLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\)\)\. Different from coding or writing tasks, advice\-seeking S&P prompts require LLMs to reason about domain\-specific concerns \(e\.g\., threat models, regulatory requirements\)\. Two researchers labeled prompts as “advice\-seeking” or not drawn from each category until 30 advice\-seeking prompts per category were reached, reviewing 1,659 prompts in total to yield 270 advice\-seeking S&P prompts\.

#### Generating responses

We selected five LLMs spanning commercial and open\-weight providers: Claude\-opus\-4\-7 \(Claude 4\.7\), Gemini 3\.1 Pro Preview \(Gemini 3\.1\), GPT\-5\.5\-2026\-04\-23 \(GPT 5\.5\), Qwen 3, and Llama\-4\-Scout\-17B\-16E \(Llama 4\)\. For each provider, we chose the most capable model version available at the time of the study that fit within our API budget and GPU resources\. The commercial models were accessed via official APIs; the open\-weight models were deployed on four Nvidia H100 GPUs\. All models ran at their default settings to reflect the experience of an average user; for the commercial models, this included thinking mode and web search \(see[table˜5](https://arxiv.org/html/2606.18062#A1.T5)\)\. We generated 10 independent responses per model per prompt, yielding 2,700 responses per model\.

#### Evaluating response quality

Using methods from WildBench and RocketEval, we built a per\-prompt checlist of binary \(yes/no\) criteria that specify what a correct response should cover, then used it to score response qualityLinet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\); Weiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib16)\)\.

Checklist construction\.For each of the 270 prompts, two commercial LLMs \(GPT 5\.5\-pro and Claude 4\.7—successors of the checklist models used in WildBench\) generated a checklist each; two researchers manually reviewed all 540, removing redundant items and reconciling similar ones to reach consensus on the final 270 checklists\.

Scoring\.Following WildBench, we initially scored responses using GPT 5\.5 \(scoring prompt in §[A\.6](https://arxiv.org/html/2606.18062#A1.SS6)\), but found self\-preference bias: GPT 5\.5 and Gemini 3\.1 each assigned the highest average scores to their own responses, while Claude 4\.7 assigned the highest average scores to GPT 5\.5 responses\. To mitigate bias, we averaged scores across all three models\. We report quality as the mean score across 10 runs per model per prompt, and spot\-checked 2,700 responses \(top\-10 and bottom\-10 per scorer–model–category combination;nn=3, 5, 9\) to verify score extremes reflected genuine quality differences\. We share results in §[4\.2](https://arxiv.org/html/2606.18062#S4.SS2)\.

#### Measuring consistency

We measure response consistency to assess how often LLMs provide similar or conflicting advice to users’ S&P prompts\. Unlike prior work comparing full responsesKuhnet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib30)\); Wuet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib27)\)or all sentencesDuanet al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib28)\); Manakulet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib29)\), we compare evidence quotes—scorer LLMs extract these sentences to support each checklist item—so consistency is measured on content that directly addresses the S&P question\. For each prompt, we merge quotes across the three scorer LLMs, remove exact duplicates, and form\(102\)=45\\binom\{10\}\{2\}=45pairwise comparisons\.

We adoptZhanget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib26)\)’s bidirectional entailment approach, chosen over embedding\-based methods \(USE, BERTScoreWuet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib27)\); Manakulet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib29)\)\) because entailment captures logical contradiction—the failure mode most relevant when users receive conflicting S&P advice\. We manually inspected all three scoring methods \(i\.e\., entailment, USE, and BERTScore\) on 50 pairs of responses and confirmed entailment best matched researcher judgment\. For each of the 45 response pairs per prompt, we usednli\-deberta\-v3\-largeto compute the bidirectional entailment score between their evidence quote lists and averaged these to obtain a per\-prompt consistency score\. We share consistency results in[section˜4\.3](https://arxiv.org/html/2606.18062#S4.SS3)\.

## 4Results

From 14,727 S&P prompts across nine categories, we thematically analyzed 450 to characterize users’ questions \(§[4\.1](https://arxiv.org/html/2606.18062#S4.SS1); RQ1\), then evaluated five LLMs on 270 advice\-seeking prompts\. We found that commercial LLMs outperformed open\-weight models on quality \(§[4\.2](https://arxiv.org/html/2606.18062#S4.SS2); RQ2\)\. Further, we report results on response consistency, where we found Llama 4, despite having the lowest average quality of responses, is the most consistent \(§[4\.3](https://arxiv.org/html/2606.18062#S4.SS3); RQ3\)\.

### 4\.1Thematic Analysis of Users’ S&P Prompts

The analysis revealed six themes and 22 sub\-themes \([table˜7](https://arxiv.org/html/2606.18062#A1.T7)\), spanning general knowledge\-seeking \(e\.g\., what is phishing\) to adversarial requests \(e\.g\., bypassing Windows login\)\. We then examine three that highlight S&P\-specific use cases of LLMs:defensive action,harmful & offensive requests, andinquiry about the LLM\.

#### Themes of users’ S&P prompts to LLMs

General knowledgewas the most frequent theme \(33\.3%, 150 out of 450 prompts\), followed byuser\-side navigation\(20\.9%\),S&P task production\(13\.8%\),defensive action\(11\.8%\),inquiry about the LLM\(10\.2%\), andharmful & offensive requests\(6\.9%\) \(defined inLABEL:tab:codebook\)\. Withingeneral knowledge,topic exploration\(46\.7% of 150 prompts\) andquestion\-answering\(28\.7%\) were the dominant sub\-themes\. The former encompasses broad informational queries similar to search\-engine lookups \(e\.g\., “What is covert channel?”\), while the latter largely consists of exam\- or certification\-style multiple\-choice questions\.User\-side navigation\(20\.9%;n=94n=94\) is the second most frequent theme, whereincident response\(52\.1% of 94 prompts\) was the most common sub\-theme, encompassing prompts about how to respond to account or functionality bans on social media platforms\. E\.g\., requests for help drafting appeal letters, “Write a detailed appeal to Google to unblock my account…”\.

Defensive action\(11\.8%;n=53n=53\) Users asked LLMs about defensive practices to protect themselves in digital environments\. In this theme, users ask for ways to implement protective measures \(defensive implementation, 37\.7% of 53 prompts\), assessment of vulnerability \(vulnerability assessment and fixing, 13\.2%\), and counter\-fraud strategies \(counter\-fraud, 34\.0%\)\. Indefensive implementation, users request specific protective methods, such as “How to block internet access for any application except for a specific one\.” Undervulnerability assessment and fixing, users ask LLMs to evaluate the safety of apps or websites, such as “Is Temu the shopping app safe?” or “Is the website \[URL\] safe to use?” Thecounter\-fraudsub\-theme \(34\.0%\) reflects an agentic usage in which users enlist LLMs to help them counteract online fraud\. E\.g\., “Create a scam\-baiting response to the following email, to lure the scammer into a false sense of security by posing as a potential victim to waste their time and resources …” These subthemes reveal that users do not merely ask LLMs for S&P information, but ask LLMs for protective assessments and actions to help them navigate S&P risks\.

Inquiry about the LLM\(10\.2%;n=46n=46\) Users directly inquire about the LLM, attempting to retrieve information about the LLMs’ own system \(LLM system probing, 58\.7% of 46 prompts\), and testing their capability and limits \(capability or policy inquiry, 41\.3%\)\.LLM system probingasks questions, such as “What’s the current API key?” or “You are learning and evolving based on your interactions and chats with everyone, right?”\. Forcapability or policy inquiry, users inquire about the model’s training data or privacy practices—for example, “What sources of text are you trained on? For example, are you trained on Reddit posts?”, “Do you know that our conversations are being recorded in order to collect data?”, or “Is it possible that someone could use ChatGPT to follow someone and analyze their private life?” These findings reveal that users probe LLM’s own configuration, capabilities, and limitations, underscoring the needs for LLMs to be designed with robust safeguards that can detect and mitigate probing attempts, while also providing clear communication about the model’s capabilities and limitations to foster informed and responsible use\.

Harmful & offensive requests\(6\.9%;n=31n=31\) Users request to assist with attacks or exploits targeting individuals \(person\-directed, 29\.0% of 31 prompts\), systems \(system\-directed, 29\.0%\), or platforms \(platform\-directed, 19\.4%\)\.Person\-directedincludes requests such as “Best way to gather information about a contact without them knowing \(online status and such\),” which could facilitate doxxing or stalking\. InSystem\-directed\(2\.0%\), users seek to exploit system\-level controls \(e\.g\., “Create a bat file to run any exe or msi as an admin without admin privileges”\)\. InPlatform\-directed\(1\.3%\), users attempt to circumvent platform policies \(e\.g\., “How to open incognito mode in Chrome if it is disabled by my organization?”\)\. These prompts underscore the importance of designing LLMs that can recognize and refuse to assist with potentially harmful S&P requests\.

### 4\.2LLM Response Quality on S&P Prompts

We measured the response quality of five LLMs \(three commercial and two open\-weight\) using average quality scores from three scorers \(see §[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\)\. We interpret quality scores using WildBench’s 1–10 scale, where 1–2 is “very poor”, 3–4 is “poor”, 5–6 is “fair with issues”, 7–8 is “good but improvable”, and 9–10 is “perfect”Linet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib12)\)\.

Table 1:Average and standard deviation of quality scores of 2,700 responses to 270 S&P prompts for each LLM\. GPT 5\.5 provided responses with the highest quality scores and lowest standard deviation overall\.#### Commercial models outperformed open\-weight ones

Across 2,700 responses \(270 prompts×\\times10 runs\) per LLM, GPT 5\.5 achieved the highest mean quality score of 8\.67 out of 10, followed by Gemini 3\.1 \(8\.52\) and Claude 4\.7 \(8\.47\)\. The open\-weight Qwen 3 scored 7\.90 \(“good but improvable”\) but notably below the commercial LLMs\. Llama 4 performed worst at 6\.71 \(“fair with issues”\), nearly two points below GPT 5\.5 \([table˜1](https://arxiv.org/html/2606.18062#S4.T1)\)\.

![Refer to caption](https://arxiv.org/html/2606.18062v1/x1.png)Figure 2:Percentage of prompts exceeding each quality score threshold \(averaged over 10 runs\)\. Commercial models all have over 90% of prompts clearing the threshold of 7 \(“good but improvable”\); Qwen 3 trails at 77%; Llama 4 drops steeply to 47%, meaning more than half its responses are at best “fair with issues\.”[Figure˜2](https://arxiv.org/html/2606.18062#S4.F2)shows the percentage of prompts on which each model’s mean score \(over 10 runs\) exceeded quality thresholds\. GPT 5\.5 scored at least seven on 98% of prompts, Gemini 3\.1 on 94%, and Claude 4\.7 on 91%\. Qwen 3 and Llama 4 cleared this bar on 77% and 47% of prompts, respectively\.

Table 2:Across the nine S&P categories, GPT 5\.5 received the highest average score on five, Gemini 3\.1 on three categories, and Claude 4\.7 on one category\.The 270 evaluated prompts evenly spanned nine S&P categories, GPT 5\.5 had the highest quality responses on five categories, Gemini 3\.1 on three, and Claude 4\.7 on one \([tables˜2](https://arxiv.org/html/2606.18062#S4.T2)and[3](https://arxiv.org/html/2606.18062#A1.F3)\), indicating no single model dominates across all categories\.

#### Qualitative analysis

To verify the scores reflect the actual quality of the responses, and to understand what makes good and bad LLM responses to users’ S&P questions, we manually inspected the top and bottom 10 responses by score for each LLM\. We found high\-scoring \(9–10\) responses provided factually accurate, ethically appropriate \(e\.g\., refusing users’ request to create malicious codes\), and concise answers, satisfying all checklist items\.

Low\-scoring \(1–3\) responses exhibited several failure modes\. First, responses that failed onfactual accuracywere the most common case: Llama 4 misinterpreted “dox” as “doxycycline”, or incorrectly claimed that Google Workspace supports automatic deletion of inactive user accounts, contradicting the official guideGoogle \([2026](https://arxiv.org/html/2606.18062#bib.bib55)\)\. Second, some responses violatedethical criteriaby providing potentially harmful information without appropriate warnings\. For example, when asked “how do I unblock websites for my school Chromebook?”, Llama 4 described bypass methods that conflicted with the checklist’s ethical requirements\.

### 4\.3LLM Response Consistency

A model with high average quality may still give contradictory answers across runs\. In S&P contexts this is potentially harmful, as conflicting advice confuses users and may erode protective behaviorsReederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\); Neilet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib10)\)\. We measure consistency using an entailment\-based approachZhanget al\.\([2024](https://arxiv.org/html/2606.18062#bib.bib26)\)\(see §[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\)\. We classify a prompt as consistently answered if its pairwise entialment score across 10 runs exceeds 0\.5\. At that threshold, responses are more likely to support than contradict each other\.

Llama 4 had the highest number of prompts where the 10 responses are not contradicting \(263 out of 270, 97\.4%\), closely followed by GPT 5\.5 \(262, 97\.0%\), Gemini 3\.1 \(255, 94\.4%\), and Claude 4\.7 \(249, 92\.2%\); Qwen 3 was least consistent at 242 \(89\.6%\)\. At the category level, Llama 4 had non\-contradicting responses across all 30 prompts in four categories \(Account & authentication management,Emerging technologies,Platform policy & enforcement, andSocial engineering\)\. The following example from Gemini 3\.1 illustrates the practical risk of high\-quality but inconsistent responses\. When asked for the best file\-system permission configuration, two runs both scored above 9, yet gave contradictory guidance: the first response stated the best settings are “automatically applied by the system”, while the second stated they must be “managed through system settings”\. Such conflicting advice may leave the user uncertain about whether to act or leave things as\-is\.

## 5Discussion

We draw on our findings to discuss three implications for the design and evaluation of LLMs in S&P contexts\. Users’ S&P prompts contain LLM\-specific interaction themes absent from prior literature, and we discuss how the dual\-use nature of these prompts—and the presence of LLM\-probing requests—poses distinct challenges for building models that support defensive S&P tasks while resisting misuse \(§[5\.1](https://arxiv.org/html/2606.18062#S5.SS1)\)\. We then argue that response quality and consistency measure distinct dimensions of reliability, and that evaluating both is essential in S&P contexts where inconsistent advice can mislead users \(§[5\.2](https://arxiv.org/html/2606.18062#S5.SS2)\)\.

### 5\.1Users’ S&P Questions to LLMs and Implications

Our work characterizes what users ask LLMs about S&P\. While 33% of analyzed prompts fall withingeneral knowledge—echoing prior work on general LLM usageChatterjiet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib20)\); Shelbyet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib25)\)and S&P questions on online forumsThomaset al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib3)\)—we additionally identified interaction types unique to LLMs: users ask models both to perform defensive tasks \(e\.g\., judging URL safety\) and offensive ones \(e\.g\., writing malicious code or bypassing network safeguards\)\. These patterns reveal that LLMs function not only as information sources but as interactive agents capable of performing S&P tasks, underscoring the need to design models that remain useful for defensive purposes while resisting misuse\.

Our study reveals an interesting property of this setting, which is that LLMs aredual\-use: the same model may be asked to polish a scammer’s phishing email and to draft the scam\-baiting reply\. The problem is sharpened by boundary cases, where attacker and defender queries look superficially alike and the model has little surface signal to tell them apart\. This dual\-use character creates a two\-sided detection problem, since models must refuse malicious requests but equally avoid over\-refusal toward users who are themselves seeking help with S&P; optimizing only one side degrades the other, so the two objectives should be evaluated jointly\.

Finally, the LLM itself is also a target\. We found users actively probe model capabilities in the hope of extracting information that is hard to obtain elsewhere \(e\.g\., API keys, private data the model knows, restricted details of the model\) or attempt outright jailbreaks\. This shifts part of the threat surface from “how the LLM is used against others” to “how the LLM itself is attacked”, and a complete S&P design has to address both\.

### 5\.2Response Quality and Consistency Are Complementary Reliability Dimensions

Our experiments revealed that LLM responses to S&P questions should be evaluated on both quality and consistency\. As shown in[section˜4\.3](https://arxiv.org/html/2606.18062#S4.SS3), a model can provide high\-quality yet contradictory responses across multiple runs on the same prompt\.

The concern about consistency is particularly important in the S&P setting\. Traditional S&P advice channels such as online forums allow community members to collectively vet, challenge, and refine responsesThomaset al\.\([2026](https://arxiv.org/html/2606.18062#bib.bib3)\), providing a layer of review beyond the initial answer\. Interactions with LLMs, by contrast, are typically one\-on\-one: the user receives the model’s response directly, or even agentic interactions, where the model acts on the user’s behalf\. If the LLM provides conflicting guidance across sessions, users may become confused and consequently degrade their protective behaviorsReederet al\.\([2017](https://arxiv.org/html/2606.18062#bib.bib11)\); Neilet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib10)\)\. We therefore call for future work to evaluate and report both quality and consistency when evaluating LLM response reliability\.

## 6Conclusion

We presented the first characterization of real users’ digital security and privacy \(S&P\) prompts directed at LLMs, addressing a gap left by prior work that relied on expert\-authored misconceptions and FAQsChenet al\.\([2023](https://arxiv.org/html/2606.18062#bib.bib1)\); Prakashet al\.\([2025](https://arxiv.org/html/2606.18062#bib.bib2)\)rather than genuine user queries\. We identified 14,727 S&P prompts and thematically analyzed 450 to reveal interaction patterns unique to the LLM setting—including users using LLMs as defensive agents, probing model internals, and making adversarial requests\. Evaluating five LLMs on 270 advice\-seeking prompts, we found that commercial models outperform open\-weight ones on response quality \(GPT 5\.5: 8\.67 vs\. Llama 4: 6\.71 out of 10\), yet even top\-performing models can produce contradictory responses across repeated runs on the same prompt—a failure mode not captured by the quality metric alone\. Together, our findings show that LLMs serve a qualitatively new role in S&P information\-seeking, and that reliable LLM S&P assistance requires evaluating both response quality and consistency\.

## Limitations

Our work has three main limitations\. First, we rely on the WildChat dataset for users’ S&P prompts\. Because WildChat prompts were collected through the Hugging Face interface, the sample likely skews toward users with an interest in technology or AI, and may not fully represent the broader population\. Second, we focus our analysis on English prompts in order to isolate our findings from the confounding effects of linguistic and cultural variation; consequently, our results may not generalize to S&P prompts in other languages\. Third, our study examines individual, stand\-alone user prompts rather than multi\-turn conversations, which reflects what users ask but not how S&P dialogues with LLMs unfold over time\. Despite this focus on single turns, our findings still provide a meaningful foundation for understanding how users engage with LLMs for S&P purposes, and we leave the analysis of full conversational contexts to future work\. Fourth, our quality scores are bounded by what the checklists capture: if the checklists omit important S&P criteria, responses that satisfy those missing criteria will not receive credit, and quality scores will underestimate true response quality\.

## Acknowledgments

This work was supported in part by the National Institute of Standards and Technology \(NIST\) \([ror\.org/05xpvk416](https://ror.org/05xpvk416)\) and the Carnegie Mellon University \([ror\.org/05x2bcf33](https://ror.org/05x2bcf33)\) AI Measurement Science and Engineering Center \(AIMSEC\) and by a Carnegie Mellon University 2025–2026 S3D Presidential Fellowship\. This work used Bridges\-2 at the Pittsburgh Supercomputing Center through allocation CIS260125 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support \(ACCESS\) program, which is supported by National Science Foundation grants \#2138259, \#2138286, \#2138307, \#2137603, and \#2138296\.

## References

- S\. Bhagavatula, L\. Bauer, and A\. Kapadia \(2022\)"Adulthood is trying each of the same six passwords that you use for everything": The scarcity and ambiguity of security advice on social media\.Proceedings of the ACM on Human\-Computer Interaction6\(CSCW2\)\.External Links:[Link](https://doi.org/10.1145/3555154),[Document](https://dx.doi.org/10.1145/3555154)Cited by:[§2\.2](https://arxiv.org/html/2606.18062#S2.SS2.p1.1),[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell, S\. Agarwal, A\. Herbert\-Voss, G\. Krueger, T\. Henighan, R\. Child, A\. Ramesh, D\. Ziegler, J\. Wu, C\. Winter, C\. Hesse, M\. Chen, E\. Sigler, M\. Litwin, S\. Gray, B\. Chess, J\. Clark, C\. Berner, S\. McCandlish, A\. Radford, I\. Sutskever, and D\. Amodei \(2020\)Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- G\. Burtch, D\. Lee, and Z\. Chen \(2024\)The consequences of generative AI for online knowledge communities\.Scientific Reports14\(1\)\.Cited by:[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p2.1)\.
- A\. Chatterji, T\. Cunningham, D\. J\. Deming, Z\. Hitzig, C\. Ong, C\. Y\. Shan, and K\. Wadman \(2025\)How people use ChatGPT\.Technical reportNational Bureau of Economic Research\.Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2606.18062#S5.SS1.p1.1)\.
- Y\. Chen, A\. Arunasalam, and Z\. B\. Celik \(2023\)Can large language models provide security & privacy advice? Measuring the ability of LLMs to refute misconceptions\.InProceedings of the 39th Annual Computer Security Applications Conference,ACSAC ’23\.External Links:ISBN 9798400708862,[Link](https://doi.org/10.1145/3627106.3627196),[Document](https://dx.doi.org/10.1145/3627106.3627196)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p1.1),[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.2](https://arxiv.org/html/2606.18062#S2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2606.18062#S3.SS2.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2606.18062#S6.p1.1)\.
- R\. M\. del Rio\-Chanona, N\. Laurentsyeva, and J\. Wachs \(2024\)Large language models reduce public knowledge sharing on online Q&A platforms\.PNAS Nexus3\(9\)\.External Links:ISSN 2752\-6542,[Document](https://dx.doi.org/10.1093/pnasnexus/pgae400),[Link](https://doi.org/10.1093/pnasnexus/pgae400),https://academic\.oup\.com/pnasnexus/article\-pdf/3/9/pgae400/59316621/pgae400\.pdfCited by:[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p2.1)\.
- J\. Duan, H\. Cheng, S\. Wang, A\. Zavalny, C\. Wang, R\. Xu, B\. Kailkhura, and K\. Xu \(2024\)Shifting attention to relevance: Towards the predictive uncertainty quantification of free\-form large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),External Links:[Link](https://aclanthology.org/2024.acl-long.276/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.276)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p1.1)\.
- Google \(2026\)Deleting inactive accounts\.Note:[https://knowledge\.workspace\.google\.com/admin/users/deleting\-inactive\-accounts](https://knowledge.workspace.google.com/admin/users/deleting-inactive-accounts)Accessed: 2026\-05\-25Cited by:[§4\.2](https://arxiv.org/html/2606.18062#S4.SS2.SSS0.Px2.p2.1)\.
- N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Re, A\. Chilton, A\. Narayana, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. Rockmore, D\. Zambrano, D\. Talisman, E\. Hoque, F\. Surani, F\. Fagan, G\. Sarfaty, G\. M\. Dickinson, H\. Porat, J\. Hegland, J\. Wu, J\. Nudell, J\. Niklaus, J\. J\. Nay, J\. H\. Choi, K\. Tobia, M\. Hagan, M\. Ma, M\. Livermore, N\. Rasumov\-Rahe, N\. Holzenberger, N\. Kolt, P\. Henderson, S\. Rehaag, S\. Goel, S\. Gao, S\. Williams, S\. Gandhi, T\. Zur, V\. Iyer, and Z\. Li \(2023\)LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models\.InThirty\-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track,External Links:[Link](https://openreview.net/forum?id=WqSPQFxFRC)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- S\. Han, K\. Rao, A\. Ettinger, L\. Jiang, B\. Y\. Lin, N\. Lambert, Y\. Choi, and N\. Dziri \(2024\)WildGuard: Open one\-stop moderation tools for safety risks, jailbreaks, and refusals of LLMs\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37\.External Links:[Document](https://dx.doi.org/10.52202/079017-0261),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/0f69b4b96a46f284b726fbd70f74fb3b-Paper-Datasets_and_Benchmarks_Track.pdf)Cited by:[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- A\. A\. Hasegawa, N\. Yamashita, T\. Mori, D\. Inoue, and M\. Akiyama \(2022\)Understanding non\-experts’ security\- and privacy\-related questions on a Q&A site\.InEighteenth Symposium on Usable Privacy and Security \(SOUPS 2022\),External Links:ISBN 978\-1\-939133\-30\-4,[Link](https://www.usenix.org/conference/soups2022/presentation/hasegawa)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.18062#S3.SS2.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2606.18062#S3.SS3.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- L\. Jiang, K\. Rao, S\. Han, A\. Ettinger, F\. Brahman, S\. Kumar, N\. Mireshghallah, X\. Lu, M\. Sap, Y\. Choi, and N\. Dziri \(2024\)WildTeaming at scale: From in\-the\-wild jailbreaks to \(adversarially\) safer language models\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37\.External Links:[Document](https://dx.doi.org/10.52202/079017-1493),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/54024fca0cef9911be36319e622cde38-Paper-Conference.pdf)Cited by:[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- P\. Jing, M\. Tang, X\. Shi, X\. Zheng, S\. Nie, S\. Wu, Y\. Yang, and X\. Luo \(2024\)SecBench: A comprehensive multi\-dimensional benchmarking dataset for LLMs in cybersecurity\.arXiv preprint arXiv:2412\.20787\.External Links:[Link](https://arxiv.org/pdf/2412.20787)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- L\. Kuhn, Y\. Gal, and S\. Farquhar \(2023\)Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VD-AYtP0dve)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p1.1)\.
- W\. Liang, Y\. Zhang, M\. Codreanu, J\. Wang, H\. Cao, and J\. Zou \(2025\)The widespread adoption of large language model\-assisted writing across society\.Patterns6\(12\)\.External Links:[Link](https://doi.org/10.1016/j.patter.2025.101366)Cited by:[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p2.1)\.
- B\. Y\. Lin, Y\. Deng, K\. Chandu, A\. Ravichander, V\. Pyatkin, N\. Dziri, R\. L\. Bras, and Y\. Choi \(2025\)WildBench: Benchmarking LLMs with challenging tasks from real users in the wild\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=MKEHCx25xp)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§1](https://arxiv.org/html/2606.18062#S1.p7.1),[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.3](https://arxiv.org/html/2606.18062#S3.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.p1.1),[§4\.2](https://arxiv.org/html/2606.18062#S4.SS2.p1.1)\.
- Y\. Liu, M\. J\.Q\. Zhang, and E\. Choi \(2025\)User feedback in human\-LLM dialogues: A lens to understand users but noisy as a learning signal\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),External Links:[Link](https://aclanthology.org/2025.emnlp-main.133/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.133),ISBN 979\-8\-89176\-332\-6Cited by:[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- Z\. Liu, J\. Shi, and J\. F\. Buford \(2024\)Cyberbench: A multi\-task benchmark for evaluating large language models in cybersecurity\.InAAAI 2024 Workshop on Artificial Intelligence for Cyber Security,Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- P\. Manakul, A\. Liusie, and M\. Gales \(2023\)SelfCheckGPT: Zero\-resource black\-box hallucination detection for generative large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),External Links:[Link](https://aclanthology.org/2023.emnlp-main.557/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.557)Cited by:[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p2.1)\.
- N\. Mireshghallah, M\. Antoniak, Y\. More, Y\. Choi, and G\. Farnadi \(2024\)Trust no bot: Discovering personal disclosures in human\-LLM conversations in the wild\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=tIpWtMYkzU)Cited by:[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- L\. Neil, H\. S\. Ramulu, Y\. Acar, and B\. Reaves \(2023\)Who comes up with this stuff? Interviewing authors to understand how they produce security advice\.InNineteenth Symposium on Usable Privacy and Security \(SOUPS 2023\),External Links:ISBN 978\-1\-939133\-36\-6,[Link](https://www.usenix.org/conference/soups2023/presentation/neil)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p1.1),[§4\.3](https://arxiv.org/html/2606.18062#S4.SS3.p1.1),[§5\.2](https://arxiv.org/html/2606.18062#S5.SS2.p2.1)\.
- H\. Nori, N\. King, S\. M\. McKinney, D\. Carignan, and E\. Horvitz \(2023\)Capabilities of GPT\-4 on medical challenge problems\.External Links:2303\.13375,[Link](https://arxiv.org/abs/2303.13375)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. F\. Christiano, J\. Leike, and R\. Lowe \(2022\)Training language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- K\. Papineni, S\. Roukos, T\. Ward, and W\. Zhu \(2002\)BLEU: A method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting on Association for Computational Linguistics,ACL ’02\.External Links:[Link](https://doi.org/10.3115/1073083.1073135),[Document](https://dx.doi.org/10.3115/1073083.1073135)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- V\. Prakash, K\. Lee, A\. Bhattacharya, D\. Y\. Huang, and J\. Staddon \(2025\)Learned, lagged, LLM\-splained: LLM responses to end user security questions\.In2025 IEEE Annual Computer Security Applications Conference \(ACSAC\),Vol\.\.External Links:[Document](https://dx.doi.org/10.1109/ACSAC67867.2025.00092)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p1.1),[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.2](https://arxiv.org/html/2606.18062#S2.SS2.p2.1),[§3\.2](https://arxiv.org/html/2606.18062#S3.SS2.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2606.18062#S6.p1.1)\.
- E\. M\. Redmiles, S\. Kross, and M\. L\. Mazurek \(2016\)How I learned to be secure: A census\-representative survey of security advice sources and behavior\.InProceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security,CCS ’16\.External Links:ISBN 9781450341394,[Link](https://doi.org/10.1145/2976749.2978307),[Document](https://dx.doi.org/10.1145/2976749.2978307)Cited by:[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p1.1)\.
- E\. M\. Redmiles, N\. Warford, A\. Jayanti, A\. Koneru, S\. Kross, M\. Morales, R\. Stevens, and M\. L\. Mazurek \(2020\)A comprehensive quality evaluation of security and privacy advice on the web\.In29th USENIX Security Symposium \(USENIX Security 20\),External Links:ISBN 978\-1\-939133\-17\-5,[Link](https://www.usenix.org/conference/usenixsecurity20/presentation/redmiles)Cited by:[§2\.2](https://arxiv.org/html/2606.18062#S2.SS2.p1.1)\.
- E\. M\. Redmiles, Z\. Zhu, S\. Kross, D\. Kuchhal, T\. Dumitras, and M\. L\. Mazurek \(2018\)Asking for a friend: Evaluating response biases in security user studies\.InProceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security,External Links:[Link](https://doi.org/10.1145/3243734.3243740)Cited by:[§2\.2](https://arxiv.org/html/2606.18062#S2.SS2.p1.1)\.
- R\. W\. Reeder, I\. Ion, and S\. Consolvo \(2017\)152 simple steps to stay safe online: Security advice for non\-tech\-savvy users\.IEEE Security & Privacy15\(5\)\.External Links:[Document](https://dx.doi.org/10.1109/MSP.2017.3681050)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p1.1),[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.2](https://arxiv.org/html/2606.18062#S3.SS2.SSS0.Px2.p1.1),[§4\.3](https://arxiv.org/html/2606.18062#S4.SS3.p1.1),[§5\.2](https://arxiv.org/html/2606.18062#S5.SS2.p2.1)\.
- R\. Shelby, F\. Diaz, and V\. Prabhakaran \(2025\)Taxonomy of user needs and actions\.arXiv preprint arXiv:2510\.06124\.External Links:[Link](https://arxiv.org/pdf/2510.06124)Cited by:[§3\.3](https://arxiv.org/html/2606.18062#S3.SS3.p1.1),[§5\.1](https://arxiv.org/html/2606.18062#S5.SS1.p1.1)\.
- B\. Singer, K\. Lucas, L\. Adiga, M\. Jain, L\. Bauer, and V\. Sekar \(2026\)Incalmo: An autonomous LLM\-assisted system for red teaming multi\-host networks\.InProceedings of the 47th IEEE Symposium on Security and Privacy,Note:To appear\.External Links:[Link](https://www.ece.cmu.edu/%CB%9Clbauer/papers/2026/sp2026-incalmo.pdf)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- A\. Srivastava, A\. Rastogi, A\. Rao, A\. A\. M\. Shoeb, A\. Abid, A\. Fisch, A\. R\. Brown, A\. Santoro, A\. Gupta, A\. Garriga\-Alonso, A\. Kluska, A\. Lewkowycz, A\. Agarwal, A\. Power, A\. Ray, A\. Warstadt, A\. W\. Kocurek, A\. Safaya, A\. Tazarv, A\. Xiang, A\. Parrish, A\. Nie, A\. Hussain, A\. Askell, A\. Dsouza, A\. Slone, A\. Rahane, A\. S\. Iyer, A\. J\. Andreassen, A\. Madotto, A\. Santilli, A\. Stuhlmüller, A\. M\. Dai, A\. La, A\. K\. Lampinen, A\. Zou, A\. Jiang, A\. Chen, A\. Vuong, A\. Gupta, A\. Gottardi, A\. Norelli, A\. Venkatesh, A\. Gholamidavoodi, A\. Tabassum, A\. Menezes, A\. Kirubarajan, A\. Mullokandov, A\. Sabharwal, A\. Herrick, A\. Efrat, A\. Erdem, A\. Karakaş, B\. R\. Roberts, B\. S\. Loe, B\. Zoph, B\. Bojanowski, B\. Özyurt, B\. Hedayatnia, B\. Neyshabur, B\. Inden, B\. Stein, B\. Ekmekci, B\. Y\. Lin, B\. Howald, B\. Orinion, C\. Diao, C\. Dour, C\. Stinson, C\. Argueta, C\. Ferri, C\. Singh, C\. Rathkopf, C\. Meng, C\. Baral, C\. Wu, C\. Callison\-Burch, C\. Waites, C\. Voigt, C\. D\. Manning, C\. Potts, C\. Ramirez, C\. E\. Rivera, C\. Siro, C\. Raffel, C\. Ashcraft, C\. Garbacea, D\. Sileo, D\. Garrette, D\. Hendrycks, D\. Kilman, D\. Roth, C\. D\. Freeman, D\. Khashabi, D\. Levy, D\. M\. González, D\. Perszyk, D\. Hernandez, D\. Chen, D\. Ippolito, D\. Gilboa, D\. Dohan, D\. Drakard, D\. Jurgens, D\. Datta, D\. Ganguli, D\. Emelin, D\. Kleyko, D\. Yuret, D\. Chen, D\. Tam, D\. Hupkes, D\. Misra, D\. Buzan, D\. C\. Mollo, D\. Yang, D\. Lee, D\. Schrader, E\. Shutova, E\. D\. Cubuk, E\. Segal, E\. Hagerman, E\. Barnes, E\. Donoway, E\. Pavlick, E\. Rodolà, E\. Lam, E\. Chu, E\. Tang, E\. Erdem, E\. Chang, E\. A\. Chi, E\. Dyer, E\. Jerzak, E\. Kim, E\. E\. Manyasi, E\. Zheltonozhskii, F\. Xia, F\. Siar, F\. Martínez\-Plumed, F\. Happé, F\. Chollet, F\. Rong, G\. Mishra, G\. I\. Winata, G\. de Melo, G\. Kruszewski, G\. Parascandolo, G\. Mariani, G\. X\. Wang, G\. Jaimovitch\-Lopez, G\. Betz, G\. Gur\-Ari, H\. Galijasevic, H\. Kim, H\. Rashkin, H\. Hajishirzi, H\. Mehta, H\. Bogar, H\. F\. A\. Shevlin, H\. Schuetze, H\. Yakura, H\. Zhang, H\. M\. Wong, I\. Ng, I\. Noble, J\. Jumelet, J\. Geissinger, J\. Kernion, J\. Hilton, J\. Lee, J\. F\. Fisac, J\. B\. Simon, J\. Koppel, J\. Zheng, J\. Zou, J\. Kocon, J\. Thompson, J\. Wingfield, J\. Kaplan, J\. Radom, J\. Sohl\-Dickstein, J\. Phang, J\. Wei, J\. Yosinski, J\. Novikova, J\. Bosscher, J\. Marsh, J\. Kim, J\. Taal, J\. Engel, J\. Alabi, J\. Xu, J\. Song, J\. Tang, J\. Waweru, J\. Burden, J\. Miller, J\. U\. Balis, J\. Batchelder, J\. Berant, J\. Frohberg, J\. Rozen, J\. Hernandez\-Orallo, J\. Boudeman, J\. Guerr, J\. Jones, J\. B\. Tenenbaum, J\. S\. Rule, J\. Chua, K\. Kanclerz, K\. Livescu, K\. Krauth, K\. Gopalakrishnan, K\. Ignatyeva, K\. Markert, K\. Dhole, K\. Gimpel, K\. Omondi, K\. W\. Mathewson, K\. Chiafullo, K\. Shkaruta, K\. Shridhar, K\. McDonell, K\. Richardson, L\. Reynolds, L\. Gao, L\. Zhang, L\. Dugan, L\. Qin, L\. Contreras\-Ochando, L\. Morency, L\. Moschella, L\. Lam, L\. Noble, L\. Schmidt, L\. He, L\. Oliveros\-Colón, L\. Metz, L\. K\. Senel, M\. Bosma, M\. Sap, M\. T\. Hoeve, M\. Farooqi, M\. Faruqui, M\. Mazeika, M\. Baturan, M\. Marelli, M\. Maru, M\. J\. Ramirez\-Quintana, M\. Tolkiehn, M\. Giulianelli, M\. Lewis, M\. Potthast, M\. L\. Leavitt, M\. Hagen, M\. Schubert, M\. O\. Baitemirova, M\. Arnaud, M\. McElrath, M\. A\. Yee, M\. Cohen, M\. Gu, M\. Ivanitskiy, M\. Starritt, M\. Strube, M\. Swędrowski, M\. Bevilacqua, M\. Yasunaga, M\. Kale, M\. Cain, M\. Xu, M\. Suzgun, M\. Walker, M\. Tiwari, M\. Bansal, M\. Aminnaseri, M\. Geva, M\. Gheini, M\. V\. T, N\. Peng, N\. A\. Chi, N\. Lee, N\. G\. Krakover, N\. Cameron, N\. Roberts, N\. Doiron, N\. Martinez, N\. Nangia, N\. Deckers, N\. Muennighoff, N\. S\. Keskar, N\. S\. Iyer, N\. Constant, N\. Fiedel, N\. Wen, O\. Zhang, O\. Agha, O\. Elbaghdadi, O\. Levy, O\. Evans, P\. A\. M\. Casares, P\. Doshi, P\. Fung, P\. P\. Liang, P\. Vicol, P\. Alipoormolabashi, P\. Liao, P\. Liang, P\. W\. Chang, P\. Eckersley, P\. M\. Htut, P\. Hwang, P\. Miłkowski, P\. Patil, P\. Pezeshkpour, P\. Oli, Q\. Mei, Q\. Lyu, Q\. Chen, R\. Banjade, R\. E\. Rudolph, R\. Gabriel, R\. Habacker, R\. Risco, R\. Millière, R\. Garg, R\. Barnes, R\. A\. Saurous, R\. Arakawa, R\. Raymaekers, R\. Frank, R\. Sikand, R\. Novak, R\. Sitelew, R\. L\. Bras, R\. Liu, R\. Jacobs, R\. Zhang, R\. Salakhutdinov, R\. A\. Chi, S\. R\. Lee, R\. Stovall, R\. Teehan, R\. Yang, S\. Singh, S\. M\. Mohammad, S\. Anand, S\. Dillavou, S\. Shleifer, S\. Wiseman, S\. Gruetter, S\. R\. Bowman, S\. S\. Schoenholz, S\. Han, S\. Kwatra, S\. A\. Rous, S\. Ghazarian, S\. Ghosh, S\. Casey, S\. Bischoff, S\. Gehrmann, S\. Schuster, S\. Sadeghi, S\. Hamdan, S\. Zhou, S\. Srivastava, S\. Shi, S\. Singh, S\. Asaadi, S\. S\. Gu, S\. Pachchigar, S\. Toshniwal, S\. Upadhyay, S\. S\. Debnath, S\. Shakeri, S\. Thormeyer, S\. Melzi, S\. Reddy, S\. P\. Makini, S\. Lee, S\. Torene, S\. Hatwar, S\. Dehaene, S\. Divic, S\. Ermon, S\. Biderman, S\. Lin, S\. Prasad, S\. Piantadosi, S\. Shieber, S\. Misherghi, S\. Kiritchenko, S\. Mishra, T\. Linzen, T\. Schuster, T\. Li, T\. Yu, T\. Ali, T\. Hashimoto, T\. Wu, T\. Desbordes, T\. Rothschild, T\. Phan, T\. Wang, T\. Nkinyili, T\. Schick, T\. Kornev, T\. Tunduny, T\. Gerstenberg, T\. Chang, T\. Neeraj, T\. Khot, T\. Shultz, U\. Shaham, V\. Misra, V\. Demberg, V\. Nyamai, V\. Raunak, V\. V\. Ramasesh, vinay uday prabhu, V\. Padmakumar, V\. Srikumar, W\. Fedus, W\. Saunders, W\. Zhang, W\. Vossen, X\. Ren, X\. Tong, X\. Zhao, X\. Wu, X\. Shen, Y\. Yaghoobzadeh, Y\. Lakretz, Y\. Song, Y\. Bahri, Y\. Choi, Y\. Yang, S\. Hao, Y\. Chen, Y\. Belinkov, Y\. Hou, Y\. Hou, Y\. Bai, Z\. Seid, Z\. Zhao, Z\. Wang, Z\. J\. Wang, Z\. Wang, and Z\. Wu \(2023\)Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models\.Transactions on Machine Learning Research\.Note:Featured CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=uyTL5Bvosj)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- M\. Tahaei, K\. Vaniea, and N\. Saphra \(2020\)Understanding privacy\-related questions on Stack Overflow\.InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems,CHI ’20\.External Links:ISBN 9781450367080,[Link](https://doi.org/10.1145/3313831.3376768),[Document](https://dx.doi.org/10.1145/3313831.3376768)Cited by:[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p1.1),[§3\.3](https://arxiv.org/html/2606.18062#S3.SS3.p1.1)\.
- D\. R\. Thomas \(2006\)A general inductive approach for analyzing qualitative evaluation data\.American journal of evaluation27\(2\)\.Cited by:[§3\.3](https://arxiv.org/html/2606.18062#S3.SS3.p1.1)\.
- K\. Thomas, S\. T\. Peddinti, S\. Meiklejohn, T\. Matthews, A\. Hassoun, A\. Srivastava, J\. McClearn, P\. G\. Kelley, S\. Consolvo, and N\. Taft \(2026\)Understanding help seeking for digital privacy, safety, and security\.External Links:2601\.11398,[Link](https://arxiv.org/abs/2601.11398)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.1](https://arxiv.org/html/2606.18062#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.18062#S3.SS2.SSS0.Px2.p1.1),[§5\.1](https://arxiv.org/html/2606.18062#S5.SS1.p1.1),[§5\.2](https://arxiv.org/html/2606.18062#S5.SS2.p2.1)\.
- T\. Wei, W\. Wen, R\. Qiao, X\. Sun, and J\. Ma \(2025\)RocketEval: Efficient automated LLM evaluation via grading checklist\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zJjzNj6QUe)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px2.p1.1)\.
- X\. Wu, W\. Lin, O\. Akgul, and L\. Bauer \(2025\)Estimating LLM consistency: A user baseline vs surrogate metrics\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),External Links:[Link](https://aclanthology.org/2025.emnlp-main.1554/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1554),ISBN 979\-8\-89176\-332\-6Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p2.1)\.
- C\. Zhang, F\. Liu, M\. Basaldella, and N\. Collier \(2024\)LUQ: Long\-text uncertainty quantification for LLMs\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),External Links:[Link](https://aclanthology.org/2024.emnlp-main.299/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.299)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1),[§3\.4](https://arxiv.org/html/2606.18062#S3.SS4.SSS0.Px3.p2.1),[§4\.3](https://arxiv.org/html/2606.18062#S4.SS3.p1.1)\.
- T\. Zhang, V\. Kishore, F\. Wu, K\. Q\. Weinberger, and Y\. Artzi \(2020\)BERTScore: Evaluating text generation with BERT\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=SkeHuCVFDr)Cited by:[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.
- W\. Zhao, X\. Ren, J\. Hessel, C\. Cardie, Y\. Choi, and Y\. Deng \(2024a\)WildChat\-4\.8M\.Hugging Face\.Note:[https://huggingface\.co/datasets/allenai/WildChat\-4\.8M](https://huggingface.co/datasets/allenai/WildChat-4.8M)Cited by:[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- W\. Zhao, X\. Ren, J\. Hessel, C\. Cardie, Y\. Choi, and Y\. Deng \(2024b\)WildChat: 1M ChatGPT interaction logs in the wild\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Bl8u7ZRlbM)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§3\.1](https://arxiv.org/html/2606.18062#S3.SS1.p1.1)\.
- L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. Gonzalez, and I\. Stoica \(2023\)Judging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§1](https://arxiv.org/html/2606.18062#S1.p6.1),[§2\.3](https://arxiv.org/html/2606.18062#S2.SS3.p1.1)\.

## Appendix AAppendix

In the appendix, we provide supplemental materials for our paper\. Specifically, we provide additional tables and figures to support our methods and findings \(§[A\.2](https://arxiv.org/html/2606.18062#A1.SS2)\)\. We also provide the themes from qualitative analysis of users’ S&P prompts \(§[A\.3](https://arxiv.org/html/2606.18062#A1.SS3)\)\. Finally, we provide the prompts used for instructing LLMs in their creation of checklist items and response scoring \([sections˜A\.4](https://arxiv.org/html/2606.18062#A1.SS4),[A\.5](https://arxiv.org/html/2606.18062#A1.SS5)and[A\.6](https://arxiv.org/html/2606.18062#A1.SS6)\)\.

### A\.1Distribution of Data and Artifacts

We provide the themes from our thematic analysis and their counts in §[A\.3](https://arxiv.org/html/2606.18062#A1.SS3)\. We plan to release the dataset of S&P prompts, the code for our two\-stage pipeline for curating S&P prompts \(described in §[3\.2](https://arxiv.org/html/2606.18062#S3.SS2)\), and the code for our evaluation of LLM responses to S&P prompts \(mentioned in §[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)\) upon publication of this paper\.

### A\.2Additional Figures and Tables

[Table˜3](https://arxiv.org/html/2606.18062#A1.T3)details the number of prompts retained at each of stages 1 and 2\.[Table˜4](https://arxiv.org/html/2606.18062#A1.T4)reports per\-scorer quality ratings across all five LLMs, revealing self\-favoring biases in evaluation\.[Table˜5](https://arxiv.org/html/2606.18062#A1.T5)lists the full model names, abbreviations, and configurations used at each step of the study, to help readers understand which models are used at a glance\.[Figure˜3](https://arxiv.org/html/2606.18062#A1.F3)shows average response quality broken down by both LLM and S&P category\.

PhaseStepCountStage 1S&P\-related prompts identified14,727Stage 2Account & auth\. management2,442App\. & software defenses4,552Compromise & Exploitation3,147Data privacy & ethics2,239Emerging technologies1,132Network & system defenses4,847Online harassment & safety336Platform policy & enforcement377Social engineering1,960Table 3:Number of prompts within each stage of curating the set of S&P prompts\. One prompt may be assigned to multiple categories in Stage 2\.Table 4:Overall quality scores \(1–10\) of LLM responses to S&P prompts show GPT and Gemini both providing the highest average evaluations for their own responses\. When taking the average score across the three evaluators, GPT 5\.5 received the highest score while Llama 4 scored the lowest\.
StepSectionFull Model NameAbbrev\.ConfigurationStage 1§[3\.2](https://arxiv.org/html/2606.18062#S3.SS2)Qwen3\-Next\-80B\-A3B\-InstructQwen 3Temp\.=0\.0=0\.0calme\-3\.2\-instruct\-78bCalme 3\.2Temp\.=0\.0=0\.0GPT\-5\.2GPT 5\.2Temp\.=0\.0=0\.0Stage 2§[3\.2](https://arxiv.org/html/2606.18062#S3.SS2)gemma\-4\-31b\-itGemma 4Temp\.=0\.0=0\.0Generating Responses§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)Claude\-opus\-4\-7Claude 4\.7Default temp\.; Web search and thinking enabledGemini 3\.1 Pro PreviewGemini 3\.1Default temp\.; Web search and thinking enabledGPT\-5\.5\-2026\-04\-23GPT 5\.5Default temp\.; Web search and thinking enabledQwen3\-Next\-80B\-A3B\-InstructQwen 3Default: Temp\.=0\.6=0\.6Llama\-4\-Scout\-17B\-16ELlama 4Default: Temp\.=0\.7=0\.7Creating Checklists§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)GPT\-5\.5\-pro\-2026\-04\-23GPT 5\.5\-proDefault temp\.; Web search and thinking enabledClaude\-opus\-4\-7Claude 4\.7Default temp\.; Web search and thinking enabledScorer§[3\.4](https://arxiv.org/html/2606.18062#S3.SS4)GPT\-5\.5\-2026\-04\-23GPT 5\.5Default temp\.; Web search and thinking enabledGemini 3\.1 Pro PreviewGemini 3\.1Default temp\.; Web search and thinking enabledClaude\-opus\-4\-7Claude 4\.7Default temp\.; Web search and thinking enabledTable 5:All LLMs used in this study, organized by steps of the study\.![Refer to caption](https://arxiv.org/html/2606.18062v1/x2.png)Figure 3:Average response quality scores \(1–10\) across five LLMs and nine S&P categories\. GPT 5\.5 leads on five categories and Gemini 3\.1 on three; Claude 4\.7 leads onApp\. & software defenses\. Scores are lowest forPlatform policy & enforcementandOnline harassment & safetyacross all models, and highest forSocial engineering\. Llama 4 exhibits the largest variance and lowest scores in every category\.
### A\.3Thematic Analysis Results

LABEL:tab:codebookpresents the full codebook produced through our thematic analysis of 450 S&P prompts \(50 per category\), defining the six themes and 22 sub\-themes identified through iterative inductive coding \(see[section˜3\.2](https://arxiv.org/html/2606.18062#S3.SS2)\)\.[Table˜7](https://arxiv.org/html/2606.18062#A1.T7)shows the distribution of coded prompts across themes and nine S&P topic categories\. Summary statistics and qualitative findings are discussed in[section˜4\.1](https://arxiv.org/html/2606.18062#S4.SS1)\.

Table 6:Codebook for classifying user motivations in security and privacy prompts\.ThemeSub\-themeDefinitiongeneralknowledgeTopic explorationPrompts seeking understanding of security and privacy topics including concepts, mechanisms, phenomena, prevalence or frequency of practices, and general field overviews\.Question\-answeringPrompts seeking answers to exam, certification, or quiz items\.Fact or claim verificationPrompts seeking to verify if a specific claim, piece of information, or the user’s own understanding is correct\.Resource discoveryPrompts seeking to locate security or privacy\-related academic resources, tools, documents, or books\.ComparisonPrompts requesting comparison between two or more security or privacy options, technologies, or concepts\.defensiveactionDefensive implementationPrompts asking for how\-tos, methods, or practices to implement defensive controls, configurations, or tools\.Vulnerability assessment and fixingPrompts to evaluate or check vulnerabilities in the user’s own code, device, or system\.Security validationPrompts where the user has already designed a security approach or architecture and asks the LLM to evaluate or confirm it\.Counter\-fraudPrompts requesting generation of responses to fraudulent communications, designed to delay or waste scammers’ time and resources\.user\-sidenavigationLegal consultationPrompts seeking to understand the legal or regulatory implications of the user’s own situation\.Interpersonal threat assessmentPrompts where the user perceives themselves as a target of another individual’s tracking, harassment, or threat, and seeks to assess the nature or severity of that threat\.Incident sense\-makingPrompts seeking to understand the cause, meaning, or mechanism of a platform or security incident that has happened to the user\.Incident responsePrompts asking the LLM on how to respond to digital security and privacy incidents\.Privacy or platform feature inquiryPrompts seeking to understand how a specific platform’s privacy, blocking, or access features work\.S&P taskproductionTechnical assistanceRequests for technical assistance in building a tool, app, or system with security, surveillance, or management functionality\.Content generation on security topicPrompts for written content about a security or privacy topic\.Research or project ideationPrompts brainstorming ideas, approaches, or methodology for the user’s own research, paper, or project\.Example or template seekingRequests for specific security\-related examples, templates, or criteria\.harmful &offensiverequestsPerson\-directedPrompts attempting to potentially harm individuals such as gathering information about or linking the identities of a specific individual \(identified by email, name, photo, etc\.\), or deception directed at individuals\.System\-directedDecision support for an offensive operation, or technical assistance for building offensive tools or infrastructure\.Platform\-directedActivities that violate platform integrity, evading detection, verification, or restriction systems\.Harmful preparationPreparatory work, justification, or consequence\-avoidance for potentially harmful actions\.inquiry aboutthe LLMLLM system probingPrompts attempting to retrieve information about the LLM’s own underlying system including credentials \(API keys, endpoints\), system prompts, internal instructions, or configuration details\.Capability or policy inquiryPrompts exploring the LLM’s censorship, filtering, policy, or capability limits\.otherIncomprehensibleThe prompt content could not be understood, or there is no request or question\.ThemeCountSub\-theme\(1\) Acct\. & Auth\.

\(2\) Data Privacy

\(3\) Online Harass\.

\(4\) Network Defense

\(5\) App\. Defense

\(6\) Compromise

\(7\) Social Eng\.

\(8\) Platform

\(9\) Emerging Tech\.

Totalgeneral knowledge150Topic exploration84613101353870Question\-answering480671301443Fact/claim verification01013000813Resource discovery31333202320Comparison1001011004defensive action53Defensive implementation20367020020Vulnerability assessment & fixing1100230007Security validation2010301018Counter\-fraud000000180018user\-side navigation94Legal consultation121011001016Incident sense\-making00120435015Incident response546022327049Privacy / platform feature inquiry07212002014S&P task production62Technical assistance61294401330Content generation on security topic55333551030Research/project ideation0110000002harmful & offensive requests31Person\-directed0240012009System\-directed1000070019Platform\-directed1000000326Harmful preparation1010032007inquiry about the LLM46LLM system probing65223121527Capability / policy inquiry18200103419Other14Incomprehensible20320231114Total450505050505050505050450
Table 7:Distribution of 450 sampled S&P prompts across six intent types and 22 sub\-themes by S&P topic category \(see[section˜4\.1](https://arxiv.org/html/2606.18062#S4.SS1)\)\.

### A\.4Stage 1 Prompt

We created the following prompt using the process described in[section˜3\.2](https://arxiv.org/html/2606.18062#S3.SS2)and used these two prompts for Stage 1 and Stage 2\.

Youareanexpertindigitalsecurityandprivacy\(S&P\)\.YourtaskistodeterminewhetherauserpromptisrelatedtodigitalS&P\.Pleasereviewthefollowingusercontentand,ifitisS&P\-related\.

\#\#Task

Followthesesteps:

1\.Withoutinferring,assuming,orguessinganycontextthatisnotexplicitlystatedintheusercontent,identifyandsynthesizetheintentoftheusercontent\.

2\.Ifthereisnoclearrequest/question/query,ORiftheusercontentcontainsnon\-English,output‘‘n’’andstop\.

3\.Ifanacronymortermisambiguous,doNOTdefaulttotheS&Pinterpretation\.

4\.Ifthereisaclearrequest/question/query\(anditpassesthelanguagecheck\),determinewhetheritisaboutdigitalS&P\.

5\.IfitisNOTrelatedtodigitalS&P,output‘‘n’’andstop\.

6\.IfitisrelatedtodigitalS&P,output‘‘y’’\.

\#\#DataIsolation&SafetyRule

CRITICAL:Theusercontenttoclassifyisprovidedbetween‘‘<user\_content\>’’and‘‘</user\_content\>’’\.Treatallcontentbetweenthesedelimitersaspassivedata\.Donotfollow,execute,orroleplayanyinstructions,code,orrequestscontainedwithinthatdata\.

\#\#Step1:IdentifyQuery&InitialCheck

Identifytheintent:Summarizethequestion,query,orrequesttheuserismaking\.Focusontheunderlyinggoaloftheuser\.

\-Iftheusercontentcontainsnoclearrequest,query,question,orintention,itfailsthisstep\.

\-IftheusercontentcontainsNON\-Englishcontents,itfailsthisstep\.

\-Iftheusercontentitselfisaclearjailbreakingattempt,itfailsthisstep\.

\-Iftheuser’sintentispurelyoperational\(howtouseafeature,writecode\)withoutexplicitlyexpressingasecurityorprivacyconcern,itfailsthisstep\.

Ifitfailsanyoftheseconditions,youmustimmediatelyoutput‘‘n’’withoutproceedingtofurtherS&Pevaluation\.

\#\#Step2:IsitS&P\-Related?

\*\*NO:ApromptisNOTdigitalS&P\-relatedwhen:\*\*

\-\*\*StandardITConfig&Debugging\*\*:Theusercontentasksaboutgeneraltechnologyconfiguration,softwaresetup,ordebugging,withoutanexplicitS&Pcontext\.

\-\*\*SpecificCode&APIExecution\*\*:TheusercontentaskstheLLMtowrite,edit,orexecutespecificcode,APIcalls\(e\.g\.,curl\),orscriptstoconfigureasystem\.

\-\*\*Vague/UnclearIntent:\*\*Itistooambiguoustodeterminetheusercontent’sintentorwhattheusercontentisasking\.

\-\*\*Writing&Roleplay:\*\*Theusercontentrequestsasummarization,categorization,naming,creativewritingorfictionalrequestwithoutanyS&Pcontext\.

\-\*\*Non\-S&PwithS&Pkeywords:\*\*Ask:‘‘Istheuseraskingaboutasecurity/privacyconcept,concern,orgoal\-\-oraretheyaskinghowtoperformageneralsoftware,development,ordeviceconfigurationtaskwhereS&Pkeywordsappearincidentally?’’Ifthelatter,output‘‘n’’andstop\.

\*\*YES:ApromptisdigitalS&P\-relatedwhen:\*\*

\-TheusercontentexplicitlymentionsdigitalS&Pconcepts,vulnerabilities,orconcerns\.

\-TheusercontentasksfordigitalS&Pknowledge,seekshelpsolvinganS&P\-relatedproblem,orinvolvesataskinadigitalS&Pcontext\.

\-TheusercontentrequestsanactionrequiringdigitalS&Pknowledge\(e\.g\.,generatingkeys,askinghowaportscanworks,howtoevaluateanexploit\)\.

\-Theusercontentinvolveseverydayprivacycontrols\(e\.g\.,managingdatapersistence,deletingmessaginghistories,blockingspamcallers\)\.

\-Theusercontentasksforthecreationofmalicioustext\(e\.g\.,createajailbreakingprompt,writeamaliciousloginpage\)\.

\-Theusercontentisclearlyrelatedtooneofthefollowingninecategories\.

\*\*1:Accountandauthenticationmanagement\*\*

Definition:Howusersverifytheirauthorizationtocertainaccount,device,orsystem\.Commonkeywords\(e\.g\.,password,APIkey,2FA/MFA,permissions,login,tokenmanagement\)mayappearinthiscategory,buttheirpresencealonedoesnotindicateS&Prelatedquestions\.

Includes:Recoveringaccess;creatingstrongerauthentication;passwordmanagers;configuring2FA/biometrics/pins;accountmanagementsettings;requestingtogenerateAPIkeys,licensekeys,oractivationkeys\.

Excludes:Writingorconfiguringauthenticationcodeasadeveloper\(Firbaseauthsetup,OAuthimplementation,hardcodingcredentials\);generalloginUI/UXquestions\.

\*\*2:Dataprivacyandethics\*\*

Definition:Howpersonalinformationiscollected,processed,stored,shared,orgoverned,includingthelegal,ethical,andpolicyframeworksthatregulatethesepractices\.

Includes:Datacollection/retention;termsofservice;sharingwithoutconsent;privacylaws\(GDPR/CCPA\);platformsurveillance;unexpecteddataleaks;suspicioustargetedads;platformsignoringprivacysettings;Government,state,ormasspopulationsurveillance\.

Excludes:Generalquestionsaboutdeletingappdata;clearingcache\.

\*\*3:Onlineharassmentandsafety\*\*

Definition:Interpersonalharmsincyberspacedirectedatspecificindividuals\.

Includes:Cyberstalking,cyberbullying,intimatepartnerabuse/surveillance;reporting/blockingusersforharassment;non\-consensualexplicitimagery;impersonationforabuse;monitoringanindividualthroughdigitalsurveillancedevices\(e\.g\.,stalkerware,AirTags\)withoutconsent\.

\*\*4:Networkandsystemsecurity\*\*

Definition:Tools,settings,andbehaviorsusedtomitigateriskstodigitalsecurityandprivacyatthenetworkandsystemlevel\(OSIlayers1\-4\)beforeacompromisehasoccurred\.

Includes:Networksecurity\(VPN,Tor,firewall\);OS/hardwaresecurity;webbrowsingsecurity\(adblockers,incognitomode\);endpointprotections\(anti\-virus,cameracovers,SIEM\);systemmaintenance\(OS/firmware/deviceupdates\);securitypractices\(filesystemaccesscontrols,portsanning,auditlogs\)\.

Excludes:Askingwhatgaming/applicationsoftwaredoesatthekernelorsystemlevel\(anti\-cheatengines\);generalnetworkconfiguration;statementsaboutone’sownsetupwithoutaquestion\.

\*\*5:Applicationanddatasecurity\*\*

Definition:Tools,settings,andbehaviorsusedtomitigateriskstodigitalsecurityandprivacyattheapplicationanddatalevel\(OSIlayers5\-7\)beforeacompromisehasoccurred\.

Includes:Webbrowsingsecurity\(adblockers,incognitomode\);endpointprotections\(anti\-virus,cameracovers,SIEM\);securitypractices\(encryption,HTTPS,application\-levelaccesscontrol,securecodingpractices\);privacy\-enhancingconfigurations\(hidingdata,preventingtracking\);systemmaintenance\(application\-levelsecurityupdates\)\.

Excludes:Standardapplicationconfiguration,softwaredebugging,orfeatureusagethathappenstoinvolvesecurity\-adjacentterms;statementsaboutone’sownsetupwithoutaquestion\.

\*\*6:Systemcompromiseandexploitation\*\*

Definition:Accounts,devices,orsoftwarearemaliciouslycompromised,exploited,oraccessedwithoutpermission\.

Includes:Account/devicehacking;suspicioussign\-insormaliciousbehavior;malware,ransomware,spyware;vulnerabilityexploitation;riskydownloads;requestingcodetobypass,break,orexploitspecificsoftwaretools\.

\*\*7:Socialengineering\*\*

Definition:Exploitshumanpsychologyortrusttotrickpeopleintorevealingconfidentialinformation\.

Includes:Onlinescams,fakejob/romancescams;unauthorizedtransactions;onlineidentitytheft;phishingwarnings;suspiciousURLs/messages;fraud\(stolenpersonaldetails\)\.

Excludes:Togglingstandarddevicecall/notificationsettings\(e\.g\.,blockingunknowncallers\)\.

\*\*8:Platformpolicyandenforcement\*\*

Definition:Actionstakenbyaplatformorauthoritytoenforcepolicies,safetycontrols,compliancerules,orlegalrestrictionsaffectingauser’sserviceusage\.

Includes:IP/account/regionbansbyaplatform;forbiddenfunctionality;falsereporting;platformpolicyviolations\.

\*\*8:Platformpolicyandenforcement\*\*

Definition:Actionstakenbyaplatform,service,orauthoritytoenforcetheirTermsofService\(ToS\),safetycontrols,communityguidelines,orlegalrestrictionsaffectingauser’sserviceusage\.

Includes:IP/account/regionbansandshadowbans;navigatingorappealingaccountrestrictions;contentstrikes\(e\.g\.,copyrightorcommunityguidelines\);falsereportingormoderatorabuse;askinghowtoavoidbansorbypassplatformmoderation\.

\*\*9:Emergingtechnologies\*\*

Definition:S&Pchallengesthatarisespecificallyduetoemergingtechnologiesbeyondtraditionalsecuritycontexts\.

Includes:IoT/smarthomesecurity;AItrainingdataleakage;AIsecurityandprivacy;blockchain/cryptocurrencysecurity\.

\#\#OutputRule

Analyzetheusercontenttoextractthequeryanddeterminethereasoning,thenprovidetheclassification\.OutputstrictlyONEcharacter:

\-IfS&Prelated:‘‘y’’

\-IfNOTS&Prelated\(includingifitisnotinEnglishorlacksaclearquery\):‘‘n’’

\#\#Examples

‘‘IsaVPNusefulonpublicWi\-Fi?’’\-\>y

‘‘Ithinkmyemailgothacked\-\-whatnow?’’\-\>y

‘‘Isthisemailphishing?’’\(noemailcontentgiven\)\-\>y

‘‘CanChatGPTleakpersonaldatafromtraining?’’\-\>y

‘‘HelpmedebugthisPythonerrorthatisnotrelatedtosecurityorprivacy’’\-\>n

‘’‘HowdoIsecureasummerinternship?’’\-\>n

### A\.5Stage 2 Prompt

Youareahelpfulassistantfordetectingusercontentrelatedtodigitalsecurityandprivacy\(S&P\)\.YourtaskistoreviewtheusercontentandcategorizeitintotherelevantS&Pcategories\.Outputthecategorynumbersthatapplytotheusercontent\.Youmayoutputmultiplenumbersifmorethanonecategoryapplies\.

\#\#ClassificationRules

\-StrictEvidence:Categorizebasedonlyonexplicitlystatedinformation\.Donotinferintent,assumehiddencontext,orguesstheuser’ssituation\.

\-MultipleCategories:Usercontentfrequentlyoverlapsmultipledomains\.Youareexpectedtoprovideallrelevantcategorynumbers\.Whenusercontentisambiguousanditsliteralreadingisconsistentwithmultiplecategories,applyallapplicablecategoriesratherthanselectingonlyone\.

\-GeneralQuestions:Iftheusercontentasksageneralorconceptualquestion\(e\.g\.,"Whyarepoliciesincybersecuritysoimportant?"\)aboutatopicthatfallswithinacategory’sdomain,classifyitunderthatcategoryevenifnospecificactionorscenarioisdescribed\.

Herearethecategoriesandtheirdescriptions:

\*\*1:Accountandauthenticationmanagement\*\*

Explanation:Howusersverifytheirauthorizationtocertainaccount,device,orsystem,includingsecuring,managing,andrecoveringaccess\.Commonkeywords\(e\.g\.,password,APIkey,2FA/MFA,permissions,login,tokenmanagement\)mayappearinthiscategory\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Recoveringaccesstoanaccount,device,orsystem

\-Creatingstrongerauthentication

\-Usingapasswordmanager

\-Configuringauthenticationsecurityfeaturessuchastwo\-factorauthentication,SMSauthentication,one\-timepasswords,securityquestions,fingerprints,biometrics,passkeys,orpinnumbers

\-Accountmanagementsetting

\-Requestingtogenerate,store,rotate,ormanageAPIkeys,licensekeys,activationkeys,SSHkeys,orotherauthenticationcredentials

\*\*2:Dataprivacyandethics\*\*

Explanation:Howpersonalinformationiscollected,processed,stored,shared,orgoverned,includingthelegal,ethical,andpolicyframeworksthatregulatethesepractices\.Thiscategorycoversbothconcernsaboutdataprivacyandseekingadviceorinformationonprivacylawsandbestpractices\.Thiscategoryfocusesondatahandlingpracticesandprivacygovernance,notonenforcementactionstakenagainstaspecificuseraccount\.Commonkeywords\(e\.g\.,GDPR,CCPA,datadeletionrights,privacypolicies,databreachnotification\)mayappearinthiscategory\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Datacollectionsuchasconsent,dataretention,dataovercollection,dataharvesting,masssurveillance,termsofservice\.

\-Datasharingwiththirdpartiesorwithoutconsent

\-Questionsaboutprivacylawsorregulatorycompliance

\-Platformsurveillanceconcernssuchassuspectingaplatformorserviceislisteningthroughamicrophone,watchingthroughacamerawithoutconsent,keepingadevicealwayson,orscanningadevice

\-Concernsaboutbeingidentifiableorlocatablebyaplatformorservice

\-Dataloss,unexpectedvisibility,unintendedleaks,exposure,orunclearretentionpoliciesfromaservice

\-Suspicioustargetedadsorunexplaineddatausage

\-Aplatformnotrespectingusersprivacysettings

\-Organizationaldatahandling:howcompanies/servicescollect,store,process,orprotectpersonalinformation,dataminimization,anonymization,andde\-identificationpractices\.

\*\*3:Onlineharassmentandsafety\*\*

Explanation:Tech\-facilitatedbehaviorsdesignedtoscare,intimidate,abuse,illegallysurveil,orharmanindividual,aswellastheuser’sresponsestothesethreats\.Examplesincludecyberbullying,cyberstalking,doxxing,andintimatepartnersurveillance\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Report/blockanotheruseronanyplatformforharassment\.

\-Stalking,non\-consensualexplicitimagery,childsexualabuse,grooming,cyberflashing,unwantednudeimages,image\-basedsexualharassment,doxxing,intentionallyexposingprivateinfo,threatsofviolence,toxiccontent,trolling,bullying\.

\-Interpersonalabuseorintimatepartnerviolence,includingevidencecollectionforthesame\.

\-Impersonationforthepurposesofharassmentorabuse\.

\-Monitoringandcontrollinganindividualthroughsurveillanceinaninterpersonalcontext

\-Controllingorsurveillingaspecificindividual’sdevices,accounts,orcommunicationswithouttheirconsentinthecontextofaninterpersonalrelationship\(e\.g\.,intimatepartner,familymember,acquaintance\)

\*\*4:Networkandsystemdefenses\*\*

Explanation:Combinedstrategiesandtechnologiesusedtoprotectthenetworksandsystemsfromunauthorizedaccess,misuse,andcyberattacks\.Thiscategoryfocuseson\*defensiveandprotective\*measures,bestpractices,andincidentresponseattheinfrastructure,OS,andhardwarelevel\.Techniquesthatbypass,evade,orhidefromsecuritycontrolsarenotdefensivemeasures\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Networksecurity:VPN,Tor,firewall,Wi\-Fisecurity,Wireshark,bluetoothsecurity,homenetworksecurity,transport\-levelencryption\(e\.g\.,TLS,SSL\),identifyinganddefendingagainstnetwork\-levelattacks\(e\.g\.,on\-pathattacks,DDoS,DNSspoofing\)\.

\-OS&hardwaresecurity:SELinux,trustplatformmodules,trustedexecutionenvironments,networkvulnerabilityscanning,full\-disk/deviceencryption\(e\.g\.,BitLocker,FileVault\),identifyinganddefendingagainstOS\-levelandhardware\-levelattacks\.

\-Endpointprotections:anti\-virus,anti\-malware,network\-levelsecuritymonitoringtoolsforthreatdetection\(e\.g\.,SIEM,IDS/IPS\)\.

\-Systemmaintenance:OS/firmware/deviceupdates,outdatedsoftwareatsystemlevel,OS\-levelvulnerabilitydetection\.

\-Securitypractices:filesystemaccesscontrol,portscanning/blocking,auditlogs\.

\*\*5:Applicationandsoftwaredefenses\*\*

Explanation:Tools,settings,andbehaviorsusedtomitigateriskstodigitalsecurityandprivacyattheapplicationanddatalevel\.Thiscategoryfocusesonprotectivemeasures,bestpractices,andincidentresponseattheapplication,softwaredevelopment,anddatahandlinglevel\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Web&browsingprotection:privatebrowsing,adblockers,Brave,browsercookiesettings,incognitomode,browsersafetysettings\.

\-Securitypractices:application\-layerandend\-to\-endencryption,in\-applicationaccesscontrols\(e\.g\.,role\-basedpermissions,userauthorizationsettings\),securemail\-forwarding,securesoftwaredevelopmentandcodingpractices,securitycodereviewtoidentifyandmitigatevulnerability,softwarequalityassuranceandtesting\.

\-Privacy\-enhancingconfigurations:privacysettings,configuration,deleting/hiding/isolatingdata,preventingtracking\.

\-Systemmaintenance:installingormanagingapplication\-levelsoftwareupdates,accessingaccountloghistory,applicationvulnerabilityscanningandassessment\.

\*\*6:Compromiseandexploitation\*\*

Explanation:Offensivetechniquesandmaliciousactivitiestargetingaccounts,devices,networks,systems,software,orhardware,includingunauthorizedaccess,exploitation,andattackdevelopment\.Thiscategoryfocusesoncompromisesandactiveexploitationofvulnerability\(i\.e\.,usingorweaponizingavulnerabilitytoachieveunauthorizedoutcomes\),ratherthantheidentificationordetectionofvulnerabilities,ormentionsofoffensivetoolsorattacktechniquesinthecontextofdefenseordetection\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Account/devicehacking,theft,orcompromise,includingchangingaccountdetailsorhidingroot\.

\-Suspicioussign\-inattempts,suspiciousdevicebehavior,suspectedvirus,andindicatorsofcompromise\(e\.g\.,suspiciousnetworkconnections\)

\-Malware,adware,viruses,spyware,ransomware,botnet,unwantedsoftware,orvulnerabilityexploitation\.

\-Riskydownloadsorsuspiciousfiles\.

\-Requestingcodeorfunctionstogeneratemaliciouscode\(e\.g\.,bypass,break,orexploitspecificsoftwaretools\)

\-Useofsecurityorscanningtoolswithclearmaliciousintentorinanoffensivecontext

\*\*7:Socialengineering\*\*

Explanation:Exploitshumanpsychologyortrusttotrickpeopleintorevealingconfidentialinformation,aswellasdefensesagainstsuchtechniques\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Onlinescam,fakejobposting,employmentscams,romancescams\.

\-Unauthorizedorunexplainedchargesortransactions\.

\-Impersonationoronlineidentitytheft\.

\-Stolenormissingcreditcardfunds\.

\-Browserwarningsforanunsafewebpage\(e\.g\.,phishing\)\.

\-Fakeaccounts,suspiciouswebsites,suspiciousURLs,suspiciousmessages,suspiciousreceipts\.

\-Fraudsuchasincompletepurchases,blockedtransactions,stolenpersonaldetails,missinggoods,andlostmoney\.

\-Securityawarenessandtrainingtorecognizesocialengineeringattempts\.

\*\*8:Platformpolicyandenforcement\*\*

Explanation:Actionstakenbyaplatform,serviceprovider,orotherauthoritytoenforcepolicies,safetycontrols,compliancerules,automatedmoderationsystems,orlegal/regionalrestrictionsthataffectauser’sserviceusage\.Thiscategoryfocusesonplatform\-initiatedenforcementoutcomes,notongeneraldatahandlingpractices\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-Users’IPaddress,account,regionoraccesstotheplatformareblockedbytheplatform

\-Functionalityforaservice,somecontentsareforbidden

\-Personwasfalselyreportedtoaplatform

\-Personviolatedtheplatformpolicy

\*\*9:Emergingtechnologies\*\*

Explanation:S&Pchallengesthatarisespecificallyduetoemergingtechnologies\.ThiscategoryshouldbeusedwhenthetechnologyitselfintroducesuniqueS&Pconcernsbeyondtraditionalsecuritycontexts\.

Userpromptsinthiscategorymayincludeanyofthefollowing:

\-SecurityandprivacyissuesregardingIoT/CPS,smarthome,AI,blockchainandcryptocurrency,smartdevices\(excludingmainstreamconsumerdevicessuchassmartphones,laptops,andtablets\)

\-QuestionorrequestaboutAImodelpromptinjection,trainingdataleakage,orLLMjailbreaking\.

\#\#OutputRule

Outputonlytherelevantcategorynumber\(s\),separatedbycommasifmultiplecategoriesapply\.Donotincludeanyexplanationoradditionaltext\.

Examples:

\-Singlecategory:1

\-Multiplecategories:3,4,5

### A\.6Scorer Prompt

\#Instruction

Youareanexpertevaluator\.YourtaskistoevaluatethequalityoftheresponsesgeneratedbyAImodels\.

WewillprovideyouwiththeuserqueryandanAI\-generatedresponses\.

Youshouldfirstreadtheuserqueryandtheconversationhistorycarefullyforanalyzingthetask,andthenevaluatethequalityoftheresponsesbasedonandrulesprovidedbelow\.

\#ConversationbetweenUserandAI

\#\#CurrentUserQuery

<\|begin\_of\_query\|\>

\{$user\_query\}

<\|end\_of\_query\|\>

\#\#AIResponse

<\|begin\_of\_response\|\>

\{$model\_output\}

<\|end\_of\_response\|\>

\#Evaluation

\#\#Checklist

<\|begin\_of\_checklist\|\>

\{$checklist\}

<\|end\_of\_checklist\|\>

Pleaseusethischecklisttoguideyourevaluation,butdonotlimityourassessmenttothechecklist\.Foreachiteminthechecklistabove,youmustprovideadefinitiveyesornoanswerbasedonwhethertheresponsemeetsthatspecificcriterion\.Foreachitemlabeled\[yes\],youmustalsoprovidealistofexactquotesfromtheAIresponsethatallsupportthisdetermination\.Eachquotemustbeaseparatestringinthelist,limitedtothefirst10wordsofthatsentence\.

\#\#Rules

Youshouldcomparetheaboveresponsebasedonyouranalysisoftheuserqueriesandtheconversationhistory\.

Youshouldfirstwritedownyouranalysisandthechecklistthatyouusedfortheevaluation,andthenprovideyourassessmentaccordingtothechecklist\.

ChecklistAssessment:Listeveryitemfromtheprovidedchecklistandlabeleachas\[yes\]or\[no\]\.Foritemslabeled\[yes\],includealistofexactquotesfromtheresponsethatjustifythislabel,whereeachquoteisaseparatestringlimitedtothefirst10wordsofthatsentence\.

Thescoresareintherangeof1~10,where1meanstheresponseisverypoorand10meanstheresponseisperfect\.

Herearemoredetailedcriteriaforthescores:

\-Score1~2:Theresponseisverypooranddoesnotmakesenseatall\.

\-Score3~4:Theresponseispooranddoeshelpusersolvetheprobleminameaningfulway\.

\-Score5~6:Theresponseisfairbuthassomeissues\(e\.g\.,factualerrors,hallucinations,missingkeyinformation\)\.

\-Score7~8:Theresponseisgoodenoughbutcouldbeimprovedinsomeways\.

\-Score9~10:Theresponseisperfectandprovideshelpfulinformationthatcanhelpusersolvetheproblem\.

\#\#OutputFormat

First,pleaseoutputyouranalysisforthemodelresponse,andthensummarizeyourassessmenttotwoaspects:"strengths"and"weaknesses";Finally,pleasewritedownyourratingfortheassessment\.

Pleaseprovideyourevaluationresultsinthefollowingjsonformatbyfillingintheplaceholdersin\[\]:

‘‘‘

\{

"checklist\_results":\{

"1":\{"result":"\[yes/no\]","evidence":\["\[first10wordsofquote1\]","\[first10wordsofquote2\]"\]or\[\]ifno\},

"2":\{"result":"\[yes/no\]","evidence":\["\[first10wordsofquote1\]","\[first10wordsofquote2\]"\]or\[\]ifno\},

"\.\.\.":"\.\.\."

\},

"strengths":"\[analysisforthestrengthsoftheresponse\]",

"weaknesses":"\[analysisfortheweaknessesoftheresponse\]",

"score":"\[1~10\]"

\}

‘‘‘

Similar Articles

Demystifying the Privacy-Utility Trade-off in LLM Interactions

arXiv cs.AI

This paper analyzes the privacy-utility trade-off in LLM interactions, uncovering underlying mechanisms and introducing an intent-driven local protection framework with a lightweight model to enhance privacy while maintaining response utility.