Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems

arXiv cs.CL Papers

Summary

This exploratory pilot study evaluates personal information output from conversational interactions in generative AI systems, finding limited impact from model design differences and suggesting inferred profiles are constructed from contextual information.

arXiv:2609.22204v1 Announce Type: new Abstract: This exploratory pilot study evaluates the scope and perceived accuracy of personal information output from ongoing conversational interactions in generative AI systems using GPT-5.2 Instant and GPT-5.2 Thinking, categorized into three output types: Fact, Inference, and Confidence. Based on the evaluation results obtained from 15 Japanese participants, differences in model design have limited impact on personal information output tendencies. Compared with the Inference type, the Fact type shows a more conservative output pattern. Regarding attribute categories, the findings indicate that Core Personal attributes associated with identification are treated relatively conservatively, whereas Behavioral and Linguistic attributes show higher accuracy across both Fact and Inference outputs. Furthermore, Holistic Profile, Psychological and Cognitive, and Residual attributes are more readily inferred, even when not supported by explicit factual outputs. Notably, the lack of null outputs for these attributes in the Inference type suggests that such inferred profiles may be constructed from indirectly available contextual information. The findings may contribute to future discussions regarding privacy awareness and personal information inference in generative AI systems.
Original Article
View Cached Full Text

Cached at: 09/22/26, 09:07 AM

# Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems
Source: [https://arxiv.org/abs/2609.22204](https://arxiv.org/abs/2609.22204)
[View PDF](https://arxiv.org/pdf/2609.22204)

> Abstract:This exploratory pilot study evaluates the scope and perceived accuracy of personal information output from ongoing conversational interactions in generative AI systems using GPT\-5\.2 Instant and GPT\-5\.2 Thinking, categorized into three output types: Fact, Inference, and Confidence\. Based on the evaluation results obtained from 15 Japanese participants, differences in model design have limited impact on personal information output tendencies\. Compared with the Inference type, the Fact type shows a more conservative output pattern\. Regarding attribute categories, the findings indicate that Core Personal attributes associated with identification are treated relatively conservatively, whereas Behavioral and Linguistic attributes show higher accuracy across both Fact and Inference outputs\. Furthermore, Holistic Profile, Psychological and Cognitive, and Residual attributes are more readily inferred, even when not supported by explicit factual outputs\. Notably, the lack of null outputs for these attributes in the Inference type suggests that such inferred profiles may be constructed from indirectly available contextual information\. The findings may contribute to future discussions regarding privacy awareness and personal information inference in generative AI systems\.

## Submission history

From: Yosuke Seki \[[view email](https://arxiv.org/show-email/5eeb695d/2609.22204)\] **\[v1\]**Tue, 1 Sep 2026 04:42:07 UTC \(866 KB\)

Similar Articles

Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 Articles

arXiv cs.CL

This scoping review analyzes 81 articles (2022-2025) examining the use of generative AI for creating and evaluating user personas, identifying strengths in reproducibility but critical issues including lack of evaluation in 45% of studies, over-reliance on GPT models (86%), and risks of circularity where the same model generates and evaluates personas.