Tag
This paper proposes 'cooperative observation' as a framework for personal intelligence, emphasizing a feedback loop between AI systems and users to improve assistance, trust, and privacy. It reports on a preliminary case study from the Organizm prototype used over six months and outlines evaluation directions.
The paper proposes an Inverse Theory of Mind (IToM) pipeline that infers user beliefs, preferences, and decision-making traits from observed interactions, using LLM-driven counterfactual reasoning to synthesize structured user personas for adaptive content recommendation across modalities, including a VisionOS spatial banking app.
Presents SERUM, a multi-pass framework that extracts structured behavioral models of user actions and intents from raw egocentric video using hierarchical VLM annotation, reducing hallucinations and producing interpretable process models without manual annotation.
Introduces IRIS, a framework that learns dynamic user personas from implicit interaction streams without explicit feedback, outperforming static and memory-only baselines on decision prediction.
This paper argues that agents should help users construct preferences rather than assuming well-formed ones, proposing the CoPref model and CoShop benchmark. Evaluations show even frontier models achieve only 56% accuracy due to poor preference expansion.
ScaleToT proposes a method to generalize structured LLM reasoning for low-activity user modeling at billion scale, using tree-of-thought refinement and training a student model to reduce cost. An online A/B test in advertising deployment showed a 6.738% increase in LT30.
ThoughtTrace introduces a large-scale dataset pairing real-world multi-turn human-AI conversations with users' self-reported thoughts, enabling improved user behavior prediction and personalized assistant training through thought-guided rewrites.
IPQA introduces a benchmark for evaluating core intent identification in personalized question answering, addressing a gap in existing metrics that focus on response quality rather than intent understanding. The paper presents a dataset construction methodology grounded in bounded rationality and demonstrates that state-of-the-art language models struggle with identifying user-prioritized intents from answer selection patterns.