Tag
This paper investigates how providing users with transparency and control over a political news recommendation system affects filter bubbles. A user study found that the enhanced interface increased awareness of filter bubbles but had heterogeneous effects on news consumption diversity.
Malleable Prompting is a novel interactive technique that reifies natural language preferences into GUI widgets (sliders, toggles, dropdowns) for direct manipulation, with a decoding algorithm that modulates token probabilities based on widget values to enable precise control over LLM generation. A user study shows it outperforms natural language prompting in precision, controllability, and transparency.
This research fine-tunes LLMs on human survey data to serve as judgmental models for group recommender systems, dynamically selecting aggregation strategies to maximize satisfaction and consensus. A user study validates that the approach aligns with human fairness and satisfaction perceptions.
Anthropic analyzed 400K Claude Code sessions and found that domain expertise is a stronger predictor of success than coding skill, with experts achieving 28-33% verified success versus 15% for novices. The study highlights that understanding the problem matters more than coding ability.
This paper analyzes real-world user queries about digital security and privacy asked to LLMs, categorizing them into nine topics and evaluating response quality and consistency across commercial and open-weight models.
This paper presents a gamified experiment where participants write responses with AI suggestions disincentivized, analyzing when humans adopt AI assistance versus maintaining creative autonomy.
Researchers from Oxford, Cambridge, MIT, CMU and other institutions conduct a mixed-methods study examining how people integrate AI tools into mathematical proof formalization workflows, finding that participants generally achieve higher formalization accuracy with AI assistance while preferring to maintain high-level human control over the proof discovery process.
A pilot study with 24 college students examines how varying levels of LLM access (none, limited, unlimited) affect essay writing quality, behavior, and perceived authorship, finding that constrained access preserves authorship confidence while unlimited access reduces creative expression and ownership.
This paper presents PrivacyAkinator, an interactive tool that helps novice developers articulate privacy design decisions via LLM-generated multiple-choice questions, achieving 47% more key decisions in 73% less time compared to NIST's PRAM methodology.
This paper presents a multimodal emotion recognition module for proactive conversational agents, using facial recognition and linguistic analysis. A user study with 20 participants reveals a 'poker face' effect where visual cues are unreliable, while linguistic analysis proves more accurate; the study also shows agents can elicit emotions through conversational adaptation.
Introduces CoTrace, a framework for goal-level attribution in human-AI collaboration, which analyzes how large language models shape goals by contributing concrete requirements and indirect influences in dialogue turns.
The COWCORPUS project, a study of 4,200 human-AI interactions, found that agents predicting their own failures and intervention moments are more useful than those simply trying to avoid errors. Researchers identified four stable trust patterns in human-AI collaboration and developed the Perfect Timing Score (PTS) to measure intervention prediction accuracy.
Anthropic presents research on how users seek personal guidance from Claude, highlighting findings on sycophancy rates across domains. The study informed the training of Claude Opus 4.7 and Mythos Preview to better protect user wellbeing.