Tag
This paper explores how creative practitioners use language as a material interface for interacting with LLMs, through a two-week study with a tangible device called the Memetic Mixer.
This paper investigates how providing users with transparency and control over a political news recommendation system affects filter bubbles. A user study found that the enhanced interface increased awareness of filter bubbles but had heterogeneous effects on news consumption diversity.
This paper introduces a graph neural network model for real-time hand gesture recognition using surface electromyography (sEMG) signals from the forearm. The method achieves 99% classification accuracy with an average processing time of 48ms on an M1 Pro CPU, outperforming existing state-of-the-art techniques.
LUMOS introduces a semantic interaction layer that converts operating system metadata into machine-readable formats, enabling AI agents to interact with computer interfaces more efficiently by reducing dependence on screenshots and visual methods.
This paper investigates execution bottlenecks in computer-use agents, comparing screen-only GUI-based approaches with skill-mediated CLI-based methods, identifying key performance differences.
This paper proposes an indirect computing model and indirect formal method for optimizing cloud computing, using Chinese information data as an example to transition from data centers to knowledge centers.
This paper proposes a framework for evaluating LLMs' ability to generate multiple responses to scientific queries at different language complexity levels. The study finds that models often vary complexity inconsistently, with Claude Sonnet 4.5 performing best but only shifting complexity correctly 46% of the time.
J. C. R. Licklider's seminal 1960 paper introduces the concept of man-computer symbiosis, envisioning a tightly-coupled partnership where humans and computers cooperate to enhance intellectual capabilities, anticipating interactive computing and future AI development.
This paper studies how humans decide when to delegate to AI and when to adopt AI suggestions in cooperative question answering, finding that confirmation bias drives suboptimal trust decisions such as under-reliance on correct AI outputs.
The author observes that the hardest part of phone-use AI agents is tracking state changes, as mobile interfaces have more dynamic and interruptive UI changes compared to desktop, and asks for others' experience.
This paper investigates the use of LLMs to generate multimodal behaviors (verbal, vocal, gestural, facial) for trust calibration in socially interactive agents. The study finds that while LLMs can produce coherent behaviors aligned with intended trustworthiness traits, they also reproduce societal gender stereotypes.
DeepMind introduces an experimental AI-powered mouse pointer that understands visual context and intent, aiming to streamline user interactions with AI across different applications.
Google DeepMind is experimenting with reimagining the mouse pointer interface using Gemini AI, allowing users to control screens through motion, speech, and natural shorthand.
The article analyzes Andrej Karpathy's perspective on using HTML as an output format for LLMs, exploring the evolution of human-computer interaction from a neuroscience viewpoint. The author argues that although the future may shift toward neural simulation, HTML will likely remain a best practice for human-AI collaboration in the near to medium term due to its engineering maintainability and low cost.
The article analyzes the psychological 'illusion of listening' where users perceive AI as empathetic due to linguistic cues, despite the lack of genuine understanding. It proposes design guidelines to ensure transparency and prevent users from outsourcing human connection to automated systems.
This paper presents six dimensions to characterize live feedback in interactive programming systems: granularity, reactivity, velocity, moldability, bidirectionality, and materiality, aiming to map the design space.
The article discusses 'LLMorphism,' a concept where humans begin to view themselves through the lens of language models, exploring the implications for human cognition and self-perception.
The author discusses the limitations of managing AI agent workflows via chat interfaces like Telegram with OpenClaw, advocating for dedicated dashboards and standardized UIs. They highlight emerging tools like Paperclip and Multica that aim to solve agent management issues.
The COWCORPUS project, a study of 4,200 human-AI interactions, found that agents predicting their own failures and intervention moments are more useful than those simply trying to avoid errors. Researchers identified four stable trust patterns in human-AI collaboration and developed the Perfect Timing Score (PTS) to measure intervention prediction accuracy.
The author explores the affordances of a screenless writing interface, noting its suitability for first drafts and stream-of-consciousness writing while highlighting limitations in editing and context retention.