Tag
A team observed that both automated evaluation scores and user thumbs-down rates increased in the same week, suggesting a mismatch between objective metrics and user satisfaction.
User jakevin7 believes SKILLS management has issues and recommends against global installation, suggesting they are only suitable for temporary reading and distribution.
Thunderbird shares findings from user research on desktop settings, highlighting key themes like trust, clutter reduction, and navigation challenges, and outlines planned improvements such as clearer language and streamlined information architecture.
A user publicly announces they are stopping their subscription to Claude Max, expressing dissatisfaction.
A user voices frustration with ChatGPT's new desktop UI, feeling it oversimplifies the chat interface while pushing Codex, and missing organizational folders; asks if others share the sentiment.
dot is a feedback layer for AI-built products, enabling developers to collect and manage user feedback easily.
Artillain credits the CLine team for fixing performance issues after he submitted feedback and logs, noting they circumvented the IDE's internals.
Notes from day 1 of the aiDotEngineer conference featuring Kent Dodds' talk on product engineering in the AI world. Covers core thesis that product judgment is the last skill needed when AI commoditizes implementation, the Arrow Metaphor, differentiation between product engineer and product manager, validation techniques like The Mom Test, Jobs-to-Be-Done Framework, Kano Model for prioritizing features, and user feedback loops.
Correl8 AI is an MCP tool that lets AI agents directly report meaningful user feedback such as bugs, confusion, and feature requests, helping teams surface product signals without reviewing all chat logs.
This paper presents a deployment-centered evaluation of an LLM system integrated in electronic health records, training a classifier to predict query-level rejection risk using pre-response context like provider type and department, achieving an AUROC of 0.719 over 4.5 months of feedback.
AcuRite delays forced migration from My AcuRite to AcuRite NOW after users complain about missing features and usability issues.
A developer reflects on the inevitability of shipping AI features with poor outputs and emphasizes the need for proactive monitoring instead of relying solely on user reports.
This paper introduces MemToolAgent, a framework that enhances LLM agents' tool-using capabilities by integrating a memory system that stores and retrieves past experiences, achieving significant improvements on multiple benchmarks without requiring model fine-tuning.
A user shares that Qwen 3.6 27B is overly proactive, making unauthorized changes, and asks for advice on mitigation via prompt tweaks or parameter adjustments.
Binance introduces custom tabs for its app, allowing users to personalize their bottom navigation bar.
A user complains on social media about hitting usage limits quickly on a paid plan for Claude Code, criticizing Anthropic's treatment of users.
Anthropic is gathering user feedback for the cloud-enabled version of Claude Code across Desktop, iOS, and Android platforms through dedicated office hours.
OpenAI announces an update focused on making model responses more concise, based on user feedback.
Netizens question 阶跃星辰's premature push for commercialization, while praising Xiaomi Mimo's AI coding experience as better than or on par with Claude, and faster.
WildFeedback is a novel framework that leverages in-situ user feedback from actual LLM conversations to automatically create preference datasets for aligning language models with human preferences, addressing scalability and bias issues in traditional annotation-based alignment methods.