commonsense-reasoning

Tag

Cards List
#commonsense-reasoning

@rohanpaul_ai: LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters a…

X AI KOLs Following · 2026-08-02 Cached

A new paper introduces SaliTrap, a benchmark of 1,145 prompts revealing that LLMs know impossible tasks but still optimize for explicit details, a failure called salience bias. Even the best models avoid traps only ~55% of the time, and awareness often doesn't prevent compliance.

0 favorites 0 likes
#commonsense-reasoning

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Hugging Face Daily Papers · 2026-07-30 Cached

This paper introduces SaliTrap, a benchmark revealing that large language models suffer from salience bias—being distracted by explicit but irrelevant conditions—in commonsense reasoning. It shows the failure is mostly due to knowledge suppression rather than absence, and that simple prompting can close the gap.

0 favorites 0 likes
#commonsense-reasoning

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

arXiv cs.AI · 2026-06-11 Cached

TouchThinker introduces a million-scale tactile reasoning dataset and benchmark to scale tactile commonsense reasoning to open-world settings, using action-aware representation for efficient reasoning.

0 favorites 0 likes
#commonsense-reasoning

STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?

Hugging Face Daily Papers · 2026-05-07 Cached

This paper identifies a critical failure mode in LLM agents where they fail to update personalized memories when new evidence conflicts with prior beliefs. It introduces the STALE benchmark and a three-dimensional probing framework, revealing that even the best models achieve only 55.2% accuracy, and proposes CUPMem as a prototype for robust memory revision.

0 favorites 0 likes
← Back to home

Submit Feedback