compositional-reasoning

Tag

Cards List
#compositional-reasoning

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

arXiv cs.AI · 2026-07-21 Cached

Introduces RECON, a benchmark for evaluating compositional reasoning over long contexts in LLM-based agents, spanning 24 case files across criminal, medical, and financial domains. The best non-oracle system achieves only 22.4% accuracy, revealing substantial limitations in current memory architectures.

0 favorites 0 likes
#compositional-reasoning

RL Post-Training Builds Compositional Reasoning Strategies

arXiv cs.CL · 2026-07-09 Cached

This paper investigates whether reinforcement learning post-training can compose primitive skills into higher-level compositional strategies, using a fully observable rewrite-grammar environment. The authors find that RL reorganizes primitive competence through phased compositional mechanisms, while rejection fine-tuning plateaus due to producing many invalid shortcut-like rewrites.

0 favorites 0 likes
#compositional-reasoning

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

arXiv cs.AI · 2026-07-01 Cached

Introduces CDR-Bench, a benchmark with 3,462 tasks to evaluate LLMs' ability to faithfully execute compositional, order-sensitive data refinement recipes. Experiments on 10+ LLMs reveal significant performance degradation in compositional and order-sensitive settings, highlighting a lack of procedural faithfulness.

0 favorites 0 likes
#compositional-reasoning

Holographic Memory for Zero-Shot Compositional Reasoning in Knowledge Graphs: A Mechanistic Study of Where and Why It Fails

arXiv cs.LG · 2026-06-25 Cached

This paper investigates holographic reduced representations for zero-shot compositional reasoning in knowledge graphs, finding that while single-hop performance is strong, composition fails due to retrieval capacity and interference effects in the superposed memory, not the bind-unbind algebra.

0 favorites 0 likes
#compositional-reasoning

R-APS: Compositional Reasoning and In-Context Meta-Learning for Constrained Design via Reflective Adversarial Pareto Search

arXiv cs.AI · 2026-06-04 Cached

R-APS (Reflective Adversarial Pareto Search) is a novel method for constrained design tasks that addresses three structural failures in LLM-based agentic systems—error propagation, robustness evaluation, and knowledge invalidation—through reasoning-mode decomposition across three timescales, requiring no fine-tuning. Evaluated on planar mechanism synthesis, it achieves 3.5x tighter robustness certificates, 46% faster iterations-to-first-admission, and 2.1x Chamfer-distance reduction over baselines.

0 favorites 0 likes
#compositional-reasoning

MAVEN: Improving Generalization in Agentic Tool Calling

arXiv cs.AI · 2026-06-01 Cached

MAVEN is a lightweight symbolic reasoning scaffold that improves generalization in agentic tool calling by using modular verification and adaptive tool orchestration. It achieves significant accuracy gains on a new stress-test benchmark (MAVEN-Bench) and remains competitive with proprietary models at a fraction of the cost.

0 favorites 0 likes
#compositional-reasoning

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

arXiv cs.AI · 2026-05-27 Cached

This paper introduces 'composition collapse', a phenomenon where language models with stable factual knowledge still fail to compose that knowledge into correct multi-hop reasoning, and proposes a double-gate protocol to isolate composition failure from atomic knowledge instability.

0 favorites 0 likes
#compositional-reasoning

Shortcut Solutions Learned by Transformers Impair Continual Compositional Reasoning

arXiv cs.LG · 2026-05-08 Cached

This research paper investigates how shortcut solutions learned by Transformer models, specifically BERT, impair their ability to perform continual compositional reasoning. It contrasts BERT with ALBERT, finding that ALBERT's recurrent nature offers better inductive bias for continual learning tasks.

0 favorites 0 likes
#compositional-reasoning

The Amazing Agent Race: Strong Tool Users, Weak Navigators

arXiv cs.CL · 2026-04-20 Cached

The Amazing Agent Race (AAR) introduces a new benchmark with 1,400 directed acyclic graph (DAG) puzzle instances to evaluate LLM agents on fork-merge tool chains and Wikipedia navigation. Evaluations reveal agents excel at tool-use (errors <17%) but struggle with navigation (27-52% of failures), exposing a critical gap invisible to existing linear benchmarks.

0 favorites 0 likes
#compositional-reasoning

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding

Hugging Face Daily Papers · 2026-04-14 Cached

Proposes Slipform, a training framework that uses lexical concreteness to select harder negatives and a margin-based Cement loss, boosting compositional reasoning in vision-language models.

0 favorites 0 likes
← Back to home

Submit Feedback