long-context-reasoning

Tag

Cards List
#long-context-reasoning

ConvMem: Convolutional Memory for Long-Context Reasoning

arXiv cs.AI · 2026-09-11 Cached

ConvMem is a training-free, parallelizable framework that reformulates long-context reasoning in large language models as hierarchical convolution to improve efficiency, avoid overfitting, and outperform baseline methods.

0 favorites 0 likes
#long-context-reasoning

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

arXiv cs.CL · 2026-08-17 Cached

This paper presents SimpleOPD, a method for on-policy distillation from long-context reasoning teachers to short-context students, overcoming tokenizer mismatch and training instability to enhance mathematical reasoning capabilities.

0 favorites 0 likes
#long-context-reasoning

Decentralized Multi-Agent Systems with Shared Context

Hugging Face Daily Papers · 2026-06-09 Cached

This paper introduces Decentralized Language Models (DeLM), a framework for multi-agent systems that uses parallel agents with a shared verified context to improve test-time scaling and reduce costs, achieving state-of-the-art results on SWE-bench Verified and LongBench-v2.

0 favorites 0 likes
#long-context-reasoning

LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards

Hugging Face Daily Papers · 2026-05-29 Cached

LongTraceRL introduces tiered distractor construction and rubric reward design to improve long-context reasoning in language models using reinforcement learning. The method generates multi-hop questions via knowledge graph random walks and uses search agent trajectories to build challenging distractors, with a rubric reward providing entity-level process supervision.

0 favorites 0 likes
#long-context-reasoning

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading

Hugging Face Daily Papers · 2026-05-11 Cached

MemReread introduces a method for long-context reasoning that avoids intermediate retrieval by decomposing questions and rereading text to recover discarded information, achieving linear time complexity. It outperforms baseline frameworks on long-context reasoning tasks.

0 favorites 0 likes
← Back to home

Submit Feedback