reasoning-llms

Tag

Cards List
#reasoning-llms

@ddkang: New research from Bridgewater AIA Labs, UIUC, and MIT: we prove what we believe to be the first non-vacuous generalizat…

X AI KOLs Timeline · 2026-07-20 Cached

Researchers from Bridgewater AIA Labs, UIUC, and MIT prove the first non-vacuous generalization bounds for reasoning LLMs trained with RLVR, providing provable accuracy lower bounds on unseen data to guide safe deployment.

0 favorites 0 likes
#reasoning-llms

On-Policy Delta Distillation

Hugging Face Daily Papers · 2026-07-16 Cached

The paper introduces On-Policy Delta Distillation (OPD^2), a new distillation reward called the delta signal that captures the difference between a teacher model and its base model before reasoning tuning, providing a more direct signal for transferring reasoning capabilities. Experiments across math, science, and code benchmarks show OPD^2 consistently outperforms conventional on-policy distillation.

0 favorites 0 likes
#reasoning-llms

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

arXiv cs.AI · 2026-07-14 Cached

This paper investigates semantic context drift in reasoning-class LLMs within hybrid decision support systems, proposing a mathematical model and a stability metric. A two-month experiment reveals latent goal-targeting drift and formulates engineering recommendations for control stability.

0 favorites 0 likes
#reasoning-llms

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

arXiv cs.CL · 2026-07-08 Cached

PluraMath extends the PolyMath dataset to 18 underrepresented languages, providing a human-validated benchmark for evaluating multilingual mathematical reasoning in LLMs. The paper reveals a persistent performance gap between high-resource and low-resource languages across 27 models.

0 favorites 0 likes
#reasoning-llms

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing

arXiv cs.CL · 2026-06-17 Cached

Introduces STATEWITNESS, an activation explainer for auditing deception in reasoning LLMs, achieving significant improvements over existing monitors and providing human-inspectable evidence.

0 favorites 0 likes
#reasoning-llms

How Do Answer Tokens Read Reasoning Traces? Self-Reading Patterns in Thinking LLMs for Quantitative Reasoning

arXiv cs.CL · 2026-04-22 Cached

Study reveals that answer tokens in thinking LLMs follow a structured self-reading pattern—forward drift plus focus on key anchors—during quantitative reasoning, and proposes a training-free SRQ steering method to exploit this for accuracy gains.

0 favorites 0 likes
← Back to home

Submit Feedback