tool-integrated-reasoning

Tag

Cards List
#tool-integrated-reasoning

Contrastive Branch Policy Optimization

arXiv cs.LG · 3d ago Cached

CBPO introduces a contrastive branch policy optimization method for fine-grained credit assignment in reinforcement learning with verifiable rewards, enhancing language model performance in tool-integrated reasoning tasks across multiple benchmarks.

0 favorites 0 likes
#tool-integrated-reasoning

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

arXiv cs.CL · 4d ago Cached

Proposes HiDiffTIR, a hierarchical difficulty-aware policy optimization framework for multi-turn tool-integrated reasoning in LLM agents, improving performance and tool invocation accuracy through fine-grained credit assignment.

0 favorites 0 likes
#tool-integrated-reasoning

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Hugging Face Daily Papers · 2026-08-04 Cached

TurnSight introduces a turn-level hindsight self-distillation framework for tool-integrated reasoning, providing dense supervision via execution-conditioned hindsight and adaptive RL advantage modulation.

0 favorites 0 likes
#tool-integrated-reasoning

ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

arXiv cs.AI · 2026-07-20 Cached

ToolVerse is a framework that automatically builds massive executable agent training environments from 422 real-world MCP environments containing 4438 tools, and proposes a task design strategy using Dynamic Unlocking Sampling Algorithm to generate long-horizon tasks, along with a Turn-Aware Relative Advantage algorithm for credit assignment in agentic reinforcement learning.

0 favorites 0 likes
#tool-integrated-reasoning

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

Hugging Face Daily Papers · 2026-06-04 Cached

The ToolMaze benchmark evaluates LLM agents' ability to handle real-world tool failures, revealing that implicit semantic failures cause the largest performance drops and that dynamic replanning remains a critical bottleneck not addressed by scaling or prompting.

0 favorites 0 likes
#tool-integrated-reasoning

Teaching Language Models to Think in Code

arXiv cs.CL · 2026-05-11 Cached

This paper introduces ThinC (Thinking in Code), a framework where language models use code blocks exclusively for reasoning after a brief natural language planning step, outperforming existing tool-integrated reasoning baselines on math benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback