Tag
Lookahead-R 将工具检索重构为受资源约束的序贯决策问题,用轻量级执行感知代理世界模型驱动成本敏感的 MCTS,在 ToolBench 的 I3 分割上取得 91.40% 的 NDCG@5,超过 ToolGen (90.16%)。
This paper proposes Graph-Based Stochastic-Power-UCT, a graph-based Monte-Carlo search algorithm that improves sample efficiency in stochastic MDPs by sharing states across trajectories, with theoretical convergence guarantees and experimental validation.
This paper introduces an autonomous AI coding agent that combines Monte Carlo Tree Search with Gemini LLM to enhance code generation, achieving a 92% success rate on complex logical prompts.
This paper evaluates AlphaZero-inspired reinforcement learning for topological control in power networks, achieving 98.43% survivability and emphasizing the effectiveness of minimalist integration with domain heuristics.
FlowScout is a framework that automatically generates tool-integrated agentic workflows from historical task-solving records, using Monte Carlo tree search guided by execution feedback. Experiments show it improves tool invocation correctness and execution score over baselines.
Introduces CEDAR, an autonomous method that uses LLM agents with Monte Carlo Tree Search to discover complex systems satisfying user-specified behavioral goals, reducing human effort and enabling goal-directed design.
Introduces MCTS-Report, a Monte Carlo Tree Search framework for generating multimodal reports from tabular data, along with the MMRBench benchmark. It outperforms strong baselines across structural completeness, numerical accuracy, chart-text alignment, and insight novelty.
This paper proposes MARS, a Monte-Carlo Tree Search-based framework for autonomously repairing multi-agent systems, along with StateMAS, a benchmark of 1,310 failure trajectories. Experiments show MARS consistently outperforms existing methods with improved repair accuracy and comparable token costs.
This paper presents a Belief-Guided architecture for Computer Go that replaces heavy MCTS with a Belief head and gating mechanism to reduce hallucination and enable professional-level play on consumer hardware.
Kernel Forge is an open-source agent harness that uses LLMs and Monte Carlo Tree Search to automatically generate and optimize CUDA kernels for any unmodified PyTorch model, achieving up to 2.83× speedup on softmax in Gemma 4 E2B.
Sakana AI introduces AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, significantly improving performance on the ARC-AGI-2 benchmark.
This paper presents EXPLORE, a framework that integrates simulator-guided Monte Carlo Tree Search with transformer-based decoding for analog topology generation, achieving a 65% success rate on a 6-component benchmark, significantly outperforming one-shot generation and sampling-and-filter baselines.
TOFFEE is a system that uses Monte Carlo Tree Search with adaptive model selection and cross-task prefix reuse to synthesize high-quality data agent trajectories at scale. These trajectories can be used for fine-tuning or in-context learning to improve data agent performance in heterogeneous enterprise environments.
This paper introduces variable-delay real-time RL, where agents decide how long to deliberate in environments that progress during decision-making, and proposes a lightweight gating policy to select state-dependent planning budgets, outperforming fixed-budget and heuristic baselines in several real-time games.
KernelPro is a closed-loop multi-agent system that uses LLMs and micro-profiling tools to automatically optimize GPU kernel code, achieving geomean speedups of 2.42×/4.69×/5.30× on KernelBench and demonstrating a measured 11.6% energy reduction at matched speed.
This paper presents a Geometry-Aware Monte Carlo Tree Search framework for solving extremal combinatorial geometry problems on n×n grids, achieving new best-known results on five out of six tested problems, including improvements for the No-Three-in-Line problem.
Introduces MODE-RAG, a multi-agent system using Variational Free Energy and Monte Carlo Tree Search to dynamically gate interventions for mitigating hallucinations in Multimodal Retrieval-Augmented Generation systems, along with the ModeVent evaluation dataset.
This paper presents Delta-Star, a deep reinforcement learning approach using AlphaZero-style self-play to discover superior lattice reduction strategies by interacting with the primitive actions of the LLL algorithm. The learned policy generalizes to higher dimensions and unseen moduli without retraining.
StarOR proposes a framework that synergizes Monte Carlo Tree Search with test-time reinforcement learning for automated optimization modeling, achieving state-of-the-art performance across multiple benchmarks.
COMET is a model-based reinforcement learning algorithm that combines a frozen object-centric encoder with a transformer-based world model and Monte Carlo Tree Search, using causal attention to focus on task-relevant objects, achieving higher scores on visual RL benchmarks.