monte-carlo-tree-search

Tag

Cards List
#monte-carlo-tree-search

Lookahead-R: Budget-Aware Tool Retrieval via Execution-Centric Planning

arXiv cs.CL ↗ · 2d ago Cached

Lookahead-R 将工具检索重构为受资源约束的序贯决策问题,用轻量级执行感知代理世界模型驱动成本敏感的 MCTS,在 ToolBench 的 I3 分割上取得 91.40% 的 NDCG@5,超过 ToolGen (90.16%)。

0 favorites 0 likes
#monte-carlo-tree-search

Graph-Based Stochastic Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper proposes Graph-Based Stochastic-Power-UCT, a graph-based Monte-Carlo search algorithm that improves sample efficiency in stochastic MDPs by sharing states across trajectories, with theoretical convergence guarantees and experimental validation.

0 favorites 0 likes
#monte-carlo-tree-search

Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM Frameworks

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper introduces an autonomous AI coding agent that combines Monte Carlo Tree Search with Gemini LLM to enhance code generation, achieving a 92% success rate on complex logical prompts.

0 favorites 0 likes
#monte-carlo-tree-search

Learning to Run Power Networks: Effective AlphaZero-inspired Topological Control

arXiv cs.LG ↗ · 2026-08-17 Cached

This paper evaluates AlphaZero-inspired reinforcement learning for topological control in power networks, achieving 98.43% survivability and emphasizing the effectiveness of minimalist integration with domain heuristics.

0 favorites 0 likes
#monte-carlo-tree-search

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

arXiv cs.LG ↗ · 2026-08-12 Cached

FlowScout is a framework that automatically generates tool-integrated agentic workflows from historical task-solving records, using Monte Carlo tree search guided by execution feedback. Experiments show it improves tool invocation correctness and execution score over baselines.

0 favorites 0 likes
#monte-carlo-tree-search

CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems

arXiv cs.AI ↗ · 2026-08-10 Cached

Introduces CEDAR, an autonomous method that uses LLM agents with Monte Carlo Tree Search to discover complex systems satisfying user-specified behavioral goals, reducing human effort and enabling goal-directed design.

0 favorites 0 likes
#monte-carlo-tree-search

Monte Carlo Tree Search for Table-to-Multimodal Report Generation

arXiv cs.AI ↗ · 2026-08-06 Cached

Introduces MCTS-Report, a Monte Carlo Tree Search framework for generating multimodal reports from tabular data, along with the MMRBench benchmark. It outperforms strong baselines across structural completeness, numerical accuracy, chart-text alignment, and insight novelty.

0 favorites 0 likes
#monte-carlo-tree-search

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper proposes MARS, a Monte-Carlo Tree Search-based framework for autonomously repairing multi-agent systems, along with StateMAS, a benchmark of 1,310 failure trajectories. Experiments show MARS consistently outperforms existing methods with improved repair accuracy and comparable token costs.

0 favorites 0 likes
#monte-carlo-tree-search

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

arXiv cs.AI ↗ · 2026-07-31 Cached

This paper presents a Belief-Guided architecture for Computer Go that replaces heavy MCTS with a Belief head and gating mechanism to reduce hallucination and enable professional-level play on consumer hardware.

0 favorites 0 likes
#monte-carlo-tree-search

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv cs.AI ↗ · 2026-07-29 Cached

Kernel Forge is an open-source agent harness that uses LLMs and Monte Carlo Tree Search to automatically generate and optimize CUDA kernels for any unmodified PyTorch model, achieving up to 2.83× speedup on softmax in Gemma 4 E2B.

0 favorites 0 likes
#monte-carlo-tree-search

Announcing Fugu-Ultra v1.1 🐡 (1 minute read)

TLDR AI ↗ · 2026-07-24 Cached

Sakana AI introduces AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, significantly improving performance on the ARC-AGI-2 benchmark.

0 favorites 0 likes
#monte-carlo-tree-search

EXPLORE: Exploration with Guided Search for Analog Topology Generation using Language Models

arXiv cs.LG ↗ · 2026-07-16 Cached

This paper presents EXPLORE, a framework that integrates simulator-guided Monte Carlo Tree Search with transformer-based decoding for analog topology generation, achieving a 65% success rate on a 6-component benchmark, significantly outperforming one-shot generation and sampling-and-filter baselines.

0 favorites 0 likes
#monte-carlo-tree-search

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

arXiv cs.AI ↗ · 2026-07-08 Cached

TOFFEE is a system that uses Monte Carlo Tree Search with adaptive model selection and cross-task prefix reuse to synthesize high-quality data agent trajectories at scale. These trajectories can be used for fine-tuning or in-context learning to improve data agent performance in heterogeneous enterprise environments.

0 favorites 0 likes
#monte-carlo-tree-search

Finding the Time to Think: Learning Planning Budgets in Real-Time RL

arXiv cs.LG ↗ · 2026-06-26 Cached

This paper introduces variable-delay real-time RL, where agents decide how long to deliberate in environments that progress during decision-making, and proposes a lightweight gating policy to select state-dependent planning budgets, outperforming fixed-budget and heuristic baselines in several real-time games.

0 favorites 0 likes
#monte-carlo-tree-search

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

arXiv cs.LG ↗ · 2026-06-26 Cached

KernelPro is a closed-loop multi-agent system that uses LLMs and micro-profiling tools to automatically optimize GPU kernel code, achieving geomean speedups of 2.42×/4.69×/5.30× on KernelBench and demonstrating a measured 11.6% energy reduction at matched speed.

0 favorites 0 likes
#monte-carlo-tree-search

Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry

arXiv cs.AI ↗ · 2026-06-26 Cached

This paper presents a Geometry-Aware Monte Carlo Tree Search framework for solving extremal combinatorial geometry problems on n×n grids, achieving new best-known results on five out of six tested problems, including improvements for the No-Three-in-Line problem.

0 favorites 0 likes
#monte-carlo-tree-search

MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation

arXiv cs.CL ↗ · 2026-06-17 Cached

Introduces MODE-RAG, a multi-agent system using Variational Free Energy and Monte Carlo Tree Search to dynamically gate interventions for mitigating hallucinations in Multimodal Retrieval-Augmented Generation systems, along with the ModeVent evaluation dataset.

0 favorites 0 likes
#monte-carlo-tree-search

Discovering Lattice Reduction Strategies via Self-Play

arXiv cs.LG ↗ · 2026-06-16 Cached

This paper presents Delta-Star, a deep reinforcement learning approach using AlphaZero-style self-play to discover superior lattice reduction strategies by interacting with the primitive actions of the LLL algorithm. The learned policy generalizes to higher dimensions and unseen moduli without retraining.

0 favorites 0 likes
#monte-carlo-tree-search

StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling

arXiv cs.LG ↗ · 2026-06-16 Cached

StarOR proposes a framework that synergizes Monte Carlo Tree Search with test-time reinforcement learning for automated optimization modeling, achieving state-of-the-art performance across multiple benchmarks.

0 favorites 0 likes
#monte-carlo-tree-search

Causal Object-Centric Models for Planning with Monte Carlo Tree Search

arXiv cs.AI ↗ · 2026-06-15 Cached

COMET is a model-based reinforcement learning algorithm that combines a frozen object-centric encoder with a transformer-based world model and Monte Carlo Tree Search, using causal attention to focus on task-relevant objects, achieving higher scores on visual RL benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback