AlphaTransit: Learning to Design City-scale Transit Routes
Summary
AlphaTransit combines Monte Carlo Tree Search with neural policy-value networks to optimize bus route design by predicting downstream quality without simulator rollouts. It achieves significant service rate improvements on a Bloomington transit benchmark.
View Cached Full Text
Cached at: 06/01/26, 07:21 PM
Paper page - AlphaTransit: Learning to Design City-scale Transit Routes
Source: https://huggingface.co/papers/2605.28730
Abstract
AlphaTransit combines Monte Carlo Tree Search with neural policy-value networks to optimize bus route design by predicting downstream quality and enabling lookahead decisions without simulator rollouts.
Designing a transit network requires many sequential route extension decisions, but their quality is often visible only after the full network is assembled. This delayed-feedback challenge lies at the heart of the Transit Route Network Design Problem (TRNDP), where route interactions can be deceptive: an extension that appears useful locally can create transfer bottlenecks, produce redundant overlap, or reduce overall throughput. To guide route construction under delayed simulator feedback, we introduce AlphaTransit, a search-based planning framework for cityscale bus network design. AlphaTransit couplesMonte Carlo Tree Search(MCTS) with aneural policy-value network: the policy proposesroute extensions, the value estimates downstream design quality, and search uses these predictions to refine each decision. This provides decision-time lookahead during route construction without running simulator rollouts inside the search tree. We evaluate AlphaTransit on a new Bloomington TRNDP benchmark with realistic road topology and censusderived demand, under mixed and full transit demand settings. In the Bloomington network, AlphaTransit attains the highestservice ratein both demand settings, reaching 54.6% and 82.1%, respectively. Relative toreinforcement learningwithout search, these correspond to 9.9% and 11.4%service rategains; relative to MCTS without learned guidance, they correspond to 2.5% and 11.2% gains. These results suggest that coupling learned guidance with MCTS is more effective than using either approach alone fortransit network design. Our code and data are publicly available in https://github.com/poudel-bibek/AlphaTransit.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2605\.28730
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### matrix-multiply/alphatransit-checkpoints Reinforcement Learning• Updatedabout 5 hours ago
Datasets citing this paper1
#### matrix-multiply/bloomington-tndp Viewer• Updatedabout 5 hours ago • 6.12k
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.28730 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning
Researchers from the University of Amsterdam propose a tabular reinforcement learning approach to the Metro Network Expansion Problem, showing it achieves comparable performance to Deep RL while reducing training episodes by 18x and carbon emissions by 12x on average. The method also incorporates social equity criteria and is evaluated on real-world metro networks in Xi'an and Amsterdam.
Smart routes: a system for development and comparison of algorithms for solving vehicle routing problems with realistic constraints
This paper introduces Smart Routes, a platform for developing and comparing algorithms for vehicle routing problems with realistic constraints, showing that deep learning and heuristic methods can match exact solutions in quality with less time for larger problem sizes.
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
This paper introduces TransitLM, a large-scale dataset of over 13 million transit route planning records from Chinese cities, enabling map-free route generation via LLMs trained directly on planning data.
Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority
This paper presents a preference-conditioned multi-objective reinforcement learning controller for transit signal priority that allows runtime tuning of the trade-off between bus priority and overall traffic delay without retraining. Experiments show it outperforms fixed-time and rule-based baselines while maintaining feasibility constraints.
ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
ChatPlanner is a novel framework that uses fine-tuned LLMs with Retrieval-Augmented Generation (RAG) to interpret user preferences from natural language queries and integrate them into public transit routing algorithms, outperforming existing route planners.