Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry
Summary
This paper presents a deep reinforcement learning approach for solving vehicle routing problems, demonstrated through three industrial truck planning case studies. The proposed method achieves over 10% cost reduction compared to baseline results and discusses generalization to more VRP variants.
View Cached Full Text
Cached at: 08/10/26, 07:58 AM
# Vehicle routing problem using deep reinforcement learning—A case study about truck planning in the industry Source: [https://arxiv.org/html/2608.06668](https://arxiv.org/html/2608.06668) ,Lili Wu[liliawu@foxmail\.com](https://arxiv.org/html/2608.06668v1/mailto:[email protected])School of Electronic Information and Electrical EngineeringShanghaiChinaandDan Hu[hudan\.must@gmail\.com](https://arxiv.org/html/2608.06668v1/mailto:[email protected])School of Computer Science, Macau University of Science and TechnologyMacauChina \(2024\) ###### Abstract\. As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assistance of digital platforms and intelligent algorithms\. Within the field of transportation research, Vehicle Routing Problem \(VRP\) has remained a persistent and enduring challenge\. In the realm of management science, experts, and scholars from both the industrial and academic sectors have continuously explored optimization models and algorithms to effectively address routing problems, from the classical Traveling Salesman Problem to the more general Vehicle Routing Problem\. These models and algorithms are applied in real\-world industrial scenarios to achieve cost optimization and reduce carbon footprints\. However, due to the complexity of real\-world problems, numerous specific constraints are often added, and challenges such as information opacity, uncertainty, and irrational human behavior may arise\. Therefore, deploying and optimizing mathematical models for VRP in practical scenarios while maintaining optimal results poses numerous challenges\. This paper discusses and provides solutions for three different logistic use cases involving external truck network design\. Through these industrial case study, the paper introduces how deep reinforcement learning\-based vehicle routing optimization has been implemented\. As a result, it can be observed that the routes optimized by reinforcement learning agent have over 10% total cost compared to baseline results\. Furthermore, the paper proposes that in future research, DRL algorithms for vehicle routing problems could be generalized into more variations of VRP\. Vehicle routing problems, Deep reinforcement learning, Transformer network ††copyright:acmlicensed††journalyear:2024††doi:XXXXXXX\.XXXXXXX††conference:Make sure to enter the correct conference title from your rights confirmation emai; June 03–05, 2018; Woodstock, NY††isbn:978\-1\-4503\-XXXX\-X/18/06††ccs:Do Not Use This Code Generate the Correct Terms for Your Paper††ccs:Do Not Use This Code Generate the Correct Terms for Your Paper††ccs:Do Not Use This Code Generate the Correct Terms for Your Paper††ccs:Do Not Use This Code Generate the Correct Terms for Your PaperFigure 1\.Transportation in the industryVehicle Routing Problem \(VRP\)## 1\.Introduction Vehicle Routing Problem \(VRP\) typically refers to organizing and dispatching a fleet of vehicles to service a series of delivery and pickup stations by planning appropriate routes\. The goal is for the vehicles to sequentially visit these stations/points to achieve certain objectives \(like minimizing the total mileage, reducing overall transportation costs, ensuring vehicles arrive at certain times, or minimizing the number of vehicles used\) while satisfying specified constraints \(such as delivery and pickup time constraints, vehicle capacity limits, driving distance limits, driving time limits, etc\.\)\. In addition, VRP is also closely linked to corporate logistics models, such as pickup and delivery modes, heterogeneous vehicle types, multi\-warehouse mode, and multi\-trip mode\. To solve VRP, a common approach is to develop an optimization model\. However, due to complexity of real\-world problems, there may also be issues during real deployment with uncertainties\. Therefore, it is challenging to guarantee optimal results for the deployment in real\-world scenarios\. In terms of VRP, there are mainly three different solution categories, which are mixed integer programming, heuristic optimization and learning\-based optimization\. For mixed integer programming \(MIP\)\[\(Bouzidet al\.,[2017](https://arxiv.org/html/2608.06668#bib.bib102)\),Singhet al\.\([2021](https://arxiv.org/html/2608.06668#bib.bib104)\),Bachet al\.\([2016](https://arxiv.org/html/2608.06668#bib.bib105)\)\], despite such algorithms are able to achieve optimal solutions for small\-scale problems within a reasonable cost by establishing strict mathematical models\. However, for large\-scale problem, as the number of decision variables increases and the constraints of real\-world problems become more complex, there is a risk that no feasible solution may be found\. Therefore, heuristic optimization methods \[Christiaens and Vanden Berghe \([2020](https://arxiv.org/html/2608.06668#bib.bib101)\),Clarke and Wright \([1964](https://arxiv.org/html/2608.06668#bib.bib106)\),Gillett and Miller \([1974](https://arxiv.org/html/2608.06668#bib.bib107)\)\] such as simulated annealing have emerged to guarantee a near\-optimal solution by intuitive or experiential constructs and providing a feasible solution for a combinatorial optimization problem within an acceptable cost\. Even if heuristic optimization can provide a feasible solution for a specific instance, it has difficulty in generalizing the solution to other instances\. Therefore, with rapid development of learning\-based methods for combinatorial optimization, reinforcement learning has illustrated advantages over heuristic optimization regarding generalization and inference efficiency\. Hence, this paper mainly aims to discuss about a deployable model\-free transformer\-based policy network\[Liet al\.\([2021](https://arxiv.org/html/2608.06668#bib.bib103)\)\] and its application on truck planning in the industry\. As a well\-known problem in management science, this paper examines the heterogeneous capacitated VRP \(HCVRP\) with real industrial applications using deep reinforcement learning\. The paper has three contributions below: begin - •This paper extends the variation of VRP by incorporating a hybrid modes for both less\-than\-truckload \(LTL\) transportation and HCVRP for real scenarios, including constraints and objective functions\. - •A Transformer\-based deep reinforcement learning algorithm has achieved the objectives of cost reduction \(e\.g\., over 10% cost reduction\) and improved load factor \(e\.g\., no more than 25% empty load\) compared to baselines\. - •It provides a deployable solution for combinatorial optimization problem with learning\-based algorithms\. ## 2\.Methodology ### 2\.1\.Modeling The set of locations/points is defined asX=\{xi\}i=0nX=\\\{x^\{i\}\\\}\_\{i=0\}^\{n\}\. Since the starting point \(depot\) can be different in this case, each location can serve as a starting pointx0x^\{0\}, which is represented by its absolute coordinates and a demand of 0\. The remaining locations describe the coordinates of the starting and ending points for each order\. The expression for the set of remaining locations isX′=X∖\{x0\}X^\{\\prime\}=X\\setminus\\\{x^\{0\}\\\}\. Each of these remaining locations is expressed as\(si,di\)\(s^\{i\},d^\{i\}\), wheresis^\{i\}contains the coordinates of both starting and ending points, anddid^\{i\}includes the demand for each order\. Furthermore, the set of different vehicle types is defined asV=\{vi\}i=1mV=\\\{v^\{i\}\\\}\_\{i=1\}^\{m\}, whereviv^\{i\}represents the maximum capacity \(i\.e\. weight capacity, volume capacity or quantity capacity\) of each vehicle type\{Qi\}\\\{Q^\{i\}\\\}and m is the number of vehicle types\. All the aforementioned variables are non\-negative integers\. Although in the current real\-world scenario, different orders may have pickups at the starting point and deliveries at the ending point, the worst\-case scenario is considered here, where all orders need to be unloaded at the last stop of the route\. Therefore, the capacity of the vehicle chosen for a route must be no less than the total number of all orders on the route\. Besides capacity\-related constraints, time\-related constraints can added according to specific real\-world problems\. For instance, each vehicle on a route shall not drive longer than 24h\. Lastly, the cost function for the VRP in the paper should be: \(1\)min∑v∈V∑i∈X∑j∈Xfixed cost\+\(D\(xi,xj\)×unit cost\)×yi,j\\min\\sum\_\{v\\in V\}\\sum\_\{i\\in X\}\\sum\_\{j\\in X\}\\text\{fixed cost\}\+\\left\(D\(x^\{i\},x^\{j\}\)\\times\\text\{unit cost\}\\right\)\\times y\_\{i,j\}Here, the pair of fixed cost value and unit cost vary over vehicle type, which are given inputs of the cost function\. Meanwhile, the fixed cost and unit cost is calculated with the pricing model provided by domain experts according to vendor data\. In addition, becausexix^\{i\}represents the coordinates of both start and end points for an order, the distanceD\(xi,xj\)D\(x^\{i\},x^\{j\}\)represents distance betweenxix^\{i\}andxjx^\{j\}\. Additionally,yi,jy\_\{i,j\}is a binary variable representing whether these two adjacent locations exist in the route\. ### 2\.2\.States In MDP problem, each state is defined asst=\(Vt,Xt\)s\_\{t\}=\(V\_\{t\},X\_\{t\}\)whereVt=\{vt1,vt2,…,vtm\}=\{\(ot1,Tt1,Gt1\),\(ot2,Tt2,Gt2\),\(ot3,Tt3,Gt3\),…\}V\_\{t\}=\\\{v\_\{t\}^\{1\},v\_\{t\}^\{2\},\\ldots,v\_\{t\}^\{m\}\\\}=\\\{\(o\_\{t\}^\{1\},T\_\{t\}^\{1\},G\_\{t\}^\{1\}\),\(o\_\{t\}^\{2\},T\_\{t\}^\{2\},G\_\{t\}^\{2\}\),\(o\_\{t\}^\{3\},T\_\{t\}^\{3\},G\_\{t\}^\{3\}\),\\ldots\\\}, whereotio\_\{t\}^\{i\}represents the remaining capacity of the corresponding selected vehicle,TtiT\_\{t\}^\{i\}represents the accumulated time spent at timesteptt, andGti=\(g0i,g1i,…\)G\_\{t\}^\{i\}=\(g\_\{0\}^\{i\},g\_\{1\}^\{i\},\\ldots\)represents the set of coordinates passed through at timesteptt, i\.e\., the information of orders already delivered\. To maintain dimension consistency, the number of points/locations included inGtiG\_\{t\}^\{i\}for each vehicle typevtiv\_\{t\}^\{i\}at timestepttis the same\. Assuming that if only vehicle typevt1v\_\{t\}^\{1\}has a new location/point added at timesteptt, then the other vehicle types for corresponding routes will repeat adding the last point that has already been included in the corresponding route\. Therefore, in the initial state,V0=\{\(Q1,0,\{0\}\),\(Q2,0,\{0\}\),…\}V\_\{0\}=\\\{\(Q^\{1\},0,\\\{0\\\}\),\(Q^\{2\},0,\\\{0\\\}\),\\ldots\\\}is the state of vehicles departing from the depot\. At the same time, in each state, the point/location state isXt=\{xt1,xt2,…,xtm\}=\{\(st0,dt0\),\(st1,dt1\),\(st2,dt2\),…,\(stm,dtm\)\}X\_\{t\}=\\\{x\_\{t\}^\{1\},x\_\{t\}^\{2\},\\ldots,x\_\{t\}^\{m\}\\\}=\\\{\(s\_\{t\}^\{0\},d\_\{t\}^\{0\}\),\(s\_\{t\}^\{1\},d\_\{t\}^\{1\}\),\(s\_\{t\}^\{2\},d\_\{t\}^\{2\}\),\\ldots,\(s\_\{t\}^\{m\},d\_\{t\}^\{m\}\)\\\}, wherestis\_\{t\}^\{i\}represents the coordinate location anddtid\_\{t\}^\{i\}represents the order volume transported at timesteptt\. ### 2\.3\.Actions In this paper, the action space is divided into two parts: the first part is to select the optimal vehicle, and the second part is to design the route with the lowest transportation cost\. Therefore, the actionata\_\{t\}at timesteptt, denoted asat∈Aa\_\{t\}\\in A, can be expressed as\(vti,xti\)\(v\_\{t\}^\{i\},x\_\{t\}^\{i\}\), meaning that vehicleviv\_\{i\}will pass through the starting point or a specific relative coordinatexix\_\{i\}at timesteptt\. When sampling the action space, only one action is sampled at each timesteptt, including coordinates of a location/point and one vehicle\. Furthermore, for actions that do not satisfy the constraints \(such as remaining capacity less than 0\), real\-time updating masks can be established to ensure that the action obtained from the sampling space satisfies the constraints\. ### 2\.4\.Reward According to real\-world scenario, the objective function of the VRP is to minimize the delivery cost\. Therefore, the reward function is designed as the negative of the delivery cost of the truck at the current timesteptt, for example,R=−∑i=1m∑c=1CrtR=\-\\sum\_\{i=1\}^\{m\}\\sum\_\{c=1\}^\{C\}r\_\{t\}where m is the number of vehicles\. Assuming that at timestepttand timestept\+1t\+1,xjx^\{j\}andxkx^\{k\}are the location/point states at timestepttandt\+1t\+1respectively, and they are both handled by the same vehicle, then the reward function for timet\+1t\+1can be described as below: \(2\)rt\+1=r\(st\+1,at\+1\)=\{0,…,fixed\_cost×\{0,1\}\+D\(xj,xk\)×\\displaystyle r\_\{t\+1\}=r\(s\_\{t\+1\},a\_\{t\+1\}\)=\\\{0,\\ldots,\\text\{fixed\\\_cost\}\\times\\\{0,1\\\}\+D\(x^\{j\},x^\{k\}\)\\timesunit\_cost,0,…\}\\displaystyle\\text\{unit\\\_cost\},0,\\ldots\\\} For function approximation, since Vehicle Routing Problem \(VRP\) and its variants are classical graph problems with hard constraints, this neural network primarily adopts Transformer\-based architectures to train the policy network\. ## 3\.Case study In the context of lean manufacturing in the industry, a supply chain term for the way materials are transported across different factories is called ”external milk run”\. As the name implies, a milk run is a way of delivering milk where many outlets need milk while each needs only a small amount, as shown in Figure[2](https://arxiv.org/html/2608.06668#S3.F2)\. Therefore, using one vehicle for delivery covers various outlets helps cost savings and inventory reduction\. Such scenario also describes what external truck routes could look like, especially for manufacturing plants\. In order to apply the proposed learning\-based optimization method into real\-world, three different external milk\-run \(EMR\) use cases across multiple industrial plants were analyzed as testing datasets\. For those three datasets, demands come from multiple manufacturing plants across cities near Shanghai\. Moreover, same as training and validation datasets, the vehicle speed is assumed to be 35km/h\. Furthermore, with the input features of daily demand data from one point to another point across plants, vehicle type information, geographical locations of the locations, service windows \(e\.g\., from 8 a\.m\. to 6 p\.m\.\), and service time for loading/unloading service to take \(e\.g\., 0\.75 hours\), etc\., the real use cases can be transformed into mathematical models\. In terms of constraints for those three use cases, the capacity of each vehicle type in the first use case is represented using the goods quantity, which are 6, 20 , 26, 30, and 44\. However, the other two use cases use weight and volume limits\. In detail, the weight capacity of each truck type is 1\.9t, 4\.75t, 7\.6t, 9\.5t, 19\.0t and the volume capacity of each truck type is 10\.368m3, 30\.0288m3, 33\.5616m3, 40\.0896m3, 60\.0m3\. Figure 2\.Illustration about Milk\-runAdditionally, besides MR, there is another transportation mode called less\-than\-truck\-load transport \(LTL\) where each supplier individually transports goods from one point to another\. Although the transportation cost does not vary over distances traveled and LTL cost is usually higher than milk\-run cost, it is still possible LTL is a must option due to customers’ considerations or due to lower LTL cost\. Hence, in order to generate different training and validation samples, we strive to closely mimic the three real datasets\. Therefore, we use the real value range of each input feature from real use cases to generate more fake training and validation samples\. As the demand representation \(QTY\) in the first use differs from the other two \(weight & volume\), Table[1](https://arxiv.org/html/2608.06668#S3.T1)and Table[2](https://arxiv.org/html/2608.06668#S3.T2)represents training input features of two different training models, respectively\. Table 1\.Training & Validation Dataset input feature range for 1st use caseTable 2\.Training & Validation input feature range for 2nd & 3rd use casesBesides decision variables in the classical heterogeneous capacitated VRP model described in section 2, another variable about whether each order is delivered through EMR or through LTL is needed in the case studies\. To determine whether the specific order shall be delivered through LTL or through EMR, LTL mode is selected only if the route delivers just a single order and cost of LTL is lower than that of EMR\. In addition, in order to generate LTL price for each sample in the training set, we calculate the LTL price based on the real demand’s LTL calculation method provided by domain experts\. Finally, 10 training samples and 8 validation samples are generated as the training dataset and the validation dataset with each sample containing the demand of 200 orders, respectively\. We used a Quadro RTX 8000 GPU for training two models: the first model was used for the 1st use case, and the second model was used for the 2nd and 3rd use cases\. The first model’s and the second model’s training time were both approximately 5 hours for 100 epochs\. Moreover, during the training process, the graph size used by the models was 200, which corresponds to the number of demands\. During the training process, the performances with the average total transportation costs of model1 and model2 on the training dataset costs are shown in Figure[3\(a\)](https://arxiv.org/html/2608.06668#S3.F3.sf1)and Figure[3\(b\)](https://arxiv.org/html/2608.06668#S3.F3.sf2)\. As shown in the figures, both cost values has been decreased over the number of epochs\. In addition, the performances with the average rewards of model1 and model2 on the validation dataset are shown in Figure[4\(a\)](https://arxiv.org/html/2608.06668#S3.F4.sf1)and Figure[4\(b\)](https://arxiv.org/html/2608.06668#S3.F4.sf2)\. As shown in the figures, as rewards are negative values, both rewards decrease over the number of epochs, which indicates performance improvements during training\. \(a\)Model1 on training dataset \(b\)Model2 on training sets Figure 3\.Average total costs on the training dataset\(a\)Model1 on validation dataset \(b\)Model2 on validation dataset Figure 4\.Average rewards on the validation datasetTherefore, during inference, if the number of orders in the test data differs from the graph size, the graph needs to be padded with zeros\. Then, we used the sum of LTL price of each order as the baseline to compare the performance of the DRL algorithm for all these three use cases on model 1 with epoch 100 for 1st use case and model 2 with 100 for 2nd, 3rd use cases, as shown in Figure[5\(a\)](https://arxiv.org/html/2608.06668#S3.F5.sf1), Figure[5\(b\)](https://arxiv.org/html/2608.06668#S3.F5.sf2)and Figure[5\(c\)](https://arxiv.org/html/2608.06668#S3.F5.sf3), where the total number of order demands for the three use cases are 171, 360, and 504, respectively\. \(a\)First use case \(b\)Second use case \(c\)Third use case Figure 5\.Total Cost Comparisons on Different Use CasesAs a result, the DRL algorithm shows a significant decrease in total transportation cost compared to baselines for the three use cases, with reductions of 23\.48%, 18\.10%, and 18\.87%, respectively\. Besides the total transportation cost savings, the detailed results consists of each newly designed EMR route, including vehicle information and the selected route for each order\. Consequently, typical routes on a map from the three use cases are visualized below\. In the first case shown in Figure[6\(a\)](https://arxiv.org/html/2608.06668#S3.F6.sf1), the route passes 5 nodes in one trip\. In the second case shown in Figure[6\(b\)](https://arxiv.org/html/2608.06668#S3.F6.sf2), the route passes 4 nodes in one trip\. In the third case shown in Figure[6\(c\)](https://arxiv.org/html/2608.06668#S3.F6.sf3), the route passes 4 nodes in one trip\. From these routes, it can be seen that the DRL algorithm is effective to generate new routes by connecting nodes together to reduce costs\. \(a\)Visualization of the first use case \(b\)Visualization of the second use case \(c\)Visualization of the third use case Figure 6\.Total Cost Comparisons on Different Use Cases ## 4\.Discussions This paper illustrates that the hybrid HCVRP can be solved using deep reinforcement learning by applying the reinforcement learning\-based environment and algorithm, as well as analyzing the training, validation and testing results\. As a result, the inference speed of the model is much faster than classical optimization algorithm such as MIP and heuristic optimization\. Moreover, the training process requires generating a large number of training samples with similar characteristics to the test and validation data\. The results have shown that the total cost of the optimized EMR routes using reinforcement learning agents are lower than baseline results\. In addition, since research on using reinforcement learning algorithms to handle VRP and its variants is still relatively rare compared to classic heuristic optimization algorithms, several research directions are worth exploring with reinforcement learning\. For example, how to enhance the generalization capability of the agent, how to add time constraints, split frequency constraint and how to automate the number of different vehicle types selections, and how to add different capacity constraints by designing reward functions or mask functions\. Moreover, strategies about how to optimize the policy network algorithms performances \(e\.g\., PPO\) and the neural network architecture \(e\.g\. Transformer\) for more combinatorial optimization problem besides VRP are other future applied research direction in the real world\. ## 5\.Conclusion In the field of transportation research, the Vehicle Routing Problem \(VRP\) is a perennially challenging issue\. Experts and scholars in both industry and academia in the field of management science are constantly exploring optimization models and algorithms to effectively address routing problems\. These solutions are then applied in real industrial scenarios to ultimately achieve cost optimization\. The paper has illustrated the effective applications of deep reinforcement\-based planning for logistics scenarios in real life, which have outperformed the baselines\. The future work could focus on how to further generalize on different variations of VRP with a unified trained model so as to reduce training efforts\. ## References - L\. Bach, J\. Lysgaard, and S\. Wøhlk \(2016\)A branch\-and\-cut\-and\-price algorithm for the mixed capacitated general routing problem\.Networks68\(3\),pp\. 161–184\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - M\. C\. Bouzid, H\. A\. Haddadene, and S\. Salhi \(2017\)An integration of lagrangian split and vns: the case of the capacitated vehicle routing problem\.Computers & Operations Research78,pp\. 513–525\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - J\. Christiaens and G\. Vanden Berghe \(2020\)Slack induction by string removals for vehicle routing problems\.Transportation Science54\(2\),pp\. 417–433\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - G\. Clarke and J\. W\. Wright \(1964\)Scheduling of vehicles from a central depot to a number of delivery points\.Operations research12\(4\),pp\. 568–581\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - B\. E\. Gillett and L\. R\. Miller \(1974\)A heuristic algorithm for the vehicle\-dispatch problem\.Operations research22\(2\),pp\. 340–349\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - J\. Li, Y\. Ma, R\. Gao, Z\. Cao, A\. Lim, W\. Song, and J\. Zhang \(2021\)Deep reinforcement learning for solving the heterogeneous capacitated vehicle routing problem\.IEEE Transactions on Cybernetics52\(12\),pp\. 13572–13585\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\. - V\. P\. Singh, K\. Sharma, and D\. Chakraborty \(2021\)A branch\-and\-bound\-based solution method for solving vehicle routing problem with fuzzy stochastic demands\.Sādhanā46,pp\. 1–17\.Cited by:[§1](https://arxiv.org/html/2608.06668#S1.p3.1)\.
Similar Articles
Deep Reinforcement Learning solution for pickup and delivery routing problems with time window and capacity constraints
This paper presents a modified JAMPR deep reinforcement learning model to solve the Pickup and Delivery problem with Capacity and Time Window constraints (CPDPTW), offering fast optimal solutions for small to medium-sized instances and suboptimal solutions for larger ones.
Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking
The paper proposes an event-driven framework using Transformer and deep reinforcement learning for dynamic multi-depot vehicle routing, comparing it with heuristic and optimization methods.
Smart routes: a system for development and comparison of algorithms for solving vehicle routing problems with realistic constraints
This paper introduces Smart Routes, a platform for developing and comparing algorithms for vehicle routing problems with realistic constraints, showing that deep learning and heuristic methods can match exact solutions in quality with less time for larger problem sizes.
A Unified Knowledge Embedded Reinforcement Learning-based Framework for Generalized Capacitated Vehicle Routing Problems
This paper proposes a unified knowledge-embedded reinforcement learning framework for generalized capacitated vehicle routing problems, combining route-first cluster-second heuristics with dynamic programming to achieve superior solution quality and strong generalization across diverse variants.
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
This research paper proposes a simulation framework using deep reinforcement learning for controlling connected and automated vehicle platoon joining in mixed traffic, showing that PPO achieves high success rates while highlighting trade-offs between safety and efficiency.