Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
Summary
This paper presents a unified empirical evaluation of methods for revising travel itineraries under resource disruptions, comparing full replanning, plan repair, and LLM-based revision approaches in terms of effectiveness, plan stability, and computational cost.
View Cached Full Text
Cached at: 09/18/26, 09:22 AM
# Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
Source: [https://arxiv.org/html/2609.19654](https://arxiv.org/html/2609.19654)
###### Abstract
Travel\-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures\. Revising these itineraries involves full replanning, classical plan repair, and LLM\-based travel\-agent revision, whose differing task formulations and evaluation protocols hinder comparison\. We conduct a systematic empirical study using two TREK\-derived benchmark sets: 500 single\-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound\-disruption cases\. We compare LLM–Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local\-revision adapter across effectiveness, plan stability, and computational cost\. LLM–Z3 with Gemini achieved the highest observed compound\-disruption success\. IPyHOPPER nearly matched that configuration’s single\-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs\. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning\. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM–Z3 adapter used compact one\-call inference, and the iTIMO adapter consumed substantially more tokens\. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings\.
###### Keywords:
Travel planning Itinerary revision Plan repair Language agents Empirical evaluation
## 1Introduction
Travel itinerary planning coordinates transportation, accommodation, and activities under user preferences, budgets, schedules, and resource availability\. Travel\-planning agents can generate feasible itineraries, including by combining large language models \(LLMs\) with formal verification tools\[[33](https://arxiv.org/html/2609.19654#bib.bib8),[29](https://arxiv.org/html/2609.19654#bib.bib1),[10](https://arxiv.org/html/2609.19654#bib.bib14),[35](https://arxiv.org/html/2609.19654#bib.bib34),[14](https://arxiv.org/html/2609.19654#bib.bib35)\]\. Flight cancellations, hotel unavailability, or attraction closures can subsequently invalidate an accepted itinerary\[[16](https://arxiv.org/html/2609.19654#bib.bib9),[20](https://arxiv.org/html/2609.19654#bib.bib4)\], shifting the task from*initial generation*to*revision*\.
Revision starts from accepted commitments and preferences\. Restoring feasibility may require changes beyond the unavailable resource, but generating a new itinerary can also replace arrangements that remain valid\. Evaluation must therefore consider preservation of the accepted itinerary\[[7](https://arxiv.org/html/2609.19654#bib.bib20)\]alongside feasibility recovery and computational cost, as illustrated in Figure[1](https://arxiv.org/html/2609.19654#S1.F1)\. Studies of LLM evaluation and agentic systems across classification, question answering, retrieval, financial analysis, data exploration, and serving provide complementary examples of task\-specific quality, reliability, and efficiency assessment\[[28](https://arxiv.org/html/2609.19654#bib.bib29),[17](https://arxiv.org/html/2609.19654#bib.bib30),[30](https://arxiv.org/html/2609.19654#bib.bib28),[25](https://arxiv.org/html/2609.19654#bib.bib27),[19](https://arxiv.org/html/2609.19654#bib.bib25),[24](https://arxiv.org/html/2609.19654#bib.bib31),[31](https://arxiv.org/html/2609.19654#bib.bib32),[18](https://arxiv.org/html/2609.19654#bib.bib33),[12](https://arxiv.org/html/2609.19654#bib.bib26)\]\.
Itinerary Revision under Resource DisruptionsTravel requestqqAccepted itineraryPPFlightHotelAttractionResource disruptionsΔ\\DeltaFlight cancelled / hotel unavailable / attraction closedUpdated environment𝒦′\\mathcal\{K\}^\{\\prime\}PPmay become infeasibleqqunchangedRevisionIfPPis infeasibleFeasible revised itineraryP′P^\{\\prime\}Infeasibility diagnosis⊥\\botorFigure 1:Itinerary revision under resource disruptions, resulting in either a feasible revision or an infeasibility diagnosis\.Three paradigms differ in their use of the accepted itinerary:*full replanning*solves the updated problem anew,*classical plan repair*reuses plan structure, and*LLM\-based travel\-agent revision*edits the itinerary directly\[[10](https://arxiv.org/html/2609.19654#bib.bib14),[32](https://arxiv.org/html/2609.19654#bib.bib21),[13](https://arxiv.org/html/2609.19654#bib.bib5)\]\. We compare these scopes by effectiveness, preservation, and computational cost\.
Motivation\.Four research gaps motivate the comparison\.\(1\) Limited benchmark coverage for itinerary revision\.Travel\-planning benchmarks emphasize itinerary generation and broader planning capabilities\[[29](https://arxiv.org/html/2609.19654#bib.bib1),[21](https://arxiv.org/html/2609.19654#bib.bib6),[1](https://arxiv.org/html/2609.19654#bib.bib2),[2](https://arxiv.org/html/2609.19654#bib.bib3),[3](https://arxiv.org/html/2609.19654#bib.bib22)\]\. Despite work on adaptive planning and itinerary modification\[[16](https://arxiv.org/html/2609.19654#bib.bib9),[13](https://arxiv.org/html/2609.19654#bib.bib5)\], benchmark support for revising accepted itineraries under resource disruptions remains limited, particularly for evaluating both single and compound disruptions\.\(2\) Lack of a unified evaluation setting\.Full replanning, classical plan repair, and LLM\-based travel\-agent revision originate from different task formulations and evaluation settings, so previously reported results are not directly comparable\. A common revision interface, feasibility criterion, and controlled disruption setting are needed to compare them\.\(3\) Incomplete evaluation dimensions\.Success rates do not reveal how much of an accepted itinerary is preserved or at what computational cost\. High repair coverage may come with extensive changes, while local revision may incur high inference or runtime costs despite making few edits\.\(4\) Limited strategy\-selection guidance\.Existing evaluations provide limited guidance for choosing a revision paradigm when deployments prioritize feasibility recovery, plan preservation, and computational cost\.
Contributions\.We compare LLM–Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO\-based local revision adapter, with four contributions:
- •A benchmark for itinerary revision under resource disruptions\.Two fixed TREK\-derived benchmark sets\[[21](https://arxiv.org/html/2609.19654#bib.bib6)\]provide 500 single\-disruption cases, including feasible and infeasible instances, and 200 feasible compound\-disruption cases\. All methods receive the same instances\.
- •A unified evaluation setting for heterogeneous revision strategies\.A common interface takes the original request, accepted itinerary, disruptions, and updated environment, and assesses complete repairs or infeasibility outcomes against a common feasibility criterion\. Unsuccessful attempts count as failures; method\-specific restrictions and budgets are explicit\.
- •A multi\-dimensional evaluation framework\.Effectiveness covers feasible repair, overall success, and correct refusal; stability covers logical edits and preservation; computational cost covers latency, LLM calls, and tokens\.
- •Practical guidelines for revision\-strategy selection\.Observed trade\-offs characterize the strengths, limitations, and suitable deployment settings of the three paradigms for the evaluated implementations and settings\.
## 2Background and Problem Formulation
This section formalizes the travel\-itinerary revision problem considered in this study\. We define the planning environment, resource disruptions, revision outcomes, plan stability, and the feasibility\-first repair objective\.
### 2\.1Travel Itinerary Planning
A travel requestqqspecifies the user’s requirements and preferences for a trip, such as destinations, dates, and budget\. The planning environment𝒦\\mathcal\{K\}contains knowledge\-base records, resource availability, and execution conditions needed to assess executability\. An itineraryXXis a complete plan comprising flights or other transportation between destinations, accommodation, local transportation where applicable, and attractions or activities with temporal and spatial assignments\.
LetC\(q,𝒦\)C\(q,\\mathcal\{K\}\)be the set of applicable constraints\. Request constraints require the itinerary to satisfy the trip requirements inqq\. Grounding requires its resources and associated information to be supported by the planning environment\. Persona or preference constraints capture the traveller’s specified needs, while budget constraints bound the cost of the trip\. Temporal consistency concerns the timing and compatibility of scheduled components; spatial consistency concerns their locations and the connections between them\. These constraints determine the feasible set
ℱ\(q,𝒦\)=\{X∣c\(X\)=truefor everyc∈C\(q,𝒦\)\}\.\\mathcal\{F\}\(q,\\mathcal\{K\}\)=\\\{X\\mid c\(X\)=\\mathrm\{true\}\\ \\text\{for every \}c\\in C\(q,\\mathcal\{K\}\)\\\}\.\(1\)Initial itinerary generation seeks any completeX∈ℱ\(q,𝒦\)X\\in\\mathcal\{F\}\(q,\\mathcal\{K\}\)\. Once accepted, the itinerary is denoted byPP, withP∈ℱ\(q,𝒦\)P\\in\\mathcal\{F\}\(q,\\mathcal\{K\}\)before disruption\.
### 2\.2Resource Disruptions and Revision Outcomes
A resource is a travel entity used in an itinerary; this study considers disruptions to flights, hotels, and attractions\. A disruptionδ\\deltamakes a previously available resource unavailable\. LetΔ=\{δ1,…,δk\}\\Delta=\\\{\\delta\_\{1\},\\ldots,\\delta\_\{k\}\\\}denote the set of disruptions:\|Δ\|=1\|\\Delta\|=1is a*single disruption*, and\|Δ\|\>1\|\\Delta\|\>1is a*compound disruption*\. Compound disruptions in this study are presented simultaneously, rather than arriving as a sequence of online changes\. They produce the updated environment
𝒦′=Update\(𝒦,Δ\)\.\\mathcal\{K\}^\{\\prime\}=\\operatorname\{Update\}\(\\mathcal\{K\},\\Delta\)\.
The requestqqremains unchanged, but disruptions may renderPPinfeasible in𝒦′\\mathcal\{K\}^\{\\prime\}through temporal and spatial dependencies\. For example, a replacement flight may shift the arrival time and prevent a scheduled attraction visit\.
The common revision interface takes the original request, accepted itinerary, disruptions, and updated environment:
f\(q,P,Δ,𝒦′\)⟶P′,⊥,orfail\.f\(q,P,\\Delta,\\mathcal\{K\}^\{\\prime\}\)\\longrightarrow P^\{\\prime\},\\ \\bot,\\ \\text\{or\}\\ \\mathrm\{fail\}\.\(2\)HereP′P^\{\\prime\}denotes a complete revised itinerary; it is a successful repair only ifP′∈ℱ\(q,𝒦′\)P^\{\\prime\}\\in\\mathcal\{F\}\(q,\\mathcal\{K\}^\{\\prime\}\)\. The outcome⊥\\botdenotes a correct infeasibility diagnosis, requiringℱ\(q,𝒦′\)=∅\\mathcal\{F\}\(q,\\mathcal\{K\}^\{\\prime\}\)=\\varnothing\. The outcomefail\\mathrm\{fail\}records an attempt that produces neither a valid repair nor a justified infeasibility outcome within the method’s execution procedure or budget\. Examples include timeouts, execution errors, exhausted search or correction budgets, and malformed or invalid candidates\.
Failure to find a feasible repair does not establish infeasibility\. The original user requirements are not silently relaxed to make the updated problem feasible\.
### 2\.3Plan Stability
Feasible revisions can differ in how much ofPPthey retain\. Plan stability is assessed through*logical itinerary components*, or*logical slots*, each representing a semantically meaningful commitment rather than a raw JSON field or textual string\. Examples include a flight assigned to a route, a hotel booking for a city, and an attraction assignment for a day\.
The logical edit measured\(P,P′\)d\(P,P^\{\\prime\}\)counts semantic itinerary changes, not textual edits\. LetS\(P\)S\(P\)denote the logical slots ofPP\. Preservation is the fraction of slots inS\(P\)S\(P\)whose assignments remain unchanged inP′P^\{\\prime\}\. Fewer edits and higher preservation generally indicate greater stability\. The measures are complementary rather than mathematically equivalent\. Section[4\.1](https://arxiv.org/html/2609.19654#S4.SS1)specifies the benchmark’s matching and counting rules and reports stability on successful feasible repairs\.
### 2\.4Feasibility\-First Repair Objective
Feasibility is the primary requirement: retaining accepted commitments does not compensate for an invalid itinerary\. Whenℱ\(q,𝒦′\)\\mathcal\{F\}\(q,\\mathcal\{K\}^\{\\prime\}\)is nonempty, an ideal minimum\-edit revision satisfies
P∗∈argminX∈ℱ\(q,𝒦′\)d\(P,X\)\.P^\{\*\}\\in\\underset\{X\\in\\mathcal\{F\}\(q,\\mathcal\{K\}^\{\\prime\}\)\}\{\\arg\\min\}\\ d\(P,X\)\.\(3\)This conceptual preference favors fewer changes among feasible alternatives; it is neither an objective implemented by all methods nor a guarantee of globally minimum\-edit repair\.
## 3Related Work and Representative Strategy Selection
Existing work relevant to itinerary revision can be organized into three broad lines according to how an accepted plan is used after the environment changes\. First, travel\-planning and solver\-assisted approaches typically formulate the updated request as a new planning problem and generate a complete feasible itinerary, with limited emphasis on preserving the previously accepted plan\[[29](https://arxiv.org/html/2609.19654#bib.bib1),[10](https://arxiv.org/html/2609.19654#bib.bib14)\]\. Second, classical plan\-repair methods explicitly reuse an existing plan, retaining prior assignments or structural decompositions while revising the parts affected by changed goals or execution conditions\[[7](https://arxiv.org/html/2609.19654#bib.bib20),[32](https://arxiv.org/html/2609.19654#bib.bib21)\]\. Third, recent LLM\-based itinerary\-modification methods operate directly on the current itinerary through semantic editing operations rather than reconstructing a formal plan\[[13](https://arxiv.org/html/2609.19654#bib.bib5)\]\. These lines differ mainly in their use of the accepted itinerary and the scope of revision\. We therefore group the literature into*full replanning*,*classical plan repair*, and*LLM\-based travel\-agent revision*, and evaluate one representative per category\.
### 3\.1Full Replanning: LLM–Z3
Full replanning seeks a complete feasible itinerary for an updated problem without necessarily preservingPP\. TravelPlanner\[[29](https://arxiv.org/html/2609.19654#bib.bib1)\]and AgentTravel\[[34](https://arxiv.org/html/2609.19654#bib.bib12)\]address itinerary generation and knowledge\-augmented planning\. External reasoning supports constraint satisfaction through verification in LLM\-Modulo\[[9](https://arxiv.org/html/2609.19654#bib.bib13)\], mixed\-integer optimization in To the Globe\[[15](https://arxiv.org/html/2609.19654#bib.bib16)\], formalized programming in LLMFP\[[11](https://arxiv.org/html/2609.19654#bib.bib15)\], and LLM–solver integration in Personal Travel Solver\[[23](https://arxiv.org/html/2609.19654#bib.bib23)\]\. RETAIL\[[6](https://arxiv.org/html/2609.19654#bib.bib17)\]and ATLAS\[[4](https://arxiv.org/html/2609.19654#bib.bib18)\]address complex or changing requirements without centering accepted\-itinerary preservation\.
Hao et al\.\[[10](https://arxiv.org/html/2609.19654#bib.bib14)\]translate requests into executable steps and code invoking a Satisfiability Modulo Theories \(SMT\) solver\. Formal solving separates constraint satisfaction from language generation and supports full replanning without a preservation objective\.
Our*LLM–Z3 Full Replan*adapter uses a fixed Python/Z3 model\[[5](https://arxiv.org/html/2609.19654#bib.bib7)\]rather than generated solver code\. One LLM call extracts structured requirements; the model solves updated resource constraints and decodesP′P^\{\\prime\}\. It does not optimize similarity toPP, which is retained for post\-hoc stability evaluation\.
### 3\.2Classical Plan Repair: IPyHOPPER
Classical repair reuses plans under changed goals, constraints, or execution conditions\. Related work spans assignment reuse in dynamic constraint satisfaction\[[27](https://arxiv.org/html/2609.19654#bib.bib19)\], localized plan adaptation\[[8](https://arxiv.org/html/2609.19654#bib.bib24)\], planning\-based repair\[[26](https://arxiv.org/html/2609.19654#bib.bib10)\], and plan stability and minimum\-change repair\[[7](https://arxiv.org/html/2609.19654#bib.bib20),[22](https://arxiv.org/html/2609.19654#bib.bib11)\]\. Hierarchical Task Network \(HTN\) planning decomposes tasks into subtasks and primitive actions, providing structure for repair methods including SHOP\-FIXER, IPyHOPPER, and REWRITE\[[32](https://arxiv.org/html/2609.19654#bib.bib21)\]\.
IPyHOPPER\[[32](https://arxiv.org/html/2609.19654#bib.bib21)\]reuses plan structure while adapting repair scope\. Given decomposition treeTTand primitive sequenceπ=plan\(T\)\\pi=\\operatorname\{plan\}\(T\), simulation identifies a failed action and replaces an ancestor task’s decomposition\. If the result is inapplicable, backtracking expands repair to a higher\-level task\. This reuse does not guarantee minimum edits\.
Our adapter reconstructsTTfromPPbefore activatingΔ\\Delta\. Updated conditions identify the affected action, and the repaired primitive sequence is decoded intoP′P^\{\\prime\}\. Candidates target disrupted slots; the worker uses domain constraints without the benchmark scorer\. Simulation is internal and does not imply that the traveller has executed a trip prefix\.
### 3\.3LLM\-Based Travel\-Agent Revision: iTIMO
LLM\-based revision edits an itinerary semantically rather than reconstructing a formal plan or repairing an HTN hierarchy\. iTIMO\[[13](https://arxiv.org/html/2609.19654#bib.bib5)\]provides an operation\-based formulation over points of interest \(POIs\):
𝒪=\{oadd,oreplace,odelete\}\.\\mathcal\{O\}=\\\{o\_\{\\mathrm\{add\}\},o\_\{\\mathrm\{replace\}\},o\_\{\\mathrm\{delete\}\}\\\}\.\(4\)These operations add, replace, or delete a POI\. The source benchmark evaluates POI\-level perturbations based on popularity, spatial distance, and category diversity\. iTIMO’s local modification task motivates our editing baseline; its formulation is not a general resource\-disruption repair algorithm\.
Our iTIMO adapter receivesqq,PP,Δ\\Delta, and𝒦′\\mathcal\{K\}^\{\\prime\}and selects operations and targets, leaving untouched components unchanged\. Unlike the source’s single POI\-level modification, it permits corrections for simultaneous disruptions within the budget in Section[4\.1](https://arxiv.org/html/2609.19654#S4.SS1)\. Corrections modify the preceding result and do not represent new arrivals ofΔ\\Delta; the adapter remains a bounded local\-repair baseline\.
Table 1:Conceptual comparison of the evaluated revision strategies\.
### 3\.4Comparison of Representative Strategies
Table[1](https://arxiv.org/html/2609.19654#S3.T1)highlights the main differences among the three representative strategies\. Full replanning treats the updated problem globally and does not use the accepted itineraryPPas a repair structure, whereas IPyHOPPER explicitly reuses the decomposition derived fromPPand adapts the repair scope hierarchically\. The iTIMO adapter instead editsPPdirectly through bounded local operations\. These differences yield three distinct revision scopes—global, adaptive hierarchical, and bounded local—and motivate their comparative evaluation under a common revision setting\.
## 4Experimental Evaluation
Figure[2](https://arxiv.org/html/2609.19654#S4.F2)summarizes the unified evaluation setting\. We evaluate the three representative strategies using identical fixed instances under both single and simultaneous compound resource disruptions\. The evaluation covers*effectiveness*,*plan stability*, and computational cost, capturing repair coverage, preservation of accepted commitments, and execution or inference overhead\.
Unified Evaluation of Representative Revision StrategiesCommon revision instance\(q,P,Δ,𝒦′\)\(q,P,\\Delta,\\mathcal\{K\}^\{\\prime\}\)Same fixed instance for all methods; method\-specific representations500 single\-disruption cases375 feasible \+ 125 infeasible200 compound\-disruption casesAll 200 feasibleLLM–Z3Full ReplanningIPyHOPPERHierarchical RepairiTIMOLocal RevisionGLOBALADAPTIVE HIERARCHICALBOUNDED LOCALSolve updated problemComplete planP′P^\{\\prime\}Formal constraint solvingNo similarity objective onPPReconstruct hierarchy fromPPHTN repair \+ backtrackingRepair may expand upwardDirect semantic edits toPPHHAAF′F^\{\\prime\}ADD / REPLACE / DELETEPreserve untouchedcomponentsCommon EvaluationEffectivenessrepair / correct refusal /overall successPlan Stabilitylogical edits / preservationComputational Costcalls / tokens / latencyFigure 2:Unified evaluation of the three representative revision strategies across effectiveness, plan stability, and computational cost\.### 4\.1Experimental Setup
The experimental setup specifies the benchmark instances, evaluation metrics, and implementation settings used for all compared strategies\.
Benchmark and tasks\.All methods receive identical fixed TREK\-derived instances\[[21](https://arxiv.org/html/2609.19654#bib.bib6)\]: a request, accepted itinerary, and disruptions in an updated environment\. Reference repairs and feasibility labels are hidden\.
The 500 single\-disruption cases comprise 375 feasible flight, optional\-hotel, and optional\-attraction disruptions and 125 infeasible required\-hotel conflicts, testing repair and infeasibility detection under unchanged requirements\. The 200 feasible compound cases comprise 50 each of AB, AC, BC, and ABC \(A/B/C: flight/hotel/attraction disruptions\), with multiple components invalidated simultaneously\. The benchmarks share 191 source queries but are reported separately; compound groups use different itineraries without difficulty matching\.
Evaluation metrics\.We evaluate each configuration along three complementary dimensions: effectiveness, plan stability, and efficiency\. Together, these metrics capture whether a revision succeeds, how much of the accepted itinerary it preserves, and the computational cost required to obtain it\.
*Effectiveness\.*Valid repairs are complete itineraries passing applicable TREK checks for requirements, grounding, persona, budget, and spatio\-temporal consistency, and avoiding disrupted resources\. Single\-disruption*feasible repair rate*,*correct refusal rate*, and*overall success rate*are percentages: valid repairs divided by 375, correct refusals divided by 125, and their sum divided by 500, respectively\. For the 125 required\-hotel conflicts, correct refusal checks recognition of the constructed conflict, not a general\-purpose proof of infeasibility\. Errors, invalid outputs, and timeouts in retained final records count as failures in the denominators, not correct refusals\. All compound cases are feasible: overall success equals repair rate over 200 cases; refusal is inapplicable\.
*Plan stability\.*Over successful feasible repairs,*mean logical edits*averagesd\(P,P′\)d\(P,P^\{\\prime\}\);*preservation rate*averages the percentage of original logical slots retaining the same resource or value at the corresponding key\. Flights are keyed by route, hotels and cars by city, and city assignments and attraction sets by day\. Daily attraction edits count asmax\(\|removed\|,\|added\|\)\\max\(\|\\mathrm\{removed\}\|,\|\\mathrm\{added\}\|\)\. These implement Section[2\.3](https://arxiv.org/html/2609.19654#S2.SS3)’s semantic measures\. Table rows use each configuration’s successful feasible subset; paired comparisons use joint successes\.
*Efficiency\.*Final iTIMO releases replaced 283 DeepSeek single\-disruption provider\-balance errors via targeted reruns and 43 Gemini compound\-disruption execution\-error records through official\-endpoint reruns\. Metrics describe retained final outputs; discarded attempts are excluded from consumption aggregation\. Where records are available, mean calls and tokens cover all retained cases, including failures\. Qwen iTIMO figures use reported summaries\. End\-to\-end latency is summarized in seconds by mean, median, and 95th percentile \(P95\), linearly interpolated where raw records are available\. IPyHOPPER’s zero calls/tokens mean no LLM inference, not zero computation\.
Implementation settings\.LLM–Z3 and the iTIMO adapter use Gemini 3\.8 Flash, DeepSeek V4 Flash, and Qwen 3\.8 Flash service aliases\. LLM–Z3 uses one call per case, temperature 0, a 2,048\-token output limit, no retry, and disabled thinking for DeepSeek/Qwen\. It checks up to 25 resource combinations without a similarity objective\. IPyHOPPER uses no LLM, scorer\-free disrupted\-slot candidates, a 300\-second limit, and no retry\. The iTIMO adapter permits two planning rounds for one disruption andk\+1k\+1forkksimultaneous disruptions; format correction may add calls\. Documented Gemini/DeepSeek settings use temperature 0 and a 2,048\-token output limit\. Qwen iTIMO adapter raw records, sampling settings, and P95 interpolation details are unavailable\.
### 4\.2Single\-Disruption Evaluation
Design\.The single\-disruption experiment evaluates isolated resource failures using 375 feasible repair cases and 125 infeasible cases, with effectiveness, stability, and efficiency assessed jointly\.
Table 2:Single\-disruption effectiveness and plan stability\. Stability metrics use successful feasible repairs\.Effectiveness and plan stability\.Table[2](https://arxiv.org/html/2609.19654#S4.T2)gives effectiveness in its first three metric columns and stability, conditional on successful feasible repairs, in its final two\. LLM–Z3 with Gemini achieved 485/500 overall successes \(97\.0%\) versus IPyHOPPER’s 484/500 \(96\.8%\); this one\-case difference alone supports no meaningful ranking\. With DeepSeek, LLM–Z3 repaired 338/375 cases and correctly refused 85/125 conflicts, whereas the iTIMO adapter repaired 318/375 and refused all 125\. Overall success therefore reversed their feasible\-repair ordering\.
Full replanning averaged about 6\.6 edits and 65% preservation, versus one edit and 92% for successful IPyHOPPER and iTIMO adapter repairs\. On the 346 jointly successful Gemini LLM–Z3/IPyHOPPER cases, respective means were 6\.809 versus 1\.000 edits and 64\.70% versus 92\.68% preservation\. Restricting to joint successes removes differences from the methods succeeding on different case subsets; the comparison remains conditional on both returning valid repairs\.
Table 3:Single\-disruption efficiency over 500 cases\. Calls and tokens are per\-case means; latency is in seconds\.Efficiency\.Table[3](https://arxiv.org/html/2609.19654#S4.T3)reports per\-case inference consumption and latency, subject to the record\-availability qualification above\. LLM–Z3 uses one call and 938–1,369 tokens per case\. The iTIMO adapter averaged 1\.132–1\.430 calls and 19,486–24,205 tokens despite approximately one logical edit per successful repair\. IPyHOPPER required no LLM inference and recorded a 1\.98\-second median\. Few edits do not imply low inference consumption; latency describes these implementations and environments, not intrinsic algorithm speed\.
Summary\.LLM–Z3 with Gemini and IPyHOPPER achieved nearly identical overall success\. Successful hierarchical and local repairs preserved more accepted commitments than full replanning, while computational costs differed substantially across strategies\.
### 4\.3Compound\-Disruption Evaluation
Design\.The 200 feasible compound cases comprise 50 each of AB, AC, BC, and ABC, representing simultaneous flight–hotel, flight–attraction, hotel–attraction, and flight–hotel–attraction disruptions; refusal is therefore inapplicable\.
Table 4:Compound\-disruption effectiveness and plan stability over 200 feasible cases\. Stability metrics use successful repairs\.Effectiveness and plan stability\.Table[4](https://arxiv.org/html/2609.19654#S4.T4)reports overall success over 200 cases, combination\-level successful counts out of 50, and stability on successful repairs\. LLM–Z3 with Gemini repaired 186/200 cases \(93\.0%\), the highest observed success, followed by IPyHOPPER at 178/200 \(89\.0%\)\. The iTIMO adapter repaired 134, 151, and 134 cases with Gemini, DeepSeek, and Qwen\. IPyHOPPER and LLM–Z3 with Gemini both repaired 47/50 AB and 49/50 BC cases; the iTIMO adapter with Gemini repaired 46/50 AB but 19/50 ABC cases\. These differences do not establish a causal effect of disruption count because the groups use different itineraries\.
Full replanning averaged about 9\.0 edits and 57% preservation, versus 2\.1–2\.2 edits and 86%–87% for hierarchical/local repair\. On 167 jointly successful Gemini LLM–Z3/IPyHOPPER cases, respective means were 9\.132 versus 2\.222 edits and 57\.28% versus 86\.77% preservation\. The paired comparison fixes the successful case subset and does not describe stability when either method fails\.
Table 5:Compound\-disruption efficiency over 200 cases\. Calls and tokens are per\-case means; latency is in seconds\.Efficiency\.Table[5](https://arxiv.org/html/2609.19654#S4.T5)reports per\-case calls, tokens, and latency under the same provenance qualifications\. LLM–Z3 used one call and averaged 1,086–1,594 tokens; the iTIMO adapter averaged 1\.660–2\.905 calls and 34,843–67,631 tokens\. IPyHOPPER required no LLM inference, but its 2\.32\-second median accompanied a 20\.23\-second mean, 300\-second P95, and 12/200 timeouts\. Median latency alone omits this tail; execution differences preclude an intrinsic speed ranking\.
Summary\.LLM–Z3 with Gemini achieved the highest observed compound success\. Hierarchical and local repairs preserved more accepted commitments, while IPyHOPPER showed a long latency tail and the iTIMO adapter higher token consumption\.
## 5Practical Guidelines for Revision Strategy Selection
No strategy dominates all evaluated dimensions\. Table[6](https://arxiv.org/html/2609.19654#S5.T6)summarizes trade\-offs within the evaluated setting\.
Table 6:Strategy\-selection guidance within the evaluated setting\.LLM–Z3: Feasibility Recovery, Compact Inference\.LLM–Z3’s observed coverage favors feasibility recovery when broad revisions are acceptable\. Its larger changes are consistent with the absence of a preservation objective, and coverage remains configuration\-dependent\.
IPyHOPPER: Strong Preservation, No LLM Inference\.IPyHOPPER is attractive when preservation is prioritized and a usable hierarchy is available, requiring no LLM inference\. Its compound\-disruption latency tail matters under strict runtime requirements\.
iTIMO: Local Editing, Higher Inference Cost\.The iTIMO adapter supports preservation\-oriented local editing, but showed lower compound coverage and substantially higher token consumption in the evaluated configurations; locality was not independently varied\.
These guidelines are limited to the evaluated implementations and synthetic resource\-disruption setting\. Repair scope is confounded with feedback, candidate restrictions, model routing, and budgets, so differences cannot be attributed solely to paradigms; no method is shown to return globally minimum\-editP∗P^\{\*\}\. Refusal covers required\-hotel conflicts only, while stability is conditioned on successful repairs and joint successes for paired comparisons\. Compound groups use different itineraries and are not difficulty\-matched\. Runtime and token results depend on implementations, environments, and service endpoints\. Exact historical equivalence for iTIMO cannot be established because its external knowledge\-base/scorer checkout was not fingerprinted; Qwen iTIMO also lacks raw records, sampling settings, and P95 interpolation details\.
## 6Conclusion
We evaluated LLM–Z3 full replanning, IPyHOPPER hierarchical repair, and iTIMO\-based local revision under single and compound resource disruptions\. LLM–Z3 with Gemini achieved the highest observed compound success, while IPyHOPPER reached comparable single\-disruption overall success with greater preservation\. Hierarchical and local repairs made fewer edits than full replanning, while the strategies exhibited distinct computational costs\. No strategy dominates all dimensions; selection should balance feasibility recovery, preservation, and computational cost\.
## References
- \[1\]S\. Chaudhuri, P\. Purkar, R\. Raghav, S\. Mallick, M\. Gupta, A\. Jana, and S\. Ghosh\(2025\)TripCraft: a benchmark for spatio\-temporally fine grained travel planning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 17035–17064\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.834)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p4.1)\.
- \[2\]W\. Chen, S\. Wang, Z\. Gao, K\. Hu, W\. Ni, S\. Di, C\. J\. Zhang, and L\. Chen\(2026\)TravelEval: a comprehensive benchmarking framework for evaluating LLM\-powered travel planning agents\.InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\. 2,pp\. 8719–8730\.External Links:[Document](https://dx.doi.org/10.1145/3770855.3817533)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p4.1)\.
- \[3\]X\. Cheng, Y\. Hu, X\. Zhang, L\. Xu, L\. Tan, Z\. Pan, X\. Li, and Y\. Liu\(2026\)Beyond itinerary planning—a real\-world benchmark for multi\-turn and tool\-using travel tasks\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),San Diego, California, United States,pp\. 29200–29251\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1347),[Link](https://aclanthology.org/2026.acl-long.1347/)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p4.1)\.
- \[4\]J\. Choi, J\. Yoon, J\. Chen, S\. Jha, and T\. Pfister\(2026\)ATLAS: constraints\-aware multi\-agent collaboration for real\-world travel planning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=mIYGiBf9Pm)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[5\]L\. de Moura and N\. Bjørner\(2008\)Z3: an efficient SMT solver\.InTools and Algorithms for the Construction and Analysis of Systems,Lecture Notes in Computer Science, Vol\.4963,pp\. 337–340\.External Links:[Document](https://dx.doi.org/10.1007/978-3-540-78800-3%5F24)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p3.1)\.
- \[6\]B\. Deng, Y\. Feng, Z\. Liu, Q\. Wei, X\. Zhu, S\. Chen, Y\. Guo, and Y\. Wang\(2025\)RETAIL: towards real\-world travel planning for large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 14870–14902\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.752/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.752)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[7\]M\. Fox, A\. Gerevini, D\. Long, and I\. Serina\(2006\)Plan stability: replanning versus plan repair\.InProceedings of the 16th International Conference on Automated Planning and Scheduling \(ICAPS\),pp\. 212–221\.Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1),[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1),[§3](https://arxiv.org/html/2609.19654#S3.p1.1)\.
- \[8\]A\. Gerevini and I\. Serina\(2000\)Fast plan adaptation through planning graphs: local and systematic search techniques\.InProceedings of the Fifth International Conference on Artificial Intelligence Planning Systems \(AIPS 2000\),pp\. 112–121\.Cited by:[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1)\.
- \[9\]A\. Gundawar, K\. Valmeekam, M\. Verma, and S\. Kambhampati\(2024\)Robust planning with compound LLM architectures: an LLM\-Modulo approach\.Note:arXiv preprint arXiv:2411\.14484External Links:[Link](https://arxiv.org/abs/2411.14484)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[10\]Y\. Hao, Y\. Chen, Y\. Zhang, and C\. Fan\(2025\)Large language models can solve real\-world planning rigorously with formal verification tools\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 3434–3483\.External Links:[Link](https://aclanthology.org/2025.naacl-long.176/),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.176)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1),[§1](https://arxiv.org/html/2609.19654#S1.p3.1),[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p2.1),[§3](https://arxiv.org/html/2609.19654#S3.p1.1)\.
- \[11\]Y\. Hao, Y\. Zhang, and C\. Fan\(2025\)Planning anything with rigor: general\-purpose zero\-shot planning with LLM\-based formalized programming\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0K1OaL6XuK)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[12\]C\. Huan, Z\. Meng, Y\. Liu, Z\. Yang, Y\. Zhu, Y\. Yun, S\. Li, R\. Gu, X\. Wu, H\. Zhang, C\. Hong, S\. Ma, G\. Chen, and C\. Tian\(2025\)Scaling graph chain\-of\-thought reasoning: a multi\-agent framework with efficient LLM serving\.arXiv preprint arXiv:2511\.01633\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.01633),[Link](https://arxiv.org/abs/2511.01633)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[13\]Z\. Huang, Y\. Ma, H\. Zhang, H\. Ma, and Z\. Sun\(2026\)iTIMO: an LLM\-empowered synthesis dataset for travel itinerary modification\.arXiv preprint arXiv:2601\.10609\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.10609)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p3.1),[§1](https://arxiv.org/html/2609.19654#S1.p4.1),[§3\.3](https://arxiv.org/html/2609.19654#S3.SS3.p1.1),[§3](https://arxiv.org/html/2609.19654#S3.p1.1)\.
- \[14\]R\. Jiang, J\. Wang, G\. Zhao, C\. Luo, K\. Wang, and W\. Zhang\(2026\)Advancing multimodal agent reasoning with long\-term neuro\-symbolic memory\.arXiv preprint arXiv:2603\.15280\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.15280),[Link](https://arxiv.org/abs/2603.15280)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1)\.
- \[15\]D\. Ju, S\. Jiang, A\. Cohen, A\. Foss, S\. Mitts, A\. Zharmagambetov, B\. Amos, X\. Li, J\. T\. Kao, M\. Fazel\-Zarandi, and Y\. Tian\(2024\)To the Globe \(TTG\): towards language\-driven guaranteed travel planning\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,pp\. 240–249\.External Links:[Link](https://aclanthology.org/2024.emnlp-demo.25/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-demo.25)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[16\]P\. Karmakar, S\. Chaudhuri, S\. Mallick, M\. Gupta, A\. Jana, and S\. Ghosh\(2026\)TripTide: a benchmark for adaptive travel planning under disruptions\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 40269–40292\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.2002)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1),[§1](https://arxiv.org/html/2609.19654#S1.p4.1)\.
- \[17\]L\. Lai, C\. Luo, Y\. Lou, M\. Ju, and Z\. Yang\(2025\)Graphy’our data: towards end\-to\-end modeling, exploring and generating report from raw data\.InCompanion of the 2025 International Conference on Management of Data,SIGMOD/PODS ’25,pp\. 147–150\.External Links:[Document](https://dx.doi.org/10.1145/3722212.3725106),[Link](https://doi.org/10.1145/3722212.3725106)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[18\]J\. Liu, L\. Chen, Z\. Yang, C\. He, M\. Ju, B\. Han, R\. Liu, and X\. Zhou\(2026\)HyperSU: corpus\-driven semantic\-unit hypergraph for retrieval\-augmented generation\.arXiv preprint arXiv:2606\.28351\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.28351),[Link](https://arxiv.org/abs/2606.28351)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[19\]J\. Liu, Z\. Chen, S\. Qiao, M\. Ju, D\. Zhang, B\. Han, S\. Yu, X\. Shu, J\. Wu, D\. Wen, X\. Cao, G\. Liu, and Z\. Yang\(2026\)A2RAG: adaptive agentic graph retrieval for cost\-aware and reliable reasoning\.In2026 IEEE 42nd International Conference on Data Engineering Workshops \(ICDEW\),pp\. 187–196\.External Links:[Document](https://dx.doi.org/10.1109/ICDEW71238.2026.00024),[Link](https://doi.org/10.1109/ICDEW71238.2026.00024)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[20\]J\. Oh, E\. Kim, and A\. Oh\(2025\)Flex\-TravelPlanner: a benchmark for flexible planning with language agents\.arXiv preprint arXiv:2506\.04649\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.04649)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1)\.
- \[21\]J\. Qi, W\. Zhang, S\. M\. Ng, F\. Xu, Y\. Chen, Y\. Li, and I\. King\(2026\)TREK: a travel reasoning and evaluation kit for LLM agents in complex trip planning\.arXiv preprint arXiv:2607\.26977\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2607.26977)Cited by:[1st item](https://arxiv.org/html/2609.19654#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2609.19654#S1.p4.1),[§4\.1](https://arxiv.org/html/2609.19654#S4.SS1.p2.1)\.
- \[22\]A\. Saetti and E\. Scala\(2025\)Optimally stable plan repair\.The Knowledge Engineering Review40,pp\. e7\.External Links:[Document](https://dx.doi.org/10.1017/S0269888925100076)Cited by:[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1)\.
- \[23\]Z\. Shao, J\. Wu, W\. Chen, and X\. Wang\(2025\)Personal Travel Solver: a preference\-driven LLM\-solver system for travel planning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Vienna, Austria,pp\. 27622–27642\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1339),[Link](https://aclanthology.org/2025.acl-long.1339/)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[24\]X\. Shu, M\. Ju, Z\. Chen, Y\. Ding, W\. Zhang, D\. Wen, and Z\. Yang\(2026\)ForexAgent: identifying trading strategies in forex markets with large language models\.In2026 IEEE International Conference on Big Data and Smart Computing \(BigComp\),pp\. 55–62\.External Links:[Document](https://dx.doi.org/10.1109/BigComp68355.2026.00019),[Link](https://doi.org/10.1109/BigComp68355.2026.00019)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[25\]X\. Tang, L\. Chen, W\. Yang, Z\. Yang, M\. Ju, X\. Shu, Z\. Yang, and Y\. Tang\(2025\)Tabular\-textual question answering: from parallel program generation to large language models\.World Wide Web28\(4\),pp\. 42\.External Links:[Document](https://dx.doi.org/10.1007/s11280-025-01351-1),[Link](https://doi.org/10.1007/s11280-025-01351-1)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[26\]R\. P\. J\. van der Krogt and M\. M\. de Weerdt\(2005\)Plan repair as an extension of planning\.InProceedings of the Fifteenth International Conference on Automated Planning and Scheduling \(ICAPS 2005\),pp\. 161–170\.External Links:[Document](https://dx.doi.org/10.5555/3037062.3037083)Cited by:[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1)\.
- \[27\]G\. Verfaillie and T\. Schiex\(1994\)Solution reuse in dynamic constraint satisfaction problems\.InProceedings of the Twelfth National Conference on Artificial Intelligence \(AAAI\),pp\. 307–312\.Cited by:[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1)\.
- \[28\]J\. Wu, X\. Tang, Z\. Yang, K\. Hao, L\. Lai, and Y\. Liu\(2025\)An experimental evaluation of LLM on image classification\.InDatabases Theory and Applications,T\. Chen, Y\. Cao, Q\. V\. H\. Nguyen, and T\. T\. Nguyen \(Eds\.\),Singapore,pp\. 506–518\.External Links:ISBN 978\-981\-96\-1242\-0,[Document](https://dx.doi.org/10.1007/978-981-96-1242-0%5F37)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[29\]J\. Xie, K\. Zhang, J\. Chen, T\. Zhu, R\. Lou, Y\. Tian, Y\. Xiao, and Y\. Su\(2024\)TravelPlanner: a benchmark for real\-world planning with language agents\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 54590–54613\.External Links:[Link](https://proceedings.mlr.press/v235/xie24j.html)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1),[§1](https://arxiv.org/html/2609.19654#S1.p4.1),[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1),[§3](https://arxiv.org/html/2609.19654#S3.p1.1)\.
- \[30\]W\. Yang, Z\. Yang, L\. Chen, R\. Yan, Z\. Yang, L\. Zhang, and Y\. Tang\(2024\)Parallel program generation for hybrid tabular\-textual question answering\.InWeb and Big Data: 8th International Joint Conference, APWeb\-WAIM 2024, Proceedings, Part I,Lecture Notes in Computer Science, Vol\.14961,pp\. 121–137\.External Links:[Document](https://dx.doi.org/10.1007/978-981-97-7232-2%5F9)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[31\]Z\. Yang, W\. Yang, G\. Liu, and L\. Qin\(2026\)RAIDS: rethinking data systems as responsible intelligent infrastructure\.arXiv preprint arXiv:2606\.21831\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.21831),[Link](https://arxiv.org/abs/2606.21831)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p2.1)\.
- \[32\]P\. Zaidins, R\. P\. Goldman, U\. Kuter, D\. S\. Nau, and M\. Roberts\(2025\)HTN plan repair algorithms compared: strengths and weaknesses of different methods\.InProceedings of the International Conference on Automated Planning and Scheduling,Vol\.35,pp\. 297–305\.External Links:[Link](https://ojs.aaai.org/index.php/ICAPS/article/view/36131),[Document](https://dx.doi.org/10.1609/icaps.v35i1.36131)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p3.1),[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p1.1),[§3\.2](https://arxiv.org/html/2609.19654#S3.SS2.p2.1),[§3](https://arxiv.org/html/2609.19654#S3.p1.1)\.
- \[33\]C\. Zhang, X\. D\. Goh, D\. Li, H\. Zhang, and Y\. Liu\(2025\)Planning with multi\-constraints via collaborative language agents\.InProceedings of the 31st International Conference on Computational Linguistics \(COLING 2025\),pp\. 10054–10082\.External Links:[Link](https://aclanthology.org/2025.coling-main.672/)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1)\.
- \[34\]J\. Zhao, J\. Feng, and Y\. Li\(2025\)AgentTravel: knowledge\-augmented LLM agent framework for urban travel planning\.InProceedings of the 1st Workshop on Knowledge Graphs & Agentic Systems Interplay \(NORA’25\),CEUR Workshop Proceedings, Vol\.4162,pp\. 55–67\.External Links:[Link](https://ceur-ws.org/Vol-4162/paper5.pdf)Cited by:[§3\.1](https://arxiv.org/html/2609.19654#S3.SS1.p1.1)\.
- \[35\]Z\. Zhao, S\. Wang, Y\. Hou, Y\. Xu, Y\. Sheng, X\. Xie, W\. Zhang, W\. Shin, and X\. Cao\(2026\)TRACE: tourism recommendation with accountable citation evidence\.arXiv preprint arXiv:2605\.07677\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.07677),[Link](https://arxiv.org/abs/2605.07677)Cited by:[§1](https://arxiv.org/html/2609.19654#S1.p1.1)\.Similar Articles
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Introduces TREK, a benchmark for evaluating LLM agents on complex travel planning tasks with deterministic scoring, covering 800 multi-constraint tasks over a synthetic knowledge base. The benchmark reveals that even strong agents struggle with unstated user needs.
When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
The ToolMaze benchmark evaluates LLM agents' ability to handle real-world tool failures, revealing that implicit semantic failures cause the largest performance drops and that dynamic replanning remains a critical bottleneck not addressed by scaling or prompting.
AI Tour Meeting: Group Travel Planning by LLM Agents
This paper proposes AI Tour Meeting, a group travel planning framework that uses multiple LLM-based agents with distinct personas to collaboratively find itineraries through natural language discussion.
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation
The paper proposes the Plan, Learn, Adapt (PLA) framework for personalized on-device itinerary generation, combining feasibility-guaranteed combinatorial planning with human preference learning via a Bradley-Terry reward model. In deployment, it achieved a 91% increase in itinerary completion rates with low latency, outperforming frontier LLMs in feasibility.
When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations
This paper proposes a difficulty-routed service-control architecture for autonomous customer-service agents, routing routine requests to a low-cost baseline and operationally coupled sessions to an escalated workflow with conflict-aware communication and write-triggered reconsideration. Evaluated on retail and airline tasks, the approach improves reliability on conflicted requests without over-allocating resources to routine ones.