Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
Summary
This arXiv paper from Walmart Global Tech presents a context-aware multi-agent framework that automates the construction of 'Lines and Ladders' pricing taxonomies for large-scale retail catalogs, achieving strong F1 and precision scores in production.
View Cached Full Text
Cached at: 08/14/26, 09:27 AM
# Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
Source: [https://arxiv.org/html/2608.12674](https://arxiv.org/html/2608.12674)
Ravi Teja ChunduriSrikaran Reddy BoyaAffiliation:Walmart Global Tech Bentonville, AR, USA Srikaran\.Reddy\.Boya@samsclub\.comDeep Narayan MishraAffiliation:Walmart Global Tech Sunnyvale, CA, USA deep\.mishra@walmart\.comAjay Kumar BAffiliation:Walmart Global Tech Seattle, WA, USA Ajay\.Kumar10@walmart\.comKarthik KumaranAffiliation:Walmart Global Tech Sunnyvale, CA, USA karthik\.kumaran@walmart\.comPranay KonaAffiliation:Walmart Global Tech Sunnyvale, CA, USA pranay\.kona@walmart\.com
###### Abstract
Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers\. However, with catalogs spanning millions of active items, manual governance of price relationships is infeasible\. Inconsistent pricing across item variants distorts customer value perception and cannibalizes sales\. To address this, we present a scalable, context\-aware Multi\-Agent Framework designed to automate the construction of “Lines and Ladders” pricing taxonomies\. Our framework employs specialized LLM agents to construct these coherent pricing structures by identifying key attributes, extracting multi\-modal values, and applying hierarchical grouping logic\. Evaluated on real\-world enterprise data and deployed in production, our 3\-Agent system achieves an F1\-score of 0\.83 for Lines, outperforming single\-agent baselines by mitigating cognitive overload\. The system achieves\>90%\>90\\%precision and\>75%\>75\\%recall in Food & Consumables, and80\.2%80\.2\\%assignment accuracy in the unstructured General Merchandise catalog\.
###### Index Terms:
Multi\-Agent Systems, Taxonomy Construction, Multi\-modal Information Extraction, Product Clustering, E\-commerce
© 2026 IEEE\. Personal use of this material is permitted\. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works\.
## IIntroduction
Retail pricing directly dictates the core profitability of a retailer\[[13](https://arxiv.org/html/2608.12674#bib.bib8)\]\. To maintain consistency, retailers structure products into “lines and ladders”—a logical hierarchy grouping items by features, quality, and business context\[[11](https://arxiv.org/html/2608.12674#bib.bib9)\]\. Traditionally, these structures are created through a manual process led by experts, where category managers make decisions based on deep domain knowledge\[[1](https://arxiv.org/html/2608.12674#bib.bib10)\]\. However, this approach is inherently static and does not meet the dynamic nature of modern retail catalogs\.
For global retailers, price governance complexity scales non\- linearly\. With millions of active items, high assortment velocity renders manual oversight infeasible\. The core problem is not merely the volume of items, but the operational bottleneck created by the inconsistent and often non\-standardized nature of catalog data\. When relationships are ignored, illogical pricing gaps emerge, leading to customer arbitrage\[[16](https://arxiv.org/html/2608.12674#bib.bib11)\]\[[9](https://arxiv.org/html/2608.12674#bib.bib16)\]\. In un\-managed catalogs, we frequently observe two gaps:Variant Inconsistency\(e\.g\., identical bikes priced at $238 vs $199 solely due to color\) andData Discrepancy\(e\.g\., a 0\.1\-inch typo in a mirror’s dimensions creating a $7\.50 gap\)\.
To address this scalability challenge, we propose a context\-aware multi\-agent framework\. Unlike traditional rule\-based systems, our approach leverages specialized agents to automatically cluster products into meaningful “Lines & Ladders”\. This system automates the identification and extraction of key attributes, customized for each ProductType\. We utilize this architecture to simulate and replicate the decision\-making logic of a human expert at scale\[[14](https://arxiv.org/html/2608.12674#bib.bib12)\]\.
This framework serves the primary objective of automated price governance, establishing the structural foundation required for next\-generation autonomous pricing systems such as algorithmic demand forecasting and dynamic pricing models\[[7](https://arxiv.org/html/2608.12674#bib.bib13)\]\. The main contributions of this paper are summarized as follows:
- •Formalizing Taxonomy Construction:We decompose pricing governance into Similarity & Variance discovery via a 3\-agent architecture\.
- •Scalable Multi\-Modal Execution:We detail a distributed pipeline utilizing a multi\-model routing strategy to process millions of items efficiently\.
- •Validating Industrial Efficacy:We provide extensive ablation studies and post\-launch metrics demonstrating the system’s scalability and operational impact in a live enterprise environment\.
## IIRelated Work
The concept of “lines and ladders” has been indirectly established in retail strategy for decades\. Dean\[[4](https://arxiv.org/html/2608.12674#bib.bib1)\]originally defined “price lining” as the practice of offering a range of products at specific, distinct price points to signal varying levels of quality to the consumer\. While there is extensive literature on product line pricing problems, as comprehensively reviewed by Guiltinan\[[8](https://arxiv.org/html/2608.12674#bib.bib2)\], our work addresses a fundamentally different challenge\. Unlike traditional approaches that focus on the mathematical optimization of setting prices, our objective is to automatically cluster a massive catalog into coherent product lines based on quality and features\. This structural organization is a prerequisite for scalable price governance rather than price setting itself\.
The term “ladder” is a more recent term and can be interpreted in various ways\. Dobbs\[[5](https://arxiv.org/html/2608.12674#bib.bib5)\]describes a ladder as a form of wholesale pricing contract where the wholesale price is contingent on the retail price, facilitating tiered price setting by the manufacturer\. In a marketing context, price laddering aligns closely with price lining, typically involving a “good\-better\-best” strategy\[[10](https://arxiv.org/html/2608.12674#bib.bib4)\]\. Furthermore, the concept of “climbing a ladder” refers to the strategy of incentivizing customers to upgrade to higher\-quality products within the offered set, as described by\[[3](https://arxiv.org/html/2608.12674#bib.bib7)\]and\[[15](https://arxiv.org/html/2608.12674#bib.bib6)\]\.
Draganska and Jain\[[6](https://arxiv.org/html/2608.12674#bib.bib3)\]further distinguish between vertical and horizontal product differentiation\. In vertical differentiation, products vary by quality, whereas in horizontal differentiation, firms offer products that vary in characteristics such as scent, color, or flavor rather than quality\. Their empirical analysis confirms that product lines function as effective price discrimination tools, justifying the common strategy of pricing lines differently based on quality while pricing variants \(e\.g\., flavors or colors\) uniformly\. While the literature offers various definitions for ladders, in the context of large\-scale retail, we specifically define ladders in terms of volume incentives\. Therefore, for the purposes of this paper, we refer to “volume discount ladders” simply as ladders\.
Although these concepts explain the strategic “why”, they fail to address the operational “how” for massive, unstructured datasets\. Pre\-LLM \(Large Language Models\) approaches relied on structured tabular data and extensive feature engineering\[[12](https://arxiv.org/html/2608.12674#bib.bib15)\], a workflow rendered impractical by the heterogeneity of modern retail catalogs\. The manual effort required to standardize attributes across millions of items is practically impossible\. LLMs offer a paradigm shift, enabling the direct interpretation of unstructured text and images to automate feature extraction at scale\.
## IIIMethodology
To overcome the limitations of manual governance, we propose a Context\-Aware Multi\-Agent Framework\. We avoid a single\-prompt approach as it yields unstable outputs\. Instead, the system leverages specialized LLMs to perform a two\-pronged analysis, discovering key attributes from non\-standardized data\.
### III\-AProblem Formulation & Scoping
LetC=\{i1,i2,…,iN\}C=\\\{i\_\{1\},i\_\{2\},\.\.\.,i\_\{N\}\\\}be a catalog ofNNitems, each with raw metadataRiR\_\{i\}\(e\.g\., title, description, image\)\. As shown in Figure[1](https://arxiv.org/html/2608.12674#S3.F1), we map each item to aLine \(LkL\_\{k\}\): a subset of items sharing functional equivalence \(e\.g\., identical items differing only in flavor or color\), where∀ia,ib∈Lk⟹Price\(ia\)=Price\(ib\)\\forall i\_\{a\},i\_\{b\}\\in L\_\{k\}\\implies Price\(i\_\{a\}\)=Price\(i\_\{b\}\)and Lines aggregate into aLadder \(MjM\_\{j\}\): a collection of Lines linked by value\-based differentiation \(e\.g\., volume, pack size\), enforcing a tiered pricing logic whereVolume\(La\)\>Volume\(Lb\)⟹UnitPrice\(La\)<UnitPrice\(Lb\)Volume\(L\_\{a\}\)\>Volume\(L\_\{b\}\)\\implies UnitPrice\(L\_\{a\}\)<UnitPrice\(L\_\{b\}\)\.
Fig\. 1:Definition of Lines & LaddersAnalyzing millions of items via a global, top\-down approach is computationally expensive and lacks context\. Therefore, we decompose the catalog usingProductType\(e\.g\., “Candles”\) as our granular anchor\. Scoping the analysis to items with shared structural characteristics ensures the extracted attributes are highly relevant, allowing the framework to generate precise, tailored taxonomies rather than relying on generic, one\-size\-fits\-all rules\.
### III\-BDual\-Pronged Attribute Discovery
Constructing effective taxonomies requires distinguishing betweenSimilarity Attributes\(features defining functional equivalence at the same price point\) andVariance Attributes\(features driving price differentiation\)\. We achieve this using two specialized LLM agents\. To ensure the generated taxonomy is strictly based on physical attributes, these agents operate without few\-shot examples of historical taxonomies\. Because legacy merchant structures often reflect ad\-hoc operational strategies \(e\.g\., grouping all items from a single supplier\) rather than strict attribute equivalence, injecting them into the prompts would introduce strategic noise and confuse the algorithmic extraction\.
#### III\-B1Identifying Price Similarity Attributes \(Agent 1\)
This agent isolates base features common to identically priced items, defining the core value proposition\. Figure[2](https://arxiv.org/html/2608.12674#S3.F2)outlines the prompt strategy\.
- •Density\-based Sampling:FromProductTypedata setDptD\_\{pt\}, we identify theKKmost frequent tuplesTsim=\{\(Brand,MSRP\)1,…,\(Brand,MSRP\)K\}T\_\{sim\}=\\\{\(Brand,MSRP\)\_\{1\},\.\.\.,\(Brand,MSRP\)\_\{K\}\\\}whereMSRPMSRPrepresents manufacturer’s suggested retail price\. We strictly utilize base MSRP to avoid promotional noise\.
- •Corpus Generation:For each tuple, we sample a corpusCtC\_\{t\}of 15 item descriptions, yieldingKKdistinct corpora representing clusters of identically priced items\.
- •Iterative Extraction:Agent 1 processes eachCtC\_\{t\}to extract consistent core attributes, utilizing a persistent memoryPsimP\_\{sim\}to iteratively refine its logic\. For a “Candles”ProductType, outputs might include\{weight\_oz, height\_inches, burn\_time\_hours\}\.
Objective:Extract quantifiable attributes justifying identical pricing\.Constraints:1\.Quantifiable Only:Reject categorical features; focus on measurable specs \(e\.g\., weight, capacity\)\.2\.Consistency:Attribute values must be invariant across the group\.3\.Evidence Tiering:∙\\bulletTier 1 \(Explicit\):Stated in text→\\rightarrowAccept\.∙\\bulletTier 2 \(Inferred\):Logically derived→\\rightarrowAccept\.∙\\bulletTier 3 \(Speculative\):No evidence→\\rightarrowReject\.4\.Standardization:Enforce output schema and reuse names from persistent memory\.Fig\. 2:Prompt strategy to identify similarity attributes
#### III\-B2Identifying Price Variance Attributes \(Agent 2\)
This agent identifies premium features that justify price stratification\. Figure[3](https://arxiv.org/html/2608.12674#S3.F3)outlines the prompt strategy, instructing the agent to act as a Pricing Strategist\.
- •Stratified Sampling:To capture price diversity, we selectBtopB\_\{top\}, the top 5 frequent brands inDptD\_\{pt\}\. For each brandb∈Btopb\\in B\_\{top\}, we identify its top 5 frequent MSRP points, creating a stratified setTvarT\_\{var\}of up to 25 groups,Tvar=\{\(Brand1,MSRP1\),T\_\{var\}=\\\{\(Brand\_\{1\},MSRP\_\{1\}\),…,\(Brand5,MSRP5\)\}\\dots,\(Brand\_\{5\},MSRP\_\{5\}\)\\\}\. Empirical trials indicated that expanding beyond 5 brands introduces excessive nomenclature variance that obscures coreProductTypesignal\.
- •Corpus Generation:We sample 5 items for each group inTvarT\_\{var\}\. Unlike Agent 1, Agent 2 ingests the entire stratified datasetDvarD\_\{var\}as a single input to analyze a wide quality and price spectrum\.
- •Comparative Analysis:Agent 2 compares low\- and high\-priced groups acrossDvarD\_\{var\}to identify differentiating features \(e\.g\.,\{number\_of\_wicks, wax\_type, has\_decorative\_vessel\}\)\. The agent is also equipped with a memory module, allowing it to incorporate feedback and refine attributes based on specific business requirements\.
Objective:Compare price\-stratified groups to identify differentiating attributes\.Analysis Framework:1\.Inventory & Evidence:List attributes with evidence tiers as shown in Figure[2](https://arxiv.org/html/2608.12674#S3.F2); Reject speculative \(Tier 3\) evidence\.2\.Type Classification:∙\\bulletType 1 \(Differentiating\):Consistent within group, varies across price tiers \(Primary\)\.∙\\bulletType 2 \(Foundational\):Essential specs regardless of price impact\.3\.Selection:Finalize attributes that explain price deltas or define identity\.Constraints:Standardized naming; no generic terms \(e\.g\., “style”\); Follow output schema\.Fig\. 3:Prompt strategy to identify variance attributesFigure[7](https://arxiv.org/html/2608.12674#S3.F7)illustrates this end\-to\-end workflow, detailing the parallel execution and subsequent synthesis of Agents outputs\.
#### III\-B3Parameter Selection & Heuristics
The sampling constants used by Agent 1 \(30 groups / 15 items\) and Agent 2 \(25 groups / 5 items\) were derived via parameter tuning to optimize the Pareto frontier between signal\-to\-noise ratio and LLM token costs\. As shown in Table[I](https://arxiv.org/html/2608.12674#S3.T1), providing more data to the LLM degrades performance\. Expanding sample breadth \(e\.g\., 80 groups\) introduces long\-tail catalog noise and niche brands that confuse schema generation\. Expanding depth \(e\.g\., 30 items\) causes context bloat, leading the LLM to hallucinate variance attributes based on minor, irrelevant differences\. Furthermore, average pipeline runtimes were scaled up to 155 minutes for the 80/30 configuration, compared to∼\\sim65 minutes for our proposed thresholds\.
TABLE I:Impact of Sample Size \(Groups / Items per Group\)
### III\-CSynthesis Agent
The similarity and variance agents yield distinct attributes sets,AsimA\_\{sim\}andAvarA\_\{var\}, which often exhibit semantic overlap\. To resolve these dualities, a Synthesis Agent performs a mapping functionΦ\\Phito produce a canonical schemaSptS\_\{pt\}:Spt=Φ\(Asim∪Avar\)S\_\{pt\}=\\Phi\(A\_\{sim\}\\cup A\_\{var\}\)
Figure[4](https://arxiv.org/html/2608.12674#S3.F4)provides the implemented prompt logic\. The execution ofΦ\\Phiensures schema integrity through three sequential operations\.
- •Attribute Consolidation & Semantic Resolution:Raw outputs from Agents 1 and 2 are aggregated to eliminate redundancy, resolve synonymy \(e\.g\., merging “flavor” and “taste”\), and enforce standardized naming conventions\.
- •Attribute\-Unit Decoupling:To resolve UOM inconsistencies \(e\.g\., oz vs g\) and enable comparative analysis, the framework explicitly decouples numerical values from units, defining a scalar value field and a UOM field for every quantitative attribute\. This yields a streamlined, high\-fidelity schema blueprint for downstream extraction\.
Objective:Synthesize Similarity and Variance lists into a canonical schema\.Analysis Framework:1\.Consolidate:Merge lists; resolve synonyms \(e\.g\., “watts”→\\rightarrow“motor\_power”\) using similarity output as precedence\.2\.Filter:Remove global attributes \(Brand, Color\) and generic terms \(e\.g\., “features”\)\.3\.UOM Decoupling:For every physical metric, generate a companion\_uomfield\.4\.Schema Definition:Define data types, defaults, and extraction prompts\.Constraints:Standardized naming; Follow output schema; Justify high\-cardinality attributes\.Fig\. 4:Prompt strategy for Schema Synthesis
### III\-DMulti\-Modal Attribute Extraction
Following schema synthesis, the framework populates the schema for every item to transform unstructured data into structured feature vectors\.
#### III\-D1Schema Definition:
The canonical schemaSptS\_\{pt\}generated in the previous step consists ofkkattributes\{a1,a2,…,ak\}\\\{a\_\{1\},a\_\{2\},\.\.\.,a\_\{k\}\\\}\. Each attributeaja\_\{j\}is formally defined as a tupleaj=⟨name,a\_\{j\}=\\langle\\textit\{name\},data type,\\textit\{data type\},default,\\textit\{default\},prompt⟩\\textit\{prompt\}\\rangle\. Figure[5](https://arxiv.org/html/2608.12674#S3.F5)shows examples of the generated attribute schema for the “Candles”ProductType\.
∙\\bulletAttribute Name:wax\_typeAttribute Type:StringDefault value:paraffinPrompt:Identify the type of wax used in the candle \(e\.g\., soy, paraffin, beeswax, coconut\) from the image or description\.∙\\bulletAttribute Name:is\_decorativeAttribute Type:BooleanDefault value:FalsePrompt:Indicate True if the candle features specific artwork, themes, or a sculpted shape\.Fig\. 5:Sample Schema Definition for Extraction
#### III\-D2Multi\-Modal Execution:
We employ a multi\-modal functionfLLMf\_\{LLM\}that maps an item’s raw dataRiR\_\{i\}\(text \+ images\) to a feature vectorViV\_\{i\}based on the schemaSptS\_\{pt\}:Vi=fLLM\(Ri,Spt\)V\_\{i\}=f\_\{LLM\}\(R\_\{i\},S\_\{pt\}\)
This multi\-modal capability is critical for retail catalogs, where key specifications \(e\.g\., “4\-pack” or “organic”\) often appear only on packaging images\. Figure[6](https://arxiv.org/html/2608.12674#S3.F6)depicts this workflow, where the model ingests item context and schema prompts to generate structured values\. Unlike traditional methods requiring thousands of specific NER models\[[2](https://arxiv.org/html/2608.12674#bib.bib14)\], this approach utilizes a single generalized agent to extract arbitrary attributes defined by the dynamic schema, resolving the scalability bottleneck\.
Fig\. 6:Implemented methodology for attribute extractionBecause traditional pre\-sanitization \(e\.g\., regex\) fails on missing or unstructured catalog data, this Multi\-Modal Extractor acts as the sanitization engine itself\. Initial experiments with smaller open\-weight models \(e\.g\., Nemotron VLM 2B/8B/12B\) failed to achieve production accuracy without large\-scale labeled data tailored to our complex schema, necessitating our reliance on frontier models\.
Fig\. 7:Implemented methodology for eachProductType
### III\-EData Standardization and Normalization
Raw extraction yields structured data, but inconsistencies impede grouping\. We employ a post\-processing pipelineΨ\\Psito normalize the extracted feature vectorViV\_\{i\}\.
- •Numerical Standardization \(UOM\):For attributes suffixed with\_uom\(e\.g\., weight\), a conversion functionϕuom\\phi\_\{uom\}transforms the raw valuevrawv\_\{raw\}and uniturawu\_\{raw\}into a normalized valuevnormv\_\{norm\}in a standard base unitubaseu\_\{base\}: vnorm=ϕuom\(vraw,uraw\)→ubasev\_\{norm\}=\\phi\_\{uom\}\(v\_\{raw\},u\_\{raw\}\)\\rightarrow u\_\{base\}For example,\(12,lbs\)→192\(12,\\text\{lbs\}\)\\rightarrow 192\(oz\)\. We enforce a strict5%5\\%numerical tolerance to absorb floating\-point inaccuracies during UOM conversions, preventing the artificial separation of identical items\.
- •Categorical Semantic Clustering:To resolve semantic fragmentation \(e\.g\., “Soy” vs\. “Soy Blend”\), an LLM\-based clustering functionCsemC\_\{sem\}maps a raw valueuju\_\{j\}from the set of unique valuesUUto a canonical termucanonu\_\{canon\}, preventing artificial fragmentation:ucanon=Csem\(uj\|U\)u\_\{canon\}=C\_\{sem\}\(u\_\{j\}\|U\)
- •Brand Normalization:To strictly govern brand identity, a hybrid functionRbrandR\_\{brand\}reconciles the extracted brandBextB\_\{ext\}with the catalog masterBintB\_\{int\}via sub\-string mapping and a safety rail based on Jaccard Similarity \(JJ\) and brand frequency \(bfb\_\{f\}\): Bfinal=\{BintifJ\(Bext,Bint\)=0∧bf\(Bint\)≥5BextotherwiseB\_\{final\}=\\begin\{cases\}B\_\{int\}&\\text\{if \}J\(B\_\{ext\},B\_\{int\}\)=0\\land b\_\{f\}\(B\_\{int\}\)\\geq 5\\\\ B\_\{ext\}&\\text\{otherwise\}\\end\{cases\}Here,J\(Bext,Bint\)J\(B\_\{ext\},B\_\{int\}\)denotes the Jaccard Similarity between the extracted and internal brand strings, andbf\(Bint\)b\_\{f\}\(B\_\{int\}\)represents the frequency count of the internal brand within the catalog\. This prioritizes the system of recordBintB\_\{int\}when the extracted brand is disjoint from a well\-established entity, preventing hallucinated overrides\.
### III\-FAdaptive Taxonomy Construction
Grouping is scoped at the brand level to respect attribute applicability, as brands within aProductTypeoften employ distinct conventions \(e\.g\., weight vs\. burn time\)\.
#### III\-F1Feature Selection & Pruning
We apply heuristic filters on the normalized vectorViV\_\{i\}to retain discriminative features:
- •Sparsity & Imputation:Through iterative deployment and experimentation across diverse catalog samples, we established operational heuristics to drop attributes with\>65%\>65\\%nullity or\>70%\>70\\%cardinality; retaining sparser features consistently caused artificial fragmentation of product lines\. Gaps in retained attributes are imputed usingSptS\_\{pt\}defaults\.
- •Numerical Binning:To handle numerical variance, we applyϵ\\epsilonneighborhood clustering\. Dictated by strict business requirements, values within a5%5\\%tolerance \(e\.g\., 12\.0 oz vs 12\.3 oz\) are binned into a single discrete valuevbinv\_\{bin\}\.
#### III\-F2Hierarchical Grouping Logic
LetAcatA\_\{cat\}be the set of categorical attributes andAnumA\_\{num\}be the set of numerical/volume attributes\.
- •Line Generation:Items are grouped by the complete setAtotal=Acat∪AnumA\_\{total\}=A\_\{cat\}\\cup A\_\{num\}\. Each unique combination forms a Line \(LkL\_\{k\}\) assigned a persistent UUID: Lk=\{i∈C∣Vi\(Atotal\)=const\}L\_\{k\}=\\\{i\\in C\\mid V\_\{i\}\(A\_\{total\}\)=\\text\{const\}\\\}
- •Ladder Generation:Relaxing constraints by excludingAnumA\_\{num\}, we group byAcatA\_\{cat\}to cluster Lines into Ladders \(MjM\_\{j\}\), linking items differing only by size or pack quantity: Mj=\{i∈C∣Vi\(Acat\)=const\}M\_\{j\}=\\\{i\\in C\\mid V\_\{i\}\(A\_\{cat\}\)=\\text\{const\}\\\}
Consequently, a ladder is defined as the union of its constituent Lines:Mj=⋃kLkM\_\{j\}=\\bigcup\_\{k\}L\_\{k\}\. This logic automatically structures the catalog, placing functionally equivalent items into Lines and connecting them via volume\-based Ladders\.
### III\-GHuman\-in\-the\-Loop Feedback Mechanism
While the multi\-agent framework provides a robust baseline, the inherent stochasticity of LLMs and the specificity of business strategy \(e\.g\., pricing soft drinks at the parent brand level\) necessitate human intervention\. We incorporate a Human\-in\-the\-Loop \(HITL\) mechanism to capture merchant corrections as high\-quality labeled signals\. This feedback updates aProductTypespecific contextual memory,PcontextP\_\{context\}, designed to drive future pipeline refinements\.
- •Strategic Alignment:Agents 1 and 2 will learn to prioritize or ignore attributes based on merchant preferences\.
- •Schema Refinement:Agent 3 will update its merging logic to respect business\-specific naming conventions and optimize extraction prompts\.
AsPcontextP\_\{context\}accumulates, agents recall these corrections to ensure subsequent runs align with established business logic, allowing the system to progressively evolve into a domain\-adapted expert\. While this HITL architecture successfully captures immediate strategic intent, quantifying the system’s long\-term convergence toward expert\-level behavior remains an ongoing area of empirical validation within our deployed environment\.
## IVDeployment at Scale
Deployed on a distributed compute cluster for horizontal scalability and fault tolerance \(Figure[8](https://arxiv.org/html/2608.12674#S4.F8)\), the framework is managed by a refresh orchestrator supporting two execution modes: full refresh \(recomputing attribute schema\) and delta refresh \(reusing persisted schema for steady\-state updates\)\.
Fig\. 8:Lines & Ladder System Design### IV\-ADistributed Pipeline Architecture
Decomposing the pipeline by Department→\\rightarrowProductTypeenables independent failure domains\. To balance reasoning and cost at scale, we employ amulti\-model routing strategyusing secure, enterprise\-managed deployments of frontier commercial multi\-modal LLMs:
- •Multi\-Agent Attribute Identifier:Coordinates the three agents\. Reasoning\-heavy tasks \(schema generation, conflict resolution\) are routed to a proprietary, enterprise\-managed frontier LLM \(comparable in scale and capability to\>100B\>100Bparameter instruction\-tuned “thinking” models\)\. A low\-latency RDBMS implements persistent memory \(PcontextP\_\{context\}\), using deterministic SQL queries to retrieve and inject historical merchant rules into prompts\. To prevent context overflow, this injected history is strictly truncated to feedback collected since the last full refresh\. Due to enterprise confidentiality, exact commercial model names cannot be disclosed\. However, all agents utilize state\-of\-the\-art frontier models with the temperature strictly set at 0 to maximize determinism\.
- •Multi\-Modal Attribute Extractor:Executes extraction functionfLLMf\_\{LLM\}\. To process millions of items, extraction is routed to a highly efficient, cost\-optimized multi\-modal model from the same family \(comparable in scale to 7B\-13B parameter Vision\-Language Models\), designed for high\-throughput, low\-latency execution\.
- •Attribute Cleaner & Standardizer:Normalizes values via pipelineΨ\\Psi\. A SQL\-based semantic cache stores synonym mappings \(e\.g\., “Soy”→\\rightarrow“Soy Wax”\) to optimize latency and reduce redundant LLM calls during delta refreshes\.
- •Lines & Ladders Clustering:Transforms standardized vectors into Lines \(LkL\_\{k\}\) and Ladders \(MjM\_\{j\}\) using hierarchical logic \(Section 3\.6\.2\)\.
### IV\-BAsynchronous Feedback Loop
Human validation utilizes an event\-driven architecture\. Merchant UI corrections emit events to a distributed message bus, where an Intelligent Feedback Classifier routes them by error category:
- •Missing Attributes→\\rightarrowTriggers agent refinement\.
- •Extraction Errors→\\rightarrowUpdates Extractor prompts\.
- •Standardization Errors→\\rightarrowUpdates SQL Cache\.
- •Strategic Decisions→\\rightarrowUpdates memoryPcontextP\_\{context\}\.
### IV\-CPerformance and Scalability
We processed∼\\sim1,700ProductTypes\(∼\\sim1 million items\) in 30 minutes using 10 worker nodes\.
- •Throughput:Distributed implementation enables linear worker node scaling within available LLM quotas\.
- •Cost Optimization:LLM inference dominates spend \(∼\\sim4K tokens/item\)\. Micro\-batching and sampling during identification achieve significant cost reductions\.
- •Observability:Structured logging enables partial reruns andProductType\-level regression analysis\.
This architecture maintains strict catalog freshness SLAs within a sustainable cost envelope\.
## VExperimental Evaluation
To validate the framework, we evaluate architectural ablation, structural accuracy, and post\-launch business impact\.
### V\-AArchitecture Ablation
To justify the multi\-agent design, we conducted an ablation study across 14 categories\. We specifically evaluate zero\-shot LLM architectures rather than traditional supervised Named Entity Recognition \(NER\) models \(e\.g\., fine\-tuned BERT\)\. Retail catalogs require dynamic schema generation \(e\.g\., extracting “wax\_type” for candles but “motor\_power” for blenders\); therefore, training and maintaining static models for thousands of evolvingProductTypesis operationally infeasible at enterprise scale\. Furthermore, traditional dense embedding\-based clustering approaches are inadequate for this task; they rely on semantic language similarity rather than the strict physical attribute equivalence and multi\-modal business context required for pricing taxonomies\. Consequently, our baseline comparison focuses on a Single\-Agent LLM, a 2\-Agent system \(Similarity \+ Variance\), and our proposed 3\-Agent system\.
TABLE II:Comparison across Agent ArchitecturesAs Table[II](https://arxiv.org/html/2608.12674#S5.T2)shows, a single LLM tasked with simultaneous similarity and variance discovery suffers severe cognitive overload \(0\.64 Lines F1\)\. While a 2\-Agent system improves performance \(0\.77 F1\), utilizing a Synthesis Agent \(3\-Agt\) to resolve semantic overlaps was strictly necessary to achieve the\>80%\>80\\%accuracy required for production\.
### V\-BQuantitative Manual Evaluation
To measure “Assignment Accuracy” \(the percentage of items correctly grouped without requiring manual edits\) in unstructured General Merchandise, domain experts evaluated 1000 randomly sampled items across 11 diverseProductTypes\. Because evaluating placement requires reviewing multi\-modal data and historical pricing strategies, this represents tens of hours of specialized annotation\. As shown in Table[III](https://arxiv.org/html/2608.12674#S5.T3), the framework achieved an80\.0%80\.0\\%weighted average accuracy\.
TABLE III:Manual Audit Results and Qualitative FeedbackQualitative feedback reveals three distinct categories of error:
- •Technical Granularity:The model occasionally missed specific technical variants, such as “Ultrasonic” vs\. “Evaporative” \(Humidifiers\), “Down Rod” vs\. “Hugger” \(Ceiling Fans\), or “Weed\-Control” vs\. “Feeding” \(Fertilizers\)\.
- •Aesthetic & Subjective Nuance:Differentiation for Lamps or Duffel Bags often relies on intangible qualities \(e\.g\., “silhouette elegance” or “material quality”\) requiring tacit domain knowledge, underscoring the need for HITL merchant intuition\.
- •Data Completeness & Complexity:Errors in Sewing Machines correlated with sparse descriptions, while Office Boards struggled with complex bundle combinations \(board \+ markers\), confirming performance is bounded by catalog data quality\.
Synthesizing these observations, in most failure cases, the grouping logic was directionally correct but missed a single differentiating attribute\. This indicates the system does not hallucinate random groups, but rather is just one feature away from perfect alignment\.
### V\-CComparative Evaluation
In Food & Consumables, we utilized merchant\-generated pricing structures across 14ProductTypesas a directional benchmark, noting that merchants often construct lines based on ad\-hoc strategies rather than strict attribute equivalence\. We pre\-processed reference data to remove statistical outliers\.
- •Matching Logic & Metrics:Because our algorithm generates novel UUIDs, direct mapping to legacy IDs is infeasible\. We utilized a maximum F1 score optimization approach \(similar to bipartite matching\) to align generated lines with merchant lines, calculating precision and recall\.
- •Results & System Guardrails:Table[IV](https://arxiv.org/html/2608.12674#S5.T4)shows an averageLine Precision of98%98\\%andLine Recall of75%75\\%\(Ladders:92%92\\%Precision,81%81\\%Recall\)\. This high\-precision/low\-recall profile is an intentional guardrail\. High\-recall groupings risk applying price changes to unrelated items \(e\.g\., incorrectly grouping 11 distinct “Prego Sauces”\)\. Our system conservatively groups only the 4 identical Alfredo sauces, safely protecting the customer experience\.
TABLE IV:Comparative Performance against Merchants reference structuresAs detailed in Table[IV](https://arxiv.org/html/2608.12674#S5.T4), the framework consistently delivers near\-perfect precision alongside a lower, category\-dependent recall\. This dynamic highlights a fundamental divergence between algorithmic strictness and human strategic intent:
- •High Precision:The agent achieves near\-perfect precision, confirming it groups physically identical items without hallucinating relationships, ensuring safe deployment\.
- •The Recall Gap:Lower recall indicates the agent frequently over\-splits lines compared to the merchant reference due to lacking strategic context\. This explains the F1 variance: highly standardized categories \(e\.g\., Deodorants, F1=1\.00\) align perfectly, whereas categories with subjective, marketing\-driven descriptions \(e\.g\., Cat Food, F1=0\.64\) rely on nuanced terms like “Ocean Whitefish Pate” vs\. “Flaked Tuna in Sauce”\. Merchants frequently over\-group these items based on promotional strategies or broad “flavor families” rather than strict attribute equivalence\. Our algorithm’s strict physical grouping naturally splits these broad merchant groups, resulting in lower recall\.
This divergence highlights the necessity of the HITL workflow \(Section 3\.7\)\. While the agent provides a safe, high\-precision cold\-start state, continuous merchant feedback enables the persistent memory to learn these strategic nuances, allowing future iterations to automatically align with the merchant’s preferred granularity\.
### V\-DPost\-Launch Business Impact
This fully launched production system processes tens of thousands of active items, generating structured Lines and Ladders for\>90%\>90\\%of the previously un\-managed catalog for the first time\. Beyond creating this entirely new data foundation, end\-to\-end telemetry tracking merchant UI workflows over a 13\-week period indicates an86\.1%86\.1\\%reduction in time\-on\-taskfor existing workflows, calculated via:New Hours=Old Hours×Remaining Workload×Remaining Time\\text\{New Hours\}=\\text\{Old Hours\}\\times\\text\{Remaining Workload\}\\times\\text\{Remaining Time\}\.
- •Workload Reduction:Automated extraction reduces manual catalog cleanup by35−40%35\-40\\%\(Remaining Workload≈0\.625\\approx 0\.625\)\.
- •Velocity Increase:Automated grouping and UI review is 4–5x faster than manual creation \(Remaining Time≈0\.222\\approx 0\.222\)\.
Applying this mid\-case scenario \(0\.625×0\.222=0\.1380\.625\\times 0\.222=0\.138\), the workflow requires only13\.9%13\.9\\%of original hours, enabling merchants to focus entirely on high\-level pricing strategy\.
## VILimitations and Future Work
- •Limitations:Despite a temperature of 0, LLM non\-determinism complicates strict reproducibility\. The system is sensitive to sampling bias, where poor catalog quality can skew attribute discovery\. Furthermore, compute costs currently restrict the framework to batch processing\. In addition, reliance on specific LLM families necessitates re\-validation upon model updates to mitigate hallucination risks\. Finally, while we argue that traditional embedding\-based clustering lacks the contextual reasoning required for this task, future work will include empirical benchmarking against these non\-LLM baselines to rigorously quantify the multi\-agent system’s added value\.
- •Future Directions:We plan to refine anchoring fromProductTypeto the brand level to capture niche features and utilize HITL feedback to stabilize non\-deterministic outputs\. We also aim to conduct longitudinal studies measuring the system’s convergence to merchant strategy over multiple HITL feedback cycles, alongside statistical significance testing and sensitivity analysis on heuristic thresholds\. To resolve data sparsity in niche categories \(∼\\sim5%\), we will implement cross\-category transfer learning\. Finally, to enable real\-time inference, we will adopt a teacher\-student architecture, distilling large models into Small Language Models \(SLMs\) for low\-latency execution\.
## VIIConclusion
This paper presented “Lines and Ladders”, a multi\-agent framework that automates retail price governance by decomposing taxonomy construction into Similarity and Variance discovery\. Synergizing LLM\-based extraction with hierarchical grouping, the system structures heterogeneous data into coherent pricing tiers\.
Evaluations confirm\>80%\>80\\%accuracy in General Merchandise and\>90%\>90\\%precision in Food categories, with75%75\\%recall targeted for HITL optimization\. Deployed in production, the system drastically reduces manual effort, enabling merchants to focus on high\-level strategy\. Beyond efficiency, this taxonomy establishes the structural foundation for autonomous pricing and anomaly detection\. Ultimately, this work demonstrates that multi\-agent architectures are highly viable for industrial\-scale governance, offering a framework broadly applicable to other heterogeneous domains like industrial supply chains and online marketplaces\.
## References
- \[1\]S\. Basroy, M\. K\. Mantrala, and R\. G\. Walters\(2001\)The impact of category management on retailer prices and performance\.Journal of Retailing,pp\. 17–18\.External Links:[Link](https://journals.sagepub.com/doi/10.1509/jmkg.65.4.16.18382?utm_source=researchgate.net&utm_medium=article)Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p1.1)\.
- \[2\]W\. Chen, K\. Shinzato, N\. Yoshinaga, and Y\. Xia\(2023\)Does named entity recognition truly not scale up to real\-world product attribute extraction?\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track,Singapore\.External Links:[Link](https://aclanthology.org/2023.emnlp-industry.16/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-industry.16)Cited by:[§III\-D2](https://arxiv.org/html/2608.12674#S3.SS4.SSS2.p2.1)\.
- \[3\]T\. DataPricing ladders: 5 pointers for outstanding performance\.Note:TGN DataAccessed on 2025\-11\-18External Links:[Link](https://tgndata.com/pricing-ladders-5-pointers-for-outstanding-performance/)Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p2.1)\.
- \[4\]J\. Dean\(1950\)Problems of product\-line pricing\.Journal of Marketing14\(4\),pp\. 518–528\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p1.1)\.
- \[5\]I\. M\. Dobbs\(2016\)When does tiered wholesale pricing create an incentive to reduce retail prices?\.Applied Economics Letters23\(11\),pp\. 777–780\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p2.1)\.
- \[6\]M\. Draganska and D\. C\. Jain\(2006\)Consumer preferences and product\-line pricing strategies: an empirical analysis\.Marketing science25\(2\),pp\. 164–174\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p3.1)\.
- \[7\]P\. S\. Fader and B\. G\. Hardie\(1996\)Modeling consumer choice among skus\.Journal of Marketing Research33,pp\. 442–452\.External Links:[Link](https://doi.org/10.2307/3152215)Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p4.1)\.
- \[8\]J\. Guiltinan\(2011\)Progress and challenges in product line pricing\.Journal of Product Innovation Management28\(5\),pp\. 744–756\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p1.1)\.
- \[9\]D\. R\. Lehmann, H\. Yuan, A\. Krishna, and R\. Briesch\(2002\)A meta\-analysis of the impact of price presentation on perceived savings\.Journal of Retailing78\(2\),pp\. 101–118\.Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p2.1)\.
- \[10\]R\. Mohammed\(2005\)The art of pricing\.Crown Business,New York\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p2.1)\.
- \[11\]K\. B\. Monroe\(2003\)Pricing: making profitable decisions\.Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p1.1)\.
- \[12\]S\. Mudgal, H\. Li, T\. Rekatsinas, A\. Doan, Y\. Park, G\. Krishnan, R\. Deep, E\. Arcaute, and V\. Raghavendra\(2018\)Deep learning for entity matching: a design space exploration\.InProceedings of the 2018 international conference on management of data,pp\. 19–34\.Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p4.1)\.
- \[13\]T\. T\. Nagle and G\. Müller\(2016\)The strategy and tactics of pricing: a guide to growing more profitably\.6th edition,Routledge\.External Links:[Link](https://www.academia.edu/39005285/THE_STRATEGY_AND_TACTICS_OF_PRICING)Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p1.1)\.
- \[14\]J\. S\. Park, J\. C\. O’Brien, C\. J\. Cai, M\. R\. Morris, P\. Liang, and M\. S\. Bernstein\(2023\)Generative agents: interactive simulacra of human behavior\.InProceedings of the 36th Annual ACM Symposium on User Interface Software and Technology,pp\. 1–22\.External Links:[Link](https://arxiv.org/pdf/2304.03442)Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p3.1)\.
- \[15\]PriceBeam\(2017\)Price ladders in emerging markets: 4 steps to higher margins\.Note:PriceBeam BlogExternal Links:[Link](https://blog.pricebeam.com/price-ladders-in-emerging-markets-4-steps-to-higher-margins)Cited by:[§II](https://arxiv.org/html/2608.12674#S2.p2.1)\.
- \[16\]L\. Xia, K\. B\. Monroe, and J\. L\. Cox\(2004\)The price is unfair\! a conceptual framework of price fairness perceptions\.Journal of Marketing,pp\. 9\.External Links:[Link](https://www.researchgate.net/publication/228590264_The_Price_Is_Unfair_A_Conceptual_Framework_of_Price_Fairness_Perceptions)Cited by:[§I](https://arxiv.org/html/2608.12674#S1.p2.1)\.Similar Articles
An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals
Presents an emerging retail portfolio management application that uses personalized, tax-aware reinforcement learning with natural language goal input, featuring a three-phase pipeline and integration with live brokerage APIs.
Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies
This paper presents A2X, an LLM-native pipeline that recursively constructs and searches hierarchical service taxonomies to overcome the limited effective context window of LLMs for service discovery in the Internet of Agents. It significantly improves retrieval accuracy and reduces token consumption compared to full-context and embedding-based baselines.
Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching
This paper presents an autoresearch loop for generating provider preference taxonomies in service marketplaces using large language models, transitioning from legacy forms to AI-native matching, with deployment results from a major U.S. marketplace.
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
E-Commerce Bench is an open-source benchmark that evaluates LLM agents on long-horizon autonomous business operation in e-commerce, featuring multi-store negotiation and dynamic events over a simulated year.
Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
Introduces Agent Bazaar, a multi-agent simulation framework for evaluating economic alignment of LLMs, identifying failure modes like algorithmic instability and Sybil deception, and training a 9B model that outperforms frontier models using targeted reinforcement learning.