PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform's Perspective
Summary
This paper introduces PlatformBid, the first comprehensive auto-bidding benchmark designed from a unified advertising platform perspective, along with BidFlow, a novel flow-matching-based auto-bidding method. Experiments show BidFlow improves target cost by +0.68% in online tests on Kuaishou.
View Cached Full Text
Cached at: 07/31/26, 10:00 AM
# PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform’s Perspective
Source: [https://arxiv.org/html/2607.27265](https://arxiv.org/html/2607.27265)
\(2026\)
###### Abstract\.
Real\-time bidding is central to computational advertising, comprising three elements: Supply Side Platform \(SSP\) selling ad impressions, Demand Side Platform \(DSP\) bidding for advertisers, and Ad Exchange conducting auctions between them\. Traditional auto\-bidding algorithms focus solely on the DSP side, maximizing advertiser conversions by adjusting bids against competitors\. However, current big ad platforms, such as social media and e\-commerce companies, now integrate SSP, DSP, and Ad Exchange functions internally\. From such ad platforms’ perspective, the goal of the auto\-bidding algorithms is not only to maximize the advertisers’ conversions, but also the total revenue of the platform\. Given the lack of platform\-centric evaluation frameworks and the pressing need to advance auto\-bidding research, we proposePlatformBid\- the first comprehensive benchmark designed from a unified ad platform’s perspective\. To accurately reflect the real\-world auto\-bidding scenarios, we define three representative settings: \(1\) homogeneous competition with identical algorithms across advertisers, \(2\) heterogeneous competition with diverse algorithmic strategies, and \(3\) promotional competition where some advertisers surge budgets for boosting sales during promotional events likeBlack Friday\. We systematically evaluate a broad spectrum of existing auto\-bidding methods across these settings, encompassing classical control methods, RL\-based methods, and recent generative methods\. Besides these methods, we further propose a novel auto\-bidding method based on flow\-matching, termedBidFlow, which leverages the flow\-matching method’s expressive policy representation to effectively handle dynamic competitive environments\. Experimental results demonstrate the effectiveness of our proposed BidFlow on the new settings\. Online experiments on Kuaishou further show a \+0\.68% improvement in target cost, providing deployment evidence for the offline\-online consistency of PlatformBid\.111Code is available at https://github\.com/YsTvT/PlatformBid\.
Auto\-bidding, Benchmark, Flow Matching, Reinforcement Learning, Generative Model
††journalyear:2026††copyright:cc††conference:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2; August 9–13, 2026; Jeju Island, Republic of Korea\.††booktitle:Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2 \(KDD 2026\), August 9–13, 2026, Jeju Island, Republic of Korea††isbn:979\-8\-4007\-2259\-2/2026/08††doi:10\.1145/3770855\.3817573††ccs:Information systems Computational advertisingFigure 1\.Overview of the unified ad platform\. Previous benchmarks primarily focus only on the DSP side\. However, as unified ad platforms integrate both DSP and SSP functionalities, a new benchmark from the platform perspective is urgently needed\.Figure 2\.Comparison of PlatformBid settings with existing DSP\-centric benchmarks\. \(a\) DSP\-centric benchmarks optimize only individual advertiser objectives, ignoring platform\-level impacts\. \(b\) Setting 1 \(Homogeneous Competition\): N advertisers with identical strategies compete dynamically\. \(c\) Setting 2 \(Heterogeneous Competition\): N/2 baseline advertisers \(white\) compete against N/2 test advertisers \(yellow\) simultaneously\. \(d\) Setting 3 \(Promotional Competition\): Selective budget amplification for specific advertisers \(yellow\) simulates promotional campaigns\. Unlike DSP\-centric approaches, PlatformBid jointly considers both advertiser and platform\-level objectives\.## 1\.Introduction
Real\-time bidding \(RTB\) has emerged as a dominant paradigm in computational advertising\(Muthukrishnan,[2009](https://arxiv.org/html/2607.27265#bib.bib23)\), significantly enhancing the efficiency of audience engagement and sales conversion on the Internet\. As illustrated in Figure[1](https://arxiv.org/html/2607.27265#S0.F1), the RTB process operates as follows: when an impression opportunity arises from a user’s activity on a supply\-side platform \(SSP\), such as watching videos on a media app, a bid request is broadcast to all competing advertisers via an Ad Exchange\. Advertisers then leverage auto\-bidding services\(Renet al\.,[2018](https://arxiv.org/html/2607.27265#bib.bib24)\)provided by demand\-side platforms \(DSPs\) to determine bid prices in real\-time\. The Ad Exchange then selects the highest bidder to display the ad, charging the winning advertiser the market prices, typically the second\-highest bid price in a generalized second\-price \(GSP\) auction setting\(Aminet al\.,[2012](https://arxiv.org/html/2607.27265#bib.bib25)\)\. This paradigm greatly simplifies advertiser operations, as they only need to specify high\-level objectives and constraints \(e\.g\., target CPA, budget limits\) to the DSP, making auto\-bidding algorithms indispensable in modern computational advertising\. Existing auto\-bidding research has predominantly adopted a DSP\-centric perspective\(Yuanet al\.,[2014](https://arxiv.org/html/2607.27265#bib.bib12); Wanget al\.,[2017](https://arxiv.org/html/2607.27265#bib.bib13); Zhanget al\.,[2014a](https://arxiv.org/html/2607.27265#bib.bib14); Ouet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib15)\), where DSP companies provide algorithmic bidding solutions designed to optimize individual advertiser objectives only\. However, recent industry developments have fundamentally altered this landscape\. The emergence and rapid growth of big media and e\-commerce companies, such as Meta, TikTok, and Kuaishou, have driven substantial increases in intra\-platform advertising activity\. Therefore, these companies operate unified ad platforms that consolidate SSP, DSP, and Ad Exchange functionalities\. This structural transformation demands a paradigm shift in the auto\-bidding research: algorithms must now balance individual advertiser objectives\(Jinet al\.,[2018](https://arxiv.org/html/2607.27265#bib.bib16); Wenet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib17)\)with platform\-level revenue optimization\.
However, there is still a lack of evaluation benchmarks designed from a unified ad platform’s perspective\. Existing benchmarks in auto\-bidding research have primarily focused on DSP\-centric scenarios\. A pioneering effort is the iPinYou dataset\(Zhanget al\.,[2014b](https://arxiv.org/html/2607.27265#bib.bib19)\), which provides real\-world RTB logs from a DSP serving multiple advertisers across various ad exchanges\. Building upon this paradigm, subsequent works have introduced benchmarks with richer features and larger scales\. For instance, AuctionNet\(Suet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib9)\)proposes a large\-scale benchmark incorporating diverse advertiser types and competitive auction dynamics\. However, all these benchmarks are still designed for the DSP\-centric auto\-bidding research\. Given the pressing need to advance auto\-bidding research, we propose the first comprehensive benchmark, termedPlatformBid, designed from a unified ad platform’s perspective\.
To develop a benchmark that accurately and comprehensively captures the essence of platform\-level auto\-bidding scenarios, we formulate three representative settings in PlatformBid, as illustrated in Figure[2](https://arxiv.org/html/2607.27265#S0.F2)\(b\): \(1\)Homogeneous competition\. This scenario simulates the platform\-wide deployment of a new auto\-bidding algorithm, where all advertisers adopt the same strategy\. This necessitates evaluating both overall platform revenue and average advertiser performance\. We model this through all advertisers employing identical algorithms\. \(2\)Heterogeneous competition\.Platforms may ensemble multiple auto\-bidding algorithms for different purposes, requiring assessment of both aggregate platform performance and inter\-algorithm competitive dynamics\. We simulate this by having half of the advertisers adopt a baseline method and the other half adopt an alternative method\. This setting evaluates whether new algorithms can improve platform\-level performance while avoiding adverse impacts on baseline advertisers\. \(3\)Promotional competition\.Ad platforms derive substantial revenue from promotional events such asBlack Friday, during which certain advertisers significantly increase their budgets to achieve more conversions\. We simulate this by selectively amplifying budgets for specific advertisers, examining auto\-bidding performance under budget imbalance\.
We systematically evaluate a broad spectrum of auto\-bidding methods on PlatformBid across the three representative settings\. The evaluated methods encompass: \(1\) Classical control methods, including PID controllers\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1)\); \(2\) Reinforcement learning \(RL\) methods, particularly offline RL approaches such as BCQ\(Fujimotoet al\.,[2019](https://arxiv.org/html/2607.27265#bib.bib2)\)and IQL\(Kostrikovet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib3)\); and \(3\) State\-of\-the\-art generative methods, comprising Decision Transformer\-based approaches\(Chenet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib4)\)such as GAS\(Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6)\)and GAVE\(Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7)\), as well as diffusion\-based methods\(Ajayet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib36); Janneret al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib37); Lianget al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib38); Zhouet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib39)\)including CBD\(Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8)\)\. Beyond these existing approaches, we proposeBidFlow, a novel flow\-matching\-based auto\-bidding method specifically designed for dynamic competitive environments\. BidFlow leverages flow\-matching’s expressive policy representation to effectively capture the complex dynamics of ad auctions and employs Q\-value\-guided distillation to distill a multi\-step flow model into an efficient one\-step policy\. BidFlow achieves the best or competitive performance in most settings, while the heterogeneous sparse setting exposes a remaining challenge under extremely limited feedback\. Furthermore, online experiment results on the Kuaishou ad platform,e\.g\., a \+0\.68% increase of target cost, validate both BidFlow’s effectiveness in more complex industrial scenarios and PlatformBid’s offline\-online consistency, confirming the benchmark’s practical value for guiding real\-world auto\-bidding algorithm development\. In summary, the contributions of this work can be summarized as follows:
- •We propose PlatformBid, the first comprehensive benchmark designed from a unified ad platform’s perspective, formulating three representative settings—homogeneous competition, heterogeneous competition, and promotional competition\.
- •We evaluate comprehensive auto\-bidding methods including classical control approaches, RL methods, and generative approaches, and propose BidFlow, a flow\-matching\-based method that achieves superior performance through expressive policy representation and Q\-value\-guided one\-step distillation\.
- •We provide extensive analysis with diverse baselines and validate offline\-online consistency through the online experiment, especially a significant \+0\.68% increase of target cost, confirming PlatformBid’s practical value for guiding real\-world auto\-bidding algorithm development\.
## 2\.Related Work
Auto\-bidding Benchmarks\.Systematic evaluation of auto\-bidding algorithms has long relied on proprietary datasets\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1); Heet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib20); Aminet al\.,[2012](https://arxiv.org/html/2607.27265#bib.bib25)\), limiting method comparison and reproducibility due to lack of unified standards\. The iPinYou dataset\(Zhanget al\.,[2014b](https://arxiv.org/html/2607.27265#bib.bib19)\)pioneered public benchmarking with real\-world RTB logs from a DSP serving multiple advertisers across ad exchanges\. However, it primarily focuses on DSP\-centric scenarios that optimize individual advertiser objectives\. Subsequent benchmarks follow this paradigm while introducing richer features and larger scales—such as AuctionGym\(Jeunenet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib10)\)and AdCraft\(Gomrokchiet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib11)\)—but still optimize individual advertiser strategies in isolation\. AuctionNet\(Suet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib9)\)advanced the field with the first large\-scale public dataset and corresponding benchmark\. However, its evaluation protocol sequentially replaces advertisers against fixed historical bids by replaying logged ad traffic, failing to capture the competitive dynamics where multiple advertisers simultaneously adjust strategies\. In contrast, as illustrated in Figure[2](https://arxiv.org/html/2607.27265#S0.F2), PlatformBid introduces dynamic multi\-advertiser competition across three representative settings, enabling comprehensive assessment of both individual advertiser performance and collective platform revenues\.
Auto\-bidding Methods\.Auto\-bidding algorithms have evolved through three main paradigms, including classical control approaches, RL methods, and Generative methods\.Classical control approaches\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1); Zhuet al\.,[2017](https://arxiv.org/html/2607.27265#bib.bib26); Yanget al\.,[2019a](https://arxiv.org/html/2607.27265#bib.bib27); Maeharaet al\.,[2018](https://arxiv.org/html/2607.27265#bib.bib28); Yanget al\.,[2019b](https://arxiv.org/html/2607.27265#bib.bib29); Kittset al\.,[2017](https://arxiv.org/html/2607.27265#bib.bib30); Geyiket al\.,[2016](https://arxiv.org/html/2607.27265#bib.bib31); Linet al\.,[2016](https://arxiv.org/html/2607.27265#bib.bib33)\), such as PID controllers\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1)\), provide basic automation through feedback control mechanisms but lack adaptability to complex market dynamics\.Reinforcement learning methods\(Panet al\.,[2020](https://arxiv.org/html/2607.27265#bib.bib34); Sutton and Barto,[2018](https://arxiv.org/html/2607.27265#bib.bib35); Liet al\.,[2026](https://arxiv.org/html/2607.27265#bib.bib47); Yanget al\.,[2026](https://arxiv.org/html/2607.27265#bib.bib46)\)have significantly advanced the field\. Online algorithms like USCB\(Heet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib20)\)and MAAB\(Wenet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib17)\)optimize bidding strategies through direct environment interaction, while offline methods including BCQ\(Fujimotoet al\.,[2019](https://arxiv.org/html/2607.27265#bib.bib2)\)and IQL\(Kostrikovet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib3)\)improve sample efficiency by learning from historical logs without risky online exploration\.Generative methods\(Niet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib41); Fuet al\.,[2020](https://arxiv.org/html/2607.27265#bib.bib42); Donget al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib43)\)offer powerful sequence modeling capabilities\. Decision Transformer\(Chenet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib4); Liuet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib40)\)formulates bidding as conditional sequence generation, avoiding value function estimation instability\. GAS\(Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6)\)extends this with post\-training search using multiple critic networks, while GAVE\(Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7)\)introduces value\-guided exploration to address the dataset quality limitation\. Diffusion\-based methods such as DiffBid\(Guoet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib5)\)and CBD\(Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8)\)provide trajectory\-level planning\-based auto\-bidding, with CBD achieving notable improvements in sparse\-reward scenarios through its completer\-aligner framework\. However, diffusion models suffer from slow inference due to iterative denoising, limiting deployment in latency\-sensitive bidding systems\. As these methods have only been validated in DSP\-centric benchmarks, we find their performance in PlatformBid is limited due to the increased complexity of environment dynamics\. To address this challenge, we propose BidFlow, which leverages flow matching for efficient inference while incorporating Q\-learning guidance for optimization\.
## 3\.PlatformBid Benchmark
In this section, we introduce our PlatformBid benchmark for platform\-centric auto\-bidding research\. We first formulate the auto\-bidding problem from a unified platform’s perspective, highlighting key differences from traditional DSP\-centric formulations\. We then describe the benchmark’s data preparation\. Subsequently, we present three evaluation settings designed to reflect real\-world auto\-bidding scenarios\. Finally, we detail the evaluation metrics that capture both advertiser\-level and platform\-level performance\.
### 3\.1\.Problem Statement
Consider an online advertising scenario involving a sequence ofHHimpression opportunities,i\.e\., impressions traffic, each identified by indexii, where PlatformBid simulates it by replaying a traffic log of impressions\. The advertiser secures an opportunity when the bidbib\_\{i\}exceeds competing bids, resulting in an associated costcic\_\{i\}\. PlatformBid employs a Generalized Second\-Price \(GSP\) auction, thuscic\_\{i\}equals the second\-highest bid among all competing advertisers\. The optimization objective seeks to maximize aggregate value from all opportunities, formulated as∑ioivi\\sum\_\{i\}o\_\{i\}v\_\{i\}, whereviv\_\{i\}represents the value of opportunityiiandoio\_\{i\}serves as a binary indicator of winning status\. Beyond budget limitations, the system must satisfy various economic requirements, including constraints on unit costs for specific conversion events such as cost\-per\-action \(CPA\)\(Heet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib20)\)\. Consequently, the auto\-bidding optimization problem under constraints can be formulated as:
\(1\)max\\displaystyle\\max∑ioivi\\displaystyle\\sum\\nolimits\_\{i\}o\_\{i\}v\_\{i\}s\.t\.∑ioici≤B,\\displaystyle\\sum\\nolimits\_\{i\}o\_\{i\}c\_\{i\}\\leq B,∑ioicij∑ioipij≤Cj,∀j,\\displaystyle\\frac\{\\sum\\nolimits\_\{i\}o\_\{i\}c\_\{ij\}\}\{\\sum\\nolimits\_\{i\}o\_\{i\}p\_\{ij\}\}\\leq C\_\{j\},\\quad\\forall j,∑ieil≤Ol,∀l,\\displaystyle\\sum\\nolimits\_\{i\}e\_\{il\}\\leq O\_\{l\},\\quad\\forall l,oi∈\{0,1\},∀i,\\displaystyle o\_\{i\}\\in\\\{0,1\\\},\\quad\\forall i,whereBBdenotes the advertiser’s total budget,pijp\_\{ij\}represents the performance indicator associated with thejj\-th advertiser\-specified constraint with upper boundCjC\_\{j\},eile\_\{il\}quantifies the impact on thell\-th platform\-level performance metric, andOlO\_\{l\}specifies the corresponding platform objective threshold\. For instance,OlO\_\{l\}may stipulate that the aggregate platform revenue across allHHopportunities must not fall below that obtained under a baseline bidding strategy, analogous to the requirement in online A/A testing\. A critical distinction from conventional DSP\-centric formulations is the explicit incorporation of platform objectivesOlO\_\{l\}\. The subsequent sections examine three distinct problem settings that emerge from different specifications ofOlO\_\{l\}\.
Prior research\(Heet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib20); Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6); Guoet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib5)\)have predominantly adopted a simplified auto\-bidding strategy formulated as:
\(2\)bi=λ0vi\+∑j=1JλjpijCj,b\_\{i\}=\\lambda\_\{0\}v\_\{i\}\+\\sum\\nolimits\_\{j=1\}^\{J\}\\lambda\_\{j\}\{p\_\{ij\}\}\{C\_\{j\}\},wherebiib\_\{i\}^\{i\}denotes the final bid amount for opportunityii\. The coefficientsλj\\lambda\_\{j\}constitute the bidding parameters\. In real\-world ad platforms, auto\-bidding algorithms typically update these bidding parameters at relatively coarse time granularities—such as 30\-minute intervals—rather than reacting to individual impression events occurring at millisecond scales\. This temporal characteristic provides sufficient computational headroom for sophisticated auto\-bidding approaches to process parameter adjustment requests\.
### 3\.2\.Data Preparation
The auto\-bidding task can be formulated as a sequential decision\-making task\. At each timestepttwithin a bidding period, the agent observes statest∈𝒮s\_\{t\}\\in\\mathcal\{S\}that encodes the current advertising status, and subsequently selects actionat∈𝒜a\_\{t\}\\in\\mathcal\{A\}to determine bidding parameters\. The advertising environment evolves according to an unknown transition dynamic𝒯\\mathcal\{T\}, where the next state is determined by both historical context and current observations\. Upon each state transition, the environment generates rewardrtr\_\{t\}, which quantifies the value accrued toward the campaign objective during periodtt\. This process repeats until the bidding horizon concludes \(e\.g\., at the end of a day\)\. The key component data of the formulation are detailed below:
- •Statests\_\{t\}: A collection of campaign\-level features describing the current advertising status, including temporal and budget\-related information like remaining time, residual budget, expenditure rate, and KPI satisfaction ratios for each constraint\.
- •Actionata\_\{t\}: Adjustments to the bidding parametersλj\\lambda\_\{j\}forj=0,…,Jj=0,\\ldots,Jat timesteptt, represented as\(atλ0,…,atλJ\)\(a\_\{t\}^\{\\lambda\_\{0\}\},\\ldots,a\_\{t\}^\{\\lambda\_\{J\}\}\)\.
- •Rewardrtr\_\{t\}: The conversion value accumulated during the interval fromtttot\+1t\+1, serving as the optimization signal for the campaign objective\.
For benchmarking learning\-based auto\-bidding methods, we adopt the mainstream offline learning paradigm, which trains policies on historical bidding logs rather than through live online interaction, thereby providing critical safety guarantees\.
Our benchmark is compatible with any datasets containing the above essential MDP components,i\.e\., states, actions, and rewards\. Currently, PlatformBid is built upon the AuctionNet\(Suet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib9)\)dataset, which contains 480K training trajectories collected from 48 advertisers operating under diverse budget and CPA constraints, along with impression traffic logs for testing simulation\. The dataset provides two variants based on reward sparsity: Dense and Sparse\. While AuctionNet focuses on single\-advertiser optimization, we extend it with a critical distinction: simultaneous multi\-advertiser policy deployment within a shared auction environment\. To prevent advertisers from colluding to submit low bids that would harm the platform’s total revenue, a floor price mechanism is adopted\. This architectural shift—from isolated single\-agent evaluation to interactive multi\-agent simulation—enables advertisers to respond to each other’s bidding strategies in real time, thereby capturing competitive dynamics in the ad environment\. We further introduce platform\-level evaluation metrics to assess platform\-level performance beyond individual advertiser objectives\.
### 3\.3\.Evaluation Settings
PlatformBid evaluates auto\-bidding algorithms under realistic multi\-advertiser competitive dynamics from a platform\-centric perspective\. A truly effective auto\-bidding algorithm must satisfy three criteria: \(1\) strong performance under universal adoption, ensuring platform\-wide efficiency; \(2\) robustness when competing against diverse strategies without crowding out other participants; and \(3\) generalization under non\-stationary conditions such as promotional events\. These criteria motivate our three evaluation settings\.
#### Setting 1: Homogeneous Competition\.
The first setting, depicted in Figure[2](https://arxiv.org/html/2607.27265#S0.F2)\(b\)\-left, evaluates platform\-wide performance when allNN=48 advertisers deploy identical bidding policies\. Each advertiser uses the same strategy and competes for impressions through the GSP auction mechanism\. Performance is measured by averaging scores across all advertisers, capturing aggregate platform outcomes rather than individual advertiser success\.
This setting directly mirrors the online experiments where a new auto\-bidding algorithm is simultaneously deployed to all advertisers\. In such deployments, the platform must evaluate whether the new algorithm improves performance at both the advertiser level and platform level compared to the in\-production baseline method\. Under homogeneous strategy adoption, emergent competitive dynamics reveal whether the algorithm achieves stable market equilibrium or triggers systemic inefficiencies such as bid inflation or collusive behavior\. This homogeneous competition scenario thus provides a controlled testbed for evaluating platform\-wide algorithm launch, measuring both aggregate revenue generation and average advertiser welfare under identical strategy adoption\.
#### Setting 2: Heterogeneous Competition\.
The second setting, illustrated in Figure[2](https://arxiv.org/html/2607.27265#S0.F2)\(b\)\-middle, evaluates robustness under strategic diversity by partitioning advertisers into two groups:N/2=24N/2=24baseline advertisers running a fixed reference policy \(vanilla Decision Transformer\(Chenet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib4)\), a widely adopted method in industry\) andN/2=24N/2=24test advertisers running the algorithm under evaluation\. This setting captures the deployment reality where new algorithms are rolled out to a subset of advertisers while others continue using existing solutions\.
This setting directly mirrors the online test with an ensemble of strategies\. In practice, platforms may not deploy a single algorithm to all advertisers simultaneously, as different scenarios have unique requirements\. Instead, new algorithms are typically tested on a subset of target advertisers while others continue using existing solutions\. This approach enables controlled performance comparison under real competitive dynamics\. Only algorithms demonstrating gains for test advertisers without degrading baseline advertiser outcomes satisfy the requirements for production deployment\.
#### Setting 3: Promotional Competition\.
The third setting, shown in Figure[2](https://arxiv.org/html/2607.27265#S0.F2)\(b\)\-right, evaluates generalization under non\-stationary conditions by simulating promotional events\. A subset of advertisers \(7 out of 48 advertisers in this benchmark\) receive doubled budgets, mimicking the budget surges accompanying sales promotions\. All advertisers deploy the same algorithm, with metrics reported separately for promotional advertisers, non\-promotional advertisers, and platform\-wide aggregates\.
This setting simulates promotional campaigns that constitute a major revenue source for ad platforms, such asBlack Friday, where participating advertisers dramatically increase budgets to boost sales\. The setting tests two critical capabilities: promotional advertisers must effectively scale their bidding strategies to utilize expanded budgets without violating CPA constraints, as optimal behavior at doubled budget diverges significantly from training distributions; meanwhile, non\-promotional advertisers must maintain competitive performance despite intensified competition from budget\-rich rivals\. This validates whether algorithms can adapt to the non\-stationary market dynamics characteristic of major promotional events while ensuring fairness across advertisers with different economic resources\.
### 3\.4\.Evaluation Metrics
Our evaluation framework employs comprehensive metrics widely adopted in industrial advertising platforms, encompassing measurements at both platform and advertiser levels\.
Platform\-centric metrics measure aggregate system outcomes:
- •Conversion: Average conversion value generated across all advertisers\.
- •Budget Utilization: Average fraction of allocated budgets spent, indicating market liquidity and inventory efficiency\.
Advertiser\-centric metrics assess individual advertiser experience and constraint satisfaction:
- •CPA Ratio: Average ratio of realized CPA to target CPA across advertisers, where lower values indicate better ROI\.
- •CPA Ratio Variance: Variance of CPA ratios, measuring consistency of constraint satisfaction across advertisers\.
- •Exceed Rate: Fraction of advertisers violating CPA constraints \(realized CPA exceeding target CPA\)\.
- •Qualified Rate: Fraction of advertisers achieving CPA ratios within the acceptable range \[0\.8, 1\.2\]\.
A comprehensive metric balances platform revenue with advertiser satisfaction could be theScore, calculated as
\(3\)Scorei=\{ConversioniifCPAi≤TargetiConversioni×\(TargetiCPAi\)2otherwise\\text\{Score\}\_\{i\}=\\begin\{cases\}\\text\{Conversion\}\_\{i\}&\\text\{if \}\\text\{CPA\}\_\{i\}\\leq\\text\{Target\}\_\{i\}\\\\ \\text\{Conversion\}\_\{i\}\\times\\left\(\\frac\{\\text\{Target\}\_\{i\}\}\{\\text\{CPA\}\_\{i\}\}\\right\)^\{2\}&\\text\{otherwise\}\\end\{cases\}This formulation reflects the platform’s preference for sustainable growth—high conversions achieved through constraint violations are penalized, as such patterns erode advertiser trust and long\-term platform health\.
Morebenchmark implementation detailsare in Appendix[5\.2](https://arxiv.org/html/2607.27265#S5.SS2)\.
Figure 3\.Three categories of auto\-bidding algorithms\.Figure 4\.Overview of the BidFlow framework\. Training \(left\): The BC flow policy learns a velocity field to transform noise into bids via flow matching\. The one\-step policy is trained to distill the BC flow policy while maximizing Q\-values\. Inference \(right\): Only the one\-step policy is used, enabling efficient single\-pass action generation without iterative sampling\.
## 4\.Methods
### 4\.1\.Auto\-bidding Baselines
As illustrated in Figure[3](https://arxiv.org/html/2607.27265#S3.F3), existing auto\-bidding methods can be categorized into three paradigms based on their inference mechanisms\. We select representative methods from each for benchmarking:
\(1\) Classical Control Methodsemploy feedback control mechanisms without learning from data\. We select PID controllers\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1)\), which compute bidding actions by combining proportional, integral, and derivative terms of the error between actual and target CPA\. Given statests\_\{t\}and process variance, the controller outputs actionata\_\{t\}through a fixed mathematical formula\.
\(2\) Reinforcement Learning Methodslearn a policyπ\(a\|s\)\\pi\(a\|s\)that directly maps states to actions of maximum accumulative value\. While RL methods can capture complex state\-action relationships in a single forward pass, they typically assume unimodal action distributions \(e\.g\., Gaussian policies\) and the MDP assumption, limiting their ability to represent multi\-modal bidding strategies required in dynamic auction environments\. Following the offline learning paradigm, we include BCQ\(Fujimotoet al\.,[2019](https://arxiv.org/html/2607.27265#bib.bib2)\)and IQL\(Kostrikovet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib3)\), representing conservative and implicit Q\-learning approaches, respectively\.
\(3\) Generative Methodsrepresent the current state\-of\-the\-art in auto\-bidding by formulating the problem as conditional generation\. These methods fall into two sub\-paradigms:
Decision\-Transformer\-based approachescapture the joint distribution of return\-to\-goRR, statess, and actionaathrough causal transformer architectures, autoregressively generating actions conditioned on target returns and historical sequences\. We testify Decision Transformer \(DT\)\(Chenet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib4)\)in both standard and score\-conditioned \(DT score\) variants, GAS\(Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6)\)with post\-training critic search, and GAVE\(Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7)\)with value\-guided exploration\. However, these methods typically model action distributions as deterministic or unimodal Gaussian, limiting their expressiveness\.
Diffusion\-based approachesleverage the distributional modeling capacity of diffusion models to generate future states, then project them to corresponding actions through inverse dynamics models\. We include CBD\(Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8)\)with its completer\-aligner framework\.
More implementation details can be seen in Appendix[A](https://arxiv.org/html/2607.27265#A1)\.
Table 1\.Comparison under PlatformBid setting 1: Homogeneous Competition\. Arrow direction “↑\\uparrow” and “↓\\downarrow” denote the performance improvement direction\. Best results are bolded\.DatasetMetricPIDIQLBCQDTDT scoreGASGAVECBDBidFlowConversion↑\\uparrow183\.98331\.79236\.48347\.79323\.06353\.00334\.52293\.98358\.90Budget Util↑\\uparrow0\.540\.900\.840\.960\.830\.931\.000\.980\.88CPA Ratio↓\\downarrow1\.391\.001\.241\.040\.910\.971\.141\.280\.88DenseCPA Ratio Var↓\\downarrow0\.690\.070\.060\.060\.030\.030\.120\.160\.03Exceed Rate↓\\downarrow0\.600\.480\.810\.560\.310\.420\.670\.750\.25Qualified Rate↑\\uparrow0\.270\.600\.420\.600\.730\.670\.440\.380\.67Score↑\\uparrow151\.74289\.74169\.60287\.62311\.77314\.27241\.57189\.34348\.04Conversion↑\\uparrow11\.2128\.5621\.5615\.6527\.8330\.1521\.5632\.9833\.06Budget Util↑\\uparrow0\.620\.870\.440\.450\.580\.651\.000\.820\.78CPA Ratio↓\\downarrow1\.941\.010\.871\.350\.810\.801\.110\.970\.85SparseCPA Ratio Var↓\\downarrow0\.620\.090\.320\.890\.150\.100\.090\.160\.11Exceed Rate↓\\downarrow0\.880\.420\.230\.460\.190\.270\.500\.440\.25Qualified Rate↑\\uparrow0\.150\.380\.310\.210\.270\.310\.460\.400\.31Score↑\\uparrow6\.8222\.7620\.9813\.7226\.6228\.4515\.5228\.2730\.59
### 4\.2\.BidFlow: A Flow\-matching\-based Auto\-bidding Method
However, we observe that existing auto\-bidding baselines exhibit limited performance in these new settings\. This degradation stems from the increasingly critical challenge of modeling complex, multi\-modal bidding distributions—a challenge amplified by more dynamic market fluctuations and heightened competitive complexity at the platform level\. To address this, we propose advancing auto\-bidding methods through flow matching, a state\-of\-the\-art generative modeling paradigm that has demonstrated remarkable success in capturing complex data distributions across multiple domains\(Lipmanet al\.,[2023](https://arxiv.org/html/2607.27265#bib.bib22); Parket al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib18); Genget al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib21)\)\. Therefore, we propose a novel flow\-matching\-based auto\-bidding method, termed BidFlow\.
As shown in Figure[4](https://arxiv.org/html/2607.27265#S3.F4), BidFlow follows the typical flow\-matching framework for offline RL\(Parket al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib18)\), comprising three key components: \(1\) a critic networkQϕ\(s,a\)Q\_\{\\phi\}\(s,a\)that estimates state\-action values; \(2\) a behavioral cloning \(BC\) flow policyμθ\(s,ε\)\\mu\_\{\\theta\}\(s,\\varepsilon\)trained via flow matching to capture the behavioral distribution in the offline dataset; and \(3\) a one\-step policyμω\(s,ε\)\\mu\_\{\\omega\}\(s,\\varepsilon\)that maximizes Q\-values while being regularized through distillation from the BC flow policy\. This one\-step policy enables faster inference compared to the multi\-step BC flow while achieving superior performance through Q\-value guidance\.
The critic is trained via standard temporal difference learning\. Given transitions\(s,a,r,s′\)\(s,a,r,s^\{\\prime\}\)sampled from the offline dataset𝒟\\mathcal\{D\}, the critic minimizes:
\(4\)ℒQ\(ϕ\)=𝔼\(s,a,r,s′\)∼𝒟\[\(Qϕ\(s,a\)−r−γQ¯\(s′,μω\(s′,ε′\)\)\)2\],\\mathcal\{L\}\_\{Q\}\(\\phi\)=\\mathbb\{E\}\_\{\(s,a,r,s^\{\\prime\}\)\\sim\\mathcal\{D\}\}\\left\[\\left\(Q\_\{\\phi\}\(s,a\)\-r\-\\gamma\\bar\{Q\}\(s^\{\\prime\},\\mu\_\{\\omega\}\(s^\{\\prime\},\\varepsilon^\{\\prime\}\)\)\\right\)^\{2\}\\right\],whereQ¯\\bar\{Q\}denotes the target network updated via exponential moving average, and next actions are sampled from the one\-step policy withε′∼𝒩\(0,I\)\\varepsilon^\{\\prime\}\\sim\\mathcal\{N\}\(0,I\)\.
The BC flow policy learns a velocity fieldvθ\(t,s,at\)v\_\{\\theta\}\(t,s,a^\{t\}\)that transforms Gaussian noise into actions through an ODE\. It is trained exclusively with the flow matching objective:
\(5\)ℒBC\(θ\)=𝔼s,a∼𝒟,a0∼𝒩\(0,I\),t∼Unif\(\[0,1\]\)\[‖vθ\(t,s,at\)−\(a−a0\)‖22\],\\mathcal\{L\}\_\{\\text\{BC\}\}\(\\theta\)=\\mathbb\{E\}\_\{\\begin\{subarray\}\{c\}s,a\\sim\\mathcal\{D\},a^\{0\}\\sim\\mathcal\{N\}\(0,I\),\\\\ t\\sim\\text\{Unif\}\(\[0,1\]\)\\end\{subarray\}\}\\left\[\\\|v\_\{\\theta\}\(t,s,a^\{t\}\)\-\(a\-a^\{0\}\)\\\|\_\{2\}^\{2\}\\right\],whereat=\(1−t\)a0\+taa^\{t\}=\(1\-t\)a^\{0\}\+tais the linear interpolation\. At inference, actions are generated by numerically solving the ODE via the Euler method overMMsteps\.
The one\-step policyμω\(s,ε\)\\mu\_\{\\omega\}\(s,\\varepsilon\)learns to directly map noise to actions in a single forward pass\. It is trained with a combined objective that balances behavioral regularization and value maximization:
ℒπ\(ω\)=𝔼s,ε\[‖μω\(s,ε\)−μθ\(s,ε\)‖22\]\+α𝔼s,ε\[−Qϕ\(s,μω\(s,ε\)\)\],\\mathcal\{L\}\_\{\\pi\}\(\\omega\)=\\mathbb\{E\}\_\{s,\\varepsilon\}\\left\[\\\|\\mu\_\{\\omega\}\(s,\\varepsilon\)\-\\mu\_\{\\theta\}\(s,\\varepsilon\)\\\|\_\{2\}^\{2\}\\right\]\+\\alpha\\mathbb\{E\}\_\{s,\\varepsilon\}\\left\[\-Q\_\{\\phi\}\(s,\\mu\_\{\\omega\}\(s,\\varepsilon\)\)\\right\],where the first term distills knowledge from the BC flow policy, the second term maximizes Q\-values for outperforming the BC Flow, andα\\alphacontrols the regularization strength\.
All three components are trained jointly at each iteration\. The complete procedure is summarized in Algorithm[1](https://arxiv.org/html/2607.27265#alg1)\. At deployment, only the one\-step policyμω\\mu\_\{\\omega\}is used for action selection, enabling efficient inference without iterative sampling\.
Algorithm 1Training Procedure of BidFlow0:Offline dataset
𝒟\\mathcal\{D\}, BC coefficient
α\\alpha
1:Initialize critic
QϕQ\_\{\\phi\}, target critic
Q¯ϕ¯\\bar\{Q\}\_\{\\bar\{\\phi\}\}, BC flow policy
vθv\_\{\\theta\}, one\-step policy
μω\\mu\_\{\\omega\}
2:whilenot convergeddo
3:Sample batch from
𝒟\\mathcal\{D\}
4:Update
QϕQ\_\{\\phi\}by minimizing TD loss \(Eq\. 4\)
5:Update
vθv\_\{\\theta\}by minimizing flow matching loss \(Eq\. 5\)
6:Compute distillation target
μθ\(s,ε\)\\mu\_\{\\theta\}\(s,\\varepsilon\)via Euler method
7:Update
μω\\mu\_\{\\omega\}by minimizing
ℒπ\(ω\)\\mathcal\{L\}\_\{\\pi\}\(\\omega\)
8:Soft update:
ϕ¯←τϕ\+\(1−τ\)ϕ¯\\bar\{\\phi\}\\leftarrow\\tau\\phi\+\(1\-\\tau\)\\bar\{\\phi\}
9:endwhile
10:returnOne\-step policy
μω\\mu\_\{\\omega\}
Table 2\.Key metrics comparison under PlatformBid setting 2: Heterogeneous Competition\. Each result pair \(Target, Average\) shows the average performance of only the target advertisers by the evaluated method and the average performance of all advertisers\. Result pairs are bolded if “Average” is the best\.DatasetMetricPIDIQLBCQDT scoreGASGAVECBDBidFlowDenseConversion↑\\uparrow170\.54, 254\.17309\.17, 331\.55231\.04, 265\.90333\.17, 345\.44340\.38, 347\.55355\.71, 330\.07313\.29, 303\.46346\.88, 347\.96CPA Ratio↓\\downarrow2\.78, 1\.910\.99, 1\.031\.19, 1\.190\.85, 0\.960\.91, 0\.970\.99, 1\.061\.20, 1\.130\.93, 0\.93Score↑\\uparrow143\.21, 216\.74285\.24, 293\.86165\.21, 190\.03321\.96, 321\.92315\.81, 315\.93287\.78, 262\.78229\.86, 233\.34312\.69, 308\.83SparseConversion↑\\uparrow11\.88, 21\.6538\.46, 32\.5937\.92, 34\.4244\.38, 37\.5036\.96, 33\.8629\.30, 32\.3242\.71, 31\.8438\.08, 28\.77CPA Ratio↓\\downarrow2\.08, 1\.280\.94, 0\.770\.63, 0\.690\.63, 0\.650\.72, 0\.701\.26, 0\.840\.89, 0\.770\.83, 1\.13Score↑\\uparrow9\.52, 20\.4734\.42, 30\.3137\.78, 33\.5943\.91, 36\.2636\.47, 32\.7721\.20, 28\.2938\.71, 29\.6035\.97, 26\.51
Table 3\.Key metric comparison under PlatformBid Setting 3: Promotional Competition\. “Pro” denotes promotional advertisers, “Non” denotes non\-promotional advertisers, and “All” denotes the average performance across all advertisers\.DatasetTypeMetricPIDIQLBCQDTDT scoreGASGAVECBDBidFlowConversion↑\\uparrow550\.71674\.00543\.14673\.14682\.57690\.43657\.71456\.86727\.29ProCPA Ratio↓\\downarrow1\.491\.101\.151\.060\.980\.951\.151\.320\.91Score↑\\uparrow519\.80587\.70429\.46555\.49634\.28661\.49456\.13347\.37701\.48Conversion↑\\uparrow124\.93294\.10234\.17307\.85316\.43319\.95304\.32252\.17310\.61DenseNonCPA Ratio↓\\downarrow1\.711\.021\.201\.121\.010\.951\.201\.370\.89Score↑\\uparrow111\.46247\.66171\.42230\.79271\.85277\.90203\.71169\.46297\.25Conversion↑\\uparrow187\.02349\.50279\.23361\.12371\.29373\.98355\.85282\.02371\.38AllCPA Ratio↓\\downarrow1\.681\.031\.191\.111\.010\.951\.201\.360\.90Score↑\\uparrow171\.01297\.25209\.05278\.14321\.73333\.84240\.52195\.41356\.20Conversion↑\\uparrow24\.8671\.0054\.5726\.0064\.7178\.8648\.5760\.8671\.86ProCPA Ratio↓\\downarrow0\.470\.950\.670\.960\.650\.711\.490\.750\.68Score↑\\uparrow24\.8564\.2454\.5725\.3864\.7178\.8524\.7059\.5771\.86Conversion↑\\uparrow5\.9527\.4417\.4915\.0523\.9023\.9322\.9329\.2728\.80SparseNonCPA Ratio↓\\downarrow1\.370\.980\.811\.120\.861\.141\.640\.990\.92Score↑\\uparrow5\.8622\.5016\.9313\.0023\.1021\.0811\.7924\.5825\.67Conversion↑\\uparrow8\.7133\.7922\.9016\.6529\.8531\.9426\.6733\.8835\.08AllCPA Ratio↓\\downarrow1\.240\.980\.791\.100\.831\.081\.620\.950\.88Score↑\\uparrow8\.6328\.5922\.4214\.8129\.1729\.5113\.6729\.6832\.41
## 5\.Experiments
### 5\.1\.Experimental Setup
We evaluate representative methods spanning the classical control approach, offline RL, and generative approaches\. PID\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1)\)serves as the classical control baseline\. For offline RL, we include BCQ\(Fujimotoet al\.,[2019](https://arxiv.org/html/2607.27265#bib.bib2)\)and IQL\(Kostrikovet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib3)\), which represent conservative and implicit Q\-learning paradigms respectively\. The generative model family includes Decision Transformer \(DT\)\(Chenet al\.,[2021](https://arxiv.org/html/2607.27265#bib.bib4)\)in both standard and reward\-conditioned \(DT score\) variants, GAS\(Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6)\)with post\-training critic search, GAVE\(Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7)\)with value\-guided exploration, and CBD\(Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8)\)with its completer\-aligner framework\. All methods are trained on the AuctionNet datasets and evaluated on both dense and sparse variants of impression traffic test data\. Implementation follows the original papers\(Chenet al\.,[2011](https://arxiv.org/html/2607.27265#bib.bib1),[2021](https://arxiv.org/html/2607.27265#bib.bib4); Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6); Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7); Kostrikovet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib3); Fujimotoet al\.,[2019](https://arxiv.org/html/2607.27265#bib.bib2); Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8); Parket al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib18)\)with hyperparameters tuned on a held\-out validation set\. For BidFlow, the BC flow policy usesM=10M=10Euler steps during training for distillation target computation\. All results are averaged over 5 random seeds\. More details about implementation are shown in Appendix[A](https://arxiv.org/html/2607.27265#A1)\.
### 5\.2\.Benchmark Implementation Details
PlatformBid is implemented as an extension to the AuctionNet simulation environment\. The simulator processes impression opportunities chronologically, soliciting bids from all active advertisers and resolving auctions according to GSP rules with floor price enforcement\.
To prevent collusive bid suppression, which is a phenomenon observed particularly in sparse\-reward environments where advertisers may collectively lower bids to reduce costs\. To tackle this issue, we introduce a floor price mechanism\. The floor price is set at 80% of the original least winning cost from the AuctionNet dataset\. If a winning bid falls below the floor price, the impression is not awarded and no transaction occurs\. This mechanism preserves market integrity by ensuring that strategic bid reduction cannot extract value below a reasonable threshold, aligning individual incentives with platform revenues\.
PlatformBid tackles its modularity goals via defining unifying abstractions over the auto\-bidding evaluation pipeline\. As depicted in Figure[5](https://arxiv.org/html/2607.27265#S5.F5), components including agent strategies \(agent\.py\), utility functions \(utls\.py\), and bidding strategies \(strategy\.py\) are gathered into evaluation scripts that are agnostic of the specific auto\-bidding algorithm implementations\. Structured command\-line configurations allow users to easily specify the evaluation setting, dataset variant, competing algorithms, and auction parameters \(e\.g\., floor price rate\), making it possible to go directly from one\-line inputs to comprehensive evaluation outputs\.
Figure 5\.PlatformBid execution diagram\. Users run evaluations as configurable experiments, where each setting loads its components from the respective Python modules\.
### 5\.3\.Result Analysis
Tables[1](https://arxiv.org/html/2607.27265#S4.T1),[2](https://arxiv.org/html/2607.27265#S4.T2), and[3](https://arxiv.org/html/2607.27265#S4.T3)present the results under key metrics across our three evaluation settings\.
BidFlow outperforms baselines\.BidFlow demonstrates superior performance in Setting 1, attaining the highest scores across both dense and sparse conversion configurations while maintaining the lowest CPA exceed rate, indicating effective alignment between advertiser objectives and platform welfare\. In Setting 3, BidFlow exhibits advantages for both promotional advertisers and non\-promotional advertisers, demonstrating robust generalization under extreme budget imbalance conditions\. Nevertheless, our analysis also identifies specific limitations\. Under Setting 2 with sparse rewards, BidFlow demonstrates performance degradation, with the baseline DT method struggling to acquire sufficient conversions\. These findings provide important insights for developing auto\-bidding methods in Setting 2 by explicitly modeling the influence between various auto\-bidding methods, highlighting a promising direction for future research\.
Algorithmic evolution brings benefits\.From simple controller methods like PID to complex approaches like DT\-based methods, we observe consistent performance improvements across all settings\. For instance, in Setting 1, BidFlow achieves a 95% improvement in conversion while maintaining superior cost efficiency compared to PID\. The sparse dataset shows similar trends with 3× better performance\. In Settings 2 and 3, advanced methods demonstrate superior adaptability\. While PID and RL methods exhibit severe performance imbalance, sophisticated algorithms like DT score maintain balanced, high performance\. These empirical results demonstrate that algorithmic evolution yields compounding benefits, not only in raw conversion metrics but also in cost efficiency and robustness across diverse competitive environments\.
Competitive and cooperative insights\.The results in Table[1](https://arxiv.org/html/2607.27265#S4.T1)demonstrates a strong negative correlation between CPA Ratio Variance and overall Score across both Dense and Sparse datasets\. BidFlow, with the lowest variance, achieves the highest Scores, while PID, with the highest variance, shows the poorest performance\. This pattern reveals that low\-variance strategies exhibit cooperative characteristics rather than aggressive competition, enabling multiple advertisers to achieve their targets simultaneously\. Conversely, high\-variance strategies reflect destructive over\-competition, causing budget exhaustion and reduced platform revenue\. This confirms that variance serves as a reliable proxy for competition intensity, where stability promotes collective welfare\.
### 5\.4\.Online Experiment Validation
To verify the effectiveness of BidFlow in practice, we deployed it on the Kuaishou live streaming ad platform\. Under this platform, advertisers set budgets with CPA/ROI constraints, and we compared BidFlow against the in\-production and heavily tuned DT method with dedicated reward design, which is the strongest baseline\. The online A/B test was conducted in two isolated environments with identical and fair traffic allocations \(each has 33% of the whole traffic\)\. This configuration corresponds to Setting 2, and Table[4](https://arxiv.org/html/2607.27265#S5.T4)presents the results based on a 13\-day online A/B test\.
The significant improvements in platform revenue \(Cost\) and conversions \(Target Cost\) not only demonstrate BidFlow’s superiority but also validate the benchmark’s practical value for maintaining offline\-online consistency\. Additionally,BidFlow has been fully deployed in Kuaishou live streaming ad platform\.
Table 4\.Online A/B test results on Kuaishou live streaming ad platform\. Performance improvement of BidFlow compared to DT\.Impression↑\\uparrowCost↑\\uparrowTarget Cost↑\\uparrowimprove\+0\.27%\+0\.30%\+0\.68%
### 5\.5\.Cross\-Dataset Validation on iPinYou
To verify that PlatformBid generalizes beyond AuctionNet, we instantiate the benchmark on the iPinYou dataset\(Zhanget al\.,[2014b](https://arxiv.org/html/2607.27265#bib.bib19)\), which differs structurally from AuctionNet in scale, advertiser pool, and auction characteristics\. We adapt all three settings to the iPinYou multi\-advertiser pool while keeping the pipeline unchanged, and report the average number of clicks per advertiser \(higher is better\), following standard iPinYou practice\. We compare BidFlow with two representative generative baselines, DT score and CBD\.
Table 5\.Cross\-dataset validation on iPinYou \(average clicks per advertiser, higher is better\)\.SettingMethodAvg Clicks↑\\uparrowS1 \(Homogeneous\)DT score309\.7CBD302\.5BidFlow \(ours\)324\.1S2 \(Heterogeneous\)DT score299\.5CBD291\.6BidFlow \(ours\)306\.2S3 \(Promotional\)DT score316\.5CBD308\.9BidFlow \(ours\)331\.8As shown in Table[5](https://arxiv.org/html/2607.27265#S5.T5), BidFlow ranks first across all three settings on iPinYou, with the heterogeneous setting \(Setting 2\) confirming that improvements on target advertisers do not come at the cost of degrading baseline advertisers\. Combined with AuctionNet results and the Kuaishou deployment, this provides cross\-platform validation across three structurally distinct environments—e\-commerce, display RTB, and short\-video—confirming that neither PlatformBid nor BidFlow is tied to a specific data source\. Additional public datasets will be incorporated as they become available\.
It is also worth noting that AuctionNet currently remains the only publicly available large\-scale dataset capturing realistic competitive bidding dynamics, and recent state\-of\-the\-art auto\-bidding works\(Liet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib6); Gaoet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib7); Guoet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib5); Liet al\.,[2025](https://arxiv.org/html/2607.27265#bib.bib8)\)are likewise evaluated on a single dataset for this reason\. PlatformBid is designed to be dataset\-agnostic and we will incorporate additional public datasets into the benchmark as they become available\.
## 6\.Limitations and Conclusion
There are still several limitations of this research\. First, we focus exclusively on auto\-bidding algorithms, while auction mechanism design also significantly impacts bidding performance and warrants further investigation\. Second, we do not explicitly model inter\-advertiser relationships—a challenging but important research direction given that real\-world platforms may host millions of advertisers with complex competitive dynamics\. Our benchmark could also be a testbed for these important research directions\.
This paper addresses the gap between existing DSP\-centric auto\-bidding benchmarks and current unified ad platforms’ requirements\. We propose PlatformBid, the first platform\-level auto\-bidding benchmark that formulates various representative competitive settings and enables systematic evaluation of diverse methods\. We further introduce BidFlow, which leverages expressive flow\-matching policies to handle dynamic multi\-advertiser competition\. BidFlow achieves state\-of\-the\-art performance in most PlatformBid settings, with online experiments confirming strong offline\-online consistency,e\.g\., a \+0\.68% improvement of target cost\. With the continued development of unified ad platforms, this work establishes a critical foundation for advancing platform\-centric auto\-bidding research\.
## Acknowledgements
This research is supported by the Big Data Computing Center of Southeast University and Kuaishou Technology\. We would also like to thank Pengfei Lyu for his valuable contribution to the online experiments\.
## References
- A\. Ajay, Y\. Du, A\. Gupta, J\. B\. Tenenbaum, T\. S\. Jaakkola, and P\. Agrawal \(2023\)Is conditional generative modeling all you need for decision making?\.InICLR ’23,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1)\.
- K\. Amin, M\. J\. Kearns, P\. B\. Key, and A\. Schwaighofer \(2012\)Budget optimization for sponsored search: censored learning in mdps\.InUAI ’12,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1),[§2](https://arxiv.org/html/2607.27265#S2.p1.1)\.
- L\. Chen, K\. Lu, A\. Rajeswaran, K\. Lee, A\. Grover, M\. Laskin, P\. Abbeel, A\. Srinivas, and I\. Mordatch \(2021\)Decision transformer: reinforcement learning via sequence modeling\.InNeurIPS ’21,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§3\.3](https://arxiv.org/html/2607.27265#S3.SS3.SSS0.Px2.p1.2),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p5.3),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1)\.
- Y\. Chen, P\. Berkhin, B\. Anderson, and N\. R\. Devanur \(2011\)Real\-time bidding algorithms for performance\-based display ad allocation\.InKDD ’11,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p1.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p2.2),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1)\.
- Z\. Dong, J\. Hao, Y\. Yuan, F\. Ni, Y\. Wang, P\. Li, and Y\. Zheng \(2024\)DiffuserLite: towards real\-time diffusion planning\.InNeurIPS ’24,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- J\. Fu, A\. Kumar, O\. Nachum, G\. Tucker, and S\. Levine \(2020\)D4RL: datasets for deep data\-driven reinforcement learning\.CoRR\.Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- S\. Fujimoto, D\. Meger, and D\. Precup \(2019\)Off\-policy deep reinforcement learning without exploration\.InICML ’19,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p3.1),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1)\.
- J\. Gao, Y\. Li, S\. Mao, P\. Jiang, N\. Jiang, Y\. Wang, Q\. Cai, F\. Pan, P\. Jiang, K\. Gai, B\. An, and X\. Zhao \(2025\)Generative auto\-bidding with value\-guided explorations\.InSIGIR 2025,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p5.3),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1),[§5\.5](https://arxiv.org/html/2607.27265#S5.SS5.p3.1)\.
- Z\. Geng, M\. Deng, X\. Bai, J\. Z\. Kolter, and K\. He \(2025\)Mean flows for one\-step generative modeling\.CoRR\.Cited by:[§4\.2](https://arxiv.org/html/2607.27265#S4.SS2.p1.1)\.
- S\. C\. Geyik, S\. Faleev, J\. Shen, S\. O’Donnell, and S\. Kolay \(2016\)Joint optimization of multiple performance metrics in online video advertising\.InKDD ’16,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- M\. Gomrokchi, O\. Levin, J\. Roach, and J\. White \(2023\)AdCraft: an advanced reinforcement learning benchmark environment for search engine marketing optimization\.CoRR\.Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p1.1)\.
- J\. Guo, Y\. Huo, Z\. Zhang, T\. Wang, C\. Yu, J\. Xu, B\. Zheng, and Y\. Zhang \(2024\)AIGB: generative auto\-bidding via conditional diffusion modeling\.InKDD ’24,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§3\.1](https://arxiv.org/html/2607.27265#S3.SS1.p2.4),[§5\.5](https://arxiv.org/html/2607.27265#S5.SS5.p3.1)\.
- Y\. He, X\. Chen, D\. Wu, J\. Pan, Q\. Tan, C\. Yu, J\. Xu, and X\. Zhu \(2021\)A unified solution to constrained bidding in online display advertising\.InKDD ’21,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p1.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§3\.1](https://arxiv.org/html/2607.27265#S3.SS1.p1.9),[§3\.1](https://arxiv.org/html/2607.27265#S3.SS1.p2.4)\.
- M\. Janner, Y\. Du, J\. B\. Tenenbaum, and S\. Levine \(2022\)Planning with diffusion for flexible behavior synthesis\.InICML ’22,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1)\.
- O\. Jeunen, S\. Murphy, and B\. Allison \(2023\)Off\-policy learning\-to\-bid with auctiongym\.InKDD ’23,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p1.1)\.
- J\. Jin, C\. Song, H\. Li, K\. Gai, J\. Wang, and W\. Zhang \(2018\)Real\-time bidding with multi\-agent reinforcement learning in display advertising\.InCIKM ’18,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- B\. Kitts, M\. Krishnan, I\. Yadav, Y\. Zeng, G\. Badeau, A\. Potter, S\. Tolkachov, E\. Thornburg, and S\. R\. Janga \(2017\)Ad serving with multiple kpis\.InKDD ’17,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- I\. Kostrikov, A\. Nair, and S\. Levine \(2022\)Offline reinforcement learning with implicit q\-learning\.InICLR ’22,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p3.1),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1)\.
- Y\. Li, J\. Gao, N\. Jiang, S\. Mao, R\. An, F\. Pan, X\. Zhao, B\. An, Q\. Cai, and P\. Jiang \(2025\)Generative auto\-bidding in large\-scale competitive auctions via diffusion completer\-aligner\.CoRR\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p6.1),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1),[§5\.5](https://arxiv.org/html/2607.27265#S5.SS5.p3.1)\.
- Y\. Li, S\. Mao, J\. Gao, N\. Jiang, Y\. Xu, Q\. Cai, F\. Pan, P\. Jiang, and B\. An \(2024\)GAS: generative auto\-bidding with post\-training search\.CoRR\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1),[§3\.1](https://arxiv.org/html/2607.27265#S3.SS1.p2.4),[§4\.1](https://arxiv.org/html/2607.27265#S4.SS1.p5.3),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1),[§5\.5](https://arxiv.org/html/2607.27265#S5.SS5.p3.1)\.
- Y\. Li, G\. Cai, S\. Yang, H\. Luo, S\. Han, X\. He, D\. Li, and L\. Feng \(2026\)Phgpo: pheromone\-guided policy optimization for long\-horizon tool planning\.arXiv preprint arXiv:2602\.13691\.Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- Z\. Liang, Y\. Mu, M\. Ding, F\. Ni, M\. Tomizuka, and P\. Luo \(2023\)AdaptDiffuser: diffusion models as adaptive self\-evolving planners\.InICML ’23,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1)\.
- C\. Lin, K\. Chuang, W\. C\. Wu, and M\. Chen \(2016\)Combining powers of two predictors in optimizing real\-time bidding strategy under constrained budget\.InCIKM ’16,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2023\)Flow matching for generative modeling\.InICLR ’23,Cited by:[§4\.2](https://arxiv.org/html/2607.27265#S4.SS2.p1.1)\.
- Z\. Liu, Z\. Guo, Y\. Yao, Z\. Cen, W\. Yu, T\. Zhang, and D\. Zhao \(2023\)Constrained decision transformer for offline safe reinforcement learning\.InICML ’23,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- T\. Maehara, A\. Narita, J\. Baba, and T\. Kawabata \(2018\)Optimal bidding strategy for brand advertising\.InIJCAI ’18,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- S\. Muthukrishnan \(2009\)Ad exchanges: research issues\.InWINE ’09,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- F\. Ni, J\. Hao, Y\. Mu, Y\. Yuan, Y\. Zheng, B\. Wang, and Z\. Liang \(2023\)MetaDiffuser: diffusion model as conditional planner for offline meta\-rl\.InICML ’23,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- W\. Ou, B\. Chen, X\. Dai, W\. Zhang, W\. Liu, R\. Tang, and Y\. Yu \(2024\)A survey on bid optimization in real\-time bidding display advertising\.ACM Trans\. Knowl\. Discov\. Data\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- L\. Pan, Q\. Cai, and L\. Huang \(2020\)Softmax deep double deterministic policy gradients\.InNeurIPS ’20,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- S\. Park, Q\. Li, and S\. Levine \(2025\)Flow q\-learning\.CoRR\.Cited by:[§4\.2](https://arxiv.org/html/2607.27265#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2607.27265#S4.SS2.p2.3),[§5\.1](https://arxiv.org/html/2607.27265#S5.SS1.p1.1)\.
- K\. Ren, W\. Zhang, K\. Chang, Y\. Rong, Y\. Yu, and J\. Wang \(2018\)Bidding machine: learning to bid for directly optimizing profits in display advertising\.IEEE Trans\. Knowl\. Data Eng\.\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- K\. Su, Y\. Huo, Z\. Zhang, S\. Dou, C\. Yu, J\. Xu, Z\. Lu, and B\. Zheng \(2024\)AuctionNet: a novel benchmark for decision\-making in large\-scale games\.InNeurIPS ’24,Cited by:[Appendix C](https://arxiv.org/html/2607.27265#A3.p1.1),[§1](https://arxiv.org/html/2607.27265#S1.p2.1),[§2](https://arxiv.org/html/2607.27265#S2.p1.1),[§3\.2](https://arxiv.org/html/2607.27265#S3.SS2.p2.1)\.
- R\. S\. Sutton and A\. G\. Barto \(2018\)Reinforcement learning \- an introduction, 2nd edition\.Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- J\. Wang, W\. Zhang, and S\. Yuan \(2017\)Display advertising with real\-time bidding \(RTB\) and behavioural targeting\.Found\. Trends Inf\. Retr\.\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- C\. Wen, M\. Xu, Z\. Zhang, Z\. Zheng, Y\. Wang, X\. Liu, Y\. Rong, D\. Xie, X\. Tan, C\. Yu, J\. Xu, F\. Wu, G\. Chen, X\. Zhu, and B\. Zheng \(2022\)A cooperative\-competitive multi\-agent framework for auto\-bidding in online advertising\.InWSDM ’22,Cited by:[Appendix B](https://arxiv.org/html/2607.27265#A2.p5.1),[Appendix B](https://arxiv.org/html/2607.27265#A2.p6.3),[§1](https://arxiv.org/html/2607.27265#S1.p1.1),[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- S\. Yang, Y\. Li, S\. He, Y\. Li, Q\. Cai, P\. Jiang, and L\. Feng \(2026\)Phase\-aware mixture of experts for agentic reinforcement learning\.InICML ’26,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- X\. Yang, D\. Sun, R\. Zhu, T\. Deng, Z\. Guo, Z\. Ding, S\. Qin, and Y\. Zhu \(2019a\)AiAds: automated and intelligent advertising system for sponsored search\.InKDD ’19,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- X\. Yang, Y\. Li, H\. Wang, D\. Wu, Q\. Tan, J\. Xu, and K\. Gai \(2019b\)Bid optimization by multivariable control in display advertising\.InKDD ’19,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
- Y\. Yuan, F\. Wang, J\. Li, and R\. Qin \(2014\)A survey on real time bidding advertising\.InProceedings of 2014 IEEE International Conference on Service Operations and Logistics, and Informatics,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- W\. Zhang, S\. Yuan, and J\. Wang \(2014a\)Optimal real\-time bidding for display advertising\.InKDD ’14,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p1.1)\.
- W\. Zhang, S\. Yuan, and J\. Wang \(2014b\)Real\-time bidding benchmarking with ipinyou dataset\.CoRR\.Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p2.1),[§2](https://arxiv.org/html/2607.27265#S2.p1.1),[§5\.5](https://arxiv.org/html/2607.27265#S5.SS5.p1.1)\.
- S\. Zhou, Y\. Du, S\. Zhang, M\. Xu, Y\. Shen, W\. Xiao, D\. Yeung, and C\. Gan \(2023\)Adaptive online replanning with diffusion models\.InNeurIPS ’23,Cited by:[§1](https://arxiv.org/html/2607.27265#S1.p4.1)\.
- H\. Zhu, J\. Jin, C\. Tan, F\. Pan, Y\. Zeng, H\. Li, and K\. Gai \(2017\)Optimized cost per click in taobao display advertising\.InKDD ’17,Cited by:[§2](https://arxiv.org/html/2607.27265#S2.p2.1)\.
## Appendix
## Appendix AMethod Details
For the PID controller, we setKp=0\.1K\_\{p\}=0\.1,Ki=0\.01K\_\{i\}=0\.01, andKd=0\.03K\_\{d\}=0\.03in all settings\. For the DT reweight variant, the reweighting coefficient is set tow=0\.2w=0\.2across all settings\. For GAS, we employ three critic networks for both dense and sparse datasets across all settings\. For BidFlow, the BC coefficientα\\alphais set as follows:α=10\\alpha=10for dense data in Settings 1 and 3,α=0\.3\\alpha=0\.3for dense data in Setting 2,α=1\\alpha=1for sparse data in Setting 2, andα=0\.03\\alpha=0\.03for sparse data in Settings 1 and 3\.
Table 6\.Key metrics comparison under PlatformBid setting 2: Heterogeneous Competition\. Each result pair \(Target, Average\) shows the average performance of only the target advertisers by the evaluated method and the average performance of all advertisers\. Result pairs are bold if “Average” is the best\.DatasetMetricPIDIQLBCQDT scoreGASGAVECBDBidFlowDenseConversion↑\\uparrow170\.54, 254\.17309\.17, 331\.55231\.04, 265\.90333\.17, 345\.44340\.38, 347\.55355\.71, 330\.07313\.29, 303\.46346\.88, 347\.96Budget Util↑\\uparrow0\.34, 0\.650\.88, 0\.930\.85, 0\.900\.85, 0\.890\.89, 0\.930\.99, 0\.970\.98, 0\.940\.93, 0\.95CPA Ratio↓\\downarrow2\.78, 1\.910\.99, 1\.031\.19, 1\.190\.85, 0\.960\.91, 0\.970\.99, 1\.061\.20, 1\.130\.93, 0\.93Qualified Rate↑\\uparrow0\.17, 0\.400\.83, 0\.690\.54, 0\.520\.63, 0\.630\.67, 0\.670\.46, 0\.480\.50, 0\.570\.50, 0\.57Score↑\\uparrow143\.21, 216\.74285\.24, 293\.86165\.21, 190\.03321\.96, 321\.92315\.81, 315\.93287\.78, 262\.78229\.86, 233\.34312\.69, 308\.83SparseConversion↑\\uparrow11\.88, 21\.6538\.46, 32\.5937\.92, 34\.4244\.38, 37\.5036\.96, 33\.8629\.30, 32\.3242\.71, 31\.8438\.08, 28\.77Budget Util↑\\uparrow0\.51, 0\.490\.99, 0\.690\.68, 0\.660\.77, 0\.690\.75, 0\.681\.00, 0\.700\.98, 0\.670\.86, 0\.66CPA Ratio↓\\downarrow2\.08, 1\.280\.94, 0\.770\.63, 0\.690\.63, 0\.650\.72, 0\.701\.26, 0\.840\.89, 0\.770\.83, 1\.13Qualified Rate↑\\uparrow0\.25, 0\.150\.42, 0\.250\.25, 0\.290\.13, 0\.130\.33, 0\.330\.33, 0\.190\.17, 0\.190\.33, 0\.29Score↑\\uparrow9\.52, 20\.4734\.42, 30\.3137\.78, 33\.5943\.91, 36\.2636\.47, 32\.7721\.20, 28\.2938\.71, 29\.6035\.97, 26\.51
Table 7\.Key metric comparison under PlatformBid Setting 3: Promotional Competition\. “Pro” denotes promotional advertisers, “Non” denotes non\-promotional advertisers, and “All” denotes the average performance across all advertisers\.DatasetTypeMetricPIDIQLBCQDTDT scoreGASGAVECBDBidFlowDenseProConversion↑\\uparrow550\.71674\.00543\.14673\.14682\.57690\.43657\.71456\.86727\.29CPA Ratio↓\\downarrow1\.491\.101\.151\.060\.980\.951\.151\.320\.91Exceed↓\\downarrow0\.570\.430\.710\.710\.420\.280\.860\.710\.29Score↑\\uparrow519\.80587\.70429\.46555\.49634\.28661\.49456\.13347\.37701\.48NonConversion↑\\uparrow124\.93294\.10234\.17307\.85316\.43319\.95304\.32252\.17310\.61CPA Ratio↓\\downarrow1\.711\.021\.201\.121\.010\.951\.201\.370\.89Exceed↓\\downarrow0\.680\.490\.850\.610\.510\.410\.660\.710\.34Score↑\\uparrow111\.46247\.66171\.42230\.79271\.85277\.90203\.71169\.46297\.25AllConversion↑\\uparrow187\.02349\.50279\.23361\.12371\.29373\.98355\.85282\.02371\.38CPA Ratio↓\\downarrow1\.681\.031\.191\.111\.010\.951\.201\.360\.90Exceed↓\\downarrow0\.670\.480\.830\.630\.490\.400\.690\.710\.69Score↑\\uparrow171\.01297\.25209\.05278\.14321\.73333\.84240\.52195\.41356\.20SparseProConversion↑\\uparrow24\.8671\.0054\.5726\.0064\.7178\.8648\.5760\.8671\.86CPA Ratio↓\\downarrow0\.470\.950\.670\.960\.650\.711\.490\.750\.68Exceed↓\\downarrow0\.380\.140\.000\.290\.000\.000\.860\.140\.00Score↑\\uparrow24\.8564\.2454\.5725\.3864\.7178\.8524\.7059\.5771\.86NonConversion↑\\uparrow5\.9527\.4417\.4915\.0523\.9023\.9322\.9329\.2728\.80CPA Ratio↓\\downarrow1\.370\.980\.811\.120\.861\.141\.640\.990\.92Exceed↓\\downarrow0\.440\.390\.200\.370\.320\.410\.780\.440\.29Score↑\\uparrow5\.8622\.5016\.9313\.0023\.1021\.0811\.7924\.5825\.67AllConversion↑\\uparrow8\.7133\.7922\.9016\.6529\.8531\.9426\.6733\.8835\.08CPA Ratio↓\\downarrow1\.240\.980\.791\.100\.831\.081\.620\.950\.88Exceed↓\\downarrow0\.380\.350\.170\.350\.270\.350\.790\.400\.25Score↑\\uparrow8\.6328\.5922\.4214\.8129\.1729\.5113\.6729\.6832\.41
## Appendix BAdditional Analysis on Setting 2 Sparse
Under Setting 2 with sparse rewards, BidFlow underperforms DT score on the target group \(35\.97 vs\. 43\.91 in Table[6](https://arxiv.org/html/2607.27265#A1.T6)\)\. We provide two complementary ablations to disentangle the underlying cause, and clarify how this case relates to deployment\-time conditions\.
Context\.In production, the Kuaishou advertising environment contains both dense conversion types \(e\.g\., app invocations, where a user jumps from Kuaishou to an external app and produces frequent reward signals\) and sparse conversion types \(e\.g\., lead submissions, which are rare events\)\. Setting 2 Sparse isolates the purely sparse case as a stress test, which does not reflect the mixed conversion distribution encountered at deployment\. Setting 2 Sparse should therefore be read in conjunction with the dense counterpart, where BidFlow is competitive\.
Ablation 1: Effect of inference architecture and Q\-guidance\.To test whether the one\-step inference architecture limits BidFlow’s expressiveness under sparse rewards, we compare three variants that share the same training data and BC flow backbone: the paper’s one\-step model withα=1\\alpha=1, a one\-step variant with zero Q\-guidance \(α=0\\alpha=0\), and the multi\-step 50\-step teacher used to produce distillation targets\.
Table 8\.Ablation of inference architecture and Q\-guidance on Setting 2 Sparse\.VariantScore \(Target\)Score \(Avg\)CPA Ratio1\-step,α=1\\alpha=1\(ours\)35\.9726\.510\.831\-step,α=0\\alpha=029\.9222\.150\.8850\-step teacher30\.4422\.691\.01The 1\-stepα=0\\alpha=0variant \(29\.92\) and the 50\-step teacher \(30\.44\) yield nearly identical scores while differing only in inference depth\. This rules out inference architecture as the primary cause of the gap to DT score\. The lift fromα=0\\alpha=0toα=1\\alpha=1\(\+6 points\) further confirms that Q\-value guidance contributes meaningfully, but does not fully close the gap, motivating a separate explanation\.
Ablation 2: Floor price sensitivity and collusive equilibrium\.We note that DT score’s CPA ratio in Setting 2 Sparse is only 0\.63, well below the constraint, which indicates that it is not competing aggressively for impressions\. Since DT score and the DT baseline share the same transformer backbone and differ only mildly in their reward design, they tend to produce correlated bidding behavior and form a low\-price equilibrium that benefits DT\-based strategies\. This is consistent with the collusive bid suppression phenomenon documented in MAAB\(Wenet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib17)\), where agents with similar architectures trained on shared data develop correlated low\-bid strategies that suppress market prices\.
To counter collusive bid suppression, floor price mechanisms are standard practice\(Wenet al\.,[2022](https://arxiv.org/html/2607.27265#bib.bib17)\)\. In PlatformBid we apply a uniform floor price ratio \(FPR\) of0\.80×0\.80\\timesthe historical market price across all settings, analogous to MAAB\-fix\. To test whether DT score’s advantage is sensitive to this floor, we increase FPR from0\.850\.85to0\.900\.90while holding all other components fixed\.
Table 9\.Floor price sensitivity on Setting 2 Sparse \(Heterogeneous\)\. Score reported as Target / Avg\.FPRMethodScore \(Target\)Score \(Avg\)CPA Ratio0\.85DT score25\.7920\.070\.750\.85BidFlow31\.9624\.350\.900\.90DT score24\.7120\.330\.870\.90BidFlow27\.9020\.710\.93Table 10\.Floor price sensitivity on Setting 1 and Setting 3 Sparse\.SettingMethod \(FPR=0\.85 / 0\.90\)ScoreCPA RatioS1 \(Homogeneous\)DT score \(0\.85\)25\.750\.92BidFlow \(0\.85\)28\.570\.88DT score \(0\.90\)25\.330\.83BidFlow \(0\.90\)26\.230\.87S3 \(Promotional\)DT score \(0\.85\)26\.13—BidFlow \(0\.85\)30\.780\.86DT score \(0\.90\)24\.390\.85BidFlow \(0\.90\)28\.880\.93The results show a clear contrast between Setting 2 and the other two\. In Setting 2, DT score’s Avg Score degrades faster than BidFlow’s as FPR is raised \(DT score:20\.07→20\.3320\.07\\rightarrow 20\.33; BidFlow:24\.35→20\.7124\.35\\rightarrow 20\.71\), and BidFlow leads at both FPR levels\. In Settings 1 and 3, BidFlow leads DT score at every FPR value\. This indicates that DT score’s advantage under sparse rewards is exclusive to the heterogeneous setting, where correlated bidding with the DT baseline population produces a low\-price equilibrium, rather than a general superiority over BidFlow\. Adaptive floor price mechanisms such as MAAB’s learned bar agent are orthogonal to PlatformBid’s evaluation focus and constitute a natural direction for future work\.
## Appendix CDiscussion on Data Realism
The realism of PlatformBid’s underlying data is a natural concern for a benchmark intended to guide deployment decisions\. PlatformBid is currently built on AuctionNet\(Suet al\.,[2024](https://arxiv.org/html/2607.27265#bib.bib9)\), which has been validated against real\-world traffic from multiple complementary perspectives\. At the feature level, the offline and online data distributions overlap substantially and exhibit the same cluster structure\. At the value level, predicted pCVR distributions are consistent with those observed online\. At the behavioral level, consumption distributions follow the same long\-tail patterns as live traffic\. AuctionNet is therefore a representative source for constructing platform\-side auto\-bidding benchmarks\.
A second concern is whether external bidders with budgets and strategies outside the auto\-bidding service affect the conclusions\. On modern large\-scale platforms, such bidders typically correspond to Real\-Time API \(RTA\) advertisers, who interact with the platform via API on a per\-impression basis\. On platforms such as Kuaishou, the majority of advertisers adopt the auto\-bidding service due to its effectiveness, and the RTA share remains small\. This minority is unlikely to alter benchmark\-level conclusions about auto\-bidding methods\.
Finally, the strongest evidence for real\-world effectiveness is the online deployment in Section 5\.3: BidFlow improves target cost by\+0\.68%\+0\.68\\%on Kuaishou’s production system, where the realized traffic distribution differs from the offline training data\. This offline–online consistency supports the benchmark as a meaningful proxy for production decisions\.Similar Articles
OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
OneBid is a unified auto-bidding foundation model that addresses challenges in diverse oCPX advertising scenarios by using Mixture-of-Experts architecture and CROP optimization, with deployment at Kuaishou showing significant improvements like +13.1% in ROAS.
Generative Auto-Bidding with Unified Modeling and Exploration
This paper introduces Guide, a framework that combines a Decision Transformer with Q-value guidance and an inverse dynamics module to balance exploration and safety in automated bidding for digital advertising, demonstrating effectiveness on public datasets and simulated auctions.
Beyond Single Slot: Joint Optimization for Multi-Slot Guaranteed Display Advertising
Proposes a joint optimization framework for multi-slot guaranteed display advertising, addressing slot-level redundancy and contract imbalance via bipartite matching and contract roulette. Online A/B tests on Meituan show significant improvements in revenue and contract fulfillment.
HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA proposes a hierarchical reinforcement learning framework for online advertising that uses a large language model for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% improvement in a large-scale A/B test.
FlowMarket
FlowMarket is a new product launched on Product Hunt described as a social network for AI agents designed to generate B2B deals.