LoopCTR: Unlocking the Loop Scaling Power for Click-Through Rate Prediction
Summary
LoopCTR introduces loop scaling to recommendation models, using MoE-based expert mixing and hyper-connected residuals to boost CTR prediction while allowing train-deep/infer-shallow deployment for low-latency serving.
View Cached Full Text
Cached at: 04/22/26, 06:17 AM
Paper page - LoopCTR: Unlocking the Loop Scaling Power for Click-Through Rate Prediction
Source: https://huggingface.co/papers/2604.19550 🔥 Recently, OpenMythos has been making waves in the AI community with itsRecurrent-Depth Transformer, showing that scaling does not have to rely solely on stacking more layers or adding more parameters. Instead, recursive computation with shared parameters can also effectively enhance a model’s reasoning capability.
Interestingly, we have just completed a new study in the recommender systems domain:LoopCTR. To the best of our knowledge,this is the first work to systematically explore loop scaling 🔁 in recommendation models.
However, loops in recommendation scenarios cannot simply reuse “the same layer over and over” ♻️. Naively sharing parameters may lead to limited expressiveness, while a fixed computation flow struggles to adapt to different samples and different loop depths.
To address this, we introduce two key designs in the Loop Block:
🧩MoE-based expert mixing: expands the expressive capacity of the shared layer, allowing a single layer to carry richer parameter capacity.
🕸️Hyper-connected residual structure: enables input-aware dynamic computation allocation, breaking the limitations of fixed residual information flow.
On top of this, LoopCTR incorporates intermediate supervision 🔍, which implicitly strengthens self-distillation while significantly reducing online inference latency.
🚀Train deep, infer shallow: the model can be trained with multiple loops but deployed with fewer loops, or even zero-loop inference, making it highly suitable for the strict latency constraints of industrial recommendation systems.
🔭Train shallow, infer deep: somewhat counterintuitively, our oracle analysis shows that models trained with shallow loops can achieve even higher performance ceilings under deeper inference settings.
This also suggests that different samples may require different computation depths. Adaptive implicit loop inference remains a highly promising direction, although our attempts with various strategies have not yet led to a fully effective solution 😢.
The experimental results are strong 💥. Even under the zero-loop inference setting, LoopCTR consistently outperforms baseline models, achieving significant performance gains with extremely low online serving overhead ⚡. This makes it highly practical for industrial deployment 🏭.
In short, LoopCTR moves beyond the traditional “just add more parameters” scaling paradigm in recommendation systems. It opens upa new dimension of loop scaling, leveraging shared-parameter architectures with better inductive biases to enable deeper, more flexible implicit reasoning.
Our oracle experiments further show that there is still substantial untapped potential in existing approaches. How to achieveadaptive and efficient latent reasoning, and fully unlock the upper bound of loop scaling, remains an exciting open problem worth exploring.🤔
Similar Articles
Mitigating Early Training Collapse in CTR Models
This paper analyzes the early training collapse phenomenon in deep neural models for click-through rate prediction and proposes mitigation strategies such as sparse feature removal and value filtering, demonstrating improvements on large-scale industrial datasets.
From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation
This paper presents a decision-centric causal optimization framework for large-scale targeting and recommendation, combining a causal Transformer, Bayesian bandit layer, and dual-based linear programming. It reports a statistically significant +7.20% lift in LinkedIn Feed marketing traffic via online A/B testing.
DeepLoop: Depth Scaling for Looped Transformers
DeepLoop introduces a residual scaling method for looped Transformers that adjusts for parameter visits, improving stability and performance when physical blocks are reused across multiple rounds.
MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction
MARCO is a Meta AI framework that decomposes clicks by intent to improve ads conversion prediction, correcting per-intent calibration bias and lifting conversions per click by +2.80% and topline metrics by +0.98% in production.
20B Looping model (paper) matches or beats Qwen3 Coder 30B at 10% of pre-training tokens
Loopie models use a looped transformer architecture to match or exceed Qwen3 Coder 30B performance with only 10% of the pre-training tokens, demonstrating strong reasoning abilities and efficient scaling.