Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Summary
This paper introduces Knowledge-Geometry Decoupling (KGD), a method for pretrain-then-transfer in streaming recommendation systems. It separates pretrained behavioral knowledge from task-specific geometry, enabling continual model refresh without interference, and reports 4-12% improvements over baselines plus successful deployment at Shopee.
View Cached Full Text
Cached at: 08/05/26, 05:43 AM
Paper page - Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
Source: https://huggingface.co/papers/2608.02738 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
Industrialrecommendersincreasinglyadoptthepretrain-then-transferparadigm,yetbehavioraldistributiondriftraisestwoquestions:whattolearnfrombehaviorsequences,andhowtotransferthelearnedknowledgewhilethepretrainedmodeliscontinuallyrefreshed.Toresolvethem,weproposeKnowledge-GeometryDecoupling(KGD).Forwhattolearn,conventionalnext-tokenpredictiontreatsadjacencyasdependencyandmayencodespurioustransitionsacrossunrelatedsessions.WeintroduceBehavioralMulti-TokenPrediction(BMTP)toretainonlycollaborativelyorsemanticallyrelatedfutureitemsassupervision,yieldingcleanerandmoretransferablebehavioralknowledge.Forhowtotransfer,pretrainedknowledgeandtask-specificgeometryimposeconflictingoptimizationdemandsonsharedparameters.Tohandleit,KGDassignsthemtoseparateparametersets:arefreshableencoderownsbehavioralknowledge,whileatasklearnerreadscontextualizedencoderstatesthroughread-onlycross-attentionandwritestask-specificgeometrythroughAnchoredCalibrationResidual(ACR)orthogonaltothepretrainedembedding.Thedecoupledownershipenablescontinualknowledgerefreshwithouttask-gradientinterferenceorinvalidatingdownstreamadaptation.KGDimprovesoverstrongpretrain-transferbaselinesby4-12%oneightpublicbenchmarksandsustainsitsadvantageovera90-dayproductionstreamwherebaselinesshownogains.KGDhasbeenfullydeployedinShopee.InaliveA/BtestonShopeeHomepageSearch,itincreasesGMVperuserby1.75%andadvertisingrevenueby1.53%,demonstratingitshighpracticalvalue.WeprovidethecoreimplementationofKGDathttps://github.com/FuCongResearchSquad/KGD4REC.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2608\.02738
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.02738 in a model README.md to link it from this page.
Datasets citing this paper1
#### PIIR/KGD-dataset Viewer• Updatedabout 1 hour ago • 18.1M • 21
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.02738 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation
This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
KDFlow is a novel knowledge distillation framework for large language models that uses a decoupled architecture with SGLang for teacher inference and FSDP2 for student training, achieving 1.44x to 6.36x speedup over existing frameworks.
SNAP-KG: Streaming Node Assignment via Projection for Knowledge Graph Entity Integration
SNAP-KG is a framework that enables efficient integration of streaming entities into knowledge graphs via multi-view clustering and inductive inference, reducing inference time and candidate search space for downstream tasks like entity resolution and link prediction.
GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data Attribution
GRASP introduces a geometry-aware, interaction-based method for scalable pretraining data attribution that models subset dynamics, outperforming existing additive approaches by over double the task-level rank correlation while reducing computation costs.
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
Stream3D-VLM is an online 3D vision-language model that enables real-time spatial understanding from streaming video by incrementally integrating geometry priors and using geometry-adaptive voxel compression, outperforming existing models on 3D spatial understanding tasks.