Brain-Inspired Hierarchical Modularity for General Continual Learning
Summary
The paper proposes a brain-inspired hierarchical modular approach for continual learning to handle online and uncertain data streams, achieving significant performance gains in tasks like embodied manipulation by leveraging pretrained foundation models.
View Cached Full Text
Cached at: 09/23/26, 09:27 AM
# Brain-inspired hierarchical modularity for general continual learning Source: [https://arxiv.org/html/2609.25146](https://arxiv.org/html/2609.25146) 2021 Hongwei YanAffiliation:School of Life Sciences, Tsinghua University, Beijing, ChinaAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Tsinghua\-Peking Center for Life Sciences, Beijing, ChinaKanglei ZhouAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, ChinaQi ChengAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, ChinaWeiyi DongAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, ChinaChunyan LanAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, ChinaGuanglong SunAffiliation:School of Life Sciences, Tsinghua University, Beijing, ChinaAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Tsinghua\-Peking Center for Life Sciences, Beijing, ChinaJun ZhouAffiliation:School of Life Sciences, Tsinghua University, Beijing, ChinaAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Tsinghua\-Peking Center for Life Sciences, Beijing, ChinaQian LiYi ZhongAffiliation:School of Life Sciences, Tsinghua University, Beijing, ChinaAffiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Tsinghua\-Peking Center for Life Sciences, Beijing, ChinaLiyuan WangEmail:[liyuanwang@tsinghua\.edu\.cn](mailto:[email protected])Affiliation:IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, ChinaAffiliation:Department of Psychological and Cognitive Sciences, Tsinghua University, Beijing, China ###### Abstract Continual learning, the ability to learn from sequential experience while retaining and adapting prior knowledge, is central to intelligent systems operating in changing environments\. However, conventional continual learning is typically studied with offline task\-wise training and clear task boundaries, leaving a substantial gap from general continual learning under online, uncertain, and evolving data streams\. In this regime, intelligent systems must separate conflicting experience to reduce interference while integrating compatible experience to promote generalization\. Inspired by the organization of the*Drosophila*learning and memory system, we identify a hierarchical modular principle that coordinates both functions through expert specialization and ensemble integration\. We instantiate this principle as lightweight modular adaptation of pretrained foundation models, combining brain\-inspired random expansion for expert routing and diversified modular integration across spatial and temporal scales\. Across visual recognition, vision\-language understanding, ego\-exo video understanding, and embodied vision\-language\-action learning, our method consistently improves learning under online and uncertain data streams, with gains exceeding 50 percentage points over replay\-free alternatives in embodied manipulation\. These findings support hierarchical modularity as a biologically grounded path for learning from dynamic experience\. ###### keywords neuro\-inspired learning, continual learning, learning and memory, catastrophic forgetting, adaptability ††equal\-contributors:These authors contributed equally to this work\.††equal\-contributors:These authors contributed equally to this work\.## 1Introduction Continual learning \(CL\)[Wang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib13);[De Lange et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib63)is a defining process through which intelligence learns, develops, and accumulates knowledge over time\. In biological organisms[Davis \(2023\)](https://arxiv.org/html/2609.25146#bib.bib31);[Li et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib32), learning from sequential experience supports immediate responses to environmental change and progressive development throughout the lifespan, enabling long\-term adaptation to changing conditions\. A similar capability is increasingly central to artificial intelligence \(AI\): moving beyond intelligence acquired primarily from static, human\-curated data requires systems that can continue to learn from their own experience, despite catastrophic forgetting[McClelland et al\. \(1995\)](https://arxiv.org/html/2609.25146#bib.bib3);[Wang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib13)and loss of plasticity[Wang et al\. \(2021a\)](https://arxiv.org/html/2609.25146#bib.bib6);[Dohare et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib53)\. Recent perspectives on an “era of experience”[Silver and Sutton \(2025\)](https://arxiv.org/html/2609.25146#bib.bib59);[LeCun \(2022\)](https://arxiv.org/html/2609.25146#bib.bib60);[Hughes et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib61)envision increasingly general agents whose capabilities emerge through persistent interaction with the external world\. Emerging directions on self\-improving agents[Zhang et al\. \(2026\)](https://arxiv.org/html/2609.25146#bib.bib69)and test\-time training[Zweiger et al\. \(2026\)](https://arxiv.org/html/2609.25146#bib.bib70);[Behrouz et al\. \(2026\)](https://arxiv.org/html/2609.25146#bib.bib71)similarly point towards systems that continue to refine their behaviour and internal knowledge after deployment\. Most existing AI studies, however, formulate CL in a simplified conventional regime, typically within a narrow task setting and with largely offline training, clear task boundaries, and auxiliary task identities[Zhou et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib33);[Wang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib19);[Wang et al\. \(2022a\)](https://arxiv.org/html/2609.25146#bib.bib11)\. These assumptions have enabled substantial progress through synaptic regularization[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1);[Wang et al\. \(2021a\)](https://arxiv.org/html/2609.25146#bib.bib6), memory replay[Buzzega et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib29);[Zhou et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib34), and dynamic architecture[Wang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib19);[Wang et al\. \(2023a\)](https://arxiv.org/html/2609.25146#bib.bib35)\. Real\-world experience instead arrives online under uncertain, overlapping, and evolving distributions across diverse models, modalities, and application scenarios\. We consider this broader regime as*general continual learning*\(GCL\)[De Lange et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib63);[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)\. Its central challenge extends beyond retaining past knowledge to a more fundamental question:how should learning be organized as data distributions evolve over time?Conflicting experience should be separated to reduce interference, whereas compatible experience should be integrated to exploit shared structure and promote generalization\. Biological organisms naturally learn under such dynamic conditions, providing a useful reference for GCL\. Among model organisms,*Drosophila*is particularly tractable because its learning circuits combine rich adaptive behaviour with increasingly detailed anatomical characterization from whole\-brain connectomes and cross\-connectome cell typing[Modi et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib9);[Davis \(2023\)](https://arxiv.org/html/2609.25146#bib.bib31);[Li et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib32);[Winding et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib50);[Lin et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib48);[Schlegel et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib51)\. In the olfactory learning and memory system, sparse, largely random projections expand sensory representations in Kenyon cells and support pattern separation[Caron et al\. \(2013\)](https://arxiv.org/html/2609.25146#bib.bib14);[Honegger et al\. \(2011\)](https://arxiv.org/html/2609.25146#bib.bib41);[Aso et al\. \(2014b\)](https://arxiv.org/html/2609.25146#bib.bib8);[Dasgupta et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib37), while downstream learning and memory are distributed across differentiated compartments with distinct spatial and temporal characteristics[Aso et al\. \(2014a\)](https://arxiv.org/html/2609.25146#bib.bib45);[Aso and Rubin \(2016\)](https://arxiv.org/html/2609.25146#bib.bib7);[Cohn et al\. \(2015\)](https://arxiv.org/html/2609.25146#bib.bib10);[Handler et al\. \(2019\)](https://arxiv.org/html/2609.25146#bib.bib12);[Cervantes\-Sandoval et al\. \(2013\)](https://arxiv.org/html/2609.25146#bib.bib38)\. Computationally, we relate these biological mechanisms to two classical paradigms of modular machine learning:*mixture\-of\-experts*\(MoE\)[Mu and Lin \(2025\)](https://arxiv.org/html/2609.25146#bib.bib39);[Jacobs et al\. \(1991\)](https://arxiv.org/html/2609.25146#bib.bib43)and*ensemble learning*\(EL\)[Dong et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib40);[Hansen and Salamon \(2002\)](https://arxiv.org/html/2609.25146#bib.bib44)\. MoE promotes specialization across dissimilar distributions to reduce interference, whereas EL integrates diversified information over related distributions to promote generalization\. Importantly, the*Drosophila*system coordinates these MoE\- and EL\-like functions through a hierarchical learning and memory organization rather than deploying them independently[Wang et al\. \(2023b\)](https://arxiv.org/html/2609.25146#bib.bib52);[Wang and Li \(2025\)](https://arxiv.org/html/2609.25146#bib.bib42)\. This hierarchical coordination provides a biological reference for jointly organizing specialization and integration in GCL\. Here we propose FlyGCL, a unified brain\-inspired framework for GCL with pretrained foundation models\. In*Drosophila*, learning and memory operate downstream of relatively stable sensory processing\. Analogously, FlyGCL retains the pretrained backbone as a stable representational substrate and organizes lightweight, parameter\-efficient learning downstream\. Brain\-inspired random expansion of pretrained representations improves instance\-level routing among specialized experts, while differentiated adaptive modules and multi\-timescale predictions introduce complementary diversity across spatial and temporal dimensions for ensemble integration\. These components realize hierarchical coordination between routing\-based specialization and spatial\-temporal integration\. This modular design accommodates different lightweight learning modules and is broadly applicable across pretrained backbones and learning settings\. Computational analyses further characterize the complementary gains of specialization and integration, and the benefit of organizing them over stable pretrained representations \(Methods\)\. We evaluate FlyGCL across diverse forms of real\-world continual experience, spanning visual recognition, vision\-language understanding, ego\-exo video understanding, and embodied vision\-language\-action learning, all under online and uncertain data streams \(Supplementary[Tabs\.S1](https://arxiv.org/html/2609.25146#A2.T1)and[S2](https://arxiv.org/html/2609.25146#A4.T2)\)\. FlyGCL consistently improves CL performance across these scenarios and pretrained models\. The gains are particularly pronounced in embodied vision\-language\-action learning\. Across spatial, object\-centric, goal\-conditioned, and long\-horizon manipulation, FlyGCL achieves final average success rates of 83\.1%, 86\.1%, 94\.5%, and 79\.1%, respectively, exceeding the strongest replay\-free CL baseline on each benchmark by 49\.5–58\.8 percentage points\. Together, these results support hierarchical modularity as a biologically grounded principle for organizing learning from dynamic experience\. abcFigure 1:Brain\-inspired hierarchical modular framework for general continual learning \(GCL\)\.[1](https://arxiv.org/html/2609.25146#S1.F1): Real\-world environments induce online and uncertain streams, where data distributions evolve and previously seen concepts may recur over time\.[1](https://arxiv.org/html/2609.25146#S1.F1): Biological inspiration from theDrosophilaolfactory learning and memory system\. Olfactory receptor neurons \(ORNs\) project to projection neurons \(PNs\) and are sparsely expanded into Kenyon cells \(KCs\), motivating hierarchical modular learning\. The upper branch illustrates ensemble learning, where multiple readoutsA1,A2,A3A\_\{1\},A\_\{2\},A\_\{3\}share the same distributionDAD\_\{A\}, while the lower branch illustrates mixture\-of\-experts, where different distributionsDA,DB,DCD\_\{A\},D\_\{B\},D\_\{C\}are assigned to specialized experts\.[1](https://arxiv.org/html/2609.25146#S1.F1): Brain\-inspired framework for GCL\. Online data streams are processed by a pretrained model and adapted through sparse expansion with nonlinear activation, expert routing, and multi\-timescale expert integration, where routing is solved by closed\-form ridge regression in the latent space\. The resulting representations support diverse downstream tasks\. ## 2Results We study*general continual learning*\(GCL\) under online, uncertain, and evolving data streams across diverse models, modalities, and application scenarios\. Its central challenge is to organize incoming experience by separating conflicting distributions to reduce interference while integrating compatible ones to promote generalization \([Fig\.1](https://arxiv.org/html/2609.25146#S1.F1)\)\. FlyGCL addresses this challenge through brain\-inspired hierarchical modularity\. We first examine its biological and computational basis, and then evaluate its generality under perceptual and embodied learning scenarios\. ### 2\.1Biological and computational basis of hierarchical modular framework  abcdefghi Figure 2:Biologically grounded evaluation of brain\-inspired hierarchical modularity\.[2](https://arxiv.org/html/2609.25146#S2.F2):Drosophila\-inspired architecture with ORN\-PN compression, sparse PN\-KC expansion, spatially differentiated experts, and temporal ensemble learning\. ORNs, olfactory receptor neurons; PNs, projection neurons; KCs, Kenyon cells; MoE, mixture\-of\-experts; EL, ensemble learning\.[2](https://arxiv.org/html/2609.25146#S2.F2): Spatial partitioning of 100 classes into five regions according to prototype proximity\.[2](https://arxiv.org/html/2609.25146#S2.F2): Area under the exposed\-class anytime accuracy curve \(AUC\) of the baseline, EL, MoE, and MoE\+EL under disjoint, blurry, and joint settings\.[2](https://arxiv.org/html/2609.25146#S2.F2): Exposed\-class accuracy of the four methods over the blurry setting\.[2](https://arxiv.org/html/2609.25146#S2.F2): Final accuracy of individual experts, baseline, and MoE across five spatial regions\.[2](https://arxiv.org/html/2609.25146#S2.F2)and[2](https://arxiv.org/html/2609.25146#S2.F2): Anytime AUC with a single head, equal\-rate ensemble, and temporal EL, with or without MoE, under disjoint and blurry settings, respectively\.[2](https://arxiv.org/html/2609.25146#S2.F2): FlyWire ORN\-PN connection weights and their truncated log\-normal approximation; inset, 50 groups with 26 ORNs converging onto one PN\.[2](https://arxiv.org/html/2609.25146#S2.F2): FlyWire PN\-KC connection weights and their truncated Gaussian approximation; inset, connection sparsity\. All results are averaged over five independent runs; error bars and shaded areas denote 95% confidence intervals\.The*Drosophila*olfactory system provides a compact biological model of learning from continuously varying sensory experience \([Fig\.1](https://arxiv.org/html/2609.25146#S1.F1)\)\. Odor signals are encoded by around 1300 olfactory receptor neurons \(ORNs\) of 50 groups and around 150 projection neurons \(PNs\) of 50 types[Wang et al\. \(2021b\)](https://arxiv.org/html/2609.25146#bib.bib47), and then transmitted through sparse, largely random PN\-KC projections into a substantially expanded population of around 2,000 Kenyon cells \(KCs\)[Caron et al\. \(2013\)](https://arxiv.org/html/2609.25146#bib.bib14);[Honegger et al\. \(2011\)](https://arxiv.org/html/2609.25146#bib.bib41);[Fulton et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib64)\. This transformation produces sparse, distributed representations that reduce overlap between odor patterns and support pattern separation[Aso et al\. \(2014b\)](https://arxiv.org/html/2609.25146#bib.bib8);[Dasgupta et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib37)\. Downstream of this relatively stable sensory representation, learning and memory are organized through theγ\\gamma,α′/β′\\alpha^\{\\prime\}/\\beta^\{\\prime\}, andα/β\\alpha/\\betaKC lobes, each partitioned into five anatomically differentiated compartments \(a total of 15 modules across three lobes\) regulated by distinct dopaminergic and output pathways[Aso et al\. \(2014a\)](https://arxiv.org/html/2609.25146#bib.bib45);[Aso and Rubin \(2016\)](https://arxiv.org/html/2609.25146#bib.bib7);[Cohn et al\. \(2015\)](https://arxiv.org/html/2609.25146#bib.bib10)\. These compartments provide spatial diversification within each lobe, while the three lobes contribute preferentially to short\-term memory, intermediate memory consolidation, and long\-term memory, respectively[Handler et al\. \(2019\)](https://arxiv.org/html/2609.25146#bib.bib12);[Cervantes\-Sandoval et al\. \(2013\)](https://arxiv.org/html/2609.25146#bib.bib38)\. The mushroom body therefore organizes learning and memory along two complementary axes: spatial compartmentalization within KC lobes and temporal differentiation across them\. We interpret this biological organization computationally as a hierarchy of two complementary paradigms in modular machine learning \([Figs\.1](https://arxiv.org/html/2609.25146#S1.F1)and[1](https://arxiv.org/html/2609.25146#S1.F1)\)\. At the first level, sparse random expansion separates sensory representations and enables selective recruitment of downstream pathways, providing a biological analogue of*mixture\-of\-experts*\(MoE\) routing for reducing interference between dissimilar distributions\. At the second level, differentiated compartments provide parallel spatial memory pathways, while theγ\\gamma,α′/β′\\alpha^\{\\prime\}/\\beta^\{\\prime\}, andα/β\\alpha/\\betalobes span progressively longer memory timescales\. Coordinating information across these spatial and temporal dimensions resembles*ensemble learning*\(EL\), which exploits diversity and integration to improve generalization over related distributions\. Computationally, MoE\-like specialization separates conflicting experience, whereas EL\-like integration combines compatible information across spatial and temporal memory components\. Their hierarchical coordination yields the central principle of GCL: separating conflicting experience while integrating compatible experience\. FlyGCL instantiates this biological organization in pretrained foundation models \([Fig\.1](https://arxiv.org/html/2609.25146#S1.F1), Methods\)\. The pretrained backbone serves as a relatively stable representational substrate, analogous to upstream sensory processing, while brain\-inspired random expansion of its representations enables instance\-level routing among specialized experts\. Diversified adaptive components provide spatial variation, whereas prediction heads with different effective memory windows provide temporal variation in each expert\. Fast predictions emphasize recent experience, slow predictions preserve information over longer timescales, and intermediate heads bridge the two, forming a computational analogue of short\-term, consolidation\-related, and long\-term memory\. This hierarchical design can be implemented with lightweight adaptation such as prompts, adapters, or LoRA[Lester et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib73);[Rebuffi et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib75);[Hu et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib72), allowing FlyGCL to operate across different pretrained models and learning scenarios\. Computational analysis further supports the complementary roles of specialization and integration, and the benefit of coordinating them on stable pretrained representations \([Fig\.1](https://arxiv.org/html/2609.25146#S1.F1), Methods\)\. To examine the hierarchical modular principle in a controlled biologically grounded setting, we construct a continual olfactory learning model that follows the population scale and modular organization of the*Drosophila*olfactory system \([Fig\.2](https://arxiv.org/html/2609.25146#S2.F2), Supplementary[Sec\.B\.1](https://arxiv.org/html/2609.25146#A2.SS1)\)\. Following prior task\-driven models of olfactory learning[Wang et al\. \(2021b\)](https://arxiv.org/html/2609.25146#bib.bib47);[Shen et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib49), the network contains 1,300 ORNs, 50 types of PNs, and 2,000 KCs\. ORN\-PN and PN\-KC connections are sampled using statistics derived from FlyWire[Dorkenwald et al\. \(2022\)](https://arxiv.org/html/2609.25146#bib.bib68)\([Figs\.2](https://arxiv.org/html/2609.25146#S2.F2)and[2](https://arxiv.org/html/2609.25146#S2.F2)\), and only the 5% most active KCs are retained for each odor\. The sensory pathway remains fixed, restricting CL to downstream learning and memory components\. We organize these components along the same spatial\-temporal hierarchy: five parallel experts represent memory pathways differentiated by spatial regions, while three temporal heads capture fast, intermediate, and slow effective memory scales\. We generate 100 classes from prototypes in a 50\-dimensional sensory space and assign nearby prototypes to five equally sized spatial regions\. Each online data stream contains 50,000 unique samples presented over five stages, ranging from strictly separated classes in Disjoint to increasing cross\-stage overlap in Blurry and Joint \(50% and 100% classes overlap, respectively\)\. We compare four targeted baselines that isolate the two dimensions of this organization: a naive baseline with a single adaptive pathway; MoE with five spatially differentiated experts and an expert router based on accumulated stage prototypes; EL with three prediction heads operating at distinct effective timescales; and the hierarchical model combining five experts with three temporal heads each \([Fig\.2](https://arxiv.org/html/2609.25146#S2.F2)\)\. Across stream configurations, MoE and EL provide complementary benefits, while their hierarchical combination consistently performs best \([Figs\.2](https://arxiv.org/html/2609.25146#S2.F2)and[2](https://arxiv.org/html/2609.25146#S2.F2)\)\. Under the disjoint setting, the area under the curve \(AUC\) performance increases from 28\.9% for the baseline to 39\.9% with MoE and 43\.2% with MoE\+EL\. Under the blurry setting, MoE\+EL reaches 30\.1%, compared with 19\.1% for the baseline, and remains consistently stronger over the course of learning\. The same ordering holds under the joint setting, where MoE\+EL also outperforms either component alone\. The expert analysis provides direct evidence of specialization \([Fig\.2](https://arxiv.org/html/2609.25146#S2.F2)\)\. Each expert is most accurate in its corresponding region, whereas its accuracy is low elsewhere\. Routing these specialized experts produces a mean final accuracy of 29\.1% across regions, compared with 19\.2% for the shared baseline\. Temporal diversity provides an additional consistent gain \([Figs\.2](https://arxiv.org/html/2609.25146#S2.F2)and[2](https://arxiv.org/html/2609.25146#S2.F2)\)\. For example, under the blurry setting with MoE, anytime AUC increases from 27\.4% with a single head to 28\.4% with three equal\-rate heads and 30\.1% when the heads use different learning rates\. This progression holds both with and without MoE in the disjoint and blurry settings, supporting distinct contributions from expert specialization and temporal integration\. ### 2\.2General Continual Learning for Visual and Vision\-Language Perception Real\-world intelligent systems continuously encounter changing perceptual experience, from evolving visual concepts to multimodal observations grounded in language\. Continual visual recognition and continual vision\-language learning therefore provide representative scenarios for studying how models preserve shared structure while adapting to distribution\-specific changes over time\. Although both scenarios have been widely studied in conventional CL, existing efforts largely rely on offline and disjoint task sequences\. We revisit them under more realistic online and blurry data streams\.  abcd Figure 3:Results on continual visual recognition benchmarks\.[3](https://arxiv.org/html/2609.25146#S2.F3): Illustration of real\-world vision scenarios, where agents continuously encounter diverse objects across evolving sessions\.[3](https://arxiv.org/html/2609.25146#S2.F3): Overall performance comparison with state\-of\-the\-art baselines\.[3](https://arxiv.org/html/2609.25146#S2.F3): Unified performance summary across datasets, pretrained models, and evaluation metrics\.[3](https://arxiv.org/html/2609.25146#S2.F3): Ablation analysis of mixture\-of\-experts \(MoE\) and ensemble learning \(EL\)\. In[3](https://arxiv.org/html/2609.25146#S2.F3)–[3](https://arxiv.org/html/2609.25146#S2.F3), FlyGCL uses LoRA modules as adaptive experts\. All results are averaged over five independent runs; error bars indicate the standard error of the mean\. Complete results for prompt\-, adapter\-, and LoRA\-based instantiations are reported in Supplementary[Tabs\.S3](https://arxiv.org/html/2609.25146#A6.T3)and[S4](https://arxiv.org/html/2609.25146#A6.T4)\.Continual Visual Recognition\.We first evaluate FlyGCL on continual visual recognition under online and blurry data streams, where samples from newly introduced and previously observed classes are probabilistically interleaved[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18);[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17), producing uncertain and evolving distributions over time \([Fig\.3](https://arxiv.org/html/2609.25146#S2.F3)\)\. We consider CIFAR\-100[Krizhevsky et al\. \(2009\)](https://arxiv.org/html/2609.25146#bib.bib4), ImageNet\-R[Hendrycks et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib26), and CUB\-200[Wah et al\. \(2011\)](https://arxiv.org/html/2609.25146#bib.bib27)datasets, spanning generic object recognition, distribution\-shifted concepts, and fine\-grained categories\. Comparisons include representative pretrained\-based CL methods, such as L2P[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20), DualPrompt[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21), and CODA\-Prompt[Smith et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib46), as well as online CL methods MVP[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)and MISA[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17)\. We report average anytime performance \(AaucA\_\{\\mathrm\{auc\}\}\) and final average performance \(AlastA\_\{\\mathrm\{last\}\}\)\. Across these benchmarks, FlyGCL consistently achieves the strongest final and anytime performance \([Figs\.3](https://arxiv.org/html/2609.25146#S2.F3)and[3](https://arxiv.org/html/2609.25146#S2.F3), Supplementary[Tab\.S3](https://arxiv.org/html/2609.25146#A6.T3)\): in the primary comparison, itsAaucA\_\{\\mathrm\{auc\}\}/AlastA\_\{\\mathrm\{last\}\}exceed the strongest baseline by 3\.5%/6\.9%, 13\.8%/18\.7%, and 12\.6%/24\.6% on CIFAR\-100, ImageNet\-R, and CUB\-200, respectively\. The performance gains are particularly clear on ImageNet\-R and CUB\-200, where distribution shifts and fine\-grained distinctions place greater demands on selective adaptation\. We next test whether this advantage depends on the pretrained representation or the adaptation interface\. The primary comparison covers three backbone settings: a model pretrained on ImageNet\-21K \(Sup\-21K\), a model pretrained on ImageNet\-21K and subsequently adapted to ImageNet\-1K \(Sup\-21K/1K\)[Russakovsky et al\. \(2015\)](https://arxiv.org/html/2609.25146#bib.bib2);[Ridnik et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib76);[Dosovitskiy et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib77), and a self\-supervised iBOT model pretrained on ImageNet\-21K \(iBOT\-21K\)[Zhou et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib55)\([Fig\.3](https://arxiv.org/html/2609.25146#S2.F3)\)\. We additionally evaluate self\-supervised checkpoints pretrained on ImageNet\-1K, including iBOT[Zhou et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib55), DINO[Caron et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib56), and MoCo v3[Chen et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib57)\(Supplementary[Tab\.S3](https://arxiv.org/html/2609.25146#A6.T3)\)\. Across these settings, FlyGCL supports prompt\-, adapter\-, and LoRA\-based adaptation[Lester et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib73);[Li and Liang \(2021\)](https://arxiv.org/html/2609.25146#bib.bib74);[Rebuffi et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib75);[Hu et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib72), and retains strong performance across the resulting combinations\. This consistency shows that the proposed hierarchical modularity is not tied to a particular pretrained representation or tuning interface\. Finally, we isolate the roles of the two components of FlyGCL design\. Removing either MoE or EL reduces performance across datasets and pretrained representations, whereas their combination performs best \([Fig\.3](https://arxiv.org/html/2609.25146#S2.F3), Supplementary[Tab\.S4](https://arxiv.org/html/2609.25146#A6.T4)\)\. Expert routing provides differentiated adaptation paths for heterogeneous visual distributions, while temporal ensemble integration stabilizes predictions as related experience recurs\. These results support that hierarchical specialization and integration provide complementary gains for continual visual recognition\.  abcdefg Figure 4:Results on continual vision\-language benchmarks\.[4](https://arxiv.org/html/2609.25146#S2.F4): Illustration of real\-world vision\-language scenarios, where an embodied agent continuously encounters evolving visual contexts and language queries\.[4](https://arxiv.org/html/2609.25146#S2.F4): Unified performance summary on CIFAR\-100 and ImageNet\-R usingAaucA\_\{\\rm auc\},AlastA\_\{\\rm last\}, forgetting, and backward transfer\.[4](https://arxiv.org/html/2609.25146#S2.F4): Performance comparison on CIFAR\-100 and ImageNet\-R\.[4](https://arxiv.org/html/2609.25146#S2.F4): Principal component analysis \(PCA\) projection of representations on CIFAR\-100\.[4](https://arxiv.org/html/2609.25146#S2.F4): Vision\-language feature drift relative to frozen CLIP; bubble size denotesAlastA\_\{\\rm last\}\.[4](https://arxiv.org/html/2609.25146#S2.F4): Hard\-negative separation by positive\-pair and hard\-negative similarities\.[4](https://arxiv.org/html/2609.25146#S2.F4): Class\-wise vision\-language margin, computed as positive\-pair similarity minus hard\-negative similarity\. All results are averaged over three independent runs; error bars indicate the standard error of the mean\. Detailed numerical results are reported in Supplementary[Tab\.S5](https://arxiv.org/html/2609.25146#A6.T5)\.Continual Vision\-Language Learning\.We next extend GCL to vision\-language models, where CL must accommodate evolving visual concepts while preserving the cross\-modal semantic structure acquired during large\-scale pretraining \([Fig\.4](https://arxiv.org/html/2609.25146#S2.F4)\)\. Compared with visual recognition, CL introduces an additional challenge: changes in visual representations may disrupt their correspondence with language and weaken the shared semantic space that supports multimodal generalization\. We evaluate FlyGCL with pretrained CLIP[Radford et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib54)on CIFAR\-100 and ImageNet\-R, comparing with classical CL methods such as EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)and LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5), as well as CLIP\-based CL methods including CLAP4CLIP[Jha et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib22)and MG\-CLIP[Huang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib23)\. FlyGCL achieves the strongest overall performance across both benchmarks \([Figs\.4](https://arxiv.org/html/2609.25146#S2.F4)and[4](https://arxiv.org/html/2609.25146#S2.F4), Supplementary[Tab\.S5](https://arxiv.org/html/2609.25146#A6.T5)\), reachingAaucA\_\{\\mathrm\{auc\}\}/AlastA\_\{\\mathrm\{last\}\}of 85\.2/79\.6% on CIFAR\-100 and 84\.7/79\.3% on ImageNet\-R, while also maintaining favorable forgetting and backward transfer\. These results show that hierarchical modularity remains effective when GCL extends from visual prediction to continually evolving cross\-modal representations\. We examine how different continual learners reshape the pretrained image\-text representation space\. Principal component analysis \(PCA\) visualization of the representation residuals reveals distinct adaptation trajectories relative to frozen CLIP \([Fig\.4](https://arxiv.org/html/2609.25146#S2.F4)\)\. FlyGCL exhibits a more constrained representation shift than competing approaches, while retaining the highest final performance\. The joint decomposition of image\- and text\-feature shifts shows the same pattern \([Fig\.4](https://arxiv.org/html/2609.25146#S2.F4)\): FlyGCL remains closer to the pretrained cross\-modal space without sacrificing adaptation accuracy\. This balance is consistent with the hierarchical design of FlyGCL, where selective expert adaptation accommodates distribution\-specific changes while temporal integration limits unnecessary displacement of the pretrained representation\. We further analyze whether this preservation extends to the semantic relationships that underpin image\-text recognition\. For each image, we compare its similarity to the matched text with that of the hardest negative text \([Fig\.4](https://arxiv.org/html/2609.25146#S2.F4)\)\. CL generally narrows this separation relative to frozen CLIP, whereas FlyGCL better preserves the positive\-to\-hard\-negative margin\. The class\-wise analysis confirms that this advantage is broadly distributed across semantic categories rather than driven by a small subset of classes \([Fig\.4](https://arxiv.org/html/2609.25146#S2.F4)\)\. The representation\- and similarity\-level analyses indicate that FlyGCL adapts to evolving visual distributions while better retaining the pretrained image\-text geometry, providing a mechanistic explanation for its stronger continual vision\-language learning performance\.  abcdefg Figure 5:Results on continual ego\-exo video understanding benchmarks\.[5](https://arxiv.org/html/2609.25146#S2.F5): Illustration of continual ego\-exo video understanding, where embodied agents encounter evolving activities and viewpoints over time\.[5](https://arxiv.org/html/2609.25146#S2.F5): Unified performance summary on EgoExoLearn and EgoExo\-Fitness usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\. For skill assessment on EgoExoLearn and EgoExo\-Fitness, results are averaged over the Relation Network \(RN\)\- and Triplet Loss \(TL\)\-based ego\-exo model variants[Huang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib66)\.[5](https://arxiv.org/html/2609.25146#S2.F5): Performance comparison for continual skill assessment on EgoExoLearn and EgoExo\-Fitness\.[5](https://arxiv.org/html/2609.25146#S2.F5),[5](https://arxiv.org/html/2609.25146#S2.F5): Stage\-wise accuracy of baseline and FlyGCL on EgoExoLearn\.[5](https://arxiv.org/html/2609.25146#S2.F5),[5](https://arxiv.org/html/2609.25146#S2.F5): Local loss landscapes under parameter perturbations for baseline and FlyGCL\. All results are averaged over three independent runs; error bars indicate the standard error of the mean\. Detailed numerical results are reported in Supplementary[Tabs\.S6](https://arxiv.org/html/2609.25146#A6.T6),[S8](https://arxiv.org/html/2609.25146#A6.T8),[S7](https://arxiv.org/html/2609.25146#A6.T7)and[S9](https://arxiv.org/html/2609.25146#A6.T9)\. ### 2\.3General Continual Learning for Embodied Perception and Action Real\-world intelligent systems must perceive changing environments while learning continuously from embodied and human\-centered experience\. Continual ego\-exo video understanding and continual vision\-language\-action learning represent two important settings in this direction, spanning the interpretation of evolving human activities and the acquisition of action policies through multimodal interaction\. These scenarios remain relatively underexplored in conventional CL, despite their direct relevance to long\-running intelligent agents\. Their video and interaction streams are inherently online, temporally continuous, and distributionally blurry, closely matching the uncertain and evolving experience targeted by GCL\. We therefore examine whether FlyGCL can extend from perception to embodied understanding and action\. Continual Ego\-Exo Video Understanding\.We extend GCL to continual ego\-exo video understanding[Yan et al\. \(2026b\)](https://arxiv.org/html/2609.25146#bib.bib65), where agents must continuously learn evolving activities from heterogeneous first\- and third\-person observations \([Fig\.5](https://arxiv.org/html/2609.25146#S2.F5)\)\. Compared with image\-based GCL, continual video learning introduces additional temporal and viewpoint variations: the same activity can exhibit substantially different visual dynamics across egocentric and exocentric views, while subject, skill, and motion patterns evolve continuously throughout the data stream\. We evaluate FlyGCL on EgoExoLearn[Huang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib66)and EgoExo\-Fitness[Li et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib67), covering continual skill assessment, action anticipation, and action classification\. Across these settings, FlyGCL achieves consistently strong overall performance on bothAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\([Figs\.5](https://arxiv.org/html/2609.25146#S2.F5)and[5](https://arxiv.org/html/2609.25146#S2.F5), Supplementary[Tabs\.S6](https://arxiv.org/html/2609.25146#A6.T6),[S8](https://arxiv.org/html/2609.25146#A6.T8),[S7](https://arxiv.org/html/2609.25146#A6.T7)and[S9](https://arxiv.org/html/2609.25146#A6.T9)\), showing that the brain\-inspired hierarchical modularity remains effective when GCL extends from static visual recognition to temporally evolving and cross\-view video representations\. The performance gains are consistent across distinct forms of embodied video understanding\. On EgoExoLearn skill assessment, FlyGCL reachesAaucA\_\{\\rm auc\}/AlastA\_\{\\rm last\}of81\.3%81\.3\\%/82\.6%82\.6\\%, outperforming replay\-, regularization\-, and prompt\-based methods\. On EgoExo\-Fitness skill assessment, FlyGCL reaches62\.1%62\.1\\%/62\.7%62\.7\\%, compared with58\.0%58\.0\\%/57\.9%57\.9\\%for the strongest competing method in the main comparison \([Fig\.5](https://arxiv.org/html/2609.25146#S2.F5)\)\. Additional evaluations on sequence verification and guidance\-based execution verification \(Supplementary[Tabs\.S10](https://arxiv.org/html/2609.25146#A6.T10)and[S11](https://arxiv.org/html/2609.25146#A6.T11)\) further demonstrate that the benefit of hierarchical modularity extends across distinct video tasks and output spaces\. We further examine how different continual learners evolve as new ego\-exo sessions arrive\. The stage\-wise trajectories \([Figs\.5](https://arxiv.org/html/2609.25146#S2.F5)and[5](https://arxiv.org/html/2609.25146#S2.F5)\) reveal substantially different forgetting patterns: the sequential fine\-tuning baseline progressively degrades on previously learned sessions, exhibiting substantial forgetting on EgoExoLearn skill assessment, whereas FlyGCL largely preserves earlier performance and maintains stable learning throughout the continual stream\. The local loss landscapes provide complementary evidence \([Figs\.5](https://arxiv.org/html/2609.25146#S2.F5)and[5](https://arxiv.org/html/2609.25146#S2.F5)\): FlyGCL converges to a flatter basin and exhibits lower sensitivity to parameter perturbations than the sequential fine\-tuning baseline\. These results suggest that hierarchical modularity improves not only average continual performance but also the robustness of the learned solution, with specialized experts accommodating heterogeneous temporal and viewpoint\-specific patterns while temporal integration limits destructive interference as related embodied experience recurs\.  abcde Figure 6:Results on continual vision\-language\-action benchmarks\.[6](https://arxiv.org/html/2609.25146#S2.F6): Illustration of continual embodied interaction, where agents follow evolving language instructions across changing environments and tasks\.[6](https://arxiv.org/html/2609.25146#S2.F6): Unified performance summary across LIBERO benchmarks and CL metrics\.[6](https://arxiv.org/html/2609.25146#S2.F6): Performance comparison with state\-of\-the\-art baselines on LIBERO\-Spatial and LIBERO\-Object usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\.[6](https://arxiv.org/html/2609.25146#S2.F6): Qualitative rollouts on LIBERO\-Spatial, comparing task execution immediately after learning and after the final CL session for EWC, DualPrompt\+, and FlyGCL\.[6](https://arxiv.org/html/2609.25146#S2.F6): Qualitative rollouts on LIBERO\-Object under the same protocol\. Results are averaged over three runs; error bars denote the standard error of the mean\.Continual Vision\-Language\-Action Learning\.We further extend GCL to continual vision\-language\-action learning, where embodied agents must connect visual observations and language instructions to actions as environments, goals, and interaction states evolve over time \([Fig\.6](https://arxiv.org/html/2609.25146#S2.F6)\)\. Unlike recognition or representation learning, continual updates in this setting affect an entire interaction policy\. The agent must preserve visuolinguistic grounding while adapting its visuomotor behaviour to new experience, even as related situations recur without clear task boundaries\. Changes in perception or action generation can alter subsequent states and compound over an interaction trajectory, placing demands on both stable representation learning and policy adaptation\. We evaluate FlyGCL on LIBERO\-Spatial, \-Object, \-Goal, and \-Long[Liu et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib28)under the primary online GCL stream and a complementary offline protocol\. In addition to EWC, LwF, L2P\+, and DualPrompt\+, we compare with sequential fine\-tuning \(SeqFT\), LoRA fine\-tuning \(SeqLoRA\), and PackNet[Mallya and Lazebnik \(2018\)](https://arxiv.org/html/2609.25146#bib.bib25)\. In the online setting, task distributions overlap and recur throughout the stream, requiring the agent to acquire new visuomotor behaviours while retaining earlier ones\. On LIBERO\-Spatial and LIBERO\-Object, FlyGCL reachesAlastA\_\{\\rm last\}/AaucA\_\{\\rm auc\}of 83\.1%/81\.6% and 86\.1%/86\.0%, exceeding the strongest replay\-free baselines by 49\.5%/37\.5% and 49\.8%/35\.9%, respectively \([Figs\.6](https://arxiv.org/html/2609.25146#S2.F6)and[6](https://arxiv.org/html/2609.25146#S2.F6)\)\. The same pattern holds for goal\-conditioned and long\-horizon manipulation \(Extended Data[Figs\.1](https://arxiv.org/html/2609.25146#Sx5.F1)and[1](https://arxiv.org/html/2609.25146#Sx5.F1)and Supplementary[Tab\.S12](https://arxiv.org/html/2609.25146#A6.T12)\)\. FlyGCL also maintains strong forward and backward transfer across the four suites, indicating that it can incorporate new interaction patterns without sacrificing previously acquired behaviours\. These results extend the benefit of hierarchical modularity from perceptual representations to instruction\-conditioned action policies\. We also evaluate the four suites under the offline protocol, in which each task is trained for multiple passes before the learner proceeds to the next one\. FlyGCL remains consistently strong across spatial, object\-centric, goal\-conditioned, and long\-horizon manipulation \(Extended Data[Fig\.2](https://arxiv.org/html/2609.25146#Sx5.F2)and Supplementary[Tab\.S13](https://arxiv.org/html/2609.25146#A6.T13)\)\. This result shows that its effectiveness is not limited to rapid online transitions\. It also applies when task changes are more structured but each task still contains continuously varying visual states, action trajectories, and interaction outcomes\. The rollout analyses further show how this stability affects task execution \([Figs\.6](https://arxiv.org/html/2609.25146#S2.F6)and[6](https://arxiv.org/html/2609.25146#S2.F6), Extended Data[Figs\.1](https://arxiv.org/html/2609.25146#Sx5.F1)and[1](https://arxiv.org/html/2609.25146#Sx5.F1)\)\. We compare FlyGCL with regularization\-based EWC and task\-expert\-based DualPrompt\+\. After subsequent updates, both methods can lose a critical part of behaviours that they execute successfully immediately after learning\. Their failures range from spatial grounding and object selection, as illustrated by DualPrompt\+ missing the target position of the black bowl and EWC selecting the wrong object instead of the chocolate pudding, to incomplete goal execution and long\-horizon action sequences\. In the latter cases, EWC pursues an incorrect goal state, while DualPrompt\+ completes the placement step but fails to close the drawer\. FlyGCL more consistently preserves complete task execution across subsequent sessions, consistent with expert routing separating task\-specific visuomotor changes and temporal integration retaining useful information across recurring experience\. ## 3Discussion This work identifies hierarchical modularity as a unified principle for GCL under online, uncertain, and evolving data streams\. Beyond preserving past knowledge, our results highlight a broader requirement: learning should be organized according to the relationships among incoming experiences\. Inspired by the organization of olfactory learning and memory in*Drosophila*, FlyGCL coordinates expert specialization and ensemble integration to separate conflicting experience while integrating compatible experience\. This design consistently improves GCL across visual recognition, vision\-language understanding, ego\-exo video understanding, and embodied vision\-language\-action learning, while remaining effective across diverse pretrained representations and parameter\-efficient adaptation mechanisms\. Some simplified brain\-inspired components were explored in our earlier conference paper[Yan et al\. \(2026a\)](https://arxiv.org/html/2609.25146#bib.bib58)\. The present work extends them into a biologically grounded hierarchical framework, supported by controlled olfactory modeling and evaluated across substantially broader models, modalities, and learning scenarios \(Supplementary Information\)\. These results establish hierarchical specialization and integration as an effective way to organize learning from dynamic experience\. From an AI perspective, GCL connects CL to a broader transition from intelligence acquired from static data to intelligence that develops through experience\. Recent perspectives on the “era of experience”[Silver and Sutton \(2025\)](https://arxiv.org/html/2609.25146#bib.bib59), autonomous machine intelligence[LeCun \(2022\)](https://arxiv.org/html/2609.25146#bib.bib60), and open\-ended learning[Hughes et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib61)similarly envision agents that continually extend their capabilities through interaction with the external world\. Such agents must not only accumulate experience, but also determine what should be reused, separated, or integrated as distributions change\. GCL provides a concrete learning paradigm for this process by bringing CL closer to the online, uncertain, and evolving conditions faced by real\-world agents\. This requirement arises in long\-running systems such as embodied robots, autonomous vehicles, personalized assistants, and healthcare or scientific monitoring systems, where perception, knowledge, and behaviour must be continually updated without predefined task boundaries\. FlyGCL provides a biologically grounded realization of this idea by organizing learning according to relationships among evolving experiences\. The biological implications are equally important\. Recent whole\-brain connectomes and cross\-connectome cell typing have provided increasingly detailed maps of the*Drosophila*nervous system[Winding et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib50);[Lin et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib48);[Schlegel et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib51), but how this anatomical organization supports learning from changing experience remains less understood\. Our computational model offers a functional interpretation of the olfactory learning and memory system by linking sparse expansion and differentiated memory pathways to specialization and integration\. Controlled olfactory experiments further show that their hierarchical coordination improves learning as experience shifts from disjoint to increasingly recurrent distributions\. These results suggest testable roles for the underlying circuit organization: differentiated pathways may reduce interference between dissimilar experiences, whereas coordinated integration may preserve shared structure and improve generalization across related experiences\. In this way, computational modelling can complement connectomics by linking anatomical organization to functional principles of learning and memory\. More broadly, our study exemplifies a bidirectional NeuroAI framework in which biological organization inspires machine\-learning principles, while computational models generate testable hypotheses for neuroscience\. The hierarchical organization of*Drosophila*olfactory learning and memory motivates a unified view of MoE and EL as complementary computational paradigms for specialization and integration in GCL\. In turn, our results suggest specific biological predictions: sparse expansion and differentiated downstream pathways should improve separation and reduce interference between dissimilar experiences; memory pathways operating across distinct spatial and temporal scales should contribute complementary information when related experiences recur; and their hierarchical coordination should be particularly beneficial when conflicting and compatible experiences coexist over time\. Extending these principles to embodied GCL also aligns with the NeuroAI vision of an “embodied Turing test”, in which intelligence develops through continuous sensorimotor interaction with the world[Zador et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib62)\. Several directions remain beyond the scope of this study\. Our biological model abstracts the overall organization of the*Drosophila*olfactory learning and memory system, leaving richer neuromodulatory dynamics, biological processes underlying memory consolidation, and behavioural feedback for future computational and experimental investigation\. On the AI side, our benchmarks extend GCL towards online, multimodal, and embodied experience, whereas longer\-term open\-ended interaction may additionally require active exploration, changing objectives, and dynamic allocation of learning resources\. FlyGCL also concentrates plasticity in lightweight modules over relatively stable pretrained representations\. Extending hierarchical specialization and integration to deeper learning within large foundation models remains an important direction\. Looking forward, extending these principles across richer biological mechanisms, open\-ended experience, and deeper model plasticity may help advance a more general science of continually developing intelligence\. ## 4Methods ### 4\.1Problem Formulation General Continual Learning\.Conventional CL[Wang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib13);[Wang et al\. \(2021a\)](https://arxiv.org/html/2609.25146#bib.bib6);[Wang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib19)often studies well\-separated tasks under largely offline training, with previous\-task data unavailable during subsequent updates\. In contrast, GCL considers single\-pass, non\-stationary streams with uncertain and blurry data distributions, without clear task boundaries or reliable task identities[Koh et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib36);[De Lange et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib63);[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)\. Formally, the data stream consists ofTTsessions,𝒮=\{𝒟1,𝒟2,…,𝒟T\}\\mathcal\{S\}=\\\{\\mathcal\{D\}\_\{1\},\\mathcal\{D\}\_\{2\},\\ldots,\\mathcal\{D\}\_\{T\}\\\}, where each session is given by𝒟t=\{\(𝐨t,i,𝐮t,i\)\}i=1Nt\\mathcal\{D\}\_\{t\}=\\\{\(\\mathbf\{o\}\_\{t,i\},\\mathbf\{u\}\_\{t,i\}\)\\\}\_\{i=1\}^\{N\_\{t\}\}\. Here,𝐨t,i∈𝒪t\\mathbf\{o\}\_\{t,i\}\\in\\mathcal\{O\}\_\{t\}denotes an observation, and𝐮t,i∈𝒰t\\mathbf\{u\}\_\{t,i\}\\in\\mathcal\{U\}\_\{t\}denotes its associated learning signal\. This notation provides a unified description of the GCL settings studied in this work:𝐨\\mathbf\{o\}may be instantiated as an image, an image\-text pair, a video clip, an ego\-exo multi\-view sequence, or an embodied visual\-language state, while𝐮\\mathbf\{u\}may correspond to a class label, semantic target, retrieval correspondence, temporal annotation, quality score, action, trajectory, or task\-success signal\. A standard model consists of a backbonefθ\(⋅\)f\_\{\\theta\}\(\\cdot\)and an output modulegψ\(⋅\)g\_\{\\psi\}\(\\cdot\)\. For notational convenience, we denote the resulting predictor asFθ,ψ=gψ∘fθF\_\{\\theta,\\psi\}=g\_\{\\psi\}\\circ f\_\{\\theta\}, i\.e\.,𝐮^=Fθ,ψ\(𝐨\)=gψ\(fθ\(𝐨\)\)\\hat\{\\mathbf\{u\}\}=F\_\{\\theta,\\psi\}\(\\mathbf\{o\}\)=g\_\{\\psi\}\(f\_\{\\theta\}\(\\mathbf\{o\}\)\)\. The learning objective is to obtain a unified mapping from𝒪=⋃t=1T𝒪t\\mathcal\{O\}=\\bigcup\_\{t=1\}^\{T\}\\mathcal\{O\}\_\{t\}to𝒰=⋃t=1T𝒰t\\mathcal\{U\}=\\bigcup\_\{t=1\}^\{T\}\\mathcal\{U\}\_\{t\}, while learning online and retaining previously acquired knowledge\. At sessiontt, samples are drawn from a local distribution\(𝐨t,i,𝐮t,i\)∼Pt\(𝐨,𝐮\)\(\\mathbf\{o\}\_\{t,i\},\\mathbf\{u\}\_\{t,i\}\)\\sim P\_\{t\}\(\\mathbf\{o\},\\mathbf\{u\}\), wherePtP\_\{t\}can evolve gradually or abruptly, overlap with previous distributions, and recur later in the stream\. In classification\-based GCL, such overlap is commonly instantiated by the Si\-Blurry setting[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17);[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18), which decomposes the global label space into a disjoint subset𝒴D\\mathcal\{Y\}^\{D\}and a blurry subset𝒴B\\mathcal\{Y\}^\{B\}, with𝒴=𝒴D∪𝒴B\\mathcal\{Y\}=\\mathcal\{Y\}^\{D\}\\cup\\mathcal\{Y\}^\{B\}and𝒴D∩𝒴B=∅\\mathcal\{Y\}^\{D\}\\cap\\mathcal\{Y\}^\{B\}=\\varnothing\. Classes in𝒴D\\mathcal\{Y\}^\{D\}are primarily associated with specific sessions, whereas classes in𝒴B\\mathcal\{Y\}^\{B\}can reappear across multiple sessions\. The disjoint ratiorD=\|𝒴D\|/\|𝒴\|r\_\{D\}=\\lvert\\mathcal\{Y\}^\{D\}\\rvert/\\lvert\\mathcal\{Y\}\\rvertcontrols the degree of session\-specific separation, while the blurry sample ratiorBr\_\{B\}controls the frequency of recurring samples\. Beyond classification, the same principle also applies more broadly, where semantic concepts, temporal states, skill levels, or action patterns may be partially shared across sessions\. Specialization and Integration\.GCL requires determining when incoming knowledge should be shared and when it should be separated\. Local distributions may share task\-relevant structure in representations, semantics, temporal dynamics, viewpoints, or behaviours, while differing in gradients, decision boundaries, temporal alignments, or action policies\. Related distributions should therefore share information to promote generalization, whereas conflicting distributions should be separated to reduce interference\. A fully shared predictorFθ,ψF\_\{\\theta,\\psi\}can exploit common structure but is vulnerable to interference, whereas fully isolated predictors preserve distribution\-specific knowledge at the cost of useful transfer\. We formalize this trade\-off with modular machine learning\. LetΩ=\{ω1,ω2,…,ωK\}\\Omega=\\\{\\omega\_\{1\},\\omega\_\{2\},\\ldots,\\omega\_\{K\}\\\}denote a set of trainable modules, such as prompts, adapters, LoRA branches, task heads, video\-specific modules, or policy modules\. When modules have their own output heads, we denote them byΨ=\{ψ1,ψ2,…,ψK\}\\Psi=\\\{\\psi\_\{1\},\\psi\_\{2\},\\ldots,\\psi\_\{K\}\\\}\. Given a representation𝐡=fθ\(𝐨\)\\mathbf\{h\}=f\_\{\\theta\}\(\\mathbf\{o\}\), a routing function produces module weightsπ\(𝐨\)=rη\(𝐡\)∈ΔK−1\\pi\(\\mathbf\{o\}\)=r\_\{\\eta\}\(\\mathbf\{h\}\)\\in\\Delta^\{K\-1\}\. We useFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}to denote the stream\-level modular predictor induced by the backbone, modules, output heads, and routing or aggregation mechanism\. MoE supports specialization by assigning inputs to adaptive modules\. For example, withk∗=argmaxkπk\(𝐨\)k^\{\*\}=\\arg\\max\_\{k\}\\pi\_\{k\}\(\\mathbf\{o\}\), the routed prediction can be written as FMoE\(𝐨\)=Fθ,ωk∗,ψk∗\(𝐨\),F\_\{\\mathrm\{MoE\}\}\(\\mathbf\{o\}\)=F\_\{\\theta,\\omega\_\{k^\{\*\}\},\\psi\_\{k^\{\*\}\}\}\(\\mathbf\{o\}\),\(1\)which reduces interference by limiting updates and predictions to the selected module\. EL instead integrates multiple predictors, FEL\(𝐨\)=𝒜\(Fθ,ω1,ψ1\(𝐨\),…,Fθ,ωK,ψK\(𝐨\)\),F\_\{\\mathrm\{EL\}\}\(\\mathbf\{o\}\)=\\mathcal\{A\}\\left\(F\_\{\\theta,\\omega\_\{1\},\\psi\_\{1\}\}\(\\mathbf\{o\}\),\\ldots,F\_\{\\theta,\\omega\_\{K\},\\psi\_\{K\}\}\(\\mathbf\{o\}\)\\right\),\(2\)where𝒜\(⋅\)\\mathcal\{A\}\(\\cdot\)denotes an aggregation function\. This improves robustness and generalization for similar distributions by combining diverse but related predictions\. However, directly combining MoE and EL is non\-trivial because they favor different distributional structures[Wang and Li \(2025\)](https://arxiv.org/html/2609.25146#bib.bib42)\. EL promotes generalization when predictors capture identical or closely related distributions[Allen\-Zhu and Li \(2023\)](https://arxiv.org/html/2609.25146#bib.bib15), whereas effective MoE routing relies on structured separation among heterogeneous distributions[Chen et al\. \(2022\)](https://arxiv.org/html/2609.25146#bib.bib16)\. In GCL, compatible and conflicting distributions may coexist: overly exclusive routing can suppress transfer among related distributions, whereas overly broad ensembling can mix incompatible modules and reintroduce interference\. GCL therefore requires hierarchical coordination between specialization and integration, separating conflicting distributions while integrating compatible ones\. ### 4\.2Theoretical Analysis Decomposing Hierarchical GCL Risk\.The preceding formulation shows that effective GCL requires coordinating specialization and integration under uncertain and evolving data streams\. We analyze this requirement using the stream\-level modular predictorFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}defined above, which comprises the backbonefθf\_\{\\theta\}, trainable modulesΩ\\Omega, output modulesΨ\\Psi, and the associated routing or aggregation operations\. LetF⋆F^\{\\star\}denote the ideal stream\-level predictor obtained if the latent structure of the local distributions\{Pt\}t=1T\\\{P\_\{t\}\\\}\_\{t=1\}^\{T\}were known\. We define the expected GCL risk as ℛGCL\(Fθ,Ω,Ψ\)=1T∑t=1T𝔼\(𝐨,𝐮\)∼Pt\[ℓ\(Fθ,Ω,Ψ\(𝐨\),𝐮\)\],\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{\(\\mathbf\{o\},\\mathbf\{u\}\)\\sim P\_\{t\}\}\\left\[\\ell\\left\(F\_\{\\theta,\\Omega,\\Psi\}\(\\mathbf\{o\}\),\\mathbf\{u\}\\right\)\\right\],\(3\)whereℓ\(⋅\)\\ell\(\\cdot\)denotes the task\-specific loss\. BecauseF⋆F^\{\\star\}is unavailable in CL, the practical objective is to approximate this ideal predictor from the observed stream𝒮\\mathcal\{S\}under non\-stationary distribution shifts\. ###### Theorem 1\(Hierarchical Decomposition of GCL Risk\)\. LetFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}be a hierarchical modular predictor learned from𝒮=\{𝒟1,…,𝒟T\}\\mathcal\{S\}=\\\{\\mathcal\{D\}\_\{1\},\\ldots,\\mathcal\{D\}\_\{T\}\\\}, where𝒟t=\{\(𝐨t,i,𝐮t,i\)\}i=1Nt\\mathcal\{D\}\_\{t\}=\\\{\(\\mathbf\{o\}\_\{t,i\},\\mathbf\{u\}\_\{t,i\}\)\\\}\_\{i=1\}^\{N\_\{t\}\}\. Under an additive excess\-risk decomposition relative toF⋆F^\{\\star\}, the expected GCL risk is upper bounded by ℛGCL\(Fθ,Ω,Ψ\)≲ℛfit\(Fθ,Ω,Ψ\)\+ℰsep\(Fθ,Ω,Ψ\)⏟separation error\+ℰint\(Fθ,Ω,Ψ\)⏟integration error\+𝒞coord\(Fθ,Ω,Ψ\)⏟coordination cost\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\\lesssim\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\+\\underbrace\{\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\}\_\{\\text\{separation error\}\}\+\\underbrace\{\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\}\_\{\\text\{integration error\}\}\+\\underbrace\{\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\}\_\{\\text\{coordination cost\}\}\.\(4\)Here, the empirical fitting term is defined on the observed stream as ℛfit\(Fθ,Ω,Ψ\)=1T∑t=1T1Nt∑i=1Ntℓ\(Fθ,Ω,Ψ\(𝐨t,i\),𝐮t,i\)\.\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\frac\{1\}\{N\_\{t\}\}\\sum\_\{i=1\}^\{N\_\{t\}\}\\ell\\left\(F\_\{\\theta,\\Omega,\\Psi\}\(\\mathbf\{o\}\_\{t,i\}\),\\mathbf\{u\}\_\{t,i\}\\right\)\.\(5\)The remaining terms are excess\-risk components:ℰsep\(Fθ,Ω,Ψ\)\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)is induced by insufficient separation of conflicting local distributions,ℰint\(Fθ,Ω,Ψ\)\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)is induced by insufficient integration of compatible predictions, and𝒞coord\(Fθ,Ω,Ψ\)\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)is the additional cost introduced by coordinating routing and integration\. The proof is provided in Supplementary[Sec\.A\.1](https://arxiv.org/html/2609.25146#A1.SS1)\. The three terms characterize distinct failure modes of hierarchical modular learning\. The separation errorℰsep\\mathcal\{E\}\_\{\\mathrm\{sep\}\}captures residual interference when samples with conflicting gradients, decision boundaries, temporal alignments, or policies share parameters\. The integration errorℰint\\mathcal\{E\}\_\{\\mathrm\{int\}\}captures the residual estimation gap when related local distributions or compatible predictors are treated in isolation\. The coordination cost𝒞coord\\mathcal\{C\}\_\{\\mathrm\{coord\}\}arises from jointly performing routing and integration, and is defined as the positive excess loss of the hierarchical modular predictor relative to the ideal stream\-level predictor: 𝒞coord\(Fθ,Ω,Ψ\)=1T∑t=1T𝔼\(𝐨,𝐮\)∼Pt\[ℓ\(Fθ,Ω,Ψ\(𝐨\),𝐮\)−ℓ\(F⋆\(𝐨\),𝐮\)\]\+,\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{\(\\mathbf\{o\},\\mathbf\{u\}\)\\sim P\_\{t\}\}\\left\[\\ell\\left\(F\_\{\\theta,\\Omega,\\Psi\}\(\\mathbf\{o\}\),\\mathbf\{u\}\\right\)\-\\ell\\left\(F^\{\\star\}\(\\mathbf\{o\}\),\\mathbf\{u\}\\right\)\\right\]\_\{\+\},\(6\)where\[⋅\]\+\[\\cdot\]\_\{\+\}denotes the positive part\. This term accounts for over\-separation of related distributions, over\-integration of incompatible modules, routing errors, and aggregation mismatch\. Effective GCL therefore requires coordinating separation and integration rather than treating routing and aggregation as independent operations\. Separation and Integration Gains\.The decomposition in[Eq\.4](https://arxiv.org/html/2609.25146#S4.E4)indicates that hierarchical modular learning is beneficial only when the gains from separation and integration outweigh the additional cost of coordinating them\. We define the routing gain and ensemble gain as Δroute=ℰsep\(Fθ,ψ\)−ℰsep\(FMoE\),Δens=ℰint\(Fsingle\)−ℰint\(FEL\),\\Delta\_\{\\mathrm\{route\}\}=\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\-\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{MoE\}\}\),\\quad\\Delta\_\{\\mathrm\{ens\}\}=\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\-\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{EL\}\}\),\(7\)whereFθ,ψF\_\{\\theta,\\psi\}is the fully shared predictor,FMoEF\_\{\\mathrm\{MoE\}\}is the routed modular predictor in[Eq\.1](https://arxiv.org/html/2609.25146#S4.E1),FsingleF\_\{\\mathrm\{single\}\}denotes a predictor using a single adaptive path, andFELF\_\{\\mathrm\{EL\}\}is the integrated predictor in[Eq\.2](https://arxiv.org/html/2609.25146#S4.E2)\.Δroute\\Delta\_\{\\mathrm\{route\}\}measures the reduction in separation error obtained by routing conflicting samples to different adaptive modules, whileΔens\\Delta\_\{\\mathrm\{ens\}\}measures the reduction in integration error obtained by aggregating compatible predictors\. ###### Proposition 1\(Separation and Integration Gains\)\. Assume that routing provides a non\-negative separation gainΔroute≥0\\Delta\_\{\\mathrm\{route\}\}\\geq 0and aggregation provides a non\-negative integration gainΔens≥0\\Delta\_\{\\mathrm\{ens\}\}\\geq 0\. Then the GCL risk of a hierarchical modular predictorFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}satisfies ℛGCL\(Fθ,Ω,Ψ\)≲\\displaystyle\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\\lesssimℛfit\(Fθ,Ω,Ψ\)\+ℰsep\(Fθ,ψ\)\+ℰint\(Fsingle\)\\displaystyle\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\+\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\(8\)−Δroute−Δens\+𝒞coord\(Fθ,Ω,Ψ\)\.\\displaystyle\-\\Delta\_\{\\mathrm\{route\}\}\-\\Delta\_\{\\mathrm\{ens\}\}\+\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\.Consequently, hierarchical modular learning improves the risk bound whenever Δroute\+Δens\>𝒞coord\(Fθ,Ω,Ψ\)\.\\Delta\_\{\\mathrm\{route\}\}\+\\Delta\_\{\\mathrm\{ens\}\}\>\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\.\(9\) The proof is provided in Supplementary[Sec\.A\.2](https://arxiv.org/html/2609.25146#A1.SS2)\.[Proposition1](https://arxiv.org/html/2609.25146#Thmproposition1)makes the trade\-off in hierarchical modular learning explicit\. Routing reduces interference by separating conflicting local distributions, whereas integration improves robustness by combining compatible predictions\. The two operations are nevertheless coupled: overly exclusive routing may separate related distributions, while overly broad aggregation may mix incompatible modules\. Effective GCL therefore requires coordinating separation and integration rather than optimizing either operation in isolation\. Pretraining\-Supported Coordination\.The condition in[Eq\.9](https://arxiv.org/html/2609.25146#S4.E9)shows that the benefit of hierarchical modular learning depends on both the gains from routing and aggregation and the cost of coordinating them\. This coordination cost is particularly relevant in GCL, where the distinction between related and conflicting local distributions may be uncertain\. We formalize this effect by separating the probability of an imperfect modular decision from its resulting prediction penalty\. ###### Proposition 2\(Pretraining\-Supported Coordination\)\. Letϵroute\\epsilon\_\{\\mathrm\{route\}\}denote the degree of routing error,ϵagg\\epsilon\_\{\\mathrm\{agg\}\}denote the mismatch introduced by aggregating modular predictions, andΔmis\(fθ\)\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)denote the prediction penalty of assigning an input to a suboptimal module under the representation induced byfθf\_\{\\theta\}\. Then the coordination cost is bounded by 𝒞coord\(Fθ,Ω,Ψ\)≤ϵrouteΔmis\(fθ\)\+ϵagg\.\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\\leq\\epsilon\_\{\\mathrm\{route\}\}\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)\+\\epsilon\_\{\\mathrm\{agg\}\}\.\(10\)Moreover, if a pretrained backbone provides a more stable and semantically organized representation than a representation learned from scratch on the online stream, then Δmis\(fθpre\)<Δmis\(fθscratch\),\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}^\{\\mathrm\{pre\}\}\)<\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}^\{\\mathrm\{scratch\}\}\),\(11\)which reduces the coordination cost in[Eq\.10](https://arxiv.org/html/2609.25146#S4.E10)\. The proof is provided in Supplementary[Sec\.A\.3](https://arxiv.org/html/2609.25146#A1.SS3)\.[Proposition2](https://arxiv.org/html/2609.25146#Thmproposition2)explains how pretrained representations reduce the cost of imperfect coordination in hierarchical GCL\. A routing error is more harmful when the selected module produces predictions that differ substantially from the appropriate one, whereas this penalty is reduced when the pretrained backbone maps related inputs into a stable semantic space\. Pretrained representations therefore contribute not only positive transfer and resistance to forgetting, but also greater robustness to imperfect routing and integration\. This supports hierarchical MoE\-EL as a lightweight modular design over stable pretrained foundation models\. Implications for FlyGCL\.The analysis above identifies three requirements for GCL: separating conflicting local distributions, integrating compatible predictions, and limiting the coordination cost between them\. These considerations motivate hierarchical modularity as the design principle of FlyGCL, particularly over stable pretrained representations\. We next instantiate this principle in a brain\-inspired GCL framework and describe its architecture and optimization\. ### 4\.3FlyGCL Model Model Overview\.FlyGCL is a unified brain\-inspired framework for GCL with pretrained foundation models\. Its design follows the organization of the*Drosophila*olfactory learning and memory system at three functional levels\. Sparse random expansion provides a distributed representation analogous to the PN\-KC pathway and supports selective recruitment of downstream pathways\. Spatially differentiated experts provide parallel memory pathways, while output heads with different update timescales capture complementary short\- and long\-term information\. Therefore, stable representation, expert routing, and temporal integration form a computational hierarchy for separating conflicting experience and integrating compatible predictions\. Pretraining\-based CL commonly introduces parameter\-efficient tuning components, such as prompts, adapters, and LoRA, as lightweight experts over a pretrained backbone\. Each expert provides a specialized learning pathway parameterized by𝝎\\bm\{\\omega\}and accumulates spatially differentiated knowledge, producing outputsf𝜽\(𝒙,𝝎\)f\_\{\\bm\{\\theta\}\}\(\\bm\{x\};\\bm\{\\omega\}\)\. Under single\-pass blurry data streams, expert\-based learning must determine both the appropriate expert for each input and how to maintain reliable predictions under limited and imbalanced online supervision\. FlyGCL addresses these challenges through random\-expanded analytic routing and temporal ensemble integration within each routed expert\. Letf𝜽\(⋅\)f\_\{\\bm\{\\theta\}\}\(\\cdot\)denote a pretrained backbone and𝒉=f𝜽\(𝒙\)∈ℝd\\bm\{h\}=f\_\{\\bm\{\\theta\}\}\(\\bm\{x\}\)\\in\\mathbb\{R\}^\{d\}its representation\. The trainable expert pool isΩ=\{𝝎1,𝝎2,…,𝝎K\}\\Omega=\\\{\\bm\{\\omega\}\_\{1\},\\bm\{\\omega\}\_\{2\},\\ldots,\\bm\{\\omega\}\_\{K\}\\\}, and the prediction of expertEkE\_\{k\}is denoted byF𝜽,𝝎k,𝝍k\(𝒙\)F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{k\},\\bm\{\\psi\}\_\{k\}\}\(\\bm\{x\}\), where𝝍k\\bm\{\\psi\}\_\{k\}is its output head\. Random\-Expanded Analytic Router\.To improve expert routing, FlyGCL uses a random\-expanded analytic router inspired by sparse expansion in the fruit fly mushroom body\. Given𝒉=f𝜽\(𝒙\)\\bm\{h\}=f\_\{\\bm\{\\theta\}\}\(\\bm\{x\}\), we apply a fixed random projection followed by nonlinear activation: 𝝋\(𝒙\)=σ\(f𝜽\(𝒙\)𝑹\)=σ\(𝒉𝑹\)∈ℝM,\\bm\{\\varphi\}\(\\bm\{x\}\)=\\sigma\\left\(f\_\{\\bm\{\\theta\}\}\(\\bm\{x\}\)\\bm\{R\}\\right\)=\\sigma\(\\bm\{h\}\\bm\{R\}\)\\in\\mathbb\{R\}^\{M\},\(12\)where𝑹∈ℝd×M\\bm\{R\}\\in\\mathbb\{R\}^\{d\\times M\}is a random matrix,M\>dM\>d, andσ\(⋅\)\\sigma\(\\cdot\)is an element\-wise activation function\. The expanded feature𝝋\(𝒙\)\\bm\{\\varphi\}\(\\bm\{x\}\)is used for instance\-level expert routing rather than final prediction, preserving the flexibility of downstream adaptive experts\. During online training, for each incoming batchℬi\\mathcal\{B\}\_\{i\}from sessiontt, we compute the expanded feature matrix𝚽i∈ℝB×M\\bm\{\\Phi\}\_\{i\}\\in\\mathbb\{R\}^\{B\\times M\}and update two statistics: 𝑮←𝑮\+𝚽i⊤𝚽i,𝑸←𝑸\+𝚽i⊤𝑪t,\\bm\{G\}\\leftarrow\\bm\{G\}\+\\bm\{\\Phi\}\_\{i\}^\{\\top\}\\bm\{\\Phi\}\_\{i\},\\quad\\bm\{Q\}\\leftarrow\\bm\{Q\}\+\\bm\{\\Phi\}\_\{i\}^\{\\top\}\\bm\{C\}\_\{t\},\(13\)where𝑮∈ℝM×M\\bm\{G\}\\in\\mathbb\{R\}^\{M\\times M\}captures second\-order feature correlations,𝑸∈ℝM×K\\bm\{Q\}\\in\\mathbb\{R\}^\{M\\times K\}stores expert\-wise feature statistics, and𝑪t∈ℝB×K\\bm\{C\}\_\{t\}\\in\\mathbb\{R\}^\{B\\times K\}denotes the expert assignment target for the current session\. The router matrix𝑼∈ℝK×M\\bm\{U\}\\in\\mathbb\{R\}^\{K\\times M\}is obtained by the closed\-form ridge solution 𝑼^⊤=\(𝑮\+λ𝑰\)−1𝑸,\\widehat\{\\bm\{U\}\}^\{\\top\}=\(\\bm\{G\}\+\\lambda\\bm\{I\}\)^\{\-1\}\\bm\{Q\},\(14\)whereλ\>0\\lambda\>0is the regularization parameter\. At inference time, the routing score and selected expert are computed as 𝒔\(𝒙\)=𝝋\(𝒙\)𝑼^⊤,E^\(𝒙\)=argmaxk≤Ksk\(𝒙\)\.\\bm\{s\}\(\\bm\{x\}\)=\\bm\{\\varphi\}\(\\bm\{x\}\)\\widehat\{\\bm\{U\}\}^\{\\top\},\\quad\\hat\{E\}\(\\bm\{x\}\)=\\arg\\max\_\{k\\leq K\}s\_\{k\}\(\\bm\{x\}\)\.\(15\)The routed prediction is Froute\(𝒙\)=F𝜽,𝝎E^,𝝍E^\(𝒙\)\.F\_\{\\rm route\}\(\\bm\{x\}\)=F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{\\hat\{E\}\},\\bm\{\\psi\}\_\{\\hat\{E\}\}\}\(\\bm\{x\}\)\.\(16\) After routing, the selected expert is updated using the task\-specific learning signal\. For a sample\(𝒙i,𝒚i\)\(\\bm\{x\}\_\{i\},\\bm\{y\}\_\{i\}\)assigned to expertE^\\hat\{E\}, the online objective is ℒi=ℓi\(F𝜽,𝝎E^,𝝍E^\(𝒙i\),𝒚i\),\\mathcal\{L\}\_\{i\}=\\ell\_\{i\}\\left\(F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{\\hat\{E\}\},\\bm\{\\psi\}\_\{\\hat\{E\}\}\}\(\\bm\{x\}\_\{i\}\),\\bm\{y\}\_\{i\}\\right\),\(17\)whereℓi\\ell\_\{i\}can be instantiated as a classification, contrastive, regression, ranking, verification, imitation, or policy\-learning loss\. Because the router is updated through accumulated statistics and solved in closed form, it avoids iterative router training and is well suited to the single\-pass dynamic data streams\. Temporal Ensemble\-based Experts\.The prediction of a routed expert depends not only on its learned representation, but also on the stability of its output head as the data distribution evolves\. FlyGCL therefore equips each expert with multiple output heads operating at different effective timescales\. For expertEkE\_\{k\}that accumulates spatially differentiated knowledge, we maintain an online head𝝍k\(0\)\\bm\{\\psi\}^\{\(0\)\}\_\{k\}andnnshadow heads\{𝝍k\(j\)\}j=1n\\\{\\bm\{\\psi\}^\{\(j\)\}\_\{k\}\\\}\_\{j=1\}^\{n\}updated with different exponential moving average \(EMA\) rates\. For a linear output head𝝍=\(𝑾,𝒃\)\\bm\{\\psi\}=\(\\bm\{W\},\\bm\{b\}\), thejj\-th EMA head of expertEkE\_\{k\}is updated as 𝑾k\(j\)←αj𝑾k\(j\)\+\(1−αj\)𝑾,𝒃k\(j\)←αj𝒃k\(j\)\+\(1−αj\)𝒃\.\\bm\{W\}\_\{k\}^\{\(j\)\}\\leftarrow\\alpha\_\{j\}\\bm\{W\}\_\{k\}^\{\(j\)\}\+\(1\-\\alpha\_\{j\}\)\\bm\{W\},\\quad\\bm\{b\}\_\{k\}^\{\(j\)\}\\leftarrow\\alpha\_\{j\}\\bm\{b\}\_\{k\}^\{\(j\)\}\+\(1\-\\alpha\_\{j\}\)\\bm\{b\}\.\(18\)Different EMA ratesαj\\alpha\_\{j\}induce distinct effective memory timescales: faster\-updating heads prioritize recent observations, resembling short\-term memory in theγ\\gammalobe of*Drosophila*, whereas progressively slower heads integrate information over longer timescales, yielding more stable decision boundaries under distribution shift and paralleling the more persistent memory supported by theα′/β′\\alpha^\{\\prime\}/\\beta^\{\\prime\}andα/β\\alpha/\\betalobes\. At inference, after the analytic router selects expertE^\\hat\{E\}, FlyGCL computes predictions from the online and EMA heads of this expert and aggregates them: FFlyGCL\(𝒙\)=𝒜\(F𝜽,𝝎E^,𝝍E^\(0\)\(𝒙\),F𝜽,𝝎E^,𝝍E^\(1\)\(𝒙\),…,F𝜽,𝝎E^,𝝍E^\(n\)\(𝒙\)\),F\_\{\\rm FlyGCL\}\(\\bm\{x\}\)=\\mathcal\{A\}\\left\(F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{\\hat\{E\}\},\\bm\{\\psi\}\_\{\\hat\{E\}\}^\{\(0\)\}\}\(\\bm\{x\}\),F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{\\hat\{E\}\},\\bm\{\\psi\}\_\{\\hat\{E\}\}^\{\(1\)\}\}\(\\bm\{x\}\),\\ldots,F\_\{\\bm\{\\theta\},\\bm\{\\omega\}\_\{\\hat\{E\}\},\\bm\{\\psi\}\_\{\\hat\{E\}\}^\{\(n\)\}\}\(\\bm\{x\}\)\\right\),\(19\)where𝒜\(⋅\)\\mathcal\{A\}\(\\cdot\)combines the head predictions, using normalized softmax weights when weighted temporal aggregation is adopted\. This temporal ensemble improves decoding robustness without removing the specialization created by expert routing\. For visual recognition settings, we further introduce a lightweight gate to calibrate the integrated temporal output\. We construct an analytic class head from the frozen pretrained representation reusing the same fixed random expansion and ridge solution described above, obtaining class evidence𝒛an\(𝒙\)\\bm\{z\}\_\{\\rm an\}\(\\bm\{x\}\)\. We calibrate the temporally aggregated class probabilities𝒑temp\(𝒙\)=FFlyGCL\(𝒙\)\\bm\{p\}\_\{\\rm temp\}\(\\bm\{x\}\)=F\_\{\\rm FlyGCL\}\(\\bm\{x\}\)according to zgate,c\(𝒙\)=logptemp,c\(𝒙\)\+λtzan,c\(𝒙\),∀c∈𝒞\(𝒙\),z\_\{\{\\rm gate\},c\}\(\\bm\{x\}\)=\\log p\_\{\{\\rm temp\},c\}\(\\bm\{x\}\)\+\\lambda\_\{t\}z\_\{\{\\rm an\},c\}\(\\bm\{x\}\),\\qquad\\forall c\\in\\mathcal\{C\}\(\\bm\{x\}\),\(20\)whereλt=λmax\(t/T\)2\\lambda\_\{t\}=\\lambda\_\{\\max\}\(t/T\)^\{2\},t/Tt/Tdenotes the normalized progress through the continual stream,λmax≥0\\lambda\_\{\\max\}\\geq 0is the final calibration strength, and𝒞\(𝒙\)\\mathcal\{C\}\(\\bm\{x\}\)denotes the category set\. The analytic evidence enhances classes supported by the stable pretrained representation and suppresses unsupported alternatives, without selecting experts or adding another temporal head\. This class\-dependent regulation may functionally resemble the modulatory role of dopamine neurons \(DANs\) over mushroom\-body output pathways[Aso et al\. \(2014b\)](https://arxiv.org/html/2609.25146#bib.bib8);[Dasgupta et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib37)\. The calibration is applied with temporal integration, while expert selection remains separately determined by the analytic router\. ### 4\.4Experimental Setups Datasets\.We evaluate FlyGCL across four CL scenarios: visual recognition, vision\-language learning, ego\-exo video understanding, and embodied vision\-language\-action learning\. For visual recognition, we use CIFAR\-100[Krizhevsky et al\. \(2009\)](https://arxiv.org/html/2609.25146#bib.bib4), containing 60,000 images from 100 classes \(50,000/10,000 training/test images\); ImageNet\-R[Hendrycks et al\. \(2021\)](https://arxiv.org/html/2609.25146#bib.bib26), containing 30,000 artistic and non\-photographic renditions from 200 ImageNet classes; and CUB\-200[Wah et al\. \(2011\)](https://arxiv.org/html/2609.25146#bib.bib27), containing 11,788 images from 200 fine\-grained bird species\. For vision\-language learning, we construct CLIP\-based continual benchmarks on CIFAR\-100 and ImageNet\-R using the same visual streams, with each class represented by the textual prompta photo of a \{class name\}\. For ego\-exo video understanding, we use EgoExoLearn[Huang et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib66), which contains 120 hours of egocentric execution and exocentric demonstration videos with gaze and multimodal annotations for cross\-view association, planning, and skill assessment, and EgoExo\-Fitness[Li et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib67), which provides synchronized ego\-exo fitness videos with two\-level temporal boundaries and interpretable action\-judgement annotations\. For embodied vision\-language\-action learning, we use LIBERO[Liu et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib28), a benchmark of language\-conditioned robotic manipulation comprising LIBERO\-Spatial, \-Object, \-Goal, and \-Long\. The first three suites each contain 10 tasks emphasizing spatial relations, object\-centric manipulation, and goal\-conditioned behaviour, respectively, while LIBERO\-Long contains 10 long\-horizon tasks derived from LIBERO\-100\. We adoptrD=50%r\_\{D\}=50\\%andrB=30%r\_\{B\}=30\\%across all scenarios \(exceptrB=10%r\_\{B\}=10\\%for visual recognition following MVP[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)and MISA[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17)\)\. Sessions are constructed using class\-, action\-semantic\-, or task\-level partitions, with category\-based partitioning used for the single\-category EgoExo\-Fitness benchmark\. Detailed benchmark setup and implementations are provided in Supplementary[AppendixB](https://arxiv.org/html/2609.25146#A2)\. Evaluation Metrics\.We report two commonly used CL metrics across all benchmarks: final average performanceAlastA\_\{\\mathrm\{last\}\}and average anytime performanceAaucA\_\{\\mathrm\{auc\}\}\. LetRt,jR\_\{t,j\}denote the task\-specific performance on sessionjjafter the model has learned through sessiontt\. Depending on the benchmark,Rt,jR\_\{t,j\}is instantiated as accuracy, ranking accuracy, Top\-5 recall, etc\. The average anytime performance is defined as Aauc=1T∑t=1T\(1t∑j=1tRt,j\),A\_\{\\mathrm\{auc\}\}=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\left\(\\frac\{1\}\{t\}\\sum\_\{j=1\}^\{t\}R\_\{t,j\}\\right\),\(21\)which measures performance throughout CL\. The final average performance is defined as Alast=1T∑j=1TRT,j,A\_\{\\mathrm\{last\}\}=\\frac\{1\}\{T\}\\sum\_\{j=1\}^\{T\}R\_\{T,j\},\(22\)which measures performance retained after learning the entire stream\. Benchmark\-specific metric instantiations and notation are detailed in Supplementary[AppendixC](https://arxiv.org/html/2609.25146#A3)\. Baseline Methods\.We compare FlyGCL with representative CL methods across four application scenarios, with sequential fine\-tuning \(SeqFT\) included as a common baseline throughout\. For continual image recognition, we additionally consider the regularization\-based methods EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)and LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5), the parameter efficiently tuning methods L2P[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20)and DualPrompt[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21), and the online CL methods MVP[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)and MISA[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17)\. For continual vision\-language learning with CLIP\-based models, we compare against EWC, LwF, L2P, DualPrompt, CLAP4CLIP[Jha et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib22), and MG\-CLIP[Huang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib23)\. For continual ego\-exo video understanding, we include EWC, LwF, L2P\+, DualPrompt\+, S\-Prompt\+[Wang et al\. \(2022b\)](https://arxiv.org/html/2609.25146#bib.bib24), and the replay\-based methods Experience Replay \(ER\)[Rolnick et al\. \(2019\)](https://arxiv.org/html/2609.25146#bib.bib30)and DER\+\+[Buzzega et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib29)\. For these adapter\-based variants, we replace the original prompt modules with adapters, with “\+” indicating this modification\. For continual embodied vision\-language\-action learning, we compare against EWC, LwF, L2P\+, DualPrompt\+, ER, and PackNet[Mallya and Lazebnik \(2018\)](https://arxiv.org/html/2609.25146#bib.bib25)\. Collectively, these baselines cover sequential fine\-tuning, regularization\-based, replay\-based, and task\-specific CL methods\. Implementation Details\.We implement FlyGCL on pretrained backbones adopted by each benchmark to ensure fair comparison with existing CL methods\. For visual recognition, we use Vision Transformer \(ViT\-B/16\) backbones with ImageNet\-based pretraining, including supervised ImageNet\-21K pretraining \(Sup\-21K\), ImageNet\-21K pretraining followed by ImageNet\-1K fine\-tuning \(Sup\-21K/1K\), and self\-supervised iBOT pretraining on ImageNet\-21K \(iBOT\-21K\)\. For vision\-language learning, we use the pretrained OpenAI CLIP model with a ViT\-B/16 image encoder\. For ego\-exo video understanding and embodied vision\-language\-action learning, we follow the backbone, input preprocessing, and evaluation pipeline of the corresponding benchmarks, replacing only the continual adaptation component across methods\. All methods use the same data stream, session order, and online update budget\. For prompt\-, adapter\-, and LoRA\-based methods, we match the number of trainable modules to FlyGCL whenever applicable, so that comparisons primarily reflect the continual coordination strategy rather than model capacity\. Hyperparameters are selected on the validation split of the first stream setting and then kept fixed across sessions\. We report the mean and standard error over multiple random seeds\. Detailed optimizer settings, learning rates, batch sizes, training budgets, and benchmark\-specific implementations are provided in Supplementary[AppendixD](https://arxiv.org/html/2609.25146#A4)\. ## Data Availability ## Code Availability ## Acknowledgments This work was supported by the NSFC Project \(No\. T2622023, No\. 62406160 and No\. 62595773\), the Beijing Natural Science Foundation \(No\. L247011\), the Beijing Nova Program \(No\. 202604841279\), and the Beijing Major Science and Technology Project \(No\. Z251100008425003\)\. ## Author Contributions Statement H\.Y\., K\.Z\. and L\.W\. conceived the project\. H\.Y\., K\.Z\. and L\.W\. designed the computational framework\. H\.Y\. and K\.Z\. performed the main experiments, assisted by Q\.C\. and W\.D\. H\.Y\., J\.Z\., G\.S\., Q\.L\., Y\.Z\., and L\.W\. contributed to the biological motivation and interpretation\. H\.Y\., K\.Z\. and L\.W\. analyzed the results\. H\.Y\., K\.Z\., and L\.W\. wrote the paper\. All authors discussed the results and revised the manuscript\. L\.W\. supervised the project\. ## Competing Interests Statement The authors declare no competing interests\. ## References - Allen\-Zhu and Li \(2023\)Z\. Allen\-Zhu and Y\. LiTowards understanding ensemble, knowledge distillation and self\-distillation in deep learning\.InInternational Conference on Learning Representations,Cited by:[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p5.1)\. - Asoet al\.\(2014a\)Y\. Aso, D\. Hattori, Y\. Yu, R\. M\. Johnston, N\. A\. Iyer, T\. Ngo, H\. Dionne, L\. Abbott, R\. Axel, H\. Tanimoto,et al\.The neuronal architecture of the mushroom body provides a logic for associative learning\.eLife3,pp\. e04577\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Aso and Rubin \(2016\)Y\. Aso and G\. M\. RubinDopaminergic neurons write and update memories with cell\-type\-specific rules\.elife5,pp\. e16135\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Asoet al\.\(2014b\)Y\. Aso, D\. Sitaraman, T\. Ichinose, K\. R\. Kaun, K\. Vogt, G\. Belliart\-Guérin, P\. Plaçais, A\. A\. Robie, N\. Yamagata, C\. Schnaitmann,et al\.Mushroom body output neurons encode valence and guide memory\-based action selection in drosophila\.eLife3,pp\. e04580\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1),[§4\.3](https://arxiv.org/html/2609.25146#S4.SS3.p8.2)\. - Behrouzet al\.\(2026\)A\. Behrouz, P\. Zhong, and V\. MirrokniTitans: learning to memorize at test time\.Advances in Neural Information Processing Systems38,pp\. 113506–113543\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1)\. - Buzzegaet al\.\(2020\)P\. Buzzega, M\. Boschini, A\. Porrello, D\. Abati, and S\. CalderaraDark experience for general continual learning: a strong, simple baseline\.Advances in Neural Information Processing Systems33,pp\. 15920–15930\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.20.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.16.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.6.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.16.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.6.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.5.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.16.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.6.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.5.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.5.1),[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Caronet al\.\(2021\)M\. Caron, H\. Touvron, I\. Misra, H\. Jégou, J\. Mairal, P\. Bojanowski, and A\. JoulinEmerging properties in self\-supervised vision transformers\.InIEEE/CVF International Conference on Computer Vision,pp\. 9650–9660\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Caronet al\.\(2013\)S\. J\. Caron, V\. Ruta, L\. F\. Abbott, and R\. AxelRandom convergence of olfactory inputs in the drosophila mushroom body\.Nature497\(7447\),pp\. 113–117\.Cited by:[§B\.1](https://arxiv.org/html/2609.25146#A2.SS1.p3.1),[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Cervantes\-Sandovalet al\.\(2013\)I\. Cervantes\-Sandoval, A\. Martin\-Peña, J\. A\. Berry, and R\. L\. DavisSystem\-like consolidation of olfactory memories in drosophila\.Journal of Neuroscience33\(23\),pp\. 9846–9854\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Chenet al\.\(2021\)X\. Chen, S\. Xie, and K\. HeAn empirical study of training self\-supervised vision transformers\.InIEEE/CVF International Conference on Computer Vision,pp\. 9640–9649\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Chenet al\.\(2022\)Z\. Chen, Y\. Deng, Y\. Wu, Q\. Gu, and Y\. LiTowards understanding the mixture\-of\-experts layer in deep learning\.Advances in Neural Information Processing Systems35,pp\. 23049–23062\.Cited by:[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p5.1)\. - Cohnet al\.\(2015\)R\. Cohn, I\. Morantte, and V\. RutaCoordinated and compartmentalized neuromodulation shapes sensory processing in drosophila\.Cell163\(7\),pp\. 1742–1755\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Dasguptaet al\.\(2017\)S\. Dasgupta, C\. F\. Stevens, and S\. NavlakhaA neural algorithm for a fundamental computing problem\.Science358\(6364\),pp\. 793–796\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1),[§4\.3](https://arxiv.org/html/2609.25146#S4.SS3.p8.2)\. - Davis \(2023\)R\. L\. DavisLearning and memory using drosophila melanogaster: a focus on advances made in the fifth decade of research\.Genetics224\(4\),pp\. iyad085\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - De Langeet al\.\(2021\)M\. De Lange, R\. Aljundi, M\. Masana, S\. Parisot, X\. Jia, A\. Leonardis, G\. Slabaugh, and T\. TuytelaarsA continual learning survey: defying forgetting in classification tasks\.IEEE Transactions on Pattern Analysis and Machine Intelligence44\(7\),pp\. 3366–3385\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1)\. - Dohareet al\.\(2024\)S\. Dohare, J\. F\. Hernandez\-Garcia, Q\. Lan, P\. Rahman, A\. R\. Mahmood, and R\. S\. SuttonLoss of plasticity in deep continual learning\.Nature632\(8026\),pp\. 768–774\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1)\. - Donget al\.\(2020\)X\. Dong, Z\. Yu, W\. Cao, Y\. Shi, and Q\. MaA survey on ensemble learning\.Frontiers of Computer Science14\(2\),pp\. 241–258\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Dorkenwaldet al\.\(2022\)S\. Dorkenwald, C\. E\. McKellar, T\. Macrina, N\. Kemnitz, K\. Lee, R\. Lu, J\. Wu, S\. Popovych, E\. Mitchell, B\. Nehoran,et al\.FlyWire: online community for whole\-brain connectomics\.Nature Methods19\(1\),pp\. 119–128\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p4.1)\. - Dosovitskiyet al\.\(2020\)A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly,et al\.An image is worth 16x16 words: transformers for image recognition at scale\.arXiv preprint arXiv:2010\.11929\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Fultonet al\.\(2024\)K\. A\. Fulton, D\. Zimmerman, A\. Samuel, K\. Vogt, and S\. R\. DattaCommon principles for odour coding across vertebrates and invertebrates\.Nature Reviews Neuroscience25\(7\),pp\. 453–472\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Handleret al\.\(2019\)A\. Handler, T\. G\. Graham, R\. Cohn, I\. Morantte, A\. F\. Siliciano, J\. Zeng, Y\. Li, and V\. RutaDistinct dopamine receptor pathways underlie the temporal sensitivity of associative learning\.Cell178\(1\),pp\. 60–75\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Hansen and Salamon \(2002\)L\. K\. Hansen and P\. SalamonNeural network ensembles\.IEEE Transactions on Pattern Analysis and Machine Intelligence12\(10\),pp\. 993–1001\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Hendryckset al\.\(2021\)D\. Hendrycks, S\. Basart, N\. Mu, S\. Kadavath, F\. Wang, E\. Dorundo, R\. Desai, T\. Zhu, S\. Parajuli, M\. Guo,et al\.The many faces of robustness: a critical analysis of out\-of\-distribution generalization\.InIEEE/CVF International Conference on Computer Vision,pp\. 8340–8349\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - Honeggeret al\.\(2011\)K\. S\. Honegger, R\. A\. Campbell, and G\. C\. TurnerCellular\-resolution population imaging reveals robust sparse coding in the drosophila mushroom body\.Journal of Neuroscience31\(33\),pp\. 11772–11785\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1)\. - Huet al\.\(2021\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLora: low\-rank adaptation of large language models\.arXiv preprint arXiv:2106\.09685\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Huanget al\.\(2025\)L\. Huang, X\. Cao, H\. Lu, Y\. Meng, F\. Yang, and X\. LiuMind the gap: preserving and compensating for the modality gap in clip\-based continual learning\.InIEEE/CVF International Conference on Computer Vision,pp\. 3777–3786\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.16.2.1.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.10.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Huanget al\.\(2024\)Y\. Huang, G\. Chen, J\. Xu, M\. Zhang, L\. Yang, B\. Pei, H\. Zhang, L\. Dong, Y\. Wang, L\. Wang,et al\.Egoexolearn: a dataset for bridging asynchronous ego\-and exo\-centric view of procedural activities in real world\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 22072–22086\.Cited by:[Figure 5](https://arxiv.org/html/2609.25146#S2.F5),[Figure 5](https://arxiv.org/html/2609.25146#S2.F5.5),[§2\.3](https://arxiv.org/html/2609.25146#S2.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - Hugheset al\.\(2024\)E\. Hughes, M\. D\. Dennis, J\. Parker\-Holder, F\. Behbahani, A\. Mavalankar, Y\. Shi, T\. Schaul, and T\. RocktäschelPosition: open\-endedness is essential for artificial superhuman intelligence\.InInternational Conference on Machine Learning,Vol\.235,pp\. 20597–20616\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§3](https://arxiv.org/html/2609.25146#S3.p2.1)\. - Jacobset al\.\(1991\)R\. A\. Jacobs, M\. I\. Jordan, S\. J\. Nowlan, and G\. E\. HintonAdaptive mixtures of local experts\.Neural Computation3\(1\),pp\. 79–87\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Jhaet al\.\(2024\)S\. Jha, D\. Gong, and L\. YaoClap4clip: continual learning with probabilistic finetuning for vision\-language models\.Advances in Neural Information Processing Systems37,pp\. 129146–129186\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.15.2.1.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.9.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Kanget al\.\(2025\)Z\. Kang, L\. Wang, X\. Zhang, and K\. AlahariAdvancing prompt\-based methods for replay\-independent general continual learning\.InInternational Conference on Learning Representations,Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.8.2.1.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Kirkpatricket al\.\(2017\)J\. Kirkpatrick, R\. Pascanu, N\. Rabinowitz, J\. Veness, G\. Desjardins, A\. A\. Rusu, K\. Milan, J\. Quan, T\. Ramalho, A\. Grabska\-Barwinska, D\. Hassabis, C\. Clopath, D\. Kumaran, and R\. HadsellOvercoming catastrophic forgetting in neural networks\.Proceedings of the National Academy of Sciences114\(13\),pp\. 3521–3526\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.11.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.21.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.3.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.31.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.17.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.7.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.17.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.7.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.6.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.6.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.6.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.6.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.6.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.6.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.6.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.6.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.4.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.6.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.17.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.7.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.6.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.6.1),[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Kohet al\.\(2021\)H\. Koh, D\. Kim, J\. Ha, and J\. ChoiOnline continual learning on class incremental blurry task configuration with anytime inference\.arXiv preprint arXiv:2110\.10031\.Cited by:[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1)\. - Krizhevskyet al\.\(2009\)A\. Krizhevsky G\. Hintonet al\.Learning multiple layers of features from tiny images\.Technical reportCiteseer\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - LeCun \(2022\)Y\. LeCunA path towards autonomous machine intelligence\.Note:OpenReviewVersion 0\.9\.2External Links:[Link](https://openreview.net/forum?id=BZ5a1r-kVsf)Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§3](https://arxiv.org/html/2609.25146#S3.p2.1)\. - Lesteret al\.\(2021\)B\. Lester, R\. Al\-Rfou, and N\. ConstantThe power of scale for parameter\-efficient prompt tuning\.arXiv preprint arXiv:2104\.08691\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Liet al\.\(2020\)F\. Li, J\. W\. Lindsey, E\. C\. Marin, N\. Otto, M\. Dreher, G\. Dempsey, I\. Stark, A\. S\. Bates, M\. W\. Pleijzier, P\. Schlegel,et al\.The connectome of the adult drosophila mushroom body provides insights into function\.eLife9,pp\. e62576\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Li and Liang \(2021\)X\. L\. Li and P\. LiangPrefix\-tuning: optimizing continuous prompts for generation\.arXiv preprint arXiv:2101\.00190\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Liet al\.\(2024\)Y\. Li, W\. Huang, A\. Wang, L\. Zeng, J\. Meng, and W\. ZhengEgoexo\-fitness: towards egocentric and exocentric full\-body action understanding\.InEuropean Conference on Computer Vision,pp\. 363–382\.Cited by:[§2\.3](https://arxiv.org/html/2609.25146#S2.SS3.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - Li and Hoiem \(2017\)Z\. Li and D\. HoiemLearning without forgetting\.IEEE Transactions on Pattern Analysis and Machine Intelligence40\(12\),pp\. 2935–2947\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.12.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.22.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.32.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.4.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.18.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.8.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.18.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.8.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.7.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.7.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.7.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.7.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.7.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.7.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.7.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.7.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.5.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.7.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.18.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.8.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.7.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.7.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Linet al\.\(2024\)A\. Lin, R\. Yang, S\. Dorkenwald, A\. Matsliah, A\. R\. Sterling, P\. Schlegel, S\. Yu, C\. E\. McKellar, M\. Costa, K\. Eichler,et al\.Network statistics of the whole\-brain connectome of drosophila\.Nature634\(8032\),pp\. 153–165\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§3](https://arxiv.org/html/2609.25146#S3.p3.1)\. - Liuet al\.\(2023\)B\. Liu, Y\. Zhu, C\. Gao, Y\. Feng, Q\. Liu, Y\. Zhu, and P\. StoneLibero: benchmarking knowledge transfer for lifelong robot learning\.Advances in Neural Information Processing Systems36,pp\. 44776–44791\.Cited by:[§2\.3](https://arxiv.org/html/2609.25146#S2.SS3.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - Mallya and Lazebnik \(2018\)A\. Mallya and S\. LazebnikPacknet: adding multiple tasks to a single network by iterative pruning\.InIEEE Conference on Computer Vision and Pattern Recognition,pp\. 7765–7773\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.29.2.1.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.4.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.4.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.4.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.4.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.4.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.4.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.4.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.4.1),[§2\.3](https://arxiv.org/html/2609.25146#S2.SS3.p5.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - McClellandet al\.\(1995\)J\. L\. McClelland, B\. L\. McNaughton, and R\. C\. O’ReillyWhy there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory\.\.Psychological Review102\(3\),pp\. 419\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1)\. - Modiet al\.\(2020\)M\. N\. Modi, Y\. Shuai, and G\. C\. TurnerThe drosophila mushroom body: from architecture to algorithm in a learning circuit\.Annual Review of Neuroscience43,pp\. 465–484\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Moonet al\.\(2023\)J\. Moon, K\. Park, J\. U\. Kim, and G\. ParkOnline class incremental learning on stochastic blurry task boundary via mask and visual prompt tuning\.InIEEE/CVF International Conference on Computer Vision,pp\. 11731–11741\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.7.2.1.1),[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Mu and Lin \(2025\)S\. Mu and S\. LinA comprehensive survey of mixture\-of\-experts: algorithms, theory, and applications\.arXiv preprint arXiv:2503\.07137\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.InInternational Conference on Machine Learning,pp\. 8748–8763\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p5.1)\. - Rebuffiet al\.\(2017\)S\. Rebuffi, H\. Bilen, and A\. VedaldiLearning multiple visual domains with residual adapters\.NeurIPS30\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Ridniket al\.\(2021\)T\. Ridnik, E\. Ben\-Baruch, A\. Noy, and L\. Zelnik\-ManorImagenet\-21k pretraining for the masses\.arXiv preprint arXiv:2104\.10972\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Rolnicket al\.\(2019\)D\. Rolnick, A\. Ahuja, J\. Schwarz, T\. P\. Lillicrap, and G\. WayneExperience replay for continual learning\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.19.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.30.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.15.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.5.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.15.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.5.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.5.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.5.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.5.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.5.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.5.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.5.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.5.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.5.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.4.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.15.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.5.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.4.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.4.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Russakovskyet al\.\(2015\)O\. Russakovsky, J\. Deng, H\. Su, J\. Krause, S\. Satheesh, S\. Ma, Z\. Huang, A\. Karpathy, A\. Khosla, M\. Bernstein,et al\.Imagenet large scale visual recognition challenge\.International Journal of Computer Vision115\(3\),pp\. 211–252\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Schlegelet al\.\(2024\)P\. Schlegel, Y\. Yin, A\. S\. Bates, S\. Dorkenwald, K\. Eichler, P\. Brooks, D\. S\. Han, M\. Gkantia, M\. Dos Santos, E\. J\. Munnelly,et al\.Whole\-brain annotation and multi\-connectome cell typing of drosophila\.Nature634\(8032\),pp\. 139–152\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§3](https://arxiv.org/html/2609.25146#S3.p3.1)\. - Shenet al\.\(2023\)Y\. Shen, S\. Dasgupta, and S\. NavlakhaReducing catastrophic forgetting with associative learning: a lesson from fruit flies\.Neural Computation35\(11\),pp\. 1797–1819\.Cited by:[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p4.1)\. - Silver and Sutton \(2025\)D\. Silver and R\. S\. SuttonWelcome to the era of experience\.InDesigning an Intelligence,Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§3](https://arxiv.org/html/2609.25146#S3.p2.1)\. - Smithet al\.\(2023\)J\. S\. Smith, L\. Karlinsky, V\. Gutta, P\. Cascante\-Bonilla, D\. Kim, A\. Arbelle, R\. Panda, R\. Feris, and Z\. KiraCoda\-prompt: continual decomposed attention\-based prompting for rehearsal\-free continual learning\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 11909–11919\.Cited by:[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.8.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1)\. - Wahet al\.\(2011\)C\. Wah, S\. Branson, P\. Welinder, P\. Perona, and S\. BelongieThe caltech\-ucsd birds\-200\-2011 dataset\.Technical reportCalifornia Institute of Technology\.Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p1.1)\. - Wang and Li \(2025\)L\. Wang and Q\. LiConvergent multi\-modular architecturefor adaptive learning in drosophila and artificial intelligence\.iScience28\(11\)\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p5.1)\. - Wanget al\.\(2023a\)L\. Wang, J\. Xie, X\. Zhang, M\. Huang, H\. Su, and J\. ZhuHierarchical decomposition of prompt\-based continual learning: rethinking obscured sub\-optimality\.Advances in Neural Information Processing Systems36,pp\. 69054–69076\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p2.1)\. - Wanget al\.\(2025\)L\. Wang, J\. Xie, X\. Zhang, H\. Su, and J\. ZhuHide\-pet: continual learning via hierarchical decomposition of parameter\-efficient tuning\.IEEE Transactions on Pattern Analysis and Machine Intelligence47\(8\),pp\. 6687–6702\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1)\. - Wanget al\.\(2021a\)L\. Wang, M\. Zhang, Z\. Jia, Q\. Li, C\. Bao, K\. Ma, J\. Zhu, and Y\. ZhongAFEC: active forgetting of negative transfer in continual learning\.InAdvances in Neural Information Processing Systems,Vol\.34\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§1](https://arxiv.org/html/2609.25146#S1.p2.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1)\. - Wanget al\.\(2023b\)L\. Wang, X\. Zhang, Q\. Li, M\. Zhang, H\. Su, J\. Zhu, and Y\. ZhongIncorporating neuro\-inspired adaptability for continual learning in artificial intelligence\.Nature Machine Intelligence5\(12\),pp\. 1356–1368\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1)\. - Wanget al\.\(2022a\)L\. Wang, X\. Zhang, Q\. Li, J\. Zhu, and Y\. ZhongCoSCL: cooperation of small continual learners is stronger than a big one\.InProceedings of the European Conference on Computer Vision,pp\. 254–271\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p2.1)\. - Wanget al\.\(2024\)L\. Wang, X\. Zhang, H\. Su, and J\. ZhuA comprehensive survey of continual learning: theory, method and application\.IEEE Transactions on Pattern Analysis and Machine Intelligence46\(8\),pp\. 5362–5383\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.25146#S4.SS1.p1.1)\. - Wanget al\.\(2021b\)P\. Y\. Wang, Y\. Sun, R\. Axel, L\. Abbott, and G\. R\. YangEvolving the olfactory system with machine learning\.Neuron109\(23\),pp\. 3879–3892\.Cited by:[§B\.1](https://arxiv.org/html/2609.25146#A2.SS1.p3.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2609.25146#S2.SS1.p4.1)\. - Wanget al\.\(2022b\)Y\. Wang, Z\. Huang, and X\. HongS\-prompts learning with pre\-trained transformers: an occam’s razor for domain incremental learning\.Advances in Neural Information Processing Systems35,pp\. 5682–5695\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.25.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.11.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.21.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.11.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.21.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.10.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.11.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.21.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.10.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.10.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Wanget al\.\(2022c\)Z\. Wang, Z\. Zhang, S\. Ebrahimi, R\. Sun, H\. Zhang, C\. Lee, X\. Ren, G\. Su, V\. Perot, J\. Dy,et al\.Dualprompt: complementary prompting for rehearsal\-free continual learning\.InEuropean Conference on Computer Vision,pp\. 631–648\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.14.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.24.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.34.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.6.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.10.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.20.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.10.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.20.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.9.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.9.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.9.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.9.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.9.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.9.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.9.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.9.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.7.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.9.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.10.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.20.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.9.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.9.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Wanget al\.\(2022d\)Z\. Wang, Z\. Zhang, C\. Lee, H\. Zhang, R\. Sun, X\. Ren, G\. Su, V\. Perot, J\. Dy, and T\. PfisterLearning to prompt for continual learning\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 139–149\.Cited by:[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.13.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.23.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.33.2.1.1),[Table S2](https://arxiv.org/html/2609.25146#A4.T2.5.1.5.2.1.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.19.1),[Table S10](https://arxiv.org/html/2609.25146#A6.T10.5.1.9.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.19.1),[Table S11](https://arxiv.org/html/2609.25146#A6.T11.5.1.9.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig1.3.1.8.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig2.3.1.8.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig3.3.1.8.1),[Table S12](https://arxiv.org/html/2609.25146#A6.T12.fig4.3.1.8.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig1.3.1.8.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig2.3.1.8.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig3.3.1.8.1),[Table S13](https://arxiv.org/html/2609.25146#A6.T13.fig4.3.1.8.1),[Table S5](https://arxiv.org/html/2609.25146#A6.T5.5.1.6.1),[Table S6](https://arxiv.org/html/2609.25146#A6.T6.5.1.8.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.19.1),[Table S7](https://arxiv.org/html/2609.25146#A6.T7.5.1.9.1),[Table S8](https://arxiv.org/html/2609.25146#A6.T8.5.1.8.1),[Table S9](https://arxiv.org/html/2609.25146#A6.T9.5.1.8.1),[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p2.1),[§4\.4](https://arxiv.org/html/2609.25146#S4.SS4.p3.1)\. - Windinget al\.\(2023\)M\. Winding, B\. D\. Pedigo, C\. L\. Barnes, H\. G\. Patsolic, Y\. Park, T\. Kazimiers, A\. Fushiki, I\. V\. Andrade, A\. Khandelwal, J\. Valdes\-Aleman,et al\.The connectome of an insect brain\.Science379\(6636\),pp\. eadd9330\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p3.1),[§3](https://arxiv.org/html/2609.25146#S3.p3.1)\. - Yanet al\.\(2026a\)H\. Yan, G\. Sun, K\. Zhou, Q\. Li, L\. Wang, and Y\. ZhongFlyPrompt: brain\-inspired random\-expanded routing with temporal\-ensemble experts for general continual learning\.InInternational Conference on Learning Representations,Cited by:[Appendix E](https://arxiv.org/html/2609.25146#A5.p1.1),[§3](https://arxiv.org/html/2609.25146#S3.p1.1)\. - Yanet al\.\(2026b\)H\. Yan, K\. Zhou, Y\. Liu, Q\. Shi, Y\. Zhong, and L\. WangCE4\{\}^\{4\}l: continual ego, exo, and ego\-exo learning\.InInternational Conference on Machine Learning,Cited by:[§2\.3](https://arxiv.org/html/2609.25146#S2.SS3.p2.1)\. - Zadoret al\.\(2023\)A\. M\. Zador, S\. Escola, B\. Richards, B\. Ölveczky, Y\. Bengio, K\. Boahen, M\. Botvinick, D\. Chklovskii, A\. Churchland, C\. Clopath, J\. J\. DiCarlo, S\. Ganguli, J\. Hawkins, K\. Kording, A\. Koulakov, Y\. LeCun, T\. Lillicrap, A\. Marblestone, B\. A\. Olshausen, A\. Pouget, C\. Savin, T\. J\. Sejnowski, E\. Simoncelli, S\. A\. Solla, D\. Sussillo, A\. S\. Tolias, and D\. TsaoCatalyzing next\-generation artificial intelligence through neuroai\.Nature Communications14,pp\. 1597\.Cited by:[§3](https://arxiv.org/html/2609.25146#S3.p4.1)\. - Zhanget al\.\(2026\)J\. Zhang, S\. Hu, C\. Lu, R\. T\. Lange, and J\. CluneDarwin gödel machine: open\-ended evolution of self\-improving agents\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1)\. - Zhouet al\.\(2021\)J\. Zhou, C\. Wei, H\. Wang, W\. Shen, C\. Xie, A\. Yuille, and T\. KongImage bert pre\-training with online tokenizer\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2609.25146#S2.SS2.p3.1)\. - Zhouet al\.\(2025\)K\. Zhou, Z\. Hao, L\. Wang, and X\. LiangAdaptive score alignment learning for continual perceptual quality assessment of 360\-degree videos in virtual reality\.IEEE Transactions on Visualization and Computer Graphics31\(5\),pp\. 2880–2890\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p2.1)\. - Zhouet al\.\(2024\)K\. Zhou, L\. Wang, X\. Zhang, H\. P\. Shum, F\. W\. Li, J\. Li, and X\. LiangMagr: manifold\-aligned graph regularization for continual action quality assessment\.InEuropean Conference on Computer Vision,Vol\.15069,pp\. 375–392\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p2.1)\. - Zweigeret al\.\(2026\)A\. Zweiger, J\. Pari, H\. Guo, Y\. Kim, and P\. AgrawalSelf\-adapting language models\.Advances in Neural Information Processing Systems38,pp\. 74084–74115\.Cited by:[§1](https://arxiv.org/html/2609.25146#S1.p1.1)\. ab cd Extended Data Fig\. 1:Extended results on continual vision\-language\-action benchmarks\.[1](https://arxiv.org/html/2609.25146#Sx5.F1): Performance comparison with state\-of\-the\-art baselines on LIBERO\-Goal usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\.[1](https://arxiv.org/html/2609.25146#Sx5.F1): Performance comparison with state\-of\-the\-art baselines on LIBERO\-Long usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\.[1](https://arxiv.org/html/2609.25146#Sx5.F1): Qualitative rollouts on LIBERO\-Goal, comparing task execution immediately after learning and after the final CL session for EWC, DualPrompt\+, and FlyGCL\.[1](https://arxiv.org/html/2609.25146#Sx5.F1): Qualitative rollouts on LIBERO\-Long under the same protocol\. Results are averaged over three independent runs; error bars denote the standard error of the mean\.abcExtended Data Fig\. 2:Offline results on continual vision\-language\-action benchmarks\.[2](https://arxiv.org/html/2609.25146#Sx5.F2): Unified performance summary across LIBERO\-Long, LIBERO\-Spatial, LIBERO\-Goal, and LIBERO\-Object under multiple continual learning metrics\.[2](https://arxiv.org/html/2609.25146#Sx5.F2): Performance comparison with state\-of\-the\-art baselines on LIBERO\-Spatial and LIBERO\-Object usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\.[2](https://arxiv.org/html/2609.25146#Sx5.F2): Performance comparison with state\-of\-the\-art baselines on LIBERO\-Goal and LIBERO\-Long usingAaucA\_\{\\rm auc\}andAlastA\_\{\\rm last\}\. Results are averaged over three independent runs; error bars denote the standard error of the mean\.## Appendix AProofs ### A\.1Proof of[Theorem1](https://arxiv.org/html/2609.25146#Thmtheorem1) ###### Proof\. LetFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}be the stream\-level modular predictor and letF⋆F^\{\\star\}be the ideal stream\-level predictor\. For brevity, we writeFFforFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}in this proof\. By definition, the expected GCL risk over the evolving stream is ℛGCL\(F\)=1T∑t=1T𝔼\(𝐨,𝐮\)∼Pt\[ℓ\(F\(𝐨\),𝐮\)\]\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{\(\\mathbf\{o\},\\mathbf\{u\}\)\\sim P\_\{t\}\}\\left\[\\ell\(F\(\\mathbf\{o\}\),\\mathbf\{u\}\)\\right\]\.\(S23\)The empirical fitting risk on the observed stream𝒮=\{𝒟1,…,𝒟T\}\\mathcal\{S\}=\\\{\\mathcal\{D\}\_\{1\},\\ldots,\\mathcal\{D\}\_\{T\}\\\}is ℛfit\(F\)=1T∑t=1T1Nt∑i=1Ntℓ\(F\(𝐨t,i\),𝐮t,i\)\.\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\frac\{1\}\{N\_\{t\}\}\\sum\_\{i=1\}^\{N\_\{t\}\}\\ell\(F\(\\mathbf\{o\}\_\{t,i\}\),\\mathbf\{u\}\_\{t,i\}\)\.\(S24\)The gap between the expected stream risk and the empirical fitting risk can be written as ℛGCL\(F\)−ℛfit\(F\)\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\)\-\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\)\.\(S25\)Under hierarchical modular learning, this residual gap has three sources\. First, if conflicting local distributions are assigned to insufficiently separated adaptive modules, their updates and predictions interfere with each other\. We denote the corresponding excess risk byℰsep\(F\)\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\)\. Second, if compatible local distributions or predictors are treated in isolation, the model loses potential transfer and suffers from larger estimation variance\. We denote this excess risk byℰint\(F\)\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\)\. Third, even if routing and integration are individually useful, their combination may be imperfect\. The positive excess loss of the hierarchical modular predictor over the ideal stream\-level predictor is 𝒞coord\(F\)=1T∑t=1T𝔼\(𝐨,𝐮\)∼Pt\[ℓ\(F\(𝐨\),𝐮\)−ℓ\(F⋆\(𝐨\),𝐮\)\]\+\.\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\)=\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\mathbb\{E\}\_\{\(\\mathbf\{o\},\\mathbf\{u\}\)\\sim P\_\{t\}\}\\left\[\\ell\(F\(\\mathbf\{o\}\),\\mathbf\{u\}\)\-\\ell\(F^\{\\star\}\(\\mathbf\{o\}\),\\mathbf\{u\}\)\\right\]\_\{\+\}\.\(S26\)By the assumed additive excess\-risk decomposition, the residual gap in[Eq\.S25](https://arxiv.org/html/2609.25146#A1.E25)is upper bounded by the sum of these three non\-negative components: ℛGCL\(F\)−ℛfit\(F\)≲ℰsep\(F\)\+ℰint\(F\)\+𝒞coord\(F\)\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\)\-\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\)\\lesssim\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\)\+\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\)\.\(S27\)Addingℛfit\(F\)\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\)to both sides gives ℛGCL\(F\)≲ℛfit\(F\)\+ℰsep\(F\)\+ℰint\(F\)\+𝒞coord\(F\)\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\)\\lesssim\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\)\+\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\)\+\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\)\.\(S28\)Substituting backF=Fθ,Ω,ΨF=F\_\{\\theta,\\Omega,\\Psi\}yields[Eq\.4](https://arxiv.org/html/2609.25146#S4.E4)\. This completes the proof\. ∎ ### A\.2Proof of[Proposition1](https://arxiv.org/html/2609.25146#Thmproposition1) ###### Proof\. For brevity, we writeFhierF\_\{\\mathrm\{hier\}\}forFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}\. From[Theorem1](https://arxiv.org/html/2609.25146#Thmtheorem1), the expected GCL risk of a hierarchical modular predictor satisfies ℛGCL\(Fhier\)≲ℛfit\(Fhier\)\+ℰsep\(Fhier\)\+ℰint\(Fhier\)\+𝒞coord\(Fhier\)\.\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\_\{\\mathrm\{hier\}\}\)\\lesssim\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\mathrm\{hier\}\}\)\+\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{hier\}\}\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{hier\}\}\)\+\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\mathrm\{hier\}\}\)\.\(S29\)By the definition of the routing gain in[Eq\.7](https://arxiv.org/html/2609.25146#S4.E7), we have Δroute=ℰsep\(Fθ,ψ\)−ℰsep\(FMoE\),\\Delta\_\{\\mathrm\{route\}\}=\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\-\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{MoE\}\}\),\(S30\)which gives ℰsep\(FMoE\)=ℰsep\(Fθ,ψ\)−Δroute\.\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{MoE\}\}\)=\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\-\\Delta\_\{\\mathrm\{route\}\}\.\(S31\)Since the hierarchical predictor uses routing as its separation mechanism, its separation error is upper bounded by that of the routed modular predictor, i\.e\., ℰsep\(Fhier\)≤ℰsep\(FMoE\)\.\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{MoE\}\}\)\.\(S32\)Combining[Eqs\.S31](https://arxiv.org/html/2609.25146#A1.E31)and[S32](https://arxiv.org/html/2609.25146#A1.E32), we obtain ℰsep\(Fhier\)≤ℰsep\(Fθ,ψ\)−Δroute\.\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\-\\Delta\_\{\\mathrm\{route\}\}\.\(S33\) Similarly, by the definition of the ensemble gain in[Eq\.7](https://arxiv.org/html/2609.25146#S4.E7), we have Δens=ℰint\(Fsingle\)−ℰint\(FEL\),\\Delta\_\{\\mathrm\{ens\}\}=\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\-\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{EL\}\}\),\(S34\)which gives ℰint\(FEL\)=ℰint\(Fsingle\)−Δens\.\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{EL\}\}\)=\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\-\\Delta\_\{\\mathrm\{ens\}\}\.\(S35\)Since the hierarchical predictor uses aggregation as its integration mechanism, its integration error is upper bounded by that of the ensemble predictor, i\.e\., ℰint\(Fhier\)≤ℰint\(FEL\)\.\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{EL\}\}\)\.\(S36\)Combining[Eqs\.S35](https://arxiv.org/html/2609.25146#A1.E35)and[S36](https://arxiv.org/html/2609.25146#A1.E36), we obtain ℰint\(Fhier\)≤ℰint\(Fsingle\)−Δens\.\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\-\\Delta\_\{\\mathrm\{ens\}\}\.\(S37\) Substituting[Eqs\.S33](https://arxiv.org/html/2609.25146#A1.E33)and[S37](https://arxiv.org/html/2609.25146#A1.E37)into[Eq\.S29](https://arxiv.org/html/2609.25146#A1.E29)yields ℛGCL\(Fhier\)≲\\displaystyle\\mathcal\{R\}\_\{\\mathrm\{GCL\}\}\(F\_\{\\mathrm\{hier\}\}\)\\lesssimℛfit\(Fhier\)\+ℰsep\(Fθ,ψ\)\+ℰint\(Fsingle\)\\displaystyle\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\mathrm\{hier\}\}\)\+\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\)\(S38\)−Δroute−Δens\+𝒞coord\(Fhier\)\.\\displaystyle\-\\Delta\_\{\\mathrm\{route\}\}\-\\Delta\_\{\\mathrm\{ens\}\}\+\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\mathrm\{hier\}\}\)\.Substituting backFhier=Fθ,Ω,ΨF\_\{\\mathrm\{hier\}\}=F\_\{\\theta,\\Omega,\\Psi\}gives[Eq\.8](https://arxiv.org/html/2609.25146#S4.E8)\. Finally, comparing this bound with the corresponding bound without routing and integration gains, ℛfit\(Fθ,Ω,Ψ\)\+ℰsep\(Fθ,ψ\)\+ℰint\(Fsingle\),\\mathcal\{R\}\_\{\\mathrm\{fit\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\+\\mathcal\{E\}\_\{\\mathrm\{sep\}\}\(F\_\{\\theta,\\psi\}\)\+\\mathcal\{E\}\_\{\\mathrm\{int\}\}\(F\_\{\\mathrm\{single\}\}\),\(S39\)shows that hierarchical modular learning improves the bound whenever Δroute\+Δens\>𝒞coord\(Fθ,Ω,Ψ\)\.\\Delta\_\{\\mathrm\{route\}\}\+\\Delta\_\{\\mathrm\{ens\}\}\>\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\theta,\\Omega,\\Psi\}\)\.\(S40\)This proves[Eq\.9](https://arxiv.org/html/2609.25146#S4.E9)and completes the proof\. ∎ ### A\.3Proof of[Proposition2](https://arxiv.org/html/2609.25146#Thmproposition2) ###### Proof\. For brevity, we writeFhierF\_\{\\mathrm\{hier\}\}forFθ,Ω,ΨF\_\{\\theta,\\Omega,\\Psi\}\. The coordination cost measures the excess loss introduced when routing and integration do not exactly match the ideal stream\-level organization\. We decompose this cost into two sources: mismatch due to imperfect routing and mismatch due to imperfect aggregation\. Letϵroute\\epsilon\_\{\\mathrm\{route\}\}denote the degree of routing error, i\.e\., the probability or expected weight that an input is assigned to a suboptimal module\. LetΔmis\(fθ\)\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)denote the maximum expected loss increase caused by such a suboptimal assignment under the representation induced byfθf\_\{\\theta\}\. Then the routing\-induced part of the coordination cost is bounded by 𝒞route\(Fhier\)≤ϵrouteΔmis\(fθ\)\.\\mathcal\{C\}\_\{\\mathrm\{route\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\epsilon\_\{\\mathrm\{route\}\}\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)\.\(S41\)Letϵagg\\epsilon\_\{\\mathrm\{agg\}\}denote the residual mismatch caused by aggregating modular predictions that are not perfectly compatible\. Then the total coordination cost is bounded by the sum of the routing\-induced mismatch and the aggregation\-induced mismatch: 𝒞coord\(Fhier\)≤𝒞route\(Fhier\)\+ϵagg\.\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\mathcal\{C\}\_\{\\mathrm\{route\}\}\(F\_\{\\mathrm\{hier\}\}\)\+\\epsilon\_\{\\mathrm\{agg\}\}\.\(S42\)Combining[Eqs\.S41](https://arxiv.org/html/2609.25146#A1.E41)and[S42](https://arxiv.org/html/2609.25146#A1.E42), we obtain 𝒞coord\(Fhier\)≤ϵrouteΔmis\(fθ\)\+ϵagg\.\\mathcal\{C\}\_\{\\mathrm\{coord\}\}\(F\_\{\\mathrm\{hier\}\}\)\\leq\\epsilon\_\{\\mathrm\{route\}\}\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)\+\\epsilon\_\{\\mathrm\{agg\}\}\.\(S43\)Substituting backFhier=Fθ,Ω,ΨF\_\{\\mathrm\{hier\}\}=F\_\{\\theta,\\Omega,\\Psi\}gives[Eq\.10](https://arxiv.org/html/2609.25146#S4.E10)\. It remains to show why pretrained representations reduce the mismatch penalty\. By definition,Δmis\(fθ\)\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}\)measures the expected prediction penalty when an input is routed to a suboptimal module\. If a representation maps related inputs closer together and makes conflicting inputs more separable, then the predictions produced by nearby or partially mismatched modules are less divergent for related inputs, while routing ambiguity is reduced for conflicting inputs\. A pretrained backbone provides such a more stable and semantically organized representation than a representation learned from scratch on the limited online stream\. Therefore, the mismatch penalty under a pretrained backbone is smaller: Δmis\(fθpre\)<Δmis\(fθscratch\)\.\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}^\{\\mathrm\{pre\}\}\)<\\Delta\_\{\\mathrm\{mis\}\}\(f\_\{\\theta\}^\{\\mathrm\{scratch\}\}\)\.\(S44\)Substituting[Eq\.S44](https://arxiv.org/html/2609.25146#A1.E44)into[Eq\.S43](https://arxiv.org/html/2609.25146#A1.E43)shows that pretraining reduces the upper bound of𝒞coord\\mathcal\{C\}\_\{\\mathrm\{coord\}\}, which proves[Eq\.11](https://arxiv.org/html/2609.25146#S4.E11)and completes the proof\. ∎ ## Appendix BAdditional Experimental Setups ### B\.1Biological Simulation Odor Classes and Spatial Regions\.We generate 100 class prototypes independently and uniformly in\[0,1\)50\[0,1\)^\{50\}\. Training and test inputs are sampled from the same space and assigned to their nearest prototype by squared Euclidean distance\. The prototypes are then grouped by proximity into five regions of 20 classes using capacity\-constrained clustering\. For each seed, the formal benchmark contains 10,000 training and 2,000 test inputs from each region\. The fixed test set therefore contains 10,000 inputs, while every training stream contains 50,000 inputs\. Continual Odor Streams\.The stream represents a learner moving through the five spatial regions and acquiring experience along this trajectory\. Each region defines the home stage of its classes\. In Disjoint, the 10,000 inputs from each region are shuffled within their home stage and the five stages are presented sequentially\. Blurry and Joint soften the boundaries between neighboring periods of experience without altering the inputs or labels\. For each home stage, a specified subset of classes remains disjoint, while 10% of the inputs from the other classes is removed, pooled across stages, globally shuffled, and reassigned under equal stage quotas\. Blurry retains 10 disjoint classes \(50%\) in each region, whereas Joint applies this procedure to all classes\. In every case, the model encounters data from all five regions in sequence, each input appears exactly once, and the total learning budget is unchanged\. Biologically Grounded Architecture\.The sensory network follows the population\-level compression and expansion described in biological and task\-driven accounts of the*Drosophila*olfactory system[Caron et al\. \(2013\)](https://arxiv.org/html/2609.25146#bib.bib14);[Wang et al\. \(2021b\)](https://arxiv.org/html/2609.25146#bib.bib47)\. Each of the 50 input channels is replicated across 26 ORNs and pooled by one PN, giving 1,300 ORNs and 50 PNs\. During training, independent Gaussian noise with standard deviation0\.10\.1is added to the ORN responses; evaluation is noise\-free\. Within each group, ORN\-PN weights are sampled from a truncated log\-normal distribution fitted to FlyWire connections between matched olfactory types and normalized to sum to one\. PN activity is rectified and projected to 2,000 KCs through a sparse PN\-KC connection matrix\. We randomly set 91\.5% of its entries to zero, sample the remaining 8\.5% from a truncated Gaussian distribution fitted to the nonzero FlyWire PN\-KC connection weights, and then keep the resulting matrix fixed\. After rectification, only the 100 largest KC responses \(top\-5% of the 2,000 KCs\) are retained for each odor\. The sensory pathway remains fixed throughout continual learning, so that all learning occurs in downstream components and the effects of learning and memory organization can be evaluated independently of changes in sensory encoding\. Baseline Variants\.MoE uses five independently initialized experts with distinct views of KC activity, representing spatially differentiated memory pathways\. During training, we associate each stage with one expert as a simple way to activate experts successively along the stream\. This association is specific to the experimental protocol rather than a requirement for known task identities or predefined stage semantics: experts could instead be activated sequentially after a fixed number of observed samples or by another online switching rule\. In parallel, the model maintains the mean KC representation of each observed stage\. At evaluation, an input is routed to the active expert whose mean representation has the highest cosine similarity to its KC activity\. The router is updated only from inputs observed so far and does not use class or region labels at inference\. The Baseline uses one bias\-free 100\-class readout trained at a learning rate of10−310^\{\-3\}\. EL replaces it with three independently initialized heads trained at learning rates10−210^\{\-2\},10−310^\{\-3\}, and10−410^\{\-4\}, representing fast, intermediate, and slow effective memory scales\. MoE uses five experts with one head each, and MoE\+EL uses three such heads within every expert\. Each head is optimized with its own cross\-entropy loss, and their class probabilities are averaged within the selected expert at inference\. The incremental EL ablation further compares these models with three independently initialized heads all trained at10−310^\{\-3\}, separating the effect of ensembling from that of different update rates\. Training and Evaluation\.All configurations share the same sensory encoding, observations, supervision, and online learning budget, without replay, so that performance differences primarily reflect the organization of the downstream learning and memory components\. Models are optimized online with Adam using a batch size of 64\. Each run comprises 50,000 inputs, with performance evaluated every 2,000 inputs at 25 post\-update checkpoints\. At each checkpoint, the exposed class set contains all classes represented by at least one input observed up to that point, and accuracy is evaluated on the corresponding test examples\. Exposed\-class anytime AUC is computed as the trapezoidal integral of this checkpoint\-wise accuracy over the number of observed inputs, normalized by the evaluated interval\. Results are averaged over five independently generated datasets, streams, and model initializations, and reported with 95% confidence intervals\. ### B\.2Continual Stream Construction Table S1:Detailed continual stream construction and evaluation metrics for all benchmarks\. Partition Unit specifies the semantic or distributional unit used to construct sessions\. Task Metric denotes the benchmark\-specific evaluation measure, while Reported CL Metrics denote its continual evaluation over the stream\.ScenarioBenchmark / TaskPartition UnitSettingTTrDr\_\{D\}/rBr\_\{B\}Task MetricReported CL MetricsVisualrecognitionCIFAR\-100Image classGCL550%50\\%/10%10\\%AccuracyAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}ImageNet\-RImage classGCL550%50\\%/10%10\\%AccuracyAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}CUB\-200Image classGCL550%50\\%/10%10\\%AccuracyAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}Vision\-languagelearningCIFAR\-100Image classGCL550%50\\%/30%30\\%AccuracyAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}ImageNet\-RImage classGCL550%50\\%/30%30\\%AccuracyAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}Ego\-exo videounderstandingEgoExoLearn: Skill AssessmentAction classGCL450%50\\%/30%30\\%Ranking accuracyAlastA\_\{\\rm last\},F¯T\\bar\{F\}\_\{T\},AaucA\_\{\\rm auc\}EgoExoLearn: Action AnticipationProcedure taskGCL850%50\\%/30%30\\%Top\-5 recallAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}EgoExo\-Fitness: Skill AssessmentAction classGCL450%50\\%/30%30\\%Ranking accuracyAlastA\_\{\\rm last\},F¯T\\bar\{F\}\_\{T\},AaucA\_\{\\rm auc\}EgoExo\-Fitness: Action ClassificationAction classGCL450%50\\%/30%30\\%AccuracyAlastA\_\{\\rm last\},F¯T\\bar\{F\}\_\{T\},AaucA\_\{\\rm auc\}EgoExo\-Fitness: Sequence VerificationAction classGCL450%50\\%/30%30\\%ROC\-AUC, mAPAlastA\_\{\\rm last\},F¯T\\bar\{F\}\_\{T\},AaucA\_\{\\rm auc\}EgoExo\-Fitness: Guidance\-based VerificationAction classGCL450%50\\%/30%30\\%Classification accuracy, F1AlastA\_\{\\rm last\},F¯T\\bar\{F\}\_\{T\},AaucA\_\{\\rm auc\}Vision\-language\-action learningLIBERO\-Spatial, \-Object, \-Goal, \-LongManipulation taskOffline CL10–Success rateAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}, FWT, NBTLIBERO\-Spatial, \-Object, \-Goal, \-LongManipulation taskGCL1050%50\\%/30%30\\%Success rateAlastA\_\{\\rm last\},AaucA\_\{\\rm auc\}, FWT, NBT Across benchmarks, sessions are constructed from the semantic unit most closely aligned with the prediction target \(Supplementary[Tab\.S1](https://arxiv.org/html/2609.25146#A2.T1)\)\. Visual recognition and vision\-language benchmarks use image classes, whereas ego\-exo video tasks use action classes or procedure\-level task groups to preserve the temporal and cross\-view structure of the original annotations\. LIBERO uses individual manipulation tasks as the partition unit\. All stream construction is performed within the official data splits, so redistribution across sessions does not transfer samples between training and evaluation sets\. For GCL evaluation, each semantic unit is assigned a home session and the stream is made blurry by allowing samples from the recurring component to appear outside that session according torBr\_\{B\}, while the disjoint component remains primarily session\-specific according torDr\_\{D\}\. Each training sample is processed under the benchmark\-specific online budget, and evaluation after sessionttcovers all sessions exposed up to that point\. The complementary offline LIBERO protocol instead presents the ten manipulation tasks sequentially and permits multiple passes over each task before moving to the next, while retaining the same task order and evaluation procedure\. ## Appendix CAdditional Evaluation Metrics This section details the task\-specific base metrics and additional CL metrics\. Consistent with the notation in the main paper, letℰj=\{\(𝐨j,n,𝐮j,n\)\}n=1Nj\\mathcal\{E\}\_\{j\}=\\\{\(\\mathbf\{o\}\_\{j,n\},\\mathbf\{u\}\_\{j,n\}\)\\\}\_\{n=1\}^\{N\_\{j\}\}denote the evaluation set of sessionjj, where𝐨j,n\\mathbf\{o\}\_\{j,n\}is an observation and𝐮j,n\\mathbf\{u\}\_\{j,n\}is its associated learning signal\. Let𝐮^j,n\(t\)\\hat\{\\mathbf\{u\}\}\_\{j,n\}^\{\(t\)\}denote the prediction obtained after the model has learned through sessiontt\. Each task\-specific metric defined below instantiates the session\-level performanceRt,jR\_\{t,j\}used in the CL aggregates in the main paper\. ### C\.1Task Metrics Accuracy\-Based Metrics\.For visual recognition, vision\-language classification, and action classification, the learning signal is a categorical labeluj,n∈𝒰ju\_\{j,n\}\\in\\mathcal\{U\}\_\{j\}, and we use classification accuracy: Acct,j=1Nj∑n=1Nj𝟏\[u^j,n\(t\)=uj,n\]\.\\mathrm\{Acc\}\_\{t,j\}=\\frac\{1\}\{N\_\{j\}\}\\sum\_\{n=1\}^\{N\_\{j\}\}\\mathbf\{1\}\\left\[\\hat\{u\}\_\{j,n\}^\{\(t\)\}=u\_\{j,n\}\\right\]\.\(S45\) Skill Assessment\.For skill assessment, the learning signal specifies the relative quality ordering of two video observations\. Let𝒫j\\mathcal\{P\}\_\{j\}denote the set of annotated ordered pairs, where\(a,b\)∈𝒫j\(a,b\)\\in\\mathcal\{P\}\_\{j\}indicates that𝐨j,a\\mathbf\{o\}\_\{j,a\}exhibits higher skill than𝐨j,b\\mathbf\{o\}\_\{j,b\}\. Given the predicted skill scoresq^j,a\(t\)\\hat\{q\}\_\{j,a\}^\{\(t\)\}andq^j,b\(t\)\\hat\{q\}\_\{j,b\}^\{\(t\)\}, ranking accuracy is defined as RAcct,j=1\|𝒫j\|∑\(a,b\)∈𝒫j𝟏\[q^j,a\(t\)\>q^j,b\(t\)\]\.\\mathrm\{RAcc\}\_\{t,j\}=\\frac\{1\}\{\\lvert\\mathcal\{P\}\_\{j\}\\rvert\}\\sum\_\{\(a,b\)\\in\\mathcal\{P\}\_\{j\}\}\\mathbf\{1\}\\left\[\\hat\{q\}\_\{j,a\}^\{\(t\)\}\>\\hat\{q\}\_\{j,b\}^\{\(t\)\}\\right\]\.\(S46\) Action Anticipation\.For action anticipation, the learning signal is the future verb or noun class, and we report class\-mean Top\-KKrecall withK=5K=5separately for verb and noun prediction\. Let𝒞\\mathcal\{C\}denote the evaluated class set,ℰj,c=\{n:uj,n=c\}\\mathcal\{E\}\_\{j,c\}=\\\{n:u\_\{j,n\}=c\\\}denote the evaluation samples belonging to classcc, andTopK\(𝐬j,n\(t\)\)\\operatorname\{TopK\}\(\\mathbf\{s\}\_\{j,n\}^\{\(t\)\}\)denote theKKclasses with the largest predicted scores\. The metric is defined as R@Kt,j=1\|𝒞\|∑c∈𝒞1\|ℰj,c\|∑n∈ℰj,c𝟏\[c∈TopK\(𝐬j,n\(t\)\)\]\.\\mathrm\{R\}@K\_\{t,j\}=\\frac\{1\}\{\\lvert\\mathcal\{C\}\\rvert\}\\sum\_\{c\\in\\mathcal\{C\}\}\\frac\{1\}\{\\lvert\\mathcal\{E\}\_\{j,c\}\\rvert\}\\sum\_\{n\\in\\mathcal\{E\}\_\{j,c\}\}\\mathbf\{1\}\\left\[c\\in\\operatorname\{TopK\}\\left\(\\mathbf\{s\}\_\{j,n\}^\{\(t\)\}\\right\)\\right\]\.\(S47\) Sequence and Guidance\-Based Verification\.For binary sequence verification, the learning signal indicates whether an observed execution sequence is valid\. We report ROC\-AUC and mAP\. Letℰj\+\\mathcal\{E\}\_\{j\}^\{\+\}andℰj−\\mathcal\{E\}\_\{j\}^\{\-\}denote the positive and negative sample sets, respectively, and lets^j,n\(t\)\\hat\{s\}\_\{j,n\}^\{\(t\)\}be the predicted verification score\. ROC\-AUC is defined as ROC\-AUCt,j=1\|ℰj\+\|\|ℰj−\|∑p∈ℰj\+∑n∈ℰj−𝟏\[s^j,p\(t\)\>s^j,n\(t\)\]\.\\mathrm\{ROC\\mbox\{\-\}AUC\}\_\{t,j\}=\\frac\{1\}\{\\lvert\\mathcal\{E\}\_\{j\}^\{\+\}\\rvert\\lvert\\mathcal\{E\}\_\{j\}^\{\-\}\\rvert\}\\sum\_\{p\\in\\mathcal\{E\}\_\{j\}^\{\+\}\}\\sum\_\{n\\in\\mathcal\{E\}\_\{j\}^\{\-\}\}\\mathbf\{1\}\\left\[\\hat\{s\}\_\{j,p\}^\{\(t\)\}\>\\hat\{s\}\_\{j,n\}^\{\(t\)\}\\right\]\.\(S48\)Following the standard ranking\-based evaluation, mAP is computed as mAPt,j=1\|𝒬j\|∑q∈𝒬jAPq\(t\),\\mathrm\{mAP\}\_\{t,j\}=\\frac\{1\}\{\\lvert\\mathcal\{Q\}\_\{j\}\\rvert\}\\sum\_\{q\\in\\mathcal\{Q\}\_\{j\}\}\\mathrm\{AP\}\_\{q\}^\{\(t\)\},\(S49\)where𝒬j\\mathcal\{Q\}\_\{j\}denotes the evaluated verification queries or groups\. For guidance\-based execution verification, we additionally report classification accuracy as defined in Eq\. \([S45](https://arxiv.org/html/2609.25146#A3.E45)\) and F1 score: F1t,j\\displaystyle\\mathrm\{F1\}\_\{t,j\}=2Prect,jRect,jPrect,j\+Rect,j,\\displaystyle=\\frac\{2\\,\\mathrm\{Prec\}\_\{t,j\}\\,\\mathrm\{Rec\}\_\{t,j\}\}\{\\mathrm\{Prec\}\_\{t,j\}\+\\mathrm\{Rec\}\_\{t,j\}\},\(S50\)Prect,j\\displaystyle\\mathrm\{Prec\}\_\{t,j\}=TPt,jTPt,j\+FPt,j,\\displaystyle=\\frac\{\\mathrm\{TP\}\_\{t,j\}\}\{\\mathrm\{TP\}\_\{t,j\}\+\\mathrm\{FP\}\_\{t,j\}\},\(S51\)Rect,j\\displaystyle\\mathrm\{Rec\}\_\{t,j\}=TPt,jTPt,j\+FNt,j\.\\displaystyle=\\frac\{\\mathrm\{TP\}\_\{t,j\}\}\{\\mathrm\{TP\}\_\{t,j\}\+\\mathrm\{FN\}\_\{t,j\}\}\.\(S52\) Embodied Vision\-Language\-Action\.For embodied vision\-language\-action, the learning signal corresponds to an action or trajectory and the evaluation target is successful task completion\. Letzj,n\(t\)∈\{0,1\}z\_\{j,n\}^\{\(t\)\}\\in\\\{0,1\\\}indicate whether the learned policy successfully completes episode observation𝐨j,n\\mathbf\{o\}\_\{j,n\}after learning through sessiontt\. The success rate is defined as SRt,j=1Nj∑n=1Njzj,n\(t\)\.\\mathrm\{SR\}\_\{t,j\}=\\frac\{1\}\{N\_\{j\}\}\\sum\_\{n=1\}^\{N\_\{j\}\}z\_\{j,n\}^\{\(t\)\}\.\(S53\) ### C\.2Continual Learning Metrics Each higher\-is\-better task metric above is instantiated asRt,jR\_\{t,j\}in the main\-paper definitions ofAlastA\_\{\\rm last\}andAaucA\_\{\\rm auc\}, whereRt,jR\_\{t,j\}denotes the performance on sessionjjafter learning through sessiontt\. For selected benchmarks, we additionally report average forgetting \(F¯T\\bar\{F\}\_\{T\}\), forward transfer \(FWT\), and negative backward transfer \(NBT\)\. Average Forgetting\.For a higher\-is\-better base metric,F¯T\\bar\{F\}\_\{T\}is defined as F¯T=1T−1∑j=1T−1\(maxt∈\{j,…,T−1\}Rt,j−RT,j\),\\bar\{F\}\_\{T\}=\\frac\{1\}\{T\-1\}\\sum\_\{j=1\}^\{T\-1\}\\left\(\\max\_\{t\\in\\\{j,\\ldots,T\-1\\\}\}R\_\{t,j\}\-R\_\{T,j\}\\right\),\(S54\)which measures the average degradation from the best historical performance on each previously observed session to its final performance\. Forward Transfermeasures how knowledge acquired from earlier continual sessions improves performance on a new session before it is learned\. Letbjb\_\{j\}denote the performance of an untrained or reference model on sessionjj\. We define FWT=1T−1∑j=2T\(Rj−1,j−bj\),\\mathrm\{FWT\}=\\frac\{1\}\{T\-1\}\\sum\_\{j=2\}^\{T\}\\left\(R\_\{j\-1,j\}\-b\_\{j\}\\right\),\(S55\)where higher values indicate stronger transfer from previous to future sessions\. Negative Backward Transfermeasures the effect of subsequent learning on previously learned sessions relative to their performance immediately after acquisition: NBT=1T−1∑j=1T−1\(Rj,j−RT,j\),\\mathrm\{NBT\}=\\frac\{1\}\{T\-1\}\\sum\_\{j=1\}^\{T\-1\}\\left\(R\_\{j,j\}\-R\_\{T,j\}\\right\),\(S56\)where lower values indicate less interference to previously learned sessions\. ## Appendix DAdditional Implementation Details Table S2:Method\-level optimization configurations across benchmark groups\. Each method is listed separately within the benchmark group where it is evaluated\. All methods within the same benchmark follow the same random seed, stream construction, session order, online update budget, and evaluation schedule\. Hyperparameters are selected on the validation split of the first stage stream and then fixed across sessions\. Opt\., optimizer\. LR, learning rate\. WD, weight decay\. Mem\., memory buffer capacity for replay\-based ER and DER\+\+\.BenchmarkMethodBackboneOpt\.LRWDBatchMem\.NotesVisualRecog\.SeqFTViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Sequentially fine\-tunes the full network\.EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses regularization based on the Fisher information matrix\.LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses output distillation from the previous model\.L2P[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses a learnable prompt pool for continual adaptation\.DualPrompt[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses task\-shared and task\-specific prompts\.MVP[Moon et al\. \(2023\)](https://arxiv.org/html/2609.25146#bib.bib18)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses GCL\-oriented prompt adaptation\.MISA[Kang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib17)ViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Initializes adaptation with pretrained prompts\.FlyGCLViT\-B/16Adam5×10−35\\times 10^\{\-3\}064–Uses adapter, prompt, or LoRA experts; the random expansion dimension and regularization parameter for regression are set to 10,000, and the temporal\-ensemble EMA decay rates are 0\.9 and 0\.99\.VisionLanguageSeqFTCLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Sequentially fine\-tunes the CLIP vision encoder\.EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Uses regularization based on the Fisher information matrix\.LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Uses output distillation from the previous model\.L2P[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Uses a learnable prompt pool for continual adaptation\.DualPrompt[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Uses task\-shared and task\-specific prompts\.CLAP4CLIP[Jha et al\. \(2024\)](https://arxiv.org/html/2609.25146#bib.bib22)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Provides a CLIP\-specific continual\-learning baseline\.MG\-CLIP[Huang et al\. \(2025\)](https://arxiv.org/html/2609.25146#bib.bib23)CLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Provides a CLIP\-specific alignment baseline\.FlyGCLCLIP ViT\-B/16AdamW5×10−45\\times 10^\{\-4\}5×10−45\\times 10^\{\-4\}64–Uses LoRA experts\.EgoExo\-Learn &\-FitnessSeqFTI3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Sequentially fine\-tunes the full network\.ER[Rolnick et al\. \(2019\)](https://arxiv.org/html/2609.25146#bib.bib30)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}3210%Uses reservoir\-based experience replay\.DER\+\+[Buzzega et al\. \(2020\)](https://arxiv.org/html/2609.25146#bib.bib29)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}3210%Combines experience replay with logit\-level distillation\.EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses regularization based on the Fisher information matrix\.LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses output distillation from the previous model\.L2P\+[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses a learnable adapter pool for continual adaptation\.DualPrompt\+[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses task\-shared and task\-specific adapters\.S\-Prompt\+[Wang et al\. \(2022b\)](https://arxiv.org/html/2609.25146#bib.bib24)I3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses adapter\-based adaptation and routing for sequential video tasks\.FlyGCLI3D / CLIP encoderAdamW10−410^\{\-4\}5×10−45\\times 10^\{\-4\}32–Uses adapter experts\.LIBEROSeqFTDiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Sequentially fine\-tunes the full network\.SeqLoRADiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Sequentially updates LoRA parameters\.PackNet[Mallya and Lazebnik \(2018\)](https://arxiv.org/html/2609.25146#bib.bib25)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Isolates parameters through model pruning\.ER[Rolnick et al\. \(2019\)](https://arxiv.org/html/2609.25146#bib.bib30)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}323210%10\\%Uses reservoir\-based experience replay\.EWC[Kirkpatrick et al\. \(2017\)](https://arxiv.org/html/2609.25146#bib.bib1)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Uses regularization based on the Fisher information matrix\.LwF[Li and Hoiem \(2017\)](https://arxiv.org/html/2609.25146#bib.bib5)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Uses output distillation from the previous model\.L2P\+[Wang et al\. \(2022d\)](https://arxiv.org/html/2609.25146#bib.bib20)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Uses a learnable adapter pool for continual adaptation\.DualPrompt\+[Wang et al\. \(2022c\)](https://arxiv.org/html/2609.25146#bib.bib21)DiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Uses task\-shared and task\-specific adapters\.FlyGCLDiT flow\-matchingAdam10−410^\{\-4\}10−610^\{\-6\}32–Uses adapter experts\. The visual\-recognition experiments use frozen ViT\-B/16 backbones with supervised or self\-supervised ImageNet pretraining, as specified in Methods and Supplementary[Tab\.S2](https://arxiv.org/html/2609.25146#A4.T2)\. The vision\-language experiments use the OpenAI CLIP ViT\-B/16 model with frozen pretrained encoders and update only the method\-specific lightweight modules and output components\. For ego\-exo video understanding, we retain each benchmark’s pretrained feature extraction and task\-specific prediction heads, changing only the continual adaptation module\. The embodied vision\-language\-action experiments initialize a DiT flow\-matching policy from a LIBERO\-90 pretrained checkpoint \(built upon DINOv2 vision encoder and CLIP text encoder\), and use the same observation and policy pipeline across methods within each suite\. Within each benchmark and task, methods share the same stream, session order, data preprocessing, mini\-batch schedule, and evaluation checkpoints\. Optimization settings that of each method are listed in Supplementary[Tab\.S2](https://arxiv.org/html/2609.25146#A4.T2)\. FlyGCL freezes the pretrained backbone and updates only the active adaptive expert and its output heads; the random\-expanded router is updated from accumulated sufficient statistics as feature covariance and mean\. Unless a method intrinsically requires stored samples, no replay memory is used\. Method\-specific regularization coefficients, prompt or LoRA dimensions, and other configurations follow the corresponding implementations and are held fixed across sessions after selection on the initial validation stream\. ## Appendix EExtensions beyond the Conference Version An earlier conference paper, FlyPrompt[Yan et al\. \(2026a\)](https://arxiv.org/html/2609.25146#bib.bib58), explored a preliminary brain\-inspired GCL framework based on random\-expanded expert routing and temporal\-ensemble prompt experts\. That study focused on continual image classification with prompt\-based adaptation over pretrained vision models\. The present work substantially extends this preliminary study in its scientific formulation, biological grounding, model generality, theoretical analysis, and empirical scope\. First, FlyGCL develops a broader hierarchical modular principle for GCL\. The present work formulates GCL around two complementary requirements: separating conflicting experience to reduce interference and integrating compatible experience to promote generalization\. We relate these requirements to MoE and EL, respectively, and study how their hierarchical coordination supports learning under online, uncertain, and evolving data distributions\. This formulation generalizes the original algorithmic design into a broader principle for organizing CL\. Second, the biological grounding is substantially expanded\. FlyPrompt focused on sparse random expansion and multi\-timescale memory in the*Drosophila*olfactory system\. FlyGCL develops a more complete correspondence with the hierarchical organization of olfactory learning and memory, including sparse sensory expansion, spatially differentiated downstream pathways, and learning and memory processes operating across distinct temporal scales\. We further introduce a controlled biologically grounded olfactory learning model that directly evaluates spatial specialization, temporal integration, and their hierarchical coordination under continual odor streams with different degrees of distributional recurrence\. Third, FlyGCL extends the model and theoretical scope beyond prompt\-based visual learning\. The hierarchical design can be instantiated with different lightweight learning modules, including prompts, adapters, LoRA branches, video\-specific modules, task heads, policy adapters, and action heads, while retaining pretrained backbones as relatively stable representational substrates\. We further develop a new theoretical analysis of hierarchical specialization and integration, decomposing GCL risk into separation, integration, and coordination terms and analyzing when the gains from routing and aggregation outweigh their coordination cost\. The analysis also characterizes how stable pretrained representations can reduce the penalty of imperfect coordination\. The empirical evaluation is also substantially expanded within visual CL\. Beyond reproducing the original image\-classification setting, FlyGCL is evaluated across multiple pretrained representations and lightweight learning interfaces, including prompts, adapters, and LoRA\. We further analyze the complementary contributions of MoE\-based specialization and EL\-based integration across different datasets and pretrained backbones\. These experiments examine whether the proposed principle remains effective beyond a particular prompt design or backbone initialization, and establish its generality across diverse pretrained visual representations\. We then extend GCL from unimodal visual recognition to multimodal and temporal understanding\. FlyGCL is evaluated on continual vision\-language learning with pretrained CLIP models, where learning must preserve cross\-modal semantic structure while accommodating evolving visual concepts\. Beyond overall performance, we analyze changes in image\-text representation spaces and the preservation of semantic relations under continual updates\. We further evaluate continual ego\-exo video understanding, where temporal dynamics, viewpoint shifts, and skill\-related variations introduce substantially more complex distribution changes than static image classification\. These settings test hierarchical specialization and integration across multimodal semantics, temporal structure, and cross\-view experience\. Finally, we extend GCL to embodied vision\-language\-action learning, where continual updates directly affect sequential decision\-making and action policies\. FlyGCL is evaluated on language\-conditioned robotic manipulation under evolving task distributions, providing a substantially more challenging setting than recognition\-oriented benchmarks\. Together with the controlled olfactory simulation, these experiments expand the empirical scope from static visual classification to biological modeling, multimodal perception, temporal multi\-view understanding, and embodied action\. They therefore validate the proposed hierarchical modular principle across a much broader range of models, modalities, and CL scenarios than the conference study\. ## Appendix FAdditional Results Table S3:Performance comparison on continual visual recognition benchmarks under the GCL setting using ViT\-B/16 backbones with different pretraining\. We report average anytime accuracyAaucA\_\{\\rm auc\}\(↑\\uparrow\), final average accuracyAlastA\_\{\\rm last\}\(↑\\uparrow\), and forgettingFF\(↓\\downarrow\)\. SL, use small backbone learning rate \(1/100 of classifier learnign rate\)\.Table S4:Ablation study of FlyGCL under the GCL setting using ViT\-B/16 backbones pretrained on ImageNet\-21K \(Sup\-21K\) and fine\-tuned on ImageNet\-1K \(Sup\-21K/1K\)\. We report average anytime accuracyAaucA\_\{\\rm auc\}\(↑\\uparrow\), final average accuracyAlastA\_\{\\rm last\}\(↑\\uparrow\), forgettingFF\(↓\\downarrow\), and negative backward transfer \(NBT,↓\\downarrow\)\.Table S5:Performance comparison on continual vision\-language benchmarks under the GCL setting\. We report final average accuracyAlastA\_\{\\rm last\}\(%,↑\\uparrow\), average anytime accuracyAaucA\_\{\\rm auc\}\(%,↑\\uparrow\), average forgettingFF\(%,↓\\downarrow\), and backward transferBWT\\mathrm\{BWT\}\(%,↑\\uparrow\)\.Table S6:Performance comparison on the continual skill assessment benchmark \(EgoExoLearn\)\. We report final average ranking accuracyAlastA\_\{\\rm last\}\(%,↑\\uparrow\), average forgettingF¯T\\bar\{F\}\_\{T\}\(%,↓\\downarrow\), and average anytime ranking accuracyAaucA\_\{\\rm auc\}\(%,↑\\uparrow\)\.Table S7:Performance comparison on the continual action anticipation benchmark \(EgoExoLearn\)\. We report final average Top\-5 recallAlastA\_\{\\rm last\}\(%,↑\\uparrow\) and average anytime Top\-5 recallAaucA\_\{\\rm auc\}\(%,↑\\uparrow\) for verb \(\-V\) and noun \(\-N\) prediction\. Avg\. denotes the average over Ego\-V, Ego\-N, Exo\-V, and Exo\-N\.Table S8:Performance comparison on the continual skill assessment benchmark \(EgoExo\-Fitness\)\. We report final average ranking accuracyA¯last\\bar\{A\}\_\{\\mathrm\{last\}\}\(%,↑\\uparrow\), average forgettingF¯T\\bar\{F\}\_\{T\}\(%,↓\\downarrow\), and average anytime ranking accuracyAaucA\_\{\\rm auc\}\(%,↑\\uparrow\)\.Table S9:Performance comparison on the continual action classification benchmark \(EgoExo\-Fitness\)\. We report final average accuracyAlastA\_\{\\rm last\}\(%,↑\\uparrow\), average forgettingF¯T\\bar\{F\}\_\{T\}\(%,↓\\downarrow\), and average anytime accuracyAaucA\_\{\\rm auc\}\(%,↑\\uparrow\)\.Table S10:Performance comparison on the continual sequence verification benchmark \(EgoExo\-Fitness\)\. We report final average performanceA¯last\\bar\{A\}\_\{\\mathrm\{last\}\}\(%,↑\\uparrow\), average forgettingF¯T\\bar\{F\}\_\{T\}\(%,↓\\downarrow\), and average anytime performanceAaucA\_\{\\rm auc\}\(%,↑\\uparrow\) in terms of ROC\-AUC and mAP\.Table S11:Performance comparison on the continual guidance\-based execution verification benchmark \(EgoExo\-Fitness\)\. We report final average performanceAlastA\_\{\\mathrm\{last\}\}\(%,↑\\uparrow\), average forgettingF¯T\\bar\{F\}\_\{T\}\(%,↓\\downarrow\), and average anytime performanceAaucA\_\{\\rm auc\}\(%,↑\\uparrow\) in terms of classification accuracy and F1 score\.Table S12:Performance comparison on continual embodied vision\-language\-action learning benchmarks under the GCL setting \(LIBERO\)\. We report final average success rateAlastA\_\{\\rm last\}\(%,↑\\uparrow\), average anytime success rateAaucA\_\{\\rm auc\}\(%,↑\\uparrow\), forward transfer \(FWT, %,↑\\uparrow\), and negative backward transfer \(NBT, %,↓\\downarrow\)\.\(a\)LIBERO\-Spatial \(b\)LIBERO\-Object \(c\)LIBERO\-Goal \(d\)LIBERO\-Long Table S13:Performance comparison on continual embodied vision\-language\-action learning benchmarks under the offline CL setting \(LIBERO\)\. We report final average success rateAlastA\_\{\\rm last\}\(%,↑\\uparrow\), average anytime success rateAaucA\_\{\\rm auc\}\(%,↑\\uparrow\), forward transfer \(FWT, %,↑\\uparrow\), and negative backward transfer \(NBT, %,↓\\downarrow\)\.\(a\)LIBERO\-Spatial \(b\)LIBERO\-Object \(c\)LIBERO\-Goal \(d\)LIBERO\-Long
Similar Articles
Homeostatic Continual Learning
This paper proposes Homeostatic Continual Learning, a method that enables AI agents to learn continuously in changing environments without catastrophic forgetting by detecting outliers and incrementally expanding world models.
The Art of Not Forgetting A Local Learning Architecture for Continual Learning
This paper introduces CMP (Cognitive Memory Primitive), a continual-learning architecture that uses sparse relational codes and local learning to reduce catastrophic forgetting, demonstrating better backward transfer than a Transformer with EWC on a byte-level language modeling protocol.
In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models
This paper proposes 4MAS, a novel neural architecture inspired by biological bilaterality and memory consolidation, to address catastrophic forgetting in lifelong learning, achieving competitive results on benchmark datasets.
Continual Learning Mechanisms Compose for Long-Horizon Memorization
This paper shows that combining complementary continual learning mechanisms enhances long-horizon memorization in language models, boosting retention by 28-fold through data, function, and weight anchors with merged LoRA.
Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory
This paper introduces MoNIM, a learnable memory module that integrates induction capabilities of attention heads with feed-forward networks to enable scalable continual learning in semiparametric language models, improving efficiency and retention of new knowledge.