A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design

arXiv cs.LG Papers

Summary

This survey comprehensively reviews resource-efficient architectures and hardware-software co-design for green AI, covering efficient model construction, training/deployment strategies, and sustainable hardware, aiming to guide sustainable large model development.

arXiv:2607.09084v1 Announce Type: new Abstract: The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability. This survey provides a comprehensive overview of the green development of large models, emphasizing resource-efficient architectures and full-stack hardware-software co-design. We systematically review recent advances in efficient model construction, including attention operator optimization, linear-complexity architectures, and model sparsification and merging, as well as training and deployment strategies such as data-efficient learning, parameter-efficient fine-tuning, and computational compression. Beyond algorithmic improvements, we explore energy-efficient AI hardware, including mainstream AI chips, memory optimization, cross-platform deployment, and sustainable infrastructure. Furthermore, we examine how large models are being applied to sustainability-critical domains such as DeepSeek, remote sensing interpretation, national-scale infrastructure, and global initiatives. Finally, we discuss key challenges and future directions, highlighting the need for continual learning paradigms, memory-centric hardware, and standardized evaluation protocols. This survey aims to offer a holistic roadmap toward sustainable, scalable, and socially responsible development of large models. Paper homepage: https://cje.ejournal.org.cn/article/doi/10.23919/cje.2025.00.438
Original Article
View Cached Full Text

Cached at: 07/13/26, 07:58 AM

# A Survey on the Green Development of Large Models: From Resource-Efficient Architectures to Hardware-Software Co-Design
Source: [https://arxiv.org/html/2607.09084](https://arxiv.org/html/2607.09084)
\\volumeyear

2026\\volumenumber35\\issuenumber5\\journalnameChinese Journal of Electronics\\startpage1\\endpage24\\DOI10\.23919/cje\.2025\.00\.438\\journaltypeReview\\corrauthXin Li, Yaowei Wang\\receivetimeManuscript Received September 30, 2025\\accepttimeAccepted January 13, 2026\\publishtimePublished Online February 7, 2026\\authorcopyFirstname1 Middlename1 Lastname1*et al\.*

Guiping Cao\\affilnums1[https://orcid.org/0000-0002-0682-2158](https://orcid.org/0000-0002-0682-2158)\{\}^\{\\lx@orcidlink\{0000\-0002\-0682\-2158\}\{\\orcidlogo\}\}Mingyue Guo\\affilnums1[https://orcid.org/0009-0005-2348-7530](https://orcid.org/0009-0005-2348-7530)\{\}^\{\\lx@orcidlink\{0009\-0005\-2348\-7530\}\{\\orcidlogo\}\}Xianchao Guan\\affilnums1,2[https://orcid.org/0009-0003-1384-0527](https://orcid.org/0009-0003-1384-0527)\{\}^\{\\lx@orcidlink\{0009\-0003\-1384\-0527\}\{\\orcidlogo\}\}Fan Yang\\affilnums1[https://orcid.org/0009-0005-0123-3505](https://orcid.org/0009-0005-0123-3505)\{\}^\{\\lx@orcidlink\{0009\-0005\-0123\-3505\}\{\\orcidlogo\}\}Ming Tao\\affilnums1[https://orcid.org/0000-0002-4662-7170](https://orcid.org/0000-0002-4662-7170)\{\}^\{\\lx@orcidlink\{0000\-0002\-4662\-7170\}\{\\orcidlogo\}\} Xin Li\\affilnums1,\*[https://orcid.org/0000-0002-1670-1368](https://orcid.org/0000-0002-1670-1368)\{\}^\{\\lx@orcidlink\{0000\-0002\-1670\-1368\}\{\\orcidlogo\}\}Yuxin Peng\\affilnums3[https://orcid.org/0000-0001-7658-3845](https://orcid.org/0000-0001-7658-3845)\{\}^\{\\lx@orcidlink\{0000\-0001\-7658\-3845\}\{\\orcidlogo\}\}Yaowei Wang\\affilnums2,1,\*[https://orcid.org/0000-0002-6110-4036](https://orcid.org/0000-0002-6110-4036)\{\}^\{\\lx@orcidlink\{0000\-0002\-6110\-4036\}\{\\orcidlogo\}\}11affiliationmark:Pengcheng Laboratory, Shenzhen 518066, China 22affiliationmark:Harbin Institute of Technology \(Shenzhen\), Shenzhen 518055, China 33affiliationmark:Peking University, Beijing 100080, China [xinlihitsz@gmail\.com, wangyw@pcl\.ac\.cn](https://arxiv.org/html/2607.09084v1/mailto:[email protected],%[email protected])

###### Abstract

The rapid expansion of large\-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability\. This survey provides a comprehensive overview of the green development of large models, emphasizing resource\-efficient architectures and full\-stack hardware\-software co\-design\. We systematically review recent advances in efficient model construction, including attention operator optimization, linear\-complexity architectures, and model sparsification and merging, as well as training and deployment strategies such as data\-efficient learning, parameter\-efficient fine\-tuning, and computational compression\. Beyond algorithmic improvements, we explore energy\-efficient AI hardware, including mainstream AI chips, memory optimization, cross\-platform deployment, and sustainable infrastructure\. Furthermore, we examine how large models are being applied to sustainability\-critical domains such as DeepSeek, remote sensing interpretation, national\-scale infrastructure, and global initiatives\. Finally, we discuss key challenges and future directions, highlighting the need for continual learning paradigms, memory\-centric hardware, and standardized evaluation protocols\. This survey aims to offer a holistic roadmap toward sustainable, scalable, and socially responsible development of large models\.

###### keywords:

Green AI, Model Efficiency, Hardware\-Software Co\-Design, Sustainable Computing, Large Models

## 1Introduction

In recent years, large\-scale Artificial Intelligence \(AI\) models, especially those built upon Transformer architectures\[[1](https://arxiv.org/html/2607.09084#bib.bib1)\], have achieved significant breakthroughs in natural language processing, computer vision, and scientific computing\. Flagship models such as BERT\[[2](https://arxiv.org/html/2607.09084#bib.bib2)\], CLIP\[[3](https://arxiv.org/html/2607.09084#bib.bib3)\], LLaVA\[[4](https://arxiv.org/html/2607.09084#bib.bib4)\], and the GPT series\[[5](https://arxiv.org/html/2607.09084#bib.bib5)\]have continuously expanded the frontier of AI\. However, this progress has come with substantial costs: the exponential increase in model parameters, training data, and input length has led to skyrocketing computational demands, resulting in massive energy consumption and environmental concerns, as shown in Table[1](https://arxiv.org/html/2607.09084#S1.T1)\. As model parameters scale from billions to trillions and input sequences extend from 1K to 100K tokens, training and inference costs have increased superlinearly, posing major challenges to theaccessibility,scalability, andsustainabilityof large\-scale AI systems\.111https://cje\.ejournal\.org\.cn/article/doi/10\.23919/cje\.2025\.00\.438

![Refer to caption](https://arxiv.org/html/2607.09084v1/x1.png)Figure 1:A triangular framework illustrating the layered relationship among model architecture, training strategies, and green AI hardware\.Table 1:Comparison of GPU hours, Energy Usage, and Carbon Footprint for training mainstream large models\.†indicates no official data available, and the values are estimated and for reference only\.ModelGPU hours\(h\)Eneregy Usage\(KWh\)Carbon Emitted\(tCO2eq\)CLIP256k V10071k24†BLOOM1,083k A100475k183GPT\-33,140k∼\\sim4,600k†V1001,287k†552†GPT\-4o∼\\sim60,000k†H10016,800k†21,660†Gemini 140,000k∼\\sim50,000k†TPUv48,000k∼\\sim12,000k†4,000∼\\sim6,000†Gemini 212,000k∼\\sim15,000k†TPUv63,000k∼\\sim5,000k†1,500∼\\sim2,500†Qwen3\-Max∼\\sim300,000k†H20023,000k†\-DeepSeek\-V32,788k H8001,087k†584LLAMA1,022k A100449k173LLAMA 21,720k A100688k†291LLAMA 330,840k H100\>\>11,000k11,390SAM18k A1007k3SAM 3172k A100 or 86k H200142k∼\\sim176k66∼\\sim78

The core of this inefficiency lies in the architectural design of dominant models\. The attention mechanism\[[1](https://arxiv.org/html/2607.09084#bib.bib1)\]in Transformers scales withO​\(N2\)O\(N^\{2\}\)time and memory complexity, creating the well\-known “quadratic wall” that severely limits the processing of long contexts\. Moreover, the dense activation paradigm requires every parameter in the model to participate in each computation, irrespective of its relevance, resulting in substantial waste of computational resources and memory\. These foundational limitations have made the training of leading\-edge models extremely resource\-intensive: for instance, training a 100K\-context model can incur nearly 100×\\timesthe cost of training a 10K\-context model\. The environmental impact is equally concerning, with recent studies estimating that training a single large model can emit as much carbon as the lifetime emissions of multiple cars\[[6](https://arxiv.org/html/2607.09084#bib.bib6)\]\.

Numerous surveys have emerged in response to these challenges\[[7](https://arxiv.org/html/2607.09084#bib.bib7),[8](https://arxiv.org/html/2607.09084#bib.bib8),[9](https://arxiv.org/html/2607.09084#bib.bib9),[10](https://arxiv.org/html/2607.09084#bib.bib10)\]\. However, most of them\[[7](https://arxiv.org/html/2607.09084#bib.bib7),[8](https://arxiv.org/html/2607.09084#bib.bib8),[9](https://arxiv.org/html/2607.09084#bib.bib9)\]concentrate on isolated technical components, such as Parameter\-Efficient Fine\-Tuning \(PEFT\), quantization, or pruning, without providing a system\-level perspective\. In particular, they often neglect the role of hardware constraints and co\-design strategies, and typically lack coverage of emerging techniques like physics\-inspired models, state\-space modeling, or model merging\. Furthermore, prior works rarely explore real\-world sustainability applications, such as remote sensing or climate modeling, where efficient AI can make a tangible environmental impact\.

To address these challenges systematically, this survey adopts a top\-down perspective that reflects both the logical hierarchy and structural dependencies of efficient AI systems\. As illustrated in Figure[1](https://arxiv.org/html/2607.09084#S1.F1), we organize thegreen development of large modelsinto atriangular frameworkspanning three tightly coupled layers:\(1\)the top layer focuses on resource\-efficient architectures, where innovations in model design shape computational patterns and compatibility with downstream hardware;\(2\)the middle layer emphasizes training and deployment optimization, bridging algorithmic techniques with system constraints; and\(3\)the bottom layer centers on hardware–software co\-design, where energy\-efficient chips, inference memory optimization, and collaborative cross\-platform deployment support upper\-layer models\. This structure is not merely hierarchical, but also bidirectional: architectural breakthroughs influence chip design, while hardware capabilities inform model structure and optimization strategy\. Unlike prior literature, we present a unified view of how efficiency gains emerge from layer\-wise collaboration rather than isolated techniques\.

As shown in Figure[2](https://arxiv.org/html/2607.09084#S1.F2), following this framework,[Sec\.2](https://arxiv.org/html/2607.09084#S2)surveys architectural\-level innovations, including physics\-inspired models like vHeat\[[11](https://arxiv.org/html/2607.09084#bib.bib11)\], state\-space models such as VMamba\[[12](https://arxiv.org/html/2607.09084#bib.bib12)\], sparsity techniques like Mixture of Experts \(MoE\), and model merging\.[Sec\.3](https://arxiv.org/html/2607.09084#S3)focuses on training and optimized computation, covering data\-efficient learning strategies, distributed training, parameter\-efficient fine\-tuning, model and computational compression, such as quantization, distillation, and speculative decoding\.[Sec\.4](https://arxiv.org/html/2607.09084#S4)examines the green AI hardware, highlighting memory optimization, cross\-platform deployment, and energy\-aware hardware system design\.[Sec\.5](https://arxiv.org/html/2607.09084#S5)extends the discussion to AI for sustainability, covering recent Chinese advances like DeepSeek and the “Aerospace·Lingmou” 3\.0 system, national\-scale infrastructure such as C2NET, and global initiatives led by Google, Meta, and others\. Finally,[Sec\.6](https://arxiv.org/html/2607.09084#S6)discusses future opportunities, including continual learning, neuromorphic computing, edge AI, and the pressing need for standardized efficiency benchmarks\. Collectively, these chapters provide a layered and interconnected perspective on the emerging landscape of green AI\.

By integrating model design, algorithm optimization, and hardware implementation into a unified framework, this survey aims to provide both theoretical insights and practical guidance for developing scalable, low\-carbon AI systems\. We emphasize the interplay of sparsity, modularity, and hardware\-awareness as recurring principles across layers, and advocate for a lifecycle\-aware approach to AI sustainability, from pretraining to deployment and beyond\.

In summary, this survey contributes to the field by offering a unified, full\-stack perspective that emphasizes the interdependence between model design, optimization strategies, and hardware implementation\. We introduce a triangular framework that captures the bidirectional influences among layers and systematically organizes recent advances within it\. Additionally, we highlight how design decisions at one layer affect constraints and opportunities at others\. By synthesizing a broad range of techniques into a cohesive view, we aim to support researchers and practitioners in navigating the complex trade\-offs involved in building efficient, scalable, and sustainable large\-scale AI systems\.

\{forest\}Figure 2:Overview of the paper structure, detailing Chapter[1](https://arxiv.org/html/2607.09084#S1)\-[7](https://arxiv.org/html/2607.09084#S7)\.
## 2Efficient Model Construction

The rapid progress in large\-scale neural networks has driven remarkable advances in AI, but at the expense of sharply increasing computational and memory costs\. As models and context lengths grow, the conventional Transformer architecture encounters fundamental scalability issues, particularly due to the quadratic complexity of its attention mechanism, which becomes a major bottleneck in long\-sequence scenarios\. This section provides a review of cutting\-edge innovations at the architectural level that are designed to mitigate the computational and memory demands of large models\. We focus on three major research directions: \(1\)attention operator optimization; \(2\)efficient model design; \(3\)model sparsification and merging\. Each of these approaches seeks to decouple model performance from resource consumption, opening pathways to sustainable and scalable AI development\.

### 2\.1Attention Operator Optimization

Computational operators are the core kernels executing tensor operations, a function crucial for enhancing computational efficiency\. A representative case is the attention mechanism in Transformers, whose quadratic complexity with respect to sequence length introduces fundamental bottlenecks in memory and speed\. To overcome these limitations, research has pivoted to hardware\-aware and efficiency\-driven operator designs\. The following section highlights several representative operator designs\.

#### 2\.1\.1Quadratic Complexity in Attention Mechanism

Traditional attention mechanisms achieve global dependency modeling by computing correlations between every position in a sequence through the scaled dot\-product attention formula:Attention⁡\(𝑸,𝑲,𝑽\)=softmax⁡\(𝑸​𝑲Tdk\)​𝑽\\operatorname\{Attention\}\(\\bm\{Q\},\\bm\{K\},\\bm\{V\}\)=\\operatorname\{softmax\}\\left\(\\frac\{\\bm\{Q\}\\bm\{K\}^\{T\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\)\\bm\{V\}, where𝑸\\bm\{Q\}\(queries\),𝑲\\bm\{K\}\(keys\), and𝑽\\bm\{V\}\(values\) are linear projections of the input sequence\. This process involves calculating𝑸​𝑲T\\bm\{Q\}\\bm\{K\}^\{T\}to generate anN×NN\\times Nsimilarity matrix, scaling it, applying softmax for normalization, and finally weighting𝑽\\bm\{V\}vectors\. Multi\-head self\-/cross\-attention extends this by running multiple attention functions in parallel to capture diverse representations\. However, the core limitation arises from the𝑸​𝑲T\\bm\{Q\}\\bm\{K\}^\{T\}computation, which exhibits quadratic complexityO​\(N2\)O\\left\(N^\{2\}\\right\)in both computation and memory requirements\. For example, when sequence length increases from 1K to 32K tokens, the attention matrix storage demand grows by a factor of 1,024, creating a severe “quadratic wall” that hinders long\-sequence processing\. Despite the success of models like BERT\[[2](https://arxiv.org/html/2607.09084#bib.bib2)\], CLIP\[[3](https://arxiv.org/html/2607.09084#bib.bib3)\], and GPT\[[57](https://arxiv.org/html/2607.09084#bib.bib57)\]in NLP and vision tasks, this complexity poses significant challenges for practical deployment in scenarios requiring long\-context understanding, such as long document analysis, high\-resolution image processing, and extended video sequences\.

#### 2\.1\.2Linear Attention

Following the quadratic complexity challenges of traditional Transformers, Linear Attention\[[13](https://arxiv.org/html/2607.09084#bib.bib13)\]reduces computational complexity fromO​\(N2\)O\\left\(N^\{2\}\\right\)to linearO​\(N\)O\(N\)through feature mapping functionsϕ​\(𝑸\)\\phi\(\\bm\{Q\}\)andϕ​\(𝑲\)\\phi\(\\bm\{K\}\)that transform Query and Key matrices into higher\-dimensional feature representations\. The key insight reformulates attention computation by applying feature mappings such asϕ​\(x\)=ReLU⁡\(x\)\\phi\(x\)=\\operatorname\{ReLU\}\(x\)orϕ​\(x\)=ELU⁡\(x\)\+1\\phi\(x\)=\\operatorname\{ELU\}\(x\)\+1to queries and keys, then leveraging matrix multiplication associativity to reorganize computation order and avoid explicit attention matrix construction\. Instead of computing the full𝑸​𝑲T\\bm\{Q\}\\bm\{K\}^\{T\}attention matrix, Linear Attention precomputes the combinations of mapped keys with values and reuses them for every query, achieving linear complexity\. While this approach theoretically enables processing of arbitrarily long sequences, Linear Attention suffers from significant performance degradation, particularly in tasks requiring precise positional information and local dependency modeling, as the linearization process loses the sharp, focused distributions that traditional Softmax attention provides\. Representative works include Linformer\[[58](https://arxiv.org/html/2607.09084#bib.bib58)\]and Linear Transformer\[[59](https://arxiv.org/html/2607.09084#bib.bib59)\], primarily applied to long document understanding and time series prediction, but practical deployment remains limited due to substantial accuracy losses that often cannot be justified by efficiency gains in core NLP tasks\.

#### 2\.1\.3Flash Attention

Flash Attention\[[14](https://arxiv.org/html/2607.09084#bib.bib14)\]addresses the memory bottleneck of standard attention mechanisms through tiled computation and memory access optimization, maintainingO​\(N2\)O\\left\(N^\{2\}\\right\)computational complexity while achieving significant memory efficiency improvements\. Unlike Linear Attention’s complexity reduction approach, Flash Attention reorganizes computation patterns by decomposing attention into blocks that fit within GPU’s SRAM, using online softmax algorithms to avoid materializing complete attention matrices and reducing memory complexity fromO​\(N2\)O\\left\(N^\{2\}\\right\)toO​\(N\)O\(N\)\. The technique minimizes expensive memory transfers between HBM \(High Bandwidth Memory\) and SRAM, achieving 2\-4×\\timesspeedups for sequences up to 8K tokens, though computational costs for ultra\-long sequences \(\>\>100K\\mathrm\{K\}tokens\) remain prohibitive due to the unchanged algorithmic complexity\. The Flash Attention family has evolved through v1/v2/v3 iterations\[[14](https://arxiv.org/html/2607.09084#bib.bib14),[15](https://arxiv.org/html/2607.09084#bib.bib15),[16](https://arxiv.org/html/2607.09084#bib.bib16)\], with v2 adding variable sequence length support and v3 enhancing hardware utilization, becoming ubiquitous in production systems including GPT\-4\[[5](https://arxiv.org/html/2607.09084#bib.bib5)\], Claude\[[60](https://arxiv.org/html/2607.09084#bib.bib60)\], and LLaMA\[[61](https://arxiv.org/html/2607.09084#bib.bib61)\]\. This memory optimization approach is complementary to Linear Attention’s complexity reduction, suggesting potential for hybrid methods that combine both algorithmic efficiency and implementation optimizations for next\-generation attention mechanisms\.

These operator\-level optimizations enable longer sequences and lower latency while bridging the gap between theoretical algorithm efficiency and real\-world performance\. Together, they form the foundation of modern deep learning engines, allowing full\-stack computation redesign without changing model semantics\. Such innovations are essential for efficient training and deployment of next\-generation large models, making AI computing more efficient and accessible\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x2.png)Figure 3:Computational complexity comparison of fundamental architectures: \(a\) Transformers\[[1](https://arxiv.org/html/2607.09084#bib.bib1)\]withO​\(N2\)O\(N^\{2\}\)self\-attention, \(b\) vHeat\[[11](https://arxiv.org/html/2607.09084#bib.bib11)\]withO​\(N1\.5\)O\(N^\{1\.5\}\)heat conduction operator, and \(c\) VMamba\[[12](https://arxiv.org/html/2607.09084#bib.bib12)\]withO​\(N\)O\(N\)cross\-scan mechanism\.

### 2\.2Efficient Model Design

To advance the computational efficiency of large\-scale models, it is essential to not only optimize computational operators for enhancing microscopic tensor computations but also adopt a holistic approach toward efficient model construction that improves end\-to\-end model efficiency\.

The following sections explore a range of efficient models and the innovative techniques behind them, highlighting their collective contribution to improved efficiency\.

#### 2\.2\.1Physics\-Inspired Models

The vHeat\[[11](https://arxiv.org/html/2607.09084#bib.bib11)\]model represents innovative exploration in physics\-inspired modeling, achieving complexity reduction fromO​\(N2\)O\\left\(N^\{2\}\\right\)toO​\(N1\.5\)O\\left\(N^\{1\.5\}\\right\)by analogizing visual information propagation to heat conduction phenomena\. As illustrated in Figure[4](https://arxiv.org/html/2607.09084#S2.F4), the model’s core innovation introduces the Heat Conduction Operator \(HCO\), designed based on general solutions of 2D heat conduction equations, simulating heat diffusion processes in the frequency domain through efficient Discrete Cosine Transform \(DCT\) and Inverse DCT \(IDCT\)\. Adaptive heat diffusion coefficients are dynamically predicted through learnable Frequency Value Embeddings \(FVEs\), achieving global receptive fields while maintaining sub\-quadratic complexity\. vHeat’s primary limitation lies in the approximation nature of physical simulation, which may not fully capture complex semantic dependencies\. The model demonstrates excellent performance in image classification and object detection tasks, with particularly notable advantages in remote sensing image processing\[[62](https://arxiv.org/html/2607.09084#bib.bib62)\]\. This establishes a critical theoretical and technical foundation for the future development of linear complexity models\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x3.png)Figure 4:Illustration of vHeat’s\[[11](https://arxiv.org/html/2607.09084#bib.bib11)\]framework\. It shows the heat conduction\-inspired mechanism and the integration of the Discrete Cosine Transform for efficient global dependency modeling\.
#### 2\.2\.2State Space Models

State Space Models \(SSMs\) have emerged as a powerful alternative to traditional attention\-based architectures, achieving breakthroughs in efficient long\-sequence modeling\. By introducing state variables to encode historical context, SSMs transform sequence modeling from quadratic attention computation to linear\-time state transitions, reducing time and space complexity fromO​\(N2\)O\\left\(N^\{2\}\\right\)toO​\(N\)O\\left\(N\\right\)\. The foundational principle of SSMs lies in their formulation as discrete\-time dynamical systems, mathematically governed by state transition matrix𝑨\\bm\{A\}, input matrix𝑩\\bm\{B\}, and output matrix𝑪\\bm\{C\}\. This formulation allows the model to capture long\-range dependencies through continuous state evolution, avoiding the need for pairwise token interactions\. Representative models include S4\[[63](https://arxiv.org/html/2607.09084#bib.bib63)\], Mamba\[[64](https://arxiv.org/html/2607.09084#bib.bib64)\], and its vision variant VMamba\[[12](https://arxiv.org/html/2607.09084#bib.bib12)\]\. S4\[[63](https://arxiv.org/html/2607.09084#bib.bib63)\]introduced a hardware\-aware efficient algorithm and demonstrated strong performance in long\-range reasoning tasks\. Mamba\[[64](https://arxiv.org/html/2607.09084#bib.bib64)\]further advanced SSMs by proposing the Selective Scan Mechanism \(SSM\), which enables state transition parameters to dynamically adapt based on input content\. This overcomes a key limitation of earlier SSMs, which struggled with content\-aware reasoning\. VMamba\[[12](https://arxiv.org/html/2607.09084#bib.bib12)\]extended these principles to computer vision, introducing a cross\-scale scanning strategy to maintain global receptive fields in visual tasks\. The architecture achieved a2\.8×2\.8\\timesspeedup and reduced GPU memory usage by86\.8%86\.8\\%in high\-resolution image processing, showcasing the versatility of state\-space methods beyond language\.

These successes highlight the role of SSMs in enabling scalable and efficient inference in data\-intensive domains, offering a promising path toward sustainable and high\-performance AI systems\.

#### 2\.2\.3Other Linear\-Complexity Models

Beyond the mainstream SSMs development trajectory, the academic community has explored numerous innovative approaches to linear complexity architectures, each addressing specific aspects of the quadratic bottleneck through distinct technical pathways\.

RWKV\(Receptance Weighted Key Value\)\[[17](https://arxiv.org/html/2607.09084#bib.bib17)\]achieves linear complexity language modeling through ingenious fusion of RNN’s recurrent mechanisms with Transformer’s parallel training advantages\. Its core innovation reformulates the attention mechanism into recursive form, maintaining global information access capabilities while achieving linear complexity and elegantly bridging the gap between sequential processing efficiency and parallel training scalability\.

RetNet\(Retentive Network\)\[[18](https://arxiv.org/html/2607.09084#bib.bib18)\]proposes the Retention Mechanism, which incorporates relative positional encoding to support parallelization during training while enabling recurrent computation during inference\. This design achieves dual optimization of both training and inference efficiency, addressing the fundamental trade\-off between training parallelizability and inference efficiency that has long challenged sequence modeling architectures\.

As illustrated in Figure[3](https://arxiv.org/html/2607.09084#S2.F3), these architectural innovations demonstrate varying levels of complexity reduction, with conventional Self\-Attention constrained byO​\(N2\)O\\left\(N^\{2\}\\right\)self\-attention complexity, while novel approaches like vHeat achieveO​\(N1\.5\)O\\left\(N^\{1\.5\}\\right\)complexity through heat conduction operators, and architectures such as VMamba reach linearO​\(N\)O\\left\(N\\right\)complexity via cross\-scan mechanisms\. These diverse explorations collectively demonstrate the rich technical ecosystem emerging around linear complexity architectures, providing multiple potential pathways for future development and driving efficient sequence modeling technology toward more mature and practical applications\.

### 2\.3Model Sparsification and Merging

Building on advances in attention operators and efficient architectures, we now examine complementary approaches focused on optimizing and consolidating existing models rather than designing new ones from scratch\. These strategies reduce the footprint of large models while preserving capabilities, directly addressing computational and memory constraints\. The following sections will delve into the technical intricacies and recent advancements within these domains\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x4.png)Figure 5:Illustration of the Mixture of Experts \(MoE\) architecture\[[65](https://arxiv.org/html/2607.09084#bib.bib65)\]\. MoE employs a gating network to dynamically route input tokens to a subset of specialized expert networks, enabling efficient sparse activation and improved scalability for large\-scale models\.#### 2\.3\.1Sparse Activation Mechanism

The core idea of sparse activation is to break through these computational bottlenecks through clever design: each input token activates only a small subset of parameters in the network, dramatically reducing actual computational requirements while maintaining or even improving model performance\. The Mixture of Experts \(MoE\) architecture\[[19](https://arxiv.org/html/2607.09084#bib.bib19),[20](https://arxiv.org/html/2607.09084#bib.bib20),[21](https://arxiv.org/html/2607.09084#bib.bib21)\]represents the most successful implementation of this principle, as shown in Figure[5](https://arxiv.org/html/2607.09084#S2.F5), replacing traditional dense FFN layers with multiple specialized expert sub\-networks and employing learnable gating mechanisms to determine which experts process each token\.

The mathematical foundation of MoE can be expressed precisely: An MoE layer containsNNexpert networks\[E1,E2,…,EN\]\[E\_\{1\},E\_\{2\},\\ldots,E\_\{N\}\], where a gating network generates logits normalized through softmax distribution\. The gating value for expertiiis:

G​\(x\)=softmax⁡\(g1​\(x\),g2​\(x\),…,gN​\(x\)\)G\(x\)=\\operatorname\{softmax\}\\left\(g\_\{1\}\(x\),g\_\{2\}\(x\),\\ldots,g\_\{N\}\(x\)\\right\)\(1\)Top\-k gating selects the experts with highest probabilities, defining active expert setτ\\tau, and the MoE output becomes:

MoE⁡\(x\)=∑i∈τG​\(x\)i⋅Ei​\(x\)\\operatorname\{MoE\}\(x\)=\\sum\_\{i\\in\\tau\}G\(x\)\_\{i\}\\cdot E\_\{i\}\(x\)\(2\)This design decomposes complex problems into specialized subtasks, with each expert handling specific input patterns, achieving dual improvements in computational efficiency and model capability while dramatically reducing computation through selective expert activation\.

Table 2:Systematic classification of MoE method paradigms\. A comprehensive taxonomy of seven distinct MoE architectures spanning from early fixed assignment approaches to modern adaptive systems, highlighting their key characteristics, representative applications, and primary advantages across different deployment scenarios and performance objectives\.ParadigmKey CharacteristicsMain ApplicationsAdvantagesStatic MoEFixed expert\-input mappingMoE\[[19](https://arxiv.org/html/2607.09084#bib.bib19)\]Early theoretical workDynamic Sparse MoELearnable gating, top\-k activationGShard\[[21](https://arxiv.org/html/2607.09084#bib.bib21)\], Switch Transformer\[[20](https://arxiv.org/html/2607.09084#bib.bib20)\]Large language modelsDense MoEAll experts activatedSoft MoE\[[66](https://arxiv.org/html/2607.09084#bib.bib66)\]High\-accuracy scenariosHierarchical MoEMulti\-level expert organizationDeep MoE\[[67](https://arxiv.org/html/2607.09084#bib.bib67)\], PLE\[[68](https://arxiv.org/html/2607.09084#bib.bib68)\]Complex structured problemsMulti\-task MoETask\-specific gatingMMoE\[[69](https://arxiv.org/html/2607.09084#bib.bib69)\], PLE\[[68](https://arxiv.org/html/2607.09084#bib.bib68)\]Recommendation systemsMulti\-modal MoEModality\-specific expertsV\- MoE\[[70](https://arxiv.org/html/2607.09084#bib.bib70)\], MoE\-LLaVA\[[4](https://arxiv.org/html/2607.09084#bib.bib4)\]Vision\-language tasksAdaptive/Evolutionary MoEDynamic expert adjustmentAdaMoE\[[71](https://arxiv.org/html/2607.09084#bib.bib71)\], EvoMoE\[[72](https://arxiv.org/html/2607.09084#bib.bib72)\]Research frontiers

As shown in Table[2](https://arxiv.org/html/2607.09084#S2.T2), the evolution of MoE technology has led to diverse architectural paradigms, each addressing specific computational challenges and application requirements\. Based on fundamental differences in routing strategies, activation patterns, and specialization mechanisms, MoE architectures can be categorized into seven distinct paradigms that represent the technical spectrum from early fixed assignment approaches to modern adaptive systems\. Each paradigm offers unique advantages for different deployment scenarios and performance objectives\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x5.png)Figure 6:Illustration of the model merging paradigm\[[22](https://arxiv.org/html/2607.09084#bib.bib22)\]\. \(a\)TTseparate models forTTtasks, \(b\) A merged model forTTtasks\.
#### 2\.3\.2Model Merging

This frontier investigates approaches to combine the knowledge and capabilities of multiple pre\-trained models, each potentially specialized for different tasks, into a single, unified model without access to the original training data\. Known as Model Merging\[[22](https://arxiv.org/html/2607.09084#bib.bib22)\], this process aims to construct a generalist model by combining several specialists\. As shown in Figure[6](https://arxiv.org/html/2607.09084#S2.F6), it represents a form of zero\-shot compression and knowledge integration, eliminating the need to store or deploy multiple separate models and thus reducing both memory overhead and inference latency\.

The techniques can be broadly categorized into four key areas each contributing unique advancements: foundational methods like Task Arithmetic\[[73](https://arxiv.org/html/2607.09084#bib.bib73)\], which establishes basic principles through task vector operations to combine or remove capabilities; efficient and scalable techniques such as Model Soups\[[74](https://arxiv.org/html/2607.09084#bib.bib74)\], which employs weight averaging to enhance generalization without increasing inference costs; specialized approaches including ZipIt\[[75](https://arxiv.org/html/2607.09084#bib.bib75)\], which tailors merging for heterogeneous architectures through feature alignment; and theoretical frameworks exemplified by Linear Mode Connectivity \(LMC\)\[[76](https://arxiv.org/html/2607.09084#bib.bib76),[77](https://arxiv.org/html/2607.09084#bib.bib77)\], which explains loss landscape connectivity to justify weight interpolation\.

Based on the aforementioned approaches, such as specifically model sparsification and merging, not only effectively tackle scalability challenges in large\-scale AI systems but also advance sustainable AI development by encouraging resource\-efficient model reuse and collaborative integration\. Collectively, they form a robust post\-design methodology that produces highly efficient, versatile, and scalable models, ultimately paving the way for more sustainable and adaptable AI ecosystems\. These framework\-level efforts enhance the overall efficiency of AI deployments while minimizing computational costs and environmental impact\.

## 3Efficient Training and Optimized Computation

The scaling of large\-scale models to billions of parameters creates a severe computational bottleneck, making training and computation exceptionally costly, time\-consuming, and environmentally unsustainable due to the substantial carbon footprint\. To address these challenges, a number of techniques have been introduced across complementary fronts for developing efficient training, includingdata\-efficient techniques,distributed training techniques,parameter\-efficient fine\-tuning, andmodel and computational compression\. This section details these essential approaches and their impacts for achieving green AI development\.

### 3\.1Data\-Efficient Techniques

The reliance of large\-scale models on massive datasets makes their training computationally intensive and environmentally costly\. Data\-efficient techniques mitigate this by maximizing the utility of every data sample, reducing resource use without compromising performance\. According to the mode of data operation, this review categorizes data\-efficient techniques into three areas:sampling,distillation, andpruning\.

#### 3\.1\.1Data\-Efficient Sampling

Non\-uniform sampling reduces training data volume without compromising model performance by selectively emphasizing the most informative examples\. As a static data scheduling strategy used in LLM pre\-training \(*e\.g*\., GPT\-3\[[78](https://arxiv.org/html/2607.09084#bib.bib78)\], LLaMA\[[79](https://arxiv.org/html/2607.09084#bib.bib79)\]\), it assigns manual sampling weights to upsample high\-quality data and downsample lower\-quality content, thereby increasing exposure to valuable information while preserving diversity\.

In contrast to this static approach, Dynamic Data Sampling is a technique that dynamically adapts the sampling strategy throughout the training process based on data characteristics, model state, or task requirements\. These methods are categorized according to their adaptation criteria\.

Difficulty\-Based Sampling\. It prioritizes samples by difficulty, often starting with easier examples to stabilize training before introducing harder ones\. Curriculum Learning \(CL\)\[[80](https://arxiv.org/html/2607.09084#bib.bib80),[81](https://arxiv.org/html/2607.09084#bib.bib81),[82](https://arxiv.org/html/2607.09084#bib.bib82)\]trains models on data ordered from easy to hard, mimicking human educational progression\. Recent extensions, such as SAI\-DPO\[[23](https://arxiv.org/html/2607.09084#bib.bib23)\], dynamically select training data for mathematical reasoning using “self\-aware difficulty” metrics across different training phases, thus enhancing both data utilization efficiency and final task performance\.

Loss\-Based Sampling\. This approach dynamically prioritizes training examples by their loss values or gradient norms to focus on informative or poorly learned samples\. Core methods include direct loss reweighting \(*e\.g*\., up\-weighting high\-loss samples\), and approximated loss sampling \(using proxy models to predict loss and reduce computation\)\. The loss value is introduced as an alternative metric in pioneering approaches\[[83](https://arxiv.org/html/2607.09084#bib.bib83),[84](https://arxiv.org/html/2607.09084#bib.bib84)\]to construct a sampling distribution that reduces gradient variance compared to uniform sampling\. Importance Sampling\[[85](https://arxiv.org/html/2607.09084#bib.bib85)\]reduces variance by weighting samples according to a tractable upper bound on the per\-sample gradient norm\. Recent dynamic loss\-based method\[[86](https://arxiv.org/html/2607.09084#bib.bib86)\]employs instance\-level data reweighting to prioritize informative samples, emphasizing challenging data while reducing focus on redundant ones\.

#### 3\.1\.2Data\-Efficient Distillation

Dataset distillation\[[87](https://arxiv.org/html/2607.09084#bib.bib87)\]aims to produce a compressed, highly informative synthetic subset that preserves the core features and distribution of original dataset, thereby drastically reducing its size\. A comprehensive review\[[88](https://arxiv.org/html/2607.09084#bib.bib88)\]outlines a general framework for data distillation\. The methods are grouped into three categories:performance matching,parameter matching, anddistribution matching\.

Performance Matching\. This method optimizes synthetic data to ensure that models trained on it achieve performance comparable to those trained on original data, as measured by minimal loss on the original dataset\. They can be further categorized into two types: \(1\) Meta learning\-based approaches\[[87](https://arxiv.org/html/2607.09084#bib.bib87),[89](https://arxiv.org/html/2607.09084#bib.bib89)\], which are computationally expensive and require substantial GPU memory; \(2\) Kernel Ridge Regression \(KRR\)\-based methods\[[24](https://arxiv.org/html/2607.09084#bib.bib24),[25](https://arxiv.org/html/2607.09084#bib.bib25),[26](https://arxiv.org/html/2607.09084#bib.bib26)\], which employ convex optimization and yield a closed\-form solution for linear models, avoiding the need for extensive inner\-loop training\.

Parameter Matching\. The core of this approach involves training an identical network architecture separately on both the synthetic and original datasets, while enforcing consistency between the model parameters derived from both data sources\. Considering the number of training steps, this approach can be further categorized into single\-step parameter matching\[[90](https://arxiv.org/html/2607.09084#bib.bib90)\]and multi\-step parameter matching\[[91](https://arxiv.org/html/2607.09084#bib.bib91),[92](https://arxiv.org/html/2607.09084#bib.bib92)\]\. Single\-step methods focus on computational efficiency by following instantaneous gradients, while multi\-step methods aim for higher accuracy by converging to optimal parameter states via multi\-step optimization\.

Distribution Matching\. This approach aims to align the distribution of synthetic data with that of real data\. Distribution matching differs from other methods by directly minimizing the distribution distance between synthetic and real data, employing metrics like Maximum Mean Discrepancy \(MMD\)\[[93](https://arxiv.org/html/2607.09084#bib.bib93)\]\. CAFE\[[94](https://arxiv.org/html/2607.09084#bib.bib94)\]constrains the feature statistics of synthetic and real samples to be consistent across all network layers except the last one\. To better capture distributional differences, NCFD\[[95](https://arxiv.org/html/2607.09084#bib.bib95)\]formulates dataset distillation as a min\-max optimization problem using the Neural Characteristic Function Discrepancy, which is a novel and theoretical metric for distribution comparison\.

#### 3\.1\.3Data\-Efficient Pruning

Data pruning strives to remove redundant samples to retain the most informative subset of dataset, known as the coreset\. Research on data pruning primarily follows two approaches:score\-basedandgeometry\-basedmethods\.

Score\-Based\. This method assigns importance metrics to data points and retain those with the highest scores\. DeepCore\[[27](https://arxiv.org/html/2607.09084#bib.bib27)\]constructs a comprehensive code library and provide an empirical study on popular coreset selection methods\. By calculating the averageL2L\_\{2\}norm of the error vector, E2LN\[[96](https://arxiv.org/html/2607.09084#bib.bib96)\]assesses the importance of training examples, identifying crucial examples very early in training\. The field has further evolved to include a variety of strategies, such as those based on “forgetting events” to measure how often each example is forgotten during training\[[97](https://arxiv.org/html/2607.09084#bib.bib97)\], prioritizing uncertain samples based on prediction variation \(Dyn\-Unc\[[98](https://arxiv.org/html/2607.09084#bib.bib98)\]\), and hybrid strategies like TDDS’s dual\-depth pruning\[[99](https://arxiv.org/html/2607.09084#bib.bib99)\]that combine difficulty and uncertainty\.

Geometry\-Based\. This approach aims to construct a coreset that better represents the underlying data distribution\[[100](https://arxiv.org/html/2607.09084#bib.bib100)\]\. To construct representative coresets, various methods have been proposed: Herding\[[101](https://arxiv.org/html/2607.09084#bib.bib101)\]minimizes distribution discrepancy; SSP\[[102](https://arxiv.org/html/2607.09084#bib.bib102)\]and Moderate\[[103](https://arxiv.org/html/2607.09084#bib.bib103)\]reduce redundancy by selecting distant or median samples, though this may harm generalization by overlooking difficult examples\. Addressing this, D2 pruning\[[104](https://arxiv.org/html/2607.09084#bib.bib104)\]uses a graph\-based message\-passing mechanism to enhance diversity and model generalization\.

In brief, data\-efficient techniques are shifting from static, rule\-based strategies toward adaptive and intelligently scheduled systems\. The trend is toward intelligent data orchestrators\. Powered by meta\- and reinforcement learning, these systems manage data selection as aself\-evolvingdata learning loop, dynamically curating the training stream for peak efficiency in large\-scale model training\.

### 3\.2Distributed Training Techniques

Distributed training leverages multiple workers to maximize computational resources, accelerate training, and improve model accuracy\. It addresses the computational and memory limitations of a single device but introduces three major challenges: \(1\) Single devices cannot meet the computational demands of modern large\-scale training; \(2\) Model parameters exceed the memory limits of a single device; \(3\) Distributed training suffers from high communication costs due to frequent synchronization\. These challenges are addressed through multipleparallelizationandmixed\-precision trainingstrategies in distributed training\.

#### 3\.2\.1Data and Model Parallelism

Data Parallelism\. As a fundamental technique for improving training throughput, data parallelism\[[28](https://arxiv.org/html/2607.09084#bib.bib28)\]replicates the model across all workers, with each processing a distinct subset of the dataset\. To maintain weight consistency, workers periodically synchronize their gradients\. ZeRO\[[29](https://arxiv.org/html/2607.09084#bib.bib29)\]\(Zero Redundancy Optimizer\), introduced by the Deep\-Speed\[[30](https://arxiv.org/html/2607.09084#bib.bib30)\], significantly reduces memory usage by optimizer states partitioning, gradients partitioning, and model parameters partitioning, eliminating redundant storage across devices\. However, its main limitation lies in the substantial communication overhead that arises with extremely large models\.

Model Parallelism\. When a model exceeds the memory capacity of a single device, model parallelism becomes necessary\. Two prominent strategies are Tensor Parallelism \(TP\)\[[31](https://arxiv.org/html/2607.09084#bib.bib31)\]and Pipeline Parallelism \(PP\)\[[32](https://arxiv.org/html/2607.09084#bib.bib32)\]\. TP decomposes weight matrices within specific layers \(*e\.g*\., attention or feed\-forward networks\) across devices, enabling parallel computation\. However, each forward and backward pass necessitates full all\-reduce operations to synchronize partial results, introducing notable communication overhead\. In contrast, PP distributes entire layers across devices arranged in a sequential pipeline, reducing per\-device memory load at the expense of potential pipeline bubbles\. While both approaches mitigate memory constraints by distributing parameters and computation, they incur additional communication costs\.

Hybrid Parallelism\. Hybrid parallelism integrates data, tensor, and pipeline parallelism to efficiently train ultra\-large\-scale models, such as those with hundreds of billions of parameters, as exemplified by 3D parallelism\[[33](https://arxiv.org/html/2607.09084#bib.bib33)\]\.

#### 3\.2\.2Mixed\-Precision Training

During large\-scale model training, GPU memory consumption primarily arises from four components: model parameters, optimizer states, intermediate activations, and temporary buffers\. To address this constraint, mixed precision training has evolved from an optional optimization to a pivotal strategy for large\-scale model development\. It effectively overcomes critical bottlenecks in GPU memory and computational throughput by utilizing lower\-precision formats for most operations \(*e\.g*\., matrix multiplications\) to reduce memory usage and accelerate computation\.

To preserve numerical stability, a master copy of the weights is maintained in a 32\-bit floating point \(FP32\) for gradient accumulation and parameter updates\. Additionally, loss scaling is applied: the loss is multiplied by a scaling factor before backpropagation to prevent small gradients from vanishing in FP16, and these gradients are unscaled before updating the master weights\. In practice, studies such as\[[105](https://arxiv.org/html/2607.09084#bib.bib105)\]have adopted FP16 or Brain Floating Point \(BF16\)\[[34](https://arxiv.org/html/2607.09084#bib.bib34)\]formats, significantly decreasing memory usage while maintaining training performance\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x6.png)Figure 7:Illustration of the model architectures of three different parameter\-efficient fine\-tuning methods: \(a\) Prompt\-based Tuning, \(b\) Adapter Tuning, and \(c\) Low\-Rank Adaptation\.

### 3\.3Parameter Efficient Fine\-Tuning

Parameter\-Efficient Fine\-Tuning \(PEFT\) refers to techniques that adapt large pretrained models to downstream tasks by updating only a small fraction of parameters or by inserting lightweight trainable modules while keeping most of the backbone frozen\. The goals are to reduce training compute and memory, minimize storage and transmission of task\-specific weights, and enable modular multi\-task deployment\. PEFT methods are widely used in language, vision, and multimodal models; they differ in where trainable capacity is placed \(input\-level prompts\[[78](https://arxiv.org/html/2607.09084#bib.bib78),[35](https://arxiv.org/html/2607.09084#bib.bib35),[106](https://arxiv.org/html/2607.09084#bib.bib106)\], modular layer inserts\[[36](https://arxiv.org/html/2607.09084#bib.bib36),[107](https://arxiv.org/html/2607.09084#bib.bib107)\], or low\-rank parameter subspaces\[[37](https://arxiv.org/html/2607.09084#bib.bib37),[108](https://arxiv.org/html/2607.09084#bib.bib108)\]\), in trade\-offs between parameter budget and final accuracy, and in deployment convenience\. Below we summarize three common PEFT families:Prompt\-based Tuning,Adapter Tuning, andLow\-Rank Adaptation, and provide representative references for each PEFT method\.

#### 3\.3\.1Prompt\-based Tuning

Prompt\-based Tuning exploits the fact that pretrained language models can often be steered toward desired downstream behavior by suitable prompts\. Early work with hand\-crafted discrete prompts demonstrated strong few\-shot capabilities in large language models\[[78](https://arxiv.org/html/2607.09084#bib.bib78)\]\. To make prompts trainable and more general, continuous \(soft\) prompt methods were proposed\. Specifically, Prompt Tuning\[[35](https://arxiv.org/html/2607.09084#bib.bib35)\]prepends a small set of learnable embeddings to the model input and trains only those embeddings, as shown in Figure[7](https://arxiv.org/html/2607.09084#S3.F7)\(a\)\. Prefix Tuning\[[106](https://arxiv.org/html/2607.09084#bib.bib106)\]injects learnable embeddings into each transformer layer’s attention mechanism so the prefix can influence internal computations without modifying model weights\. P‑tuning and P‑tuning v2\[[57](https://arxiv.org/html/2607.09084#bib.bib57),[109](https://arxiv.org/html/2607.09084#bib.bib109)\]introduce richer parameterizations and optimization strategies for continuous prompts and show effectiveness across generation and classification tasks\. Prompt\-based methods are attractive due to their minimal parameter footprint and easy per\-task storage, but their effectiveness can be sensitive to model scale, prompt length, and initialization; they may underperform other PEFT methods on smaller backbones or on tasks requiring deeper representational changes\. Recent work addresses multi\-task prompt learning, initialization heuristics, and optimization improvements to boost stability and generalization\.

#### 3\.3\.2Adapter Tuning

As shown in Figure[7](https://arxiv.org/html/2607.09084#S3.F7)\(b\), Adapter Tuning inserts compact, trainable adapters into frozen layers and trains only those modules\. Typical adapters use a bottleneck architecture \(down\-projection → nonlinearity → up\-projection\), adding a few parameters while enabling flexible transformations of layer activations\. The Series Adapter\[[36](https://arxiv.org/html/2607.09084#bib.bib36),[110](https://arxiv.org/html/2607.09084#bib.bib110)\], which places adapter blocks sequentially inside each Transformer block, was an early and widely used design\. To reduce latency and preserve parallelism, Parallel Adapters\[[107](https://arxiv.org/html/2607.09084#bib.bib107)\]place a lightweight side branch that runs in parallel with the original sublayer and fuse its output with the main branch; these designs reduce inference overhead while retaining much of serial adapters’ performance\. Extensions such as AdapterFusion\[[111](https://arxiv.org/html/2607.09084#bib.bib111)\]and related methods\[[112](https://arxiv.org/html/2607.09084#bib.bib112),[113](https://arxiv.org/html/2607.09084#bib.bib113)\]fuse multiple task adapters for transfer\. Current directions include automated adapter architecture search\[[114](https://arxiv.org/html/2607.09084#bib.bib114),[115](https://arxiv.org/html/2607.09084#bib.bib115)\], adapter compression\[[116](https://arxiv.org/html/2607.09084#bib.bib116)\], and adapting adapter ideas to vision and multimodal transformers\[[117](https://arxiv.org/html/2607.09084#bib.bib117),[118](https://arxiv.org/html/2607.09084#bib.bib118)\]\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x7.png)Figure 8:Illustration of techniques commonly employed for model compression\. Mainly include \(a\) quantization, \(b\) pruning, and \(c\) knowledge distillation\.
#### 3\.3\.3Low\-Rank Adaptation

Low\-Rank Adaptation constrains the learned update to a weight matrix to lie in a low\-dimensional subspace, expressing the update as a product of two small matrices\. As shown in Figure[7](https://arxiv.org/html/2607.09084#S3.F7)\(c\), LoRA\[[37](https://arxiv.org/html/2607.09084#bib.bib37)\]is the canonical example\. Instead of updating a full weight matrix𝑾\\bm\{W\}, LoRA freezes𝑾\\bm\{W\}and learns low\-rank matrices𝑨\\bm\{A\}and𝑩\\bm\{B\}so that the effective weight becomes𝑾\+α​𝑨​𝑩\\bm\{W\}\+\\alpha\\bm\{A\}\\bm\{B\}during training \(whereα\\alphais a scaling factor\)\. When applied to a linear projection that maps an input𝑿\\bm\{X\}via𝑾\\bm\{W\}, the LoRA update modifies this mapping to𝑾​𝑿\+α​𝑨​𝑩​𝑿\\bm\{W\}\\bm\{X\}\+\\alpha\\bm\{A\}\\bm\{B\}\\bm\{X\}\. LoRA is typically applied to critical linear projections \(*e\.g*\., query/key/value or feed\-forward projections\) in Transformer blocks\. The approach offers strong parameter–performance trade\-offs for large models, requires little additional GPU memory, and can be merged into base weights for inference, simplifying deployment\. Since its introduction, LoRA has inspired many variants: selective layer application\[[119](https://arxiv.org/html/2607.09084#bib.bib119)\], dynamic/adaptive rank selection\[[120](https://arxiv.org/html/2607.09084#bib.bib120),[121](https://arxiv.org/html/2607.09084#bib.bib121)\], and hybrid schemes combining LoRA with adapters or prompts\[[122](https://arxiv.org/html/2607.09084#bib.bib122),[123](https://arxiv.org/html/2607.09084#bib.bib123)\]\. LoRA’s engineering advantages, including compatibility, easy merging into base weights, and straightforward deployment, have driven broad industrial adoption and spurred numerous benchmarks and open\-source implementations such as Hugging Face tools\[[124](https://arxiv.org/html/2607.09084#bib.bib124)\]and the PEFT library\[[125](https://arxiv.org/html/2607.09084#bib.bib125)\]\. Beyond these widely used libraries, several open\-source ecosystems further simplify LoRA\-style fine\-tuning\. FastChat\[[126](https://arxiv.org/html/2607.09084#bib.bib126)\]provides a lightweight framework for instruction tuning and chat model serving with built\-in LoRA adapter support\. Unsloth\[[127](https://arxiv.org/html/2607.09084#bib.bib127)\]offers optimized kernels and quantized LoRA training, reducing memory and compute demands\. ColossalAI\[[128](https://arxiv.org/html/2607.09084#bib.bib128)\]integrates LoRA with system\-level optimizations such as parallelism and memory partitioning, enabling efficient fine\-tuning of large models on limited hardware\. These toolkits broaden practical PEFT adoption by lowering resource requirements and easing deployment\.

### 3\.4Model and Computational Compression

Model and computational compression aim to reduce the computational size by eliminating redundant information, thus enhancing both storage efficiency and computational performance during inference\. As illustrated in Figure[8](https://arxiv.org/html/2607.09084#S3.F8)and Figure[9](https://arxiv.org/html/2607.09084#S3.F9), key techniques encompasslow\-precision quantization,model pruning,knowledge distillation, andspeculative decoding\. The following subsections will delve deeper into these techniques and their role in enhancing model performance during inference\.

#### 3\.4\.1Low\-Precision Quantization

Low\-precision quantization has become a crucial technique for improving the training and inference efficiency of large\-scale deep learning models\[[129](https://arxiv.org/html/2607.09084#bib.bib129)\]\. Originally, 32\-bit floating point \(FP32\) was the standard for model training, but the growing size of models and datasets has made its computational cost prohibitive\[[130](https://arxiv.org/html/2607.09084#bib.bib130)\]\. To mitigate this issue, lower\-precision formats such as FP16, BF16, and more recently FP8 and INT8 have been widely adopted\. These formats offer benefits including reduced memory consumption and faster computation\[[131](https://arxiv.org/html/2607.09084#bib.bib131)\]\. Nevertheless, numerical overflow and precision degradation remain critical challenges, especially for long\-sequence training and large\-scale inference tasks\[[132](https://arxiv.org/html/2607.09084#bib.bib132)\]\.

As quantization techniques continue to advance, several methods have been proposed to improve stability and performance\. LLM\.int8\(\)\[[133](https://arxiv.org/html/2607.09084#bib.bib133)\]demonstrates that INT8 quantization can significantly reduce memory usage while preserving near\-FP16 accuracy, and LLM\-QAT\[[38](https://arxiv.org/html/2607.09084#bib.bib38)\]incorporates quantization into training to enhance robustness to low precision\. To further address instability, dynamic precision strategies such as Fallback Quantization\[[134](https://arxiv.org/html/2607.09084#bib.bib134)\]temporarily elevate precision to prevent overflow, while ShiftQuant\[[135](https://arxiv.org/html/2607.09084#bib.bib135)\]enables sub\-8\-bit training via gradient estimation and L1 normalization\. More recently, the pursuit of extreme efficiency has spurred interest in INT4 and even 1\-bit quantization\. Frameworks such as OneBit\[[136](https://arxiv.org/html/2607.09084#bib.bib136)\]and BitNet\[[137](https://arxiv.org/html/2607.09084#bib.bib137)\]illustrate the feasibility of binarizing weights with minimal accuracy loss, and techniques like ParetoQ\[[138](https://arxiv.org/html/2607.09084#bib.bib138)\]and ABQ\-LLM\[[139](https://arxiv.org/html/2607.09084#bib.bib139)\]explore arbitrary\-bit configurations to flexibly balance precision and efficiency\.

#### 3\.4\.2Model Pruning

Model pruning serves as a pivotal optimization strategy in large language models, enabling significant enhancements in inference speed by selectively removing less critical parameters while preserving core functionality\[[140](https://arxiv.org/html/2607.09084#bib.bib140)\]\. This approach reduces the computational footprint during deployment, allowing for faster processing on resource\-constrained devices without necessitating extensive retraining\[[141](https://arxiv.org/html/2607.09084#bib.bib141)\]\.

Recent surveys on efficient LLMs highlight pruning innovations addressing inference bottlenecks like high FLOPs and memory demands\[[142](https://arxiv.org/html/2607.09084#bib.bib142),[143](https://arxiv.org/html/2607.09084#bib.bib143)\]\. One\-shot methods like SparseGPT\[[39](https://arxiv.org/html/2607.09084#bib.bib39)\]induce up to 50% sparsity in models such as OPT\-175B, reducing inference time via parameter elimination without iterative fine\-tuning, aiding latency\-sensitive deployments\. Activation\-aware techniques like Wanda\[[144](https://arxiv.org/html/2607.09084#bib.bib144)\]prune based on weight\-activation interactions, yielding1\.24×1\.24\\timesspeedups by cutting redundant computations for dynamic querying\. Structured approaches such as SliceGPT\[[145](https://arxiv.org/html/2607.09084#bib.bib145)\]use PCA to remove low\-importance components, achieving1\.87×1\.87\\timesacceleration and 30% compression in Transformers, boosting scalability for mobile and distributed systems\. Semi\-structured patterns in E\-Sparse\[[146](https://arxiv.org/html/2607.09084#bib.bib146)\]align with hardware like Tensor Cores for1\.53×1\.53\\timesthroughput gains, linking sparsity to practical utilization\. These advances mitigate LLM inefficiencies, enabling broader use in resource\-constrained environments\.

#### 3\.4\.3Knowledge Distillation

Knowledge distillation emerges as a sophisticated refinement technique for large language models, facilitating accelerated inference by transferring distilled insights from a robust teacher model to a more compact student counterpart, thereby slashing computational demands during runtime\[[147](https://arxiv.org/html/2607.09084#bib.bib147),[148](https://arxiv.org/html/2607.09084#bib.bib148)\]\. This method optimizes deployment efficiency, enabling quicker response times on hardware with limited capabilities without the need for exhaustive reconfiguration\[[149](https://arxiv.org/html/2607.09084#bib.bib149)\]\.

Contemporary research highlights black\-box and white\-box strategies for boosting LLM inference velocity\. For example, Orca 2\[[150](https://arxiv.org/html/2607.09084#bib.bib150)\]uses synthetic data from teacher models to enhance reasoning in smaller students, achieving2–3×2–3\\timesgains over Vicuna on complex tasks with minimal data and 40% acceleration in reasoning pipelines\. TinyLLM\[[40](https://arxiv.org/html/2607.09084#bib.bib40)\]aggregates multi\-teacher knowledge into a small student, delivering up to 30% inference speedup on edge devices while preserving 90% accuracy across benchmarks\. DA\-KD\[[151](https://arxiv.org/html/2607.09084#bib.bib151)\]applies difficulty\-aware sampling to focus on challenging samples, yielding1\.8×1\.8\\timesthroughput improvements for efficient LLMs in real\-time applications like chatbots\. Symbolic Chain\-of\-Thought Distillation refines symbolic reasoning paths, allowing small models to emulate step\-by\-step thinking with 25% reduced latency in logical scenarios\[[152](https://arxiv.org/html/2607.09084#bib.bib152)\]\.

#### 3\.4\.4Speculative Decoding

Speculative decoding functions as an innovative speedup tactic for large language models, hastening inference by employing a lightweight drafter to propose token sequences that a primary verifier then assesses in parallel, thereby curtailing sequential computations and elevating generation rates, as shown in Figure[9](https://arxiv.org/html/2607.09084#S3.F9)\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x8.png)Figure 9:Illustration of the speculative decoding mechanism\[[153](https://arxiv.org/html/2607.09084#bib.bib153)\]\. The speculative decoding methods draft some tokens and verify them in parallel in one decoding step\.Contemporary explorations in speculative decoding for LLMs focus on draft\-verification synergies to enhance throughput in resource\-intensive inference\. For example, Confidence\-Modulated Speculative Decoding\[[154](https://arxiv.org/html/2607.09084#bib.bib154)\]applies adaptive confidence thresholds for dynamic drafting lengths, yielding up to2\.5×2\.5\\timesfaster decoding in variable\-context LLMs with noise robustness\. Recurrent Drafter\[[155](https://arxiv.org/html/2607.09084#bib.bib155)\]uses RNN\-based drafters for sequential efficiency, achieving3×3\\timesspeedups in long\-form generation by cutting verification overheads in Llama variants\. Efficient Multi\-sample Speculative Decoding\[[156](https://arxiv.org/html/2607.09084#bib.bib156)\]leverages parallel sampling to raise acceptance rates, providing2\.2×2\.2\\timesinference gains and reduced latency in NAACL\-benchmarked LLMs\. SpecEE\[[41](https://arxiv.org/html/2607.09084#bib.bib41)\]merges speculative execution with early exiting, cutting computational demands by 4% and boosting on\-device throughput for edge\-deployed LLMs\.

## 4Green Computing AI Chips and Hardware\-Software Co\-Design

As Large Language Models \(LLMs\) continue to scale in complexity and deployment, the demand for high\-performance and efficient hardware has intensified\. Traditional acceleration platforms, while effective, have exposed growing concerns over energy consumption, carbon emissions, and deployment scalability\. In response, both academia and industry have invested heavily in the development of green computing solutions—from architectural innovations in AI chips to sustainable data center infrastructure\[[157](https://arxiv.org/html/2607.09084#bib.bib157),[158](https://arxiv.org/html/2607.09084#bib.bib158)\]\. This chapter explores the evolving landscape of AI hardware, emphasizing mainstream chip ecosystems and emerging paradigms in energy\-efficient hardware\-software co\-design\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x9.png)Figure 10:Application Categories of AI Chips\.### 4\.1Mainstream AI Chips and Application Ecology

#### 4\.1\.1Classification and Comparison of AI Chips

As shown in Figure[10](https://arxiv.org/html/2607.09084#S4.F10), AI chips can be broadly categorized based on their deployment environments \(*i\.e*\., cloud or edge\), and their computational roles in training or inference\. Central Processing Units \(CPUs\), while flexible, are generally inefficient for large\-scale matrix operations\. Graphics Processing Units \(GPUs\), such as NVIDIA’s A100 and H100, have become the de facto standard for training LLMs due to their massive parallelism and rich software ecosystem \(*e\.g*\., CUDA\)\. Tensor Processing Units \(TPUs\)\[[43](https://arxiv.org/html/2607.09084#bib.bib43)\], developed by Google, are custom ASICs optimized for matrix multiplication and serve as a backbone for models like GPT and PaLM\. Field\-Programmable Gate Arrays \(FPGAs\) offer reconfigurability and are well\-suited for edge inference but require significant design effort\[[159](https://arxiv.org/html/2607.09084#bib.bib159)\]\. Digital Signal Processors \(DSPs\) and Application\-Specific Integrated Circuits \(ASICs\) provide optimized performance for specific workloads, especially in low\-power scenarios\[[42](https://arxiv.org/html/2607.09084#bib.bib42)\]\.

Table 3:Global Representative AI Chip Enterprises and Its Product\.CompanyTypical ChipsYearArchFunctionInternational AI chip enterprisesNVIDIAH1002022GPUcloud trainingAMDMI3002023GPUcloud trainingIntelGaudi 32024NPUcloud trainingGoogleTPUv52023ASICcloud trainingQualcommCloud AI1002020ASICcloud inferenceAppleM32023ARMedge inferenceSamsungExynos 21002021ARMedge inferenceIBMTrueNorth2015–edge inferenceDomestic AI chip enterprisesCambriconMLU 5902024ASICcloud trainingHUAWEIAscend 9102019NPUcloud trainingHorizonJourney 62023ARMedge inferenceT\-headHanguang 8002019FPGAcloud inferenceBaiduKunlun 22022FPGAcloud inference

#### 4\.1\.2Global AI Chip Ecosystems

As shown in[Tab\.3](https://arxiv.org/html/2607.09084#S4.T3), leading international companies such as NVIDIA, Google, AMD, and Intel dominate the AI hardware market\. Firstly, NVIDIA’s GPU platform combines rapidly evolving silicon architectures, such as Ampere, Hopper, and Blackwell, with a vertically integrated software stack, including CUDA, CUTLASS, cuBLAS/cuDNN, NCCL, and Triton Inference Server, and system\-level interconnects like NVLink and NVSwitch, enabling scalable training across tens of thousands of accelerators\. Secondly, Google’s TPU series\[[43](https://arxiv.org/html/2607.09084#bib.bib43)\], from TPU v2/v3 to v5, demonstrates the datacenter efficiency of AI\-specific ASICs by leveraging systolic matrix engines, wafer\-scale torus interconnects, and optical circuit switches for dynamic topology reconfiguration\. These innovations achieve high model FLOPs utilization in large\-scale training while minimizing energy consumption per operation\. Google integrates TPUs into its cloud infrastructure, supporting the training of cutting\-edge models such as Gemini\[[160](https://arxiv.org/html/2607.09084#bib.bib160)\]\. Meanwhile, AMD’s MI300 series and Intel’s Gaudi 3 accelerators are broadening the competitive landscape by emphasizing mixed\-precision computing and energy\-aware scheduling\. These platforms are supported by comprehensive toolchains and SDKs that enable developers to optimize LLM workloads across both training and inference stages\[[161](https://arxiv.org/html/2607.09084#bib.bib161)\]\.

#### 4\.1\.3Domestic AI Chip Development and Ecosystems

As shown in[Tab\.3](https://arxiv.org/html/2607.09084#S4.T3), China’s AI chip ecosystem has diversified across cloud and edge scenarios\[[161](https://arxiv.org/html/2607.09084#bib.bib161)\]\. Cloud\- and edge\-oriented NPUs from Cambricon\[[44](https://arxiv.org/html/2607.09084#bib.bib44)\]\(*e\.g*\., MLU series\) target training and inference with compiler toolchains and operator libraries integrated into domestic AI frameworks\. Horizon Robotics focuses on edge autonomy with high\-efficiency SoCs \(Journey/Sunrise\), combining perception and planning accelerators for automotive ADAS/AD\. Huawei’s Ascend \(*e\.g*\., 910/310\) provides a vertically integrated stack for cloud and edge with the CANN/AscendCL toolchain, MindSpore framework, and model zoo integration; the series have been deployed for general AI services and domain\-specific inference\. In special\-purpose domains, Bitmain’s experience in ultra–high\-efficiency matrix engines informs inference ASIC design practices, while T\-Head \(Alibaba\) advances RISC‑V based CPUs \(XuanTie\) and cloud inference chips \(Hanguang\) within a broader open\-source and cloud ecosystem\. These vendors are increasingly developing end\-to\-end stacks, including compilers, graph optimization, quantization, and runtimes, as software maturity has become a decisive factor in achieving effective energy efficiency at the application level

#### 4\.1\.4Energy and Performance Perspectives

The energy efficiency of AI chips is an increasingly critical metric, often measured in TOPS/W \(Tera\-operations per Second per Watt\)\. For instance, while NVIDIA’s H100 achieves state\-of\-the\-art throughput, its power envelope exceeds 700W, posing challenges for widespread deployment\. In contrast, edge ASICs like those from Bitmain or Apple’s Neural Engine offer much higher energy efficiency for inference\. The software stack plays a pivotal role in green computing\. Frameworks like TensorRT, TVM, and MindSpore enable operator\-level optimizations, while sparsity\-aware libraries reduce redundant computation\. Matching model architecture \(*e\.g*\., Transformer variants\) to chip capabilities is essential for maximizing energy\-performance trade\-offs\.

### 4\.2Inference memory optimization

Inference memory optimization serves as a critical advancement in large language models, achieved through software\-hardware co\-design that accelerates inference by dynamically reallocating storage resources via intelligent caching mechanisms and offloading inactive data to secondary storage layers, thereby reducing peak memory demands without sacrificing throughput\. Key strategies encompass KV cache optimization, memory offloading, and early exiting\.

![Refer to caption](https://arxiv.org/html/2607.09084v1/x10.png)Figure 11:Illustration of the key\-value caching mechanism\[[162](https://arxiv.org/html/2607.09084#bib.bib162)\]\.#### 4\.2\.1KV Cache Optimization

KV cache optimization emerges as a pivotal acceleration technique in large language models, realized through software\-hardware co\-design that boosts inference efficiency by compressing key\-value stores using selective retention and precision reduction methods, thereby mitigating memory bottlenecks in long sequences while preserving accuracy\[[163](https://arxiv.org/html/2607.09084#bib.bib163),[164](https://arxiv.org/html/2607.09084#bib.bib164)\]\. Fig\.[11](https://arxiv.org/html/2607.09084#S4.F11)demonstrates the key\-value cache usage in both phases\. During the initialization phase, the LLM generates the key\-value cache for each token in the input prompt\. In the subsequent decoding phase, the LLM only needs to compute the query, key, and value of one newly generated token, leveraging the precomputed key\-value cache to facilitate the process step by step\.

Research on LLM\-accelerated KV cache optimization categorizes strategies into tag\-level eviction and model\-level quantization, using hardware\-aware designs to reduce memory overhead in long\-context scenarios\. For instance, CacheGen\[[165](https://arxiv.org/html/2607.09084#bib.bib165)\]applies custom tensor encoding via distributional properties, enabling 2\-4×\\timesfaster context loading and streaming for Llama\-2\-70B in high\-throughput serving\. LaCache\[[45](https://arxiv.org/html/2607.09084#bib.bib45)\]uses ladder\-shaped caching with training\-free hierarchical retention, boosting inference speed by 1\.5\-2×\\timesand reducing cache size by 40% for Mistral\-7B in long\-sequence generation\. FastGen\[[166](https://arxiv.org/html/2607.09084#bib.bib166)\]employs LLM profiling for adaptive compression across attention heads, halving memory usage without quality loss and accelerating decoding by 30% in dynamic workloads\. Gaoet al\.propose hybrid sparsity\-quantization frameworks, achieving up to3×3\\timesthroughput gains by minimizing recomputation in extended contexts\[[167](https://arxiv.org/html/2607.09084#bib.bib167)\]\.

#### 4\.2\.2Memory Offloading

Memory offloading emerges as a pivotal technique for accelerating LLM inference on resource\-constrained hardware, realized through software\-hardware co\-design that enables the execution of massive models by dynamically swapping parameters between high\-speed GPU memory and slower but more capacious external storage such as CPU DRAM or NVMe SSDs\[[168](https://arxiv.org/html/2607.09084#bib.bib168)\]\.

Recent advancements refine offloading strategies to balance bandwidth and compute demands\. For example, Aqua\[[46](https://arxiv.org/html/2607.09084#bib.bib46)\]employs network acceleration in multi\-GPU clusters to cut paging overheads in inference state transfers, achieving up to2×2\\timesspeedups via RDMA weight prefetching\. HeadInfer\[[47](https://arxiv.org/html/2607.09084#bib.bib47)\]uses head\-wise KV cache offloading, partitioning attention heads across devices to reduce full\-layer storage and increase memory efficiency by 40% without accuracy loss\. For MoE architectures, a latency\-hiding scheme\[[169](https://arxiv.org/html/2607.09084#bib.bib169)\]overlaps expert activations with data transfers, minimizing idle times for seamless scaling on heterogeneous setups\. SpecOffload\[[170](https://arxiv.org/html/2607.09084#bib.bib170)\]applies speculative partial offloading to tap underused GPU capacity, predicting and caching low\-activation parameters on CPU for1\.51\.5\-3×3\\timesfaster iterative decoding\. Complementary methods, such as heterogeneous speculative decoding, multilevel pretraining offloads, and dynamic token pruning, illustrate the growing inference optimization ecosystem\.

#### 4\.2\.3Early Exiting for Inference Acceleration

Early exiting has become a cornerstone of conditional computation in LLM inference acceleration, dynamically halting autoregressive decoding at intermediate layers when token predictions meet predefined confidence criteria, thus curbing computational overhead and enabling real\-time applications on diverse hardware without compromising semantic integrity\[[171](https://arxiv.org/html/2607.09084#bib.bib171)\]\.

Over the past few years, innovations like AdaInfer\[[48](https://arxiv.org/html/2607.09084#bib.bib48)\]and FREE\[[172](https://arxiv.org/html/2607.09084#bib.bib172)\]have paved the way for more adaptive mechanisms, while broader efforts in speculative integration enhance throughput\. SpecEE\[[41](https://arxiv.org/html/2607.09084#bib.bib41)\]deploys a speculation\-driven engine that anticipates exit points via auxiliary predictors, verifying drafts in parallel to yield up to2\.5×2\.5\\timesfaster generation on GPUs by minimizing redundant layer traversals\. Finally, SPADE introduces a hybrid algorithm fusing confidence monitoring with space\-aligned decoding, projecting intermediate states to align with deeper representations and enabling seamless exits that preserve coherence, boosting efficiency by2×2\\timesin edge deployments\[[173](https://arxiv.org/html/2607.09084#bib.bib173)\]\.

### 4\.3Cross\-platform deployment and adaptation

Cross\-platform deployment and adaptation for large\-scale models facilitates efficient execution across heterogeneous hardware like CPUs, GPUs, and edge devices by addressing compatibility and optimization challenges through three core strategies: unified inference frameworks, automatic operator scheduling, and key operator discretization\. The following subsections explore these techniques and their impacts on model performance\.

#### 4\.3\.1Unified Inference Framework

Unified inference frameworks streamline LLM cross\-platform deployment by abstracting hardware details into a cohesive API, enabling seamless transitions across CPUs, GPUs, and edge devices while maintaining performance\.

Key approaches include LLMBox\[[49](https://arxiv.org/html/2607.09084#bib.bib49)\], which unifies training, inference, and evaluation pipelines with modular extensions for backends like PyTorch and TensorRT to reduce deployment overheads\. HERMES\[[174](https://arxiv.org/html/2607.09084#bib.bib174)\]supports multi\-stage pipelines with heterogeneous clients, dynamically batching requests across concurrent models to optimize distributed resource use\. ScaleLLM\[[175](https://arxiv.org/html/2607.09084#bib.bib175)\]emphasizes end\-to\-end efficiency via operator fusion and paged attention for scalable serving beyond standard inference\. Together, these reduce porting efforts by up to 50% in multi\-tenant settings\. Edge latency\-aware LLM customization\[[176](https://arxiv.org/html/2607.09084#bib.bib176)\]uses distillation in unified backends for privacy\-preserving inference on constrained hardware\. Overall, these frameworks boost adaptability, yielding22\-3×3\\timesthroughput improvements across ARM and x86 ecosystems\.

#### 4\.3\.2Automatic Operator Scheduling

Automatic operator scheduling optimizes LLM inference by dynamically allocating computational kernels to heterogeneous hardware, mitigating bottlenecks in cross\-platform execution through runtime profiling and predictive mapping\.

The Past\-Future Scheduler in LightLLM\[[50](https://arxiv.org/html/2607.09084#bib.bib50)\]balances historical and speculative workloads under SLA constraints, prioritizing token sequences to improve goodput by1\.5×1\.5\\timesin shared clusters\. Decentralized serving task scheduling\[[177](https://arxiv.org/html/2607.09084#bib.bib177)\]applies heuristic algorithms to distribute inference across edge nodes, enabling fault\-tolerant adaptation for large\-scale deployments\. DynamoLLM\[[178](https://arxiv.org/html/2607.09084#bib.bib178)\]uses reconfiguration loops to optimize cluster topologies for energy efficiency, aligning operator dispatch with SLOs through reinforcement learning policies\. These approaches handle multi\-tenancy by integrating scheduling with KV cache management\.

#### 4\.3\.3Key Operator Discretization

Key operator discretization enhances cross\-platform LLM adaptation by decomposing complex kernels into modular, hardware\-agnostic units, allowing fine\-grained quantization and fusion for efficient porting across accelerators\.

ClusterFusion\[[179](https://arxiv.org/html/2607.09084#bib.bib179)\]scales fusion primitives to cluster\-level operators, discretizing attention and linear layers for distributed inference with2×2\\timesmemory savings\. FlashDecoding\+\+\[[180](https://arxiv.org/html/2607.09084#bib.bib180)\]uses asynchronous kernel fusion in LLM decoding, discretizing GEMM operations for flat optimization and heuristic prefetching\. Qtile\[[51](https://arxiv.org/html/2607.09084#bib.bib51)\]accelerates quantized serving through tile\-based operator discretization, integrating low\-bit formats with fusion to minimize overhead in multi\-precision setups and mitigate variances on ARM and NVIDIA hardware\. KPerfIR\[[52](https://arxiv.org/html/2607.09084#bib.bib52)\]advances a compiler\-centric ecosystem for GPU kernel fusion, discretizing operators via 4D parallelism to balance workloads in LLM pipelines\. These methods emphasize hardware\-agnostic optimizations, enabling broader LLM accessibility in diverse computational environments\.

### 4\.4Energy\-Efficient Hardware System

Beyond AI chips, some hardware directions are seeking step\-function improvements in performance per watt by reducing energy consumption associated with data movement and cooling, and by tailoring computation to model structure\.

#### 4\.4\.1Emerging Green Chip Technologies

To overcome the limitations of the Von Neumann architecture, where data shuttling between memory and compute units incurs substantial energy overhead, In\-Memory Computing \(IMC\) has emerged as a promising alternative\[[181](https://arxiv.org/html/2607.09084#bib.bib181)\]\. IMC performs computation within the memory array itself, significantly reducing data movement\. Architectures like TranCIM, RIME, and SmartInfinity showcase how compute\-in\-memory enables energy\-efficient matrix operations, particularly for low\-bit quantized models\. Additionally, neuromorphic chips\[[53](https://arxiv.org/html/2607.09084#bib.bib53)\]such as IBM’s TrueNorth and Intel’s Loihi simulate brain\-like spiking neuron behavior\[[182](https://arxiv.org/html/2607.09084#bib.bib182)\], enabling ultra\-low\-power inference\. Optical computing and quantum accelerators also show promise, though they remain in experimental stages\[[183](https://arxiv.org/html/2607.09084#bib.bib183)\]\. These paradigm\-shifting technologies aim to support future LLMs with orders\-of\-magnitude improvements in energy efficiency\.

#### 4\.4\.2Decentralized Green Learning

Decentralized learning paradigms further reduce the environmental cost of AI by curbing data movement and right‑sizing compute\. Federated Learning \(FL\) offers a distributed training paradigm where models are trained locally on edge devices and only parameter updates are exchanged\. This approach not only enhances data privacy but also minimizes the need for large\-scale data transmission to central servers, reducing overall energy consumption\[[184](https://arxiv.org/html/2607.09084#bib.bib184)\]\. Recent studies show that federated fine\-tuning of LMs can lead to competitive performance with significantly lower compute and communication costs, especially when combined with techniques like LoRA and quantization\-aware training\. Besides, collaborative inference systems like PETALS\[[54](https://arxiv.org/html/2607.09084#bib.bib54)\]propose decentralized model hosting, where users share compute workloads, enhancing both accessibility and sustainability\.

#### 4\.4\.3Green Data Centers

Data centers hosting LLM workloads must address their rapidly growing carbon footprints\[[55](https://arxiv.org/html/2607.09084#bib.bib55)\]\. One strategy is geographical optimization, whereby data centers are located near renewable energy sources—such as hydropower in Sichuan, geothermal in Iceland, or solar farms in the American Southwest\[[6](https://arxiv.org/html/2607.09084#bib.bib6)\]\. China’s “East Data, West Computing” initiative exemplifies this approach, relocating compute\-intensive LLM training to western regions with abundant green energy\[[56](https://arxiv.org/html/2607.09084#bib.bib56)\]\. Moreover, natural cooling strategies, including the use of ambient air or seawater for thermal management, are replacing traditional air\-conditioning systems\. Innovations in liquid immersion cooling and heat reuse \(*e\.g*\., district heating\) further contribute to reducing the energy overhead of data center operations\.

## 5AI for Sustainability Applications

As large models continue to evolve in scale and complexity, their potential to contribute to global sustainability efforts is becoming increasingly evident\. Beyond concerns over their own carbon footprint, large models and AI systems are now being actively deployed to empower sustainable practices across multiple domains\. This chapter explores recent advancements in AI applications for sustainability, focusing on efficient model architectures, high\-throughput remote sensing, national\-scale computing infrastructures, and domain\-specific green innovations\.

### 5\.1Parallel Training and RL\-Driven Reasoning Paradigm

DeepSeek\[[185](https://arxiv.org/html/2607.09084#bib.bib185)\]exemplifies a dual focus on efficient large\-model training and inference\. In training, it leverages multi\-level parallelism, including tensor, pipeline, and data parallelism, coupled with latency\-hiding kernels, operator fusion, and topology\-aware communication\. These techniques maximize GPU utilization and minimize memory and communication overhead, aligning with best practices for distributed LLM training\[[8](https://arxiv.org/html/2607.09084#bib.bib8)\]\. On the inference side, DeepSeek adopts reinforcement learning at scale to improve reasoning efficiency\. By learning structured decision\-making policies and task\-specific tool usage, the model reduces redundant computation and improves output quality per token and per joule\. This approach reflects a broader trend: sustainable AI increasingly enhances energy efficiency through co\-optimized training pipelines and inference\-time policy learning, rather than relying solely on raw compute scaling\[[186](https://arxiv.org/html/2607.09084#bib.bib186)\]\.

### 5\.2Remote Sensing Interpretation: RS‑vHeat and “Aerospace·Lingmou” 3\.0

Remote sensing workloads require high\-throughput, energy\-efficient inference across multi\-sensor platforms under stringent latency and power constraints\. RS‑vHeat\[[62](https://arxiv.org/html/2607.09084#bib.bib62)\]addresses these challenges through a physics\-inspired architecture that models semantic propagation as heat diffusion\[[11](https://arxiv.org/html/2607.09084#bib.bib11)\]\. Developed around the “Aerospace·Lingmou” 3\.0 kernel, it replaces global attention with localized heat conduction operators, significantly reducing memory traffic while maintaining global receptive fields\. This physical inductive bias enhances locality, facilitates aggressive kernel fusion and cache reuse, and achieves substantial efficiency gains: 84% memory reduction, 24% lower FLOPs, and 2\.7×\\timeshigher throughput compared to attention\-based models, while maintaining state\-of\-the\-art performance across optical, SAR, thermal, and hyperspectral modalities\. The compact design enables efficient edge deployment and federated learning, reducing cloud dependency and emissions\.

### 5\.3China Computing NET \(C2NET\): A National\-Scale Sustainable Infrastructure

C2NET envisions a unified, sovereign “compute network” that interconnects heterogeneous compute \(GPU, NPU, CPU, FPGA\) via high\-speed links and programmable scheduling, providing on‑demand, policy‑aware capacity to national strategic workloads\. As the digital economy’s substrate, C2NET targets holistic sustainability in three directions\. First, siting and energy: it supports “East Data, West Computing”, placing resource‑hungry clusters near renewable resources and abundant land while serving coastal demand over high‑performance backbones, cutting carbon intensity and land use\[[56](https://arxiv.org/html/2607.09084#bib.bib56)\]\. Second, orchestration: traffic‑ and topology‑aware schedulers allocate tasks to the cleanest, closest, and least‑congested sites, maximizing MFU and minimizing network energy per job\[[8](https://arxiv.org/html/2607.09084#bib.bib8)\]\. Third, openness and resilience: a standardized interface for resource discovery, carbon\-aware scheduling, and privacy‑preserving execution \(*e\.g*\., federated analytics\) lowers barriers to sharing clean capacity nationwide\. As a result, C2NET not only decarbonizes supply \(clean power\) but also demand, turning compute into a managed utility with sustainability service‑level objectives\.

### 5\.4Google and Meta: Full\-Stack Optimization for Sustainable AI

Leading technology companies such as Google\[[187](https://arxiv.org/html/2607.09084#bib.bib187)\]and Meta\[[188](https://arxiv.org/html/2607.09084#bib.bib188)\]have pioneered full\-stack optimizations to advance sustainable AI development\. Google’s Gemini system\[[160](https://arxiv.org/html/2607.09084#bib.bib160)\]integrates accelerator\-level telemetry, dynamic model routing, and carbon\-aware data center scheduling, achieving a 44×\\timesreduction in per\-prompt emissions within one year\. Meta, formerly Facebook AI, implements end\-to\-end carbon accounting across all stages of the model lifecycle, from training to inference, while leveraging flash attention, quantization, and sparsity\-based techniques to reduce computational demands\. Beyond infrastructure optimization, both companies deploy large\-scale models for sustainability\-critical applications, including climate forecasting, water consumption prediction, and disaster response\. These initiatives illustrate how improvements in system\-level efficiency, combined with AI\-for\-good applications, can collectively enable scalable and environmentally responsible AI development\.

### 5\.5Global AI Applications for Multi\-Domain Sustainability

Artificial Intelligence is increasingly recognized as a global enabler of sustainability across multiple domains\. Inenergy systems, AI techniques are used for smart grid optimization, dynamic load balancing, and predictive maintenance of renewable infrastructure worldwide\[[189](https://arxiv.org/html/2607.09084#bib.bib189)\]\. Inmaterials science, machine learning accelerates the discovery of photovoltaic and battery materials through high\-throughput simulations and generative models, supported by initiatives from the EU and Department of Energy\.

Inclimate science, AI enhances the resolution of global climate models via downscaling, improving local policy relevance and disaster preparedness\[[190](https://arxiv.org/html/2607.09084#bib.bib190)\]\. Organizations like NASA and the UN leverage AI\-powered early warning systems for floods, wildfires, and hurricanes\. Meanwhile,AI\-for\-coderesearch, including AlphaCode\[[191](https://arxiv.org/html/2607.09084#bib.bib191)\], shows that LLMs can generate energy\-efficient software that reduces runtime energy by 10–20%, promoting greener engineering workflows\[[192](https://arxiv.org/html/2607.09084#bib.bib192)\]\. These global efforts illustrate how AI not only reduces its own footprint but also amplifies sustainability across science, infrastructure, and digital ecosystems\.

## 6Discussion and Future Outlook

While significant progress has been made in the design of efficient model architectures, the optimization of deployment pipelines, and the development of green AI hardware, many open challenges still remain unresolved\. This section provides a reflective discussion on the key limitations and outlines forward\-looking perspectives that may define the next phase of green LLM development\.

### 6\.1Learning without Starting Over: Continual, Incremental, and Federated Training

A significant proportion of the carbon footprint associated with LLMs stems from the repeated re\-training of monolithic architectures\. A more sustainable learning paradigm prioritizes “learning without starting over\.” Continual and incremental learning\[[193](https://arxiv.org/html/2607.09084#bib.bib193),[194](https://arxiv.org/html/2607.09084#bib.bib194),[195](https://arxiv.org/html/2607.09084#bib.bib195),[196](https://arxiv.org/html/2607.09084#bib.bib196)\], under conditions of model homogeneity, heterogeneity, and modality expansion, should be established as core training paradigms rather than secondary considerations\. Promising approaches include: \(i\) elastic parameter partitioning through sparse or adaptive subnetworks that localize model updates and mitigate catastrophic forgetting; \(ii\) retrieval\-augmented pretraining and fine\-tuning, which externalize knowledge updates into lightweight, mutable indices instead of relying on full\-scale weight reconfiguration; and \(iii\) federated and split learning frameworks that enable the exchange of sparsified, differentially private model updates, with provable convergence even under straggler effects and non\-IID data distributions\. Compiler and runtime systems can further support sustainability by identifying and freezing “stable” layers, scheduling updates only for “plastic” regions\. Energy\-aware Reinforcement Learning from Human Feedback \(RLHF\) and optimization techniques can penalize computationally intensive update pathways\. Collectively, these strategies reduce lifecycle emissions by replacing periodic, resource\-intensive model overhauls with targeted, verifiable, and minimal parameter updates\.

### 6\.2Beyond GPUs: Co\-Designing Models with Memory\- and Physics\-Centric Compute

Achieving step\-function gains in energy efficiency requires tighter model–hardware fusion and new substrates\. Memory\-centric compute \(*e\.g*\., computing near\-/in\-memory, CNM/CIM\) can collapse the von Neumann bottleneck in bandwidth\-limited layers \(attention/MLP\) if models adopt sparsity/low‑rank formats and compilers provide calibration and error‑aware schedules\. Neuromorphic/event\-driven units are effective for sparse perception and on\-device intelligent agents, while photonic interconnects promise ultra‑low‑energy data movement and, in the longer term, enable analog matrix computation\. Quantum accelerators remain in early exploration\. Critically, models must be co\-designed with these emerging hardware platforms in mind: state\-space models \(*e\.g*\., Mamba\) that scale linearly with sequence length, and modular routing mechanisms that minimize data movement,*etc*\. A unified Intermediate Representation \(IR\) that explicitly captures data movement metrics, such as bytes per operation and data residency levels, combined with autotuners capable of jointly optimizing precision, sparsity, and memory placement, can pave the way for scalable and energy\-efficient heterogeneous computing\.

### 6\.3Edge AI: Efficiency on Constrained Devices

Deploying large models on edge devices remains challenging due to hardware limitations, computational constraints, high energy consumption, and heterogeneous environments\. Future advancements will rely on new approaches beyond existing popular techniques like model compression, pruning, and quantization\. For example, neuromorphic computing, with brain\-inspired architectures, will deliver significantly higher energy efficiency and real\-time processing for cognitive tasks through parallel and event\-driven operation\. Furthermore, quantum neural networks and mobile quantum processing units could revolutionize edge AI by leveraging quantum properties to process massive datasets and optimize decisions at unprecedented speeds, potentially enabling hybrid quantum\-classical AI for real\-time applications\. The synergy with 6G networks will be critical, providing the ultra\-low latency and high\-speed connectivity required for synchronizing AI across large\-scale Internet of Things systems \(IoT\)\. Finally, the rise of edge\-native AI models optimized for on\-device use will reduce cloud dependency and enable more advanced, self\-sufficient edge applications\.

### 6\.4Standardized and Unified Evaluation

Equally critical is the lack of standardized, comprehensive benchmarks for evaluating the resource efficiency of large models\. Many existing metrics, such as FLOPs, parameter count, or model size, only capture isolated dimensions of efficiency and fail to reflect the complex trade\-offs between performance and resource consumption\. Moreover, differences in hardware configurations, evaluation protocols, and reporting standards make it difficult to compare methods fairly\. This fragmentation hinders progress and obscures the real\-world impact of proposed optimizations\. A more holistic evaluation framework is needed, one that incorporates not only computational and memory efficiency but also energy consumption, carbon emissions, and lifecycle sustainability\. Emerging tools such as CodeCarbon\[[197](https://arxiv.org/html/2607.09084#bib.bib197)\]and experiment\-impact\-tracker offer promising directions for measuring environmental impact, while performance\-efficiency Pareto frontiers provide a principled lens for evaluating trade\-offs\. Moving forward, the community would greatly benefit from a unified benchmark suite that supports multi\-dimensional evaluation across diverse deployment contexts and model sizes\.

## 7Conclusion

The era of large\-scale AI models has brought transformative advances across natural language processing, computer vision, and scientific discovery\. However, these advancements entail unprecedented computational and energy demands\. This survey explores the green development path of large models, covering efficient architectures, optimized training and inference, hardware\-software co\-design, and sustainability\-driven applications\.

We review advances like sparse activation, physics\-inspired modeling, dynamic data selection, and parameter\-efficient fine\-tuning, which reduce resource usage while maintaining performance\. On the hardware side, we analyze energy\-efficient AI chips, memory\-centric computing, and decentralized learning systems that lower environmental impact\. We also highlight AI’s growing role in sustainability domains, including remote sensing and green infrastructure\.

Looking ahead, we advocate for a shift toward lifecycle\-aware model development, standardized efficiency benchmarks, and co\-designed algorithm–hardware systems\. Embedding environmental responsibility at every layer of AI development will ensure that future models are not only more powerful but also more sustainable and widely accessible\.

###### Acknowledgements\.

This work was supported in part by the National Natural Science Foundation of China under Grant 62536003, and also in part by the Major Key Project of Pengcheng Laboratory under Grant PCL2025A14\.

## References

- \[1\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit,*et al\.*, “Attention is all you need”,*Advances in neural information processing systems*, in press, 2017\.
- \[2\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova, “Bert: Pre\-training of deep bidirectional transformers for language understanding”, in*Proceedings of the North American Chapter of the Association for Computational Linguistics*, 2019\.
- \[3\]A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh,*et al\.*, “Learning transferable visual models from natural language supervision”, in*International conference on machine learning*, PmLR, pp\.8748–8763, 2021\.
- \[4\]B\. Lin, Z\. Tang, Y\. Ye, J\. Cui,*et al\.*, “Moe\-llava: Mixture of experts for large vision\-language models”,*arXiv preprint arXiv:2401\.15947*, in press, 2024\.
- \[5\]J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad,*et al\.*, “Gpt\-4 technical report”,*arXiv preprint arXiv:2303\.08774*, in press, 2023\.
- \[6\]A\. Radovanović, R\. Koningstein, I\. Schneider, B\. Chen,*et al\.*, “Carbon\-aware computing for datacenters”,*IEEE Transactions on Power Systems*, vol\.38, no\.2, pp\.1270–1280, 2022\.
- \[7\]C\. Guo, F\. Cheng, Z\. Du, J\. Kiessling,*et al\.*, “A survey: Collaborative hardware and software design in the era of large language models”,*IEEE Circuits and Systems Magazine*, vol\.25, no\.1, pp\.35–57, 2025\.
- \[8\]J\. Duan, S\. Zhang, Z\. Wang, L\. Jiang,*et al\.*, “Efficient training of large language models on distributed infrastructures: A survey”,*arXiv preprint arXiv:2407\.20018*, in press, 2024\.
- \[9\]L\. Shen, Y\. Sun, Z\. Yu, L\. Ding,*et al\.*, “On efficient training of large\-scale deep learning models”,*ACM Computing Surveys*, vol\.57, no\.3, pp\.1–36, 2024\.
- \[10\]G\. Bai, Z\. Chai, C\. Ling, S\. Wang,*et al\.*, “Beyond efficiency: A systematic survey of resource\-efficient large language models”,*arXiv preprint arXiv:2401\.00625*, in press, 2024\.
- \[11\]Z\. Wang, Y\. Liu, Y\. Tian, Y\. Liu,*et al\.*, “Building vision models upon heat conduction”, in*Proceedings of the Computer Vision and Pattern Recognition Conference*, pp\.9707–9717, 2025\.
- \[12\]Y\. Liu, Y\. Tian, Y\. Zhao, H\. Yu,*et al\.*, “Vmamba: Visual state space model”,*Advances in neural information processing systems*, in press, 2024\.
- \[13\]A\. Katharopoulos, A\. Vyas, N\. Pappas, and F\. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention”, in*International conference on machine learning*, 2020\.
- \[14\]T\. Dao, D\. Fu, S\. Ermon, A\. Rudra, and C\. Ré, “Flashattention: Fast and memory\-efficient exact attention with io\-awareness”,*Advances in neural information processing systems*, in press, 2022\.
- \[15\]T\. Dao, “Flashattention\-2: Faster attention with better parallelism and work partitioning”,*arXiv preprint arXiv:2307\.08691*, in press, 2023\.
- \[16\]J\. Shah, G\. Bikshandi, Y\. Zhang, V\. Thakkar,*et al\.*, “Flashattention\-3: Fast and accurate attention with asynchrony and low\-precision”,*Advances in Neural Information Processing Systems*, in press, 2024\.
- \[17\]B\. Peng, E\. Alcaide, Q\. Anthony, A\. Albalak,*et al\.*, “Rwkv: Reinventing rnns for the transformer era”, in*Findings of the Association for Computational Linguistics: EMNLP 2023*, 2023\.
- \[18\]Y\. Sun, L\. Dong, S\. Huang, S\. Ma,*et al\.*, “Retentive network: A successor to transformer for large language models”,*arXiv preprint arXiv:2307\.08621*, in press, 2023\.
- \[19\]R\. Jacobs, M\. Jordan, S\. Nowlan, and G\. Hinton, “Adaptive mixtures of local experts”,*Neural Computation*, in press, 1991\.
- \[20\]N\. Shazeer, A\. Mirhoseini,*et al\.*, “Outrageously large neural networks: The sparsely\-gated mixture\-of\-experts layer”, in*International Conference on Learning Representations*, 2017\.
- \[21\]D\. Lepikhin, H\. Lee, Y\. Xu, D\. Chen,*et al\.*, “Gshard: Scaling giant models with conditional computation and automatic sharding”, in*International Conference on Learning Representations*, 2020\.
- \[22\]E\. Yang, L\. Shen, G\. Guo, X\. Wang,*et al\.*, “Model merging in llms, mllms, and beyond: Methods, theories, applications and opportunities”,*arXiv preprint arXiv:2408\.07666*, in press, 2024\.
- \[23\]J\. Rao, X\. Liu, H\. Deng, Z\. Lin,*et al\.*, “Dynamic sampling that adapts: Iterative dpo for self\-aware mathematical reasoning”,*arXiv preprint arXiv:2505\.16176*, in press, 2025\.
- \[24\]T\. Nguyen, Z\. Chen, and J\. Lee, “Dataset meta\-learning from kernel ridge\-regression”, in*International Conference on Learning Representations*, 2021\.
- \[25\]Y\. Zhou, E\. Nezhadarya, and J\. Ba, “Dataset distillation using neural feature regression”,*Advances in Neural Information Processing Systems*, in press, 2022\.
- \[26\]N\. Loo, R\. Hasani, A\. Amini, and D\. Rus, “Efficient dataset distillation using random feature approximation”,*Advances in Neural Information Processing Systems*, in press, 2022\.
- \[27\]C\. Guo, B\. Zhao, and Y\. Bai, “Deepcore: A comprehensive library for coreset selection in deep learning”, in*International Conference on Database and Expert Systems Applications*, pp\.181–195, 2022\.
- \[28\]S\. Li, Y\. Zhao, R\. Varma, O\. Salpekar,*et al\.*, “Pytorch distributed: Experiences on accelerating data parallel training”,*Proceedings of the VLDB Endowment*, vol\.13, no\.12, pp\.3005–3018, 2020\.
- \[29\]S\. Rajbhandari, J\. Rasley, O\. Ruwase, and Y\. He, “Zero: Memory optimizations toward training trillion parameter models”, in*SC20: International Conference for High Performance Computing, Networking, Storage and Analysis*, IEEE, pp\.1–16, 2020\.
- \[30\]J\. Rasley, S\. Rajbhandari, O\. Ruwase, and Y\. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters”, in*Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining*, pp\.3505–3506, 2020\.
- \[31\]M\. Shoeybi, M\. Patwary, R\. Puri, P\. LeGresley,*et al\.*, “Megatron\-lm: Training multi\-billion parameter language models using model parallelism”,*arXiv preprint arXiv:1909\.08053*, in press, 2019\.
- \[32\]Y\. Huang, Y\. Cheng, A\. Bapna, O\. Firat,*et al\.*, “Gpipe: Efficient training of giant neural networks using pipeline parallelism”,*Advances in neural information processing systems*, vol\.32, 2019\.
- \[33\]D\. Narayanan, M\. Shoeybi, J\. Casper, P\. LeGresley,*et al\.*, “Efficient large\-scale language model training on gpu clusters using megatron\-lm”, in*SC21: International Conference for High Performance Computing, Networking, Storage and Analysis*, IEEE, pp\.1–14, 2021\.
- \[34\]B\. Workshop, T\. L\. Scao, A\. Fan, C\. Akiki,*et al\.*, “Bloom: A 176b\-parameter open\-access multilingual language model”,*arXiv preprint arXiv:2211\.05100*, in press, 2022\.
- \[35\]B\. Lester, R\. Al\-Rfou, and N\. Constant, “The power of scale for parameter\-efficient prompt tuning”, in*Proceedings of the Conference on Empirical Methods in Natural Language Processing*, 2021\.
- \[36\]N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone,*et al\.*, “Parameter\-efficient transfer learning for nlp”, in*International conference on machine learning*, pp\.2790–2799, 2019\.
- \[37\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu,*et al\.*, “Lora: Low\-rank adaptation of large language models\.”,*International Conference on Learning Representations*, vol\.1, no\.2, artilce no\.3, 2022\.
- \[38\]Z\. Liu, B\. Oguz, C\. Zhao, E\. Chang,*et al\.*, “Llm\-qat: Data\-free quantization aware training for large language models”, in*Findings of the Association for Computational Linguistics ACL 2024*, pp\.467–484, 2024\.
- \[39\]E\. Frantar and D\. Alistarh, “Sparsegpt: Massive language models can be accurately pruned in one\-shot”, in*International conference on machine learning*, 2023\.
- \[40\]S\. V\. Kandala, P\. Medaranga, and A\. Varshney, “Tinyllm: A framework for training and deploying language models at the edge computers”,*arXiv preprint arXiv:2412\.15304*, in press, 2024\.
- \[41\]J\. Xu, J\. Pan, Y\. Zhou, S\. Chen,*et al\.*, “Specee: Accelerating large language model inference with speculative early exiting”, in*Proceedings of the 52nd Annual International Symposium on Computer Architecture*, pp\.467–481, 2025\.
- \[42\]Y\. Hu, Y\. Liu, and Z\. Liu, “A survey on convolutional neural network accelerators: Gpu, fpga and asic”, in*2022 14th International Conference on Computer Research and Development \(ICCRD\)*, IEEE, pp\.100–107, 2022\.
- \[43\]N\. P\. Jouppi, C\. Young, N\. Patil, D\. Patterson,*et al\.*, “In\-datacenter performance analysis of a tensor processing unit”, in*Proceedings of the 44th annual international symposium on computer architecture*, pp\.1–12, 2017\.
- \[44\]S\. Liu, Z\. Du, J\. Tao, D\. Han,*et al\.*, “Cambricon: An instruction set architecture for neural networks”,*ACM SIGARCH Computer Architecture News*, vol\.44, no\.3, pp\.393–405, 2016\.
- \[45\]D\. Shi, Y\. Fu, X\. Yuan, Z\. Yu,*et al\.*, “Lacache: Ladder\-shaped kv caching for efficient long\-context modeling of large language models”,*arXiv preprint arXiv:2507\.14204*, in press, 2025\.
- \[46\]A\. Vijaya Kumar, G\. Antichi, and R\. Singh, “Aqua: Network\-accelerated memory offloading for llms in scale\-up gpu domains”, in*Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2*, pp\.48–62, 2025\.
- \[47\]C\. Luo, Z\. Cai, H\. Sun, J\. Xiao,*et al\.*, “Headinfer: Memory\-efficient llm inference by head\-wise offloading”, in*ICML 2025 Workshop on Long\-Context Foundation Models*, \.
- \[48\]S\. Fan, X\. Jiang, X\. Li, X\. Meng,*et al\.*, “Not all layers of llms are necessary during inference”, in*Proceedings of the Thirty\-Fourth International Joint Conference on Artificial Intelligence*, pp\.5083–5091, 2025\.
- \[49\]T\. Tang, H\. Yiwen, B\. Li, W\. Luo,*et al\.*, “Llmbox: A comprehensive library for large language models”, in*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics*, pp\.388–399, 2024\.
- \[50\]R\. Gong, S\. Bai, S\. Wu, Y\. Fan,*et al\.*, “Past\-future scheduler for llm serving under sla guarantees”, in*Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2*, pp\.798–813, 2025\.
- \[51\]Q\. Zhang, M\. Zhai, R\. Sun, and J\. Zhai, “Qfactory: Accelerating quantized large language model serving with qtile graphs”, in*2025 USENIX Annual Technical Conference*, pp\.631–646, 2025\.
- \[52\]Y\. Guan, Y\. Fang, K\. Zhou, C\. Robeck,*et al\.*, “Kperfir: Towards an open and compiler\-centric ecosystem for gpu kernel performance tooling on modern ai workloads”,*arXiv preprint arXiv:2505\.21661*, in press, 2025\.
- \[53\]A\. Basu, L\. Deng, C\. Frenkel, and X\. Zhang, “Spiking neural network integrated circuits: A review of trends and future directions”, in*2022 IEEE custom integrated circuits conference \(CICC\)*, IEEE, pp\.1–8, 2022\.
- \[54\]A\. Borzunov, D\. Baranchuk, T\. Dettmers, M\. Riabinin,*et al\.*, “Petals: Collaborative inference and fine\-tuning of large models”, in*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\)*, pp\.558–568, 2023\.
- \[55\]Z\. Cao, X\. Zhou, X\. Wu, Z\. Zhu,*et al\.*, “Data center sustainability: Revisits and outlooks”,*IEEE Transactions on Sustainable Computing*, vol\.9, no\.3, pp\.236–248, 2023\.
- \[56\]Y\. Zhang, H\. Li, and S\. Wang, “Decarbonizing data centers through regional bits migration: A comprehensive assessment of china’s ‘eastern data, western computing’initiative and its global implications”,*Applied Energy*, vol\.392, artilce no\.126020, 2025\.
- \[57\]X\. Liu, Y\. Zheng, Z\. Du, M\. Ding,*et al\.*, “Gpt understands, too”,*AI Open*, vol\.5, pp\.208–215, 2024\.
- \[58\]S\. Wang, B\. Z\. Li, M\. Khabsa, H\. Fang, and H\. Ma, “Linformer: Self\-attention with linear complexity”,*arXiv preprint arXiv:2006\.04768*, in press, 2020\.
- \[59\]I\. Schlag, K\. Irie, and J\. Schmidhuber, “Linear transformers are secretly fast weight programmers”, in*International conference on machine learning*, pp\.9355–9366, 2021\.
- \[60\]A\. Anthropic, “The claude 3 model family: Opus, sonnet, haiku”,*Claude\-3 Model Card*, vol\.1, no\.1, artilce no\.4, 2024\.
- \[61\]H\. Touvron, L\. Martin, K\. Stone, P\. Albert,*et al\.*, “Llama 2: Open foundation and fine\-tuned chat models”,*arXiv preprint arXiv:2307\.09288*, in press, 2023\.
- \[62\]H\. Hu, P\. Wang, H\. Bi, B\. Tong,*et al\.*, “Rs\-vheat: Heat conduction guided efficient remote sensing foundation model”,*arXiv preprint arXiv:2411\.17984*, in press, 2024\.
- \[63\]A\. Gu, K\. Goel, and C\. Ré, “Efficiently modeling long sequences with structured state spaces”, in*International Conference on Learning Representations*, 2022\.
- \[64\]A\. Gu and T\. Dao, “Mamba: Linear\-time sequence modeling with selective state spaces”,*arXiv preprint arXiv:2312\.00752*, in press, 2023\.
- \[65\]D\. Dai, C\. Deng, C\. Zhao, R\. X\. Xu,*et al\.*, “Deepseekmoe: Towards ultimate expert specialization in mixture\-of\-experts language models”, in*Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics*, pp\.1280–1297, 2024\.
- \[66\]J\. Puigcerver, C\. R\. Ruiz, B\. Mustafa, and N\. Houlsby, “From sparse to soft mixtures of experts”, in*The Twelfth International Conference on Learning Representations*, 2024\.
- \[67\]X\. Wang, F\. Yu, L\. Dunlap, Y\.\-A\. Ma,*et al\.*, “Deep mixture of experts via shallow embedding”, in*Uncertainty in Artificial Intelligence*, pp\.552–562, 2020\.
- \[68\]H\. Tang, J\. Liu, M\. Zhao, and X\. Gong, “Progressive layered extraction \(ple\): A novel multi\-task learning \(mtl\) model for personalized recommendations”, in*Proceedings of the 14th ACM conference on recommender systems*, pp\.269–278, 2020\.
- \[69\]J\. Ma, Z\. Zhao, X\. Yi, J\. Chen,*et al\.*, “Modeling task relationships in multi\-task learning with multi\-gate mixture\-of\-experts”, in*Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining*, pp\.1930–1939, 2018\.
- \[70\]C\. Riquelme, J\. Puigcerver, B\. Mustafa, M\. Neumann,*et al\.*, “Scaling vision with sparse mixture of experts”,*Advances in Neural Information Processing Systems*, vol\.34, pp\.8583–8595, 2021\.
- \[71\]Z\. Zeng, Y\. Miao, H\. Gao, H\. Zhang, and Z\. Deng, “Adamoe: Token\-adaptive routing with null experts for mixture\-of\-experts language models”, in*Findings of the Association for Computational Linguistics: EMNLP*, pp\.6223–6235, 2024\.
- \[72\]L\. Jing, Y\. Gao, Z\. Wang, W\. Lan,*et al\.*, “Evomoe: Expert evolution in mixture of experts for multimodal large language models”,*arXiv preprint arXiv:2505\.23830*, in press, 2025\.
- \[73\]G\. Ilharco, M\. T\. Ribeiro, M\. Wortsman, L\. Schmidt,*et al\.*, “Editing models with task arithmetic”,*ICLR*, in press, 2023\.
- \[74\]M\. Wortsman, G\. Ilharco, S\. Y\. Gadre, R\. Roelofs,*et al\.*, “Model soups: Averaging weights of multiple fine\-tuned models improves accuracy without increasing inference time”, in*International conference on machine learning*, 2022\.
- \[75\]G\. Stoica, D\. Bolya, J\. B\. Bjorner, P\. Ramesh,*et al\.*, “Zipit\! merging models from different tasks without training”, in*The Twelfth International Conference on Learning Representations*, 2023\.
- \[76\]R\. Entezari, H\. Sedghi, O\. Saukh, and B\. Neyshabur, “The role of permutation invariance in linear mode connectivity of neural networks”, in*Sparsity in Neural Networks\-Advancing Understanding and Practice: SNN Workshop 2021*, 2021\.
- \[77\]J\. Frankle, G\. K\. Dziugaite, D\. Roy, and M\. Carbin, “Linear mode connectivity and the lottery ticket hypothesis”, in*International Conference on Machine Learning*, 2020\.
- \[78\]T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah,*et al\.*, “Language models are few\-shot learners”,*Advances in neural information processing systems*, vol\.33, pp\.1877–1901, 2020\.
- \[79\]H\. Touvron, T\. Lavril, G\. Izacard, X\. Martinet,*et al\.*, “Llama: Open and efficient foundation language models”,*arXiv preprint arXiv:2302\.13971*, in press, 2023\.
- \[80\]Y\. Bengio, J\. Louradour, R\. Collobert, and J\. Weston, “Curriculum learning”, in*Proceedings of the 26th annual international conference on machine learning*, pp\.41–48, 2009\.
- \[81\]X\. Wang, Y\. Chen, and W\. Zhu, “A survey on curriculum learning”,*IEEE transactions on pattern analysis and machine intelligence*, vol\.44, no\.9, pp\.4555–4576, 2021\.
- \[82\]L\. Xiao, X\. Yang, F\. Peng, M\. Yan,*et al\.*, “Clip\-vg: Self\-paced curriculum adapting of clip for visual grounding”,*IEEE Transactions on Multimedia*, vol\.26, pp\.4334–4347, 2023\.
- \[83\]A\. Katharopoulos and F\. Fleuret, “Biased importance sampling for deep neural network training”,*arXiv preprint arXiv:1706\.00043*, in press, 2017\.
- \[84\]A\. H\. Jiang, D\. L\.\-K\. Wong, G\. Zhou, D\. G\. Andersen,*et al\.*, “Accelerating deep learning by focusing on the biggest losers”,*arXiv preprint arXiv:1910\.00762*, in press, 2019\.
- \[85\]A\. Katharopoulos and F\. Fleuret, “Not all samples are created equal: Deep learning with importance sampling”, in*International conference on machine learning*, PMLR, pp\.2525–2534, 2018\.
- \[86\]D\. Sow, H\. Woisetschläger, S\. K\. Bulusu, S\. Wang,*et al\.*, “Dynamic loss\-based sample reweighting for improved large language model pretraining”, in*International Conference on Learning Representations*, 2025\.
- \[87\]T\. Wang, J\.\-Y\. Zhu, A\. Torralba, and A\. A\. Efros, “Dataset distillation”,*arXiv preprint arXiv:1811\.10959*, in press, 2018\.
- \[88\]R\. Yu, S\. Liu, and X\. Wang, “Dataset distillation: A comprehensive review”,*IEEE transactions on pattern analysis and machine intelligence*, vol\.46, no\.1, pp\.150–170, 2023\.
- \[89\]Z\. Deng and O\. Russakovsky, “Remember the past: Distilling datasets into addressable memories for neural networks”,*Advances in Neural Information Processing Systems*, vol\.35, pp\.34391–34404, 2022\.
- \[90\]L\. Zhang, J\. Zhang, B\. Lei, S\. Mukherjee,*et al\.*, “Accelerating dataset distillation via model augmentation”, in*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\.11950–11959, 2023\.
- \[91\]G\. Cazenavette, T\. Wang, A\. Torralba, A\. A\. Efros, and J\.\-Y\. Zhu, “Dataset distillation by matching training trajectories”, in*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\.4750–4759, 2022\.
- \[92\]G\. Li, R\. Togo, T\. Ogawa, and M\. Haseyama, “Dataset distillation using parameter pruning”,*IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences*, vol\.107, no\.6, pp\.936–940, 2024\.
- \[93\]B\. Zhao and H\. Bilen, “Dataset condensation with distribution matching”, in*Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision*, pp\.6514–6523, 2023\.
- \[94\]K\. Wang, B\. Zhao, X\. Peng, Z\. Zhu,*et al\.*, “Cafe: Learning to condense dataset by aligning features”, in*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\.12196–12205, 2022\.
- \[95\]S\. Wang, Y\. Yang, Z\. Liu, C\. Sun,*et al\.*, “Dataset distillation with neural characteristic function: A minmax perspective”, in*Proceedings of the Computer Vision and Pattern Recognition Conference*, pp\.25570–25580, 2025\.
- \[96\]M\. Paul, S\. Ganguli, and G\. K\. Dziugaite, “Deep learning on a data diet: Finding important examples early in training”,*Advances in neural information processing systems*, vol\.34, pp\.20596–20607, 2021\.
- \[97\]M\. Toneva, A\. Sordoni, R\. T\. des Combes, A\. Trischler,*et al\.*, “An empirical study of example forgetting during deep neural network learning”, in*International Conference on Learning Representations*, 2019\.
- \[98\]M\. He, S\. Yang, T\. Huang, and B\. Zhao, “Large\-scale dataset pruning with dynamic uncertainty”, in*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\.7713–7722, 2024\.
- \[99\]X\. Zhang, J\. Du, Y\. Li, W\. Xie, and J\. T\. Zhou, “Spanning training progress: Temporal dual\-depth scoring \(tdds\) for enhanced dataset pruning”, in*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pp\.26223–26232, 2024\.
- \[100\]S\. Agarwal, H\. Arora, S\. Anand, and C\. Arora, “Contextual diversity for active learning”, in*European Conference on Computer Vision*, Springer, pp\.137–153, 2020\.
- \[101\]M\. Welling, “Herding dynamical weights to learn”, in*Proceedings of the 26th annual international conference on machine learning*, pp\.1121–1128, 2009\.
- \[102\]B\. Sorscher, R\. Geirhos, S\. Shekhar, S\. Ganguli, and A\. Morcos, “Beyond neural scaling laws: Beating power law scaling via data pruning”,*Advances in Neural Information Processing Systems*, vol\.35, pp\.19523–19536, 2022\.
- \[103\]X\. Xia, J\. Liu, J\. Yu, X\. Shen,*et al\.*, “Moderate coreset: A universal method of data selection for real\-world data\-efficient deep learning”, in*The Eleventh International Conference on Learning Representations*, 2022\.
- \[104\]A\. Maharana, P\. Yadav, and M\. Bansal, “D2 pruning: Message passing for balancing diversity and difficulty in data pruning”, in*The Twelfth International Conference on Learning Representations*, 2024\.
- \[105\]P\. Micikevicius, S\. Narang, J\. Alben, G\. Diamos,*et al\.*, “Mixed precision training”, in*International Conference on Learning Representations*, 2018\.
- \[106\]X\. L\. Li and P\. Liang, “Prefix\-tuning: Optimizing continuous prompts for generation”, in*Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\)*, pp\.4582–4597, 2021\.
- \[107\]J\. He, C\. Zhou, X\. Ma, T\. Berg\-Kirkpatrick, and G\. Neubig, “Towards a unified view of parameter\-efficient transfer learning”, in*International Conference on Learning Representations*, 2022\.
- \[108\]C\. Wei, Y\. Shu, Y\. T\. He, and F\. R\. Yu, “Flexora: Flexible low rank adaptation for large language models”,*arXiv preprint arXiv:2408\.10774*, in press, 2024\.
- \[109\]X\. Liu, K\. Ji, Y\. Fu, W\. Tam,*et al\.*, “P\-tuning: Prompt tuning can be comparable to fine\-tuning across scales and tasks”, in*Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\)*, pp\.61–68, 2022\.
- \[110\]C\. Fu, H\. Huang, X\. Chen, Y\. Tian, and J\. Zhao, “Learn\-to\-share: A hardware\-friendly transfer learning framework exploiting computation and parameter sharing”, in*International Conference on Machine Learning*, PMLR, pp\.3469–3479, 2021\.
- \[111\]J\. Pfeiffer, A\. Kamath, A\. Rücklé, K\. Cho, and I\. Gurevych, “Adapterfusion: Non\-destructive task composition for transfer learning”, in*Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume*, pp\.487–503, 2021\.
- \[112\]Y\. Wang, S\. Agarwal, S\. Mukherjee, X\. Liu,*et al\.*, “Adamix: Mixture\-of\-adaptations for parameter\-efficient model tuning”, in*Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing*, pp\.5744–5760, 2022\.
- \[113\]S\. He, R\.\-Z\. Fan, L\. Ding, L\. Shen,*et al\.*, “Mera: Merging pretrained adapters for few\-shot learning”,*arXiv preprint arXiv:2308\.15982*, in press, 2023\.
- \[114\]N\. Lawton, A\. Kumar, G\. Thattai, A\. Galstyan, and G\. Ver Steeg, “Neural architecture search for parameter\-efficient fine\-tuning of large pre\-trained language models”, in*Findings of the Association for Computational Linguistics: ACL 2023*, pp\.8506–8515, 2023\.
- \[115\]H\. Zhou, X\. Wan, I\. Vulić, and A\. Korhonen, “Autopeft: Automatic configuration search for parameter\-efficient fine\-tuning”,*Transactions of the Association for Computational Linguistics*, vol\.12, pp\.525–542, 2024\.
- \[116\]S\. He, L\. Ding, D\. Dong, J\. Zhang, and D\. Tao, “Sparseadapter: An easy approach for improving the parameter\-efficiency of adapters”, in*Findings of the Association for Computational Linguistics: EMNLP 2022*, pp\.2184–2190, 2022\.
- \[117\]L\. Zhang, A\. Rao, and M\. Agrawala, “Adding conditional control to text\-to\-image diffusion models”, in*Proceedings of the IEEE/CVF international conference on computer vision*, pp\.3836–3847, 2023\.
- \[118\]M\. Tao, B\.\-K\. Bao, H\. Tang, and C\. Xu, “Galip: Generative adversarial clips for text\-to\-image synthesis”, in*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pp\.14214–14223, 2023\.
- \[119\]H\. Zhou, X\. Lu, W\. Xu, C\. Zhu,*et al\.*, “Lora\-drop: Efficient lora parameter pruning based on output evaluation”, in*Proceedings of the 31st International Conference on Computational Linguistics*, pp\.5530–5543, 2025\.
- \[120\]H\. U\. K\. Shinwari and M\. Usama, “Ard\-lora: Dynamic rank allocation for parameter\-efficient fine\-tuning of foundation models with heterogeneous adaptation needs”,*arXiv preprint arXiv:2506\.18267*, in press, 2025\.
- \[121\]H\. Lu, C\. Zhao, J\. Xue, L\. Yao,*et al\.*, “Adaptive rank, reduced forgetting: Knowledge retention in continual learning vision\-language models with dynamic rank\-selective lora”,*arXiv preprint arXiv:2412\.01004*, in press, 2024\.
- \[122\]L\. Xiao, X\. Yang, F\. Peng, Y\. Wang, and C\. Xu, “Hivg: Hierarchical multimodal fine\-grained modulation for visual grounding”, in*Proceedings of the 32nd ACM International Conference on Multimedia*, pp\.5460–5469, 2024\.
- \[123\]S\. Chen, W\. Wang, X\. Chen, P\. Lu,*et al\.*, “Llama\-lora neural prompt engineering: A deep tuning framework for automatically generating chinese text logical reasoning thinking chains”,*Data Intelligence*, vol\.6, no\.2, pp\.375–408, 2024\.
- \[124\]T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond,*et al\.*, “Transformers: State\-of\-the\-art natural language processing”, in*Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations*, Association for Computational Linguistics, Online, pp\.38–45,[https://www\.aclweb\.org/anthology/2020\.emnlp\-demos\.6](https://www.aclweb.org/anthology/2020.emnlp-demos.6), 2020\.
- \[125\]S\. Mangrulkar, S\. Gugger, L\. Debut, Y\. Belkada,*et al\.*, “PEFT: State\-of\-the\-art parameter\-efficient fine\-tuning methods”,[https://github\.com/huggingface/peft](https://github.com/huggingface/peft), 2022\.
- \[126\]L\. Zheng, W\.\-L\. Chiang, Y\. Sheng, S\. Zhuang,*et al\.*, “Judging llm\-as\-a\-judge with mt\-bench and chatbot arena”, 2023,[2306\.05685](https://arxiv.org/html/2607.09084v1/2306.05685)\.
- \[127\]M\. H\. Daniel Han and U\. team, “Unsloth”, 2023,[http://github\.com/unslothai/unsloth](http://github.com/unslothai/unsloth)\.
- \[128\]S\. Li, H\. Liu, Z\. Bian, J\. Fang,*et al\.*, “Colossal\-ai: A unified deep learning system for large\-scale parallel training”, in*Proceedings of the 52nd International Conference on Parallel Processing*, Association for Computing Machinery, New York, NY, USA, ICPP ’23, artilce no\.766–775,[10\.1145/3605573\.3605613](https://arxiv.org/doi.org/10.1145/3605573.3605613),[https://doi\.org/10\.1145/3605573\.3605613](https://doi.org/10.1145/3605573.3605613), 2023, ISBN 9798400708435\.
- \[129\]R\. Wang, Z\. Gao, L\. Zhang, S\. Yue, and Z\. Gao, “Empowering large language models to edge intelligence: A survey of edge efficient llms and techniques”,*Computer Science Review*, vol\.57, artilce no\.100755, 2025\.
- \[130\]J\. Yuan, H\. Li, X\. Ding, W\. Xie,*et al\.*, “Give me fp32 or give me death? challenges and solutions for reproducible reasoning”,*arXiv preprint arXiv:2506\.09501*, in press, 2025\.
- \[131\]M\. Van Baalen, A\. Kuzmin, S\. S\. Nair, Y\. Ren,*et al\.*, “Fp8 versus int8 for efficient deep learning inference”,*arXiv preprint arXiv:2303\.17951*, in press, 2023\.
- \[132\]H\. Peng, K\. Wu, Y\. Wei, G\. Zhao,*et al\.*, “Fp8\-lm: Training fp8 large language models”,*arXiv preprint arXiv:2310\.18313*, in press, 2023\.
- \[133\]T\. Dettmers, M\. Lewis, Y\. Belkada, and L\. Zettlemoyer, “Gpt3\. int8 \(\): 8\-bit matrix multiplication for transformers at scale”,*Advances in neural information processing systems*, vol\.35, pp\.30318–30332, 2022\.
- \[134\]P\. Zhang, J\. Wei, J\. Zhang, J\. Zhu, and J\. Chen, “Accurate int8 training through dynamic block\-level fallback”,*arXiv preprint arXiv:2503\.08040*, in press, 2025\.
- \[135\]W\. Guo, D\. Liu, W\. Xie, Y\. Li,*et al\.*, “Towards accurate and efficient sub\-8\-bit integer training”,*arXiv preprint arXiv:2411\.10948*, in press, 2024\.
- \[136\]Y\. Xu, X\. Han, Z\. Yang, S\. Wang,*et al\.*, “Onebit: Towards extremely low\-bit large language models”,*Advances in Neural Information Processing Systems*, vol\.37, pp\.66357–66382, 2024\.
- \[137\]H\. Wang, S\. Ma, L\. Dong, S\. Huang,*et al\.*, “Bitnet: Scaling 1\-bit transformers for large language models”,*arXiv preprint arXiv:2310\.11453*, in press, 2023\.
- \[138\]Z\. Liu, C\. Zhao, H\. Huang, S\. Chen,*et al\.*, “Paretoq: Improving scaling laws in extremely low\-bit llm quantization”, in*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, \.
- \[139\]C\. Zeng, S\. Liu, Y\. Xie, H\. Liu,*et al\.*, “Abq\-llm: Arbitrary\-bit quantized inference acceleration for large language models”, in*Proceedings of the AAAI Conference on Artificial Intelligence*, vol\.39, pp\.22299–22307, 2025\.
- \[140\]Y\. Liu, J\. Ning, S\. Xia, X\. Gao,*et al\.*, “Pruning large language models by identifying and preserving functional networks”,*arXiv preprint arXiv:2508\.05239*, in press, 2025\.
- \[141\]B\. Hou, Q\. Chen, J\. Wang, G\. Yin,*et al\.*, “Instruction\-following pruning for large language models”,*arXiv preprint arXiv:2501\.02086*, in press, 2025\.
- \[142\]J\. Whitmore, C\. Hastings, A\. Patel, and S\. Brody, “Efficient inference of large language models through model compression”, in press, 2025\.
- \[143\]Y\. Liu, H\. Yang, Y\. Chen, R\. Zhang,*et al\.*, “Pat: Pruning\-aware tuning for large language models”, in*Proceedings of the AAAI Conference on Artificial Intelligence*, 2025\.
- \[144\]M\. Sun, Z\. Liu, A\. Bair, and J\. Z\. Kolter, “A simple and effective pruning approach for large language models”,*arXiv preprint arXiv:2306\.11695*, in press, 2023\.
- \[145\]S\. Ashkboos, M\. L\. Croci, M\. G\. d\. Nascimento, T\. Hoefler, and J\. Hensman, “Slicegpt: Compress large language models by deleting rows and columns”,*arXiv preprint arXiv:2401\.15024*, in press, 2024\.
- \[146\]Y\. Li, L\. Niu, X\. Zhang, K\. Liu,*et al\.*, “E\-sparse: Boosting the large language model inference through entropy\-based n: M sparsity”,*arXiv preprint arXiv:2310\.15929*, in press, 2023\.
- \[147\]C\. Yang, Y\. Zhu, W\. Lu, Y\. Wang,*et al\.*, “Survey on knowledge distillation for large language models: Methods, evaluation, and application”,*ACM Transactions on Intelligent Systems and Technology*, in press, 2024\.
- \[148\]X\. Xu, M\. Li, C\. Tao, T\. Shen,*et al\.*, “A survey on knowledge distillation of large language models”,*arXiv preprint arXiv:2402\.13116*, in press, 2024\.
- \[149\]M\. Takrouri, N\. M\. Cuadrado, and M\. Takáč, “Knowledge distillation from large language models for household energy modeling”,*arXiv preprint arXiv:2502\.03034*, in press, 2025\.
- \[150\]A\. Mitra, L\. Del Corro, S\. Mahajan, A\. Codas,*et al\.*, “Orca 2: Teaching small language models how to reason”,*arXiv preprint arXiv:2311\.11045*, in press, 2023\.
- \[151\]C\. He, Y\. Ding, J\. Guo, R\. Gong,*et al\.*, “Da\-kd: Difficulty\-aware knowledge distillation for efficient large language models”, in*Forty\-second International Conference on Machine Learning*, \.
- \[152\]L\. H\. Li, J\. Hessel, Y\. Yu, X\. Ren,*et al\.*, “Symbolic chain\-of\-thought distillation: Small models can also” think” step\-by\-step”, in*The 61st Annual Meeting Of The Association For Computational Linguistics*, 2023\.
- \[153\]X\. Luo, Y\. Wang, Q\. Zhu, Z\. Zhang,*et al\.*, “Turning trash into treasure: Accelerating inference of large language models with token recycling”,*arXiv preprint arXiv:2408\.08696*, in press, 2024\.
- \[154\]J\. Sen, S\. Dasgupta, and H\. Waghela, “Confidence\-modulated speculative decoding for large language models”,*arXiv preprint arXiv:2508\.15371*, in press, 2025\.
- \[155\]Y\. Cheng, A\. Zhang, X\. Zhang, C\. Wang, and Y\. Wang, “Recurrent drafter for fast speculative decoding in large language models”,*arXiv preprint arXiv:2403\.09919*, in press, 2024\.
- \[156\]Y\. Ni, C\. Liu, Y\. Tang, K\. Han, and Y\. Wang, “Ems\-sd: Efficient multi\-sample speculative decoding for accelerating large language models”, in*Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\)*, pp\.9307–9320, 2025\.
- \[157\]S\. A\. Budennyy, V\. D\. Lazarev, N\. N\. Zakharenko, A\. N\. Korovin,*et al\.*, “Eco2ai: Carbon emissions tracking of machine learning models as the first step towards sustainable ai”, in*Doklady mathematics*, Springer, vol\.106, pp\.S118–S128, 2022\.
- \[158\]E\. Strubell, A\. Ganesh, and A\. McCallum, “Energy and policy considerations for modern deep learning research”, in*Proceedings of the AAAI conference on artificial intelligence*, vol\.34, 2020\.
- \[159\]Y\.\-H\. Chen, T\. Krishna, J\. S\. Emer, and V\. Sze, “Eyeriss: An energy\-efficient reconfigurable accelerator for deep convolutional neural networks”,*IEEE journal of solid\-state circuits*, vol\.52, no\.1, pp\.127–138, 2016\.
- \[160\]G\. Team, R\. Anil, S\. Borgeaud, J\.\-B\. Alayrac,*et al\.*, “Gemini: A family of highly capable multimodal models”,*arXiv preprint arXiv:2312\.11805*, in press, 2023\.
- \[161\]J\. Wu and D\. Ren, “Development of artificial intelligence chips in china”,*Strategic Study of Chinese Academy of Engineering*, vol\.27, no\.1, pp\.133–141, 2025\.
- \[162\]B\. Wu, Y\. Zhong, Z\. Zhang, S\. Liu,*et al\.*, “Fast distributed inference serving for large language models”,*arXiv preprint arXiv:2305\.05920*, in press, 2023\.
- \[163\]Y\. Liu, J\. Fu, S\. Liu, Y\. Zou,*et al\.*, “Kv cache compression for inference efficiency in llms: A review”,*arXiv preprint arXiv:2508\.06297*, in press, 2025\.
- \[164\]H\. Li, Y\. Li, A\. Tian, T\. Tang,*et al\.*, “A survey on large language model acceleration based on kv cache management”,*arXiv preprint arXiv:2412\.19442*, in press, 2024\.
- \[165\]Y\. Liu, H\. Li, Y\. Cheng, S\. Ray,*et al\.*, “Cachegen: Kv cache compression and streaming for fast large language model serving”, in*Proceedings of the ACM SIGCOMM 2024 Conference*, 2024\.
- \[166\]S\. Ge, Y\. Zhang, L\. Liu, M\. Zhang,*et al\.*, “Model tells you what to discard: Adaptive kv cache compression for llms”, in*12th International Conference on Learning Representations, ICLR 2024*, 2024\.
- \[167\]G\. WEI, X\. Zhou, P\. Sun, T\. Zhang, and Y\. Wen, “Rethinking key\-value cache compression techniques for large language model serving”, in*Eighth Conference on Machine Learning and Systems*, \.
- \[168\]M\. Davies, N\. Crago, K\. Sankaralingam, and C\. Kozyrakis, “Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need”,*arXiv preprint*, in press, 2025\.
- \[169\]Z\. Wang, Z\. Zhang, Y\. Zhou, Z\. Wang,*et al\.*, “Accelerating mixture\-of\-experts inference by hiding offloading latency with speculative decoding”,*arXiv preprint arXiv:2508\.21706*, in press, 2025\.
- \[170\]X\. Zhuge, X\. Shen, Z\. Wang, F\. Dang,*et al\.*, “Specoffload: Unlocking latent gpu capacity for llm inference on resource\-constrained devices”,*arXiv preprint arXiv:2505\.10259*, in press, 2025\.
- \[171\]R\. Miao, Y\. Yan, X\. Yao, and T\. Yang, “An efficient inference framework for early\-exit large language models”,*arXiv preprint arXiv:2407\.20272*, in press, 2024\.
- \[172\]D\. J\. Bajpai and M\. K\. Hanawal, “Free: Fast and robust vision language models with early exits”,*arXiv preprint arXiv:2506\.06884*, in press, 2025\.
- \[173\]B\. Zheng, M\. Ma, Z\. Lin, and T\. Yang, “A hybrid early\-exit algorithm for large language models based on space alignment decoding \(spade\)”,*arXiv preprint arXiv:2507\.17618*, in press, 2025\.
- \[174\]A\. R\. Bambhaniya, H\. Wu, S\. Subramanian, S\. Srinivasan,*et al\.*, “Understanding and optimizing multi\-stage ai inference pipelines”,*arXiv preprint arXiv:2504\.09775*, in press, 2025\.
- \[175\]Y\. Yao, H\. Jin, A\. D\. Shah, S\. Han,*et al\.*, “Scalellm: A resource\-frugal llm serving framework by optimizing end\-to\-end efficiency”, in*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track*, pp\.279–289, 2024\.
- \[176\]C\. Tian, X\. Qin, K\. Tam, L\. Li,*et al\.*, “Clone: Customizing llms for efficient latency\-aware inference at the edge”,*arXiv preprint arXiv:2506\.02847*, in press, 2025\.
- \[177\]Y\. Wu, S\. Tang, C\. Yu, B\. Yang,*et al\.*, “Task scheduling in geo\-distributed computing: A survey”,*IEEE Transactions on Parallel and Distributed Systems*, in press, 2025\.
- \[178\]J\. Stojkovic, C\. Zhang, Í\. Goiri, J\. Torrellas, and E\. Choukse, “Dynamollm: Designing llm inference clusters for performance and energy efficiency”, in*2025 IEEE International Symposium on High Performance Computer Architecture \(HPCA\)*, IEEE, 2025\.
- \[179\]I\. T\. Kurniawan and B\. R\. Trilaksono, “Clusterfusion: Leveraging radar spatial features for radar\-camera 3d object detection in autonomous vehicles”,*IEEE Access*, vol\.11, pp\.121511–121528, 2023\.
- \[180\]K\. Hong, G\. Dai, J\. Xu, Q\. Mao,*et al\.*, “Flashdecoding\+\+: Faster large language model inference with asynchronization, flat gemm optimization, and heuristics”,*Proceedings of Machine Learning and Systems*, vol\.6, pp\.148–161, 2024\.
- \[181\]W\. Wan, R\. Kubendran, C\. Schaefer, S\. B\. Eryilmaz,*et al\.*, “A compute\-in\-memory chip based on resistive random\-access memory”,*Nature*, vol\.608, no\.7923, pp\.504–512, 2022\.
- \[182\]M\. Davies, “Lessons from loihi: Progress in neuromorphic computing”, in*2021 Symposium on VLSI Circuits*, IEEE, pp\.1–2, 2021\.
- \[183\]B\. J\. Shastri, A\. N\. Tait, T\. Ferreira de Lima, W\. H\. Pernice,*et al\.*, “Photonics for artificial intelligence and neuromorphic computing”,*Nature Photonics*, vol\.15, no\.2, pp\.102–114, 2021\.
- \[184\]Y\. Sun, H\. Ochiai, and H\. Esaki, “Decentralized deep learning for multi\-access edge computing: A survey on communication efficiency and trustworthiness”,*IEEE Transactions on Artificial Intelligence*, vol\.3, no\.6, pp\.963–972, 2021\.
- \[185\]D\. Guo, D\. Yang, H\. Zhang, J\. Song,*et al\.*, “Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning”,*arXiv preprint arXiv:2501\.12948*, in press, 2025\.
- \[186\]O\. O\. Oyewole and J\. F\. Joseph, “Sustainable ai and green computing: Reducing the environmental impact of large\-scale models with energy\-efficient techniques”,*International Journal of Scientific Research in Network Security and Communication*, vol\.13, no\.3, pp\.19–26, 2025\.
- \[187\]C\. Elsworth, K\. Huang, D\. Patterson, I\. Schneider,*et al\.*, “Measuring the environmental impact of delivering ai at google scale”,*arXiv preprint arXiv:2508\.15734*, in press, 2025\.
- \[188\]C\.\-J\. Wu, R\. Raghavendra, U\. Gupta, B\. Acun,*et al\.*, “Sustainable ai: Environmental implications, challenges and opportunities”,*Proceedings of machine learning and systems*, vol\.4, pp\.795–813, 2022\.
- \[189\]G\. Gupta, A\. Tomar, S\. Pamulaparthyvenkata, and A\. Balakrishnan, “A comprehensive review of machine learning techniques for smart grid optimization”,*Secure Energy Optimization: Leveraging Internet of Things and Artificial Intelligence for Enhanced Efficiency*, in press, pp\.29–60, 2025\.
- \[190\]D\. Rolnick, P\. L\. Donti, L\. H\. Kaack, K\. Kochanski,*et al\.*, “Tackling climate change with machine learning”,*ACM Computing Surveys \(CSUR\)*, vol\.55, no\.2, pp\.1–96, 2022\.
- \[191\]Y\. Li, D\. Choi, J\. Chung, N\. Kushman,*et al\.*, “Competition\-level code generation with alphacode”,*Science*, vol\.378, no\.6624, pp\.1092–1097, 2022\.
- \[192\]L\. Solovyeva, S\. Weidmann, and F\. Castor, “Ai\-powered, but power\-hungry? energy efficiency of llm\-generated code”, in*2025 IEEE/ACM Second International Conference on AI Foundation Models and Software Engineering \(Forge\)*, IEEE, pp\.49–60, 2025\.
- \[193\]P\. Deng, J\. Zhang, X\. Sheng, C\. Yan,*et al\.*, “Multi\-granularity class prototype topology distillation for class\-incremental source\-free unsupervised domain adaptation”, in*Proceedings of the Computer Vision and Pattern Recognition Conference*, pp\.30566–30576, 2025\.
- \[194\]Z\. Yang, L\. Li, J\. Zhang, T\. Wang,*et al\.*, “Domain shared and specific prompt learning for incremental monocular depth estimation”, in*Proceedings of the 32nd ACM International Conference on Multimedia*, pp\.8306–8315, 2024\.
- \[195\]J\. Yin, L\. Li, J\. Zhang, Y\. Gao,*et al\.*, “Progressive homeostatic and plastic prompt tuning for audio\-visual multi\-task incremental learning”, in*Proceedings of the IEEE/CVF International Conference on Computer Vision*, pp\.2022–2033, 2025\.
- \[196\]X\. Tao, X\. Hong, X\. Chang, S\. Dong,*et al\.*, “Few\-shot class\-incremental learning”, in*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pp\.12183–12192, 2020\.
- \[197\]J\. Gidlund, “Green ai: Reducing the carbon footprint of automated testing”, 2025\.

Similar Articles

Is AI at this scale actually sustainable?

Reddit r/artificial

This article questions the sustainability of large-scale AI datacenters, discussing water and energy demands, and evaluating potential solutions like orbital datacenters and efficiency improvements.

Is AI ever going to become resource efficient?

Reddit r/ArtificialInteligence

A discussion questioning the long-term sustainability of AI models due to high compute costs and reliance on investor funding, pondering whether resource efficiency improvements can prevent a bubble burst.

Is This Sustainable?

Hacker News Top

A senior engineer reflects on three years of deep AI integration in software development, noting the collapse of the idea-to-demo gap and the shift of bottlenecks from engineering to coordination, while raising concerns about sustainability and unequal access to AI tools.