DxPTA: An Architecture Design Space Exploration with Optical Dataflow-guided Strategy for HW/SW Co-Design of Photonic Transformer Accelerators
Summary
This paper proposes DxPTA, a novel design space exploration methodology for efficient HW/SW co-design of photonic transformer accelerators that meet area, power, energy, and latency constraints. It achieves up to 15.2x faster searching time than exhaustive approaches, enabling efficient PTA designs for diverse transformer models.
View Cached Full Text
Cached at: 06/08/26, 09:15 AM
# DxPTA: An Architecture Design Space Exploration with Optical Dataflow-guided Strategy for HW/SW Co-Design of Photonic Transformer Accelerators
Source: [https://arxiv.org/html/2606.06515](https://arxiv.org/html/2606.06515)
Rachmad Vidya Wicaksana Putra, Solomon Micheal Serunjogi, Mahmoud Rasras, and Muhammad ShafiqueRachmad Vidya Wicaksana Putra is with eBRAIN Lab, Division of Engineering, New York University \(NYU\) Abu Dhabi, United Arab Emirates; \(e\-mail: rachmad\.putra@nyu\.edu\)\. Solomon Micheal Serunjogi is with Photonic Research Lab \(PRL\), Division of Engineering, New York University \(NYU\) Abu Dhabi, United Arab Emirates; \(e\-mail: sms10215@nyu\.edu\)\. Mahmoud Rasras is the Director of Photonic Research Lab \(PRL\), Division of Engineering, New York University \(NYU\) Abu Dhabi, United Arab Emirates \(UAE\); \(e\-mail: mrasras@nyu\.edu\)\. Muhammad Shafique is the Director of eBRAIN Lab, Division of Engineering, New York University \(NYU\) Abu Dhabi, United Arab Emirates; \(e\-mail: muhammad\.shafique@nyu\.edu\)\.
###### Abstract
Transformer\-based networks have emerged as prominent AI models with state\-of\-the\-art performance, which potentially pave the way toward artificial general intelligence \(AGI\)\. However, their large sizes still hinder their efficient implementation, thus highlighting the need for alternate solutions to enable their energy\-efficient acceleration\. Recently, state\-of\-the\-art works propose photonic transformer accelerators \(PTAs\) with significant speedup and energy efficiency improvements over the conventional electronic accelerators\. However, their PTA architectures are developed without considering the application constraints \(e\.g\., area, power, energy, and latency\)\. Moreover, their manual design approach also requires huge design time to determine a suitable architecture for the targeted application, hence making this approach not scalable\. To address these limitations, we proposeDxPTA, a novel design space exploration methodology for enabling efficient hardware/software co\-design of the appropriate PTA architecture that meets all constraints\. It is achieved by \(1\) identifying the PTA architecture parameters based on the coherent optical dataflow; \(2\) analyzing the impact/significance of the parameters; and \(3\) leveraging this analysis for devising a constraint\-aware architecture search algorithm\. Experimental results show that, our DxPTA can find the appropriate PTA architectures for different transformer\-based models \(i\.e\., DeiT\-T/S/B and BERT\-B/L\)\. It achieves up to 26mm2area, 4\.8W power, 39mJ energy, and 6ms latency, for constraints of 50mm2area, 5W power, 50mJ energy, and 10ms latency; with 15\.2x faster searching time than the exhaustive approach\. These results demonstrate the potential of DxPTA methodology for enabling efficient PTA designs for diverse AGI\-based applications\.
## IIntroduction
Transformer\-based network models\[[31](https://arxiv.org/html/2606.06515#bib.bib58)\], such as Vision Transformers \(ViTs\) and Large Language Models \(LLMs\), have emerged as prominent AI models with state\-of\-the\-art performance for solving diverse machine learning tasks, such as vision and natural language processing \(NLP\)\[[6](https://arxiv.org/html/2606.06515#bib.bib57),[30](https://arxiv.org/html/2606.06515#bib.bib33),[11](https://arxiv.org/html/2606.06515#bib.bib55),[9](https://arxiv.org/html/2606.06515#bib.bib56)\], hence potentially paving the way toward artificial general intelligence \(AGI\)\[[18](https://arxiv.org/html/2606.06515#bib.bib4)\]\[[33](https://arxiv.org/html/2606.06515#bib.bib5)\]\. However, this state\-of\-the\-art performance comes at high computational and memory requirements as shown in Fig\.[1](https://arxiv.org/html/2606.06515#S1.F1)\(a\), hence leading to huge power/energy consumption\[[9](https://arxiv.org/html/2606.06515#bib.bib56)\]\. This condition makes it difficult to obtain high performance efficiency when processing transformer models for wide\-scale implementations across diverse applications\. A potential solution is employing specialized accelerators for expediting the inference of transformers, thus minimizing power consumption and improving energy efficiency\[[17](https://arxiv.org/html/2606.06515#bib.bib17),[22](https://arxiv.org/html/2606.06515#bib.bib14),[32](https://arxiv.org/html/2606.06515#bib.bib29),[26](https://arxiv.org/html/2606.06515#bib.bib30),[36](https://arxiv.org/html/2606.06515#bib.bib27),[10](https://arxiv.org/html/2606.06515#bib.bib3),[35](https://arxiv.org/html/2606.06515#bib.bib28)\]; see Fig\.[1](https://arxiv.org/html/2606.06515#S1.F1)\(b\)\. However, conventional electronic accelerators face challenges related to their diminishing performance efficiency \(i\.e\., increased power dissipation\-per\-unit area and slower performance gains\), as transistor circuits reach the limits of Dennard scaling\[[28](https://arxiv.org/html/2606.06515#bib.bib13)\]\.
Recently, electronic\-photonic integrated circuit \(EPIC\)\-based solutions, so\-calledphotonic accelerators\[[25](https://arxiv.org/html/2606.06515#bib.bib19),[23](https://arxiv.org/html/2606.06515#bib.bib24),[8](https://arxiv.org/html/2606.06515#bib.bib26),[34](https://arxiv.org/html/2606.06515#bib.bib8)\], have been studied as an alternative for achieving significant speedup and efficiency improvements over the electronic accelerators, due to their ultra\-high speed, high bandwidth, and low energy consumption\[[8](https://arxiv.org/html/2606.06515#bib.bib26)\]\. Therefore, employments of photonic accelerators for expediting neural network \(NN\) workloads are actively being explored\. For instance, photonic tensor core \(PTC\) developments leveraging optical components such as Mach\-Zehnder Interferometer \(MZI\)\[[24](https://arxiv.org/html/2606.06515#bib.bib23)\], Micro\-Ring Resonator \(MRR\)\-based bank\[[29](https://arxiv.org/html/2606.06515#bib.bib21)\]\[[27](https://arxiv.org/html/2606.06515#bib.bib22)\], and Phase Change Material \(PCM\)\-based crossbar\[[7](https://arxiv.org/html/2606.06515#bib.bib20)\]\. However, these works still target in accelerating traditional convolutional neural networks \(CNN\) workloads\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\], hence indicating the need for further studies to enable high\-performance and energy\-efficient inference of transformer models\. Therefore, thetargeted research problemin this paper ishow can we effectively enable high performance and energy\-efficient inference of transformer models using photonic\-based accelerators? A solution to this problem may enable the efficient deployments of transformer\-based models on photonic\-based computing systems\.
Figure 1:\(a\)Transformer\-based models typically can improve the performance at the cost of larger memory size \(i\.e\., higher number of parameters\); based on the data from\[[9](https://arxiv.org/html/2606.06515#bib.bib56)\]\.\(b\)Experimental results of running Data\-efficient Image Transformer Base \(DeiT\-B\)\[[30](https://arxiv.org/html/2606.06515#bib.bib33)\]with different compute platforms: CPU, GPU, CMOS\-based accelerators \(i\.e\., AutoViT\-4bit\[[16](https://arxiv.org/html/2606.06515#bib.bib25)\]and HeatViT\-8bit\[[5](https://arxiv.org/html/2606.06515#bib.bib31)\]\) and photonic\-based Lightening\-Transformer \(LT\) accelerators \(i\.e\., LT\-Base\-4bit and LT\-Large\-4bit\); based on data from\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]\.### I\-AState\-of\-the\-art of Photonic Transformer Accelerators \(PTAs\) and Their Limitations
Recent works propose PTA designs that employ statically\-operated PTCs using MRR banks\[[14](https://arxiv.org/html/2606.06515#bib.bib7),[13](https://arxiv.org/html/2606.06515#bib.bib6),[3](https://arxiv.org/html/2606.06515#bib.bib15),[1](https://arxiv.org/html/2606.06515#bib.bib10),[2](https://arxiv.org/html/2606.06515#bib.bib12)\]and PCM crossbars\[[15](https://arxiv.org/html/2606.06515#bib.bib9)\]\. Another work proposes theLightening\-Transformer \(LT\)\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]based on dynamically\-operated PTCs\. It inspires further studies in digital\-to\-analog converter \(DAC\) design\[[12](https://arxiv.org/html/2606.06515#bib.bib11),[4](https://arxiv.org/html/2606.06515#bib.bib18)\]and reconfigurability\[[38](https://arxiv.org/html/2606.06515#bib.bib16)\]\. LT improves the performance and efficiency of transformer inference over other PTAs by enabling dynamic operations of full\-range input operands, making it the state\-of\-the\-art PTA design\. Despite their benefits, all these works still have the following limitations\.
- •Their architectures are developed without considering application constraints \(e\.g\., area, power, energy, and latency\), and hence their designs are not directly applicable for targeted applications and leading to sub\-optimal performance and efficiency gains\.
- •Their manual design approach needs a huge design time and power/energy consumption to develop a suitable architecture for the targeted applications, thus making this approach not scalable\.
To illustrate the limitations of state\-of\-the\-arts and related research challenges, we perform an experimental case study, which will be discussed in Section[I\-B](https://arxiv.org/html/2606.06515#S1.SS2)\.
### I\-BCase Study and Research Challenges
Figure 2:Experimental results considering different configurations of architecture parameters \(i\.e\.,NtN\_\{t\}andNcN\_\{c\}\) in the 4\-bit LT accelerator for\(a\)power and area; and\(b\)energy consumption and latency considering the DeiT\-Base\[[30](https://arxiv.org/html/2606.06515#bib.bib33)\]\.We explore the impact of different architecture parameters of the state\-of\-the\-art 4\-bit LT accelerator\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]\. Here, we vary the number of tiles \(NtN\_\{t\}\) and the number of cores\-per\-tile \(NcN\_\{c\}\)\. For workload, we consider the DeiT\-Base model\[[30](https://arxiv.org/html/2606.06515#bib.bib33)\]\. Details of the LT hardware architecture and the experimental setup are provided in Section[II\-B](https://arxiv.org/html/2606.06515#S2.SS2)and Section[IV](https://arxiv.org/html/2606.06515#S4), respectively\. Experimental results are shown in Fig\.[2](https://arxiv.org/html/2606.06515#S1.F2), from which we make the following key observations\.
- •Different configurations lead to different profiles of area, power, energy, and latency, thus highlighting the wide range of design choices for developing photonic accelerators\.
- •IncreasingNtN\_\{t\}orNcN\_\{c\}leads to higher power and larger area due to more complex circuitry \(seeA\), but it may reduce latency and energy consumption due to increased parallelism \(seeB\)\. This shows the need for trade\-off analysis in accelerator design\.
- •The state\-of\-the\-art LT design \(withNtN\_\{t\}=4 andNcN\_\{c\}=2\) may not meet the constraints\. For instance, in low\-power applications with max\. 5W, the design incurs significantly more power \(∼\\sim15W\) as shown byC, indicating the need for a custom architecture\.
These observations expose several research challenges in devising solutions for the targeted research problem, as outlined below\.
- •The solution should leverage the characteristics of photonic devices and optical dataflow to find the PTA architecture that meets all constraints, ensuring its applicability for diverse applications\.
- •The solution should minimize the searching time of PTA architecture, hence expediting the design time and providing a scalable design approach for diverse applications\.
### I\-COur Novel Contributions
To address the targeted research problem and related challenges, we proposeDxPTA,a novel architectureDesign space exploration methodology leveraging coherent optical dataflow\-guided strategy for efficient hardware/software \(HW/SW\) co\-design ofPhotonicTransformerAccelerators while meeting multiple constraints \(i\.e\., area, power, energy, and latency\)\. It employs the following key steps \(see an overview in Fig\.[3](https://arxiv.org/html/2606.06515#S1.F3)and details in Fig\.[4](https://arxiv.org/html/2606.06515#S1.F4)\)\.
- •Identify the architecture parameters of PTA \(Section[III\-A](https://arxiv.org/html/2606.06515#S3.SS1)\):It targets to analyze the characteristics of PTA architecture \(including its hierarchy and photonic devices\) and its coherent optical dataflow for identifying the prominent architecture parameters\.
- •Analyze the impact of architecture parameters \(Section[III\-B](https://arxiv.org/html/2606.06515#S3.SS2)\):It identifies the significance of architecture parameters \(e\.g\.,NtN\_\{t\}andNcN\_\{c\}\) by observing their impact on area, power, energy, and latency\. The information will be leveraged for architecture search\.
- •Devise the constraint\-aware search algorithm \(Section[III\-C](https://arxiv.org/html/2606.06515#S3.SS3)\):It explores architecture candidates by leveraging the coherent optical dataflow and parameter significance, evaluates their energy\-delay products \(EDPs\), and then select the one that has the lowest EDP and meets all constraints\.
Key Results:We evaluate our DxPTA methodology using Python implementation, and then run it on an Nvidia RTX 6000 Ada GPU machine, while considering diverse transformer workloads \(i\.e\., DeiT\-T/S/B111DeiT\-T/S/B refers to DeiT\-Tiny, DeiT\-Small, and DeiT\-Base, respectively\.and BERT\-B/L222BERT\-B/L refers to BERT\-Base and BERT\-Large, respectively\.\)\. Experimental results show that, DxPTA successfully finds the accelerator architectures that meet all constraints\. It achieves up to 26mm2area, 4\.8W power, 39mJ energy, and 6ms latency across all investigated models, for constraints of 50mm2area, 5W power, 50mJ energy, and 10ms latency, with 15\.2x faster searching time than the exhaustive approach\.
Figure 3:Our novel contributions in this work\.Figure 4:Our novel DxPTA methodology with its key steps: \(1\) identifying of the architecture parameters of PTA; \(2\) analyzing the impact of architecture parameters; and \(3\) devising the constraint\-aware search algorithm\.
## IIPreliminaries
### II\-ATransformer\-based Networks
A transformer\-based network typically consists of multiple identical blocks, known asencoderanddecoderblocks\. Each block consists of a multi\-head self\-attention \(MHA\) module, a feed\-forward network \(FFN\), shortcut connections, as well as a layer normalization \(LN\)\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]\. Furthermore, the decoder block also has cross\-attention and masked self\-attention modules\. The basic encoder block can be formulated as Eq\.[1](https://arxiv.org/html/2606.06515#S2.E1)\-[2](https://arxiv.org/html/2606.06515#S2.E2)\. Here,Xl\\textbf\{X\}\_\{l\}is the input sequences ofll\-th layer\.
X^l\+1=MHA\(LN\(Xl\)\)\+Xl\\begin\{split\}\\hat\{\\textbf\{X\}\}\_\{l\+1\}=\\textit\{MHA\}\(\\textit\{LN\}\(\\textbf\{X\}\_\{l\}\)\)\+\\textbf\{X\}\_\{l\}\\end\{split\}\(1\)Xl\+1=FFN\(LN\(X^l\+1\)\)\+X^l\+1\\begin\{split\}\\textbf\{X\}\_\{l\+1\}=\\textit\{FFN\}\(\\textit\{LN\}\(\\hat\{\\textbf\{X\}\}\_\{l\+1\}\)\)\+\\hat\{\\textbf\{X\}\}\_\{l\+1\}\\end\{split\}\(2\)Multi\-head self\-attention \(MHA\) module hasHHself\-attention heads, where each head transforms the input vector into separate vectors: query \(Q\), key \(K\), and value \(V\) vectors\. The attention function between these input vectors can be calculated using Eq\.[3](https://arxiv.org/html/2606.06515#S2.E3)\. Here,dkd\_\{k\}is the dimension ofQandK\.
Attention\(Q,K,V\)=softmax\(QK⊺dk\)VAttention\(\\textbf\{Q\},\\textbf\{K\},\\textbf\{V\}\)=softmax\\left\(\\frac\{\\textbf\{Q\}\\textbf\{K\}^\{\\intercal\}\}\{\\sqrt\{d\_\{k\}\}\}\\right\)\\textbf\{V\}\(3\)
### II\-BPhotonic Transformer Accelerator \(PTA\)
Figure 5:Architecture of the LT accelerator; based on\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]\.In this work,we focus on the LT accelerator\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]as the reference design since it is the state\-of\-the\-art PTA architecture, whose descriptions are provided in the following; see an overview in Fig\.[5](https://arxiv.org/html/2606.06515#S2.F5)\. The LT Accelerator consists of analog photonic computing elements for accelerating general matrix multiplication \(GEMM\), optical interconnect for data transmission, and electronics for other operations \(e\.g\., data storage, signal conversion, nonlinear functions, and softmax\); see Fig\.[5](https://arxiv.org/html/2606.06515#S2.F5)\. Following are its key design points\.
- •A single LT chip typically containsNtN\_\{t\}tiles, and each tile clustersNcN\_\{c\}dynamically\-operated photonic tensor cores \(DPTCs\)\.
- •A DPTC is known as thecoreof LT accelerator, and it contains an array ofNhN\_\{h\}×\\timesNvN\_\{v\}dynamically\-operated dot\-product engine \(DDot\)\.
- •A single DDot performs optical dot\-product operation between two full\-range vectorsx→\\vec\{x\}andy→\\vec\{y\}based on coherent interference\.
Dot\-product operation in DDot is performed with the following steps\. First, each pair of inputs \(xi,yix\_\{i\},y\_\{i\}\) is encoded in the same wavelengthλi\\lambda\_\{i\}through thewavelength\-division multiplexing \(WDM\)technique\. These signals are passed to the two arms of 50:50directional coupler \(DC\)with \-90°\\degreephase shifter \(PS\)\. The outputs of DC \(zi0,zi1z^\{0\}\_\{i\},z^\{1\}\_\{i\}\) are orthogonal in complex plane and can be computed with Eq\.[4](https://arxiv.org/html/2606.06515#S2.E4)\.
\(zi0zi1\)=12\(1jj1\)⏟DC\(100e−jπ/2\)⏟PS\(xiyi\)=12\(xi\+yij\(xi−yi\)\)\\begin\{split\}\\begin\{pmatrix\}z\_\{i\}^\{0\}\\\\ z\_\{i\}^\{1\}\\end\{pmatrix\}&=\\underbrace\{\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{pmatrix\}1&j\\\\ j&1\\end\{pmatrix\}\}\_\{DC\}\\underbrace\{\\begin\{pmatrix\}1&0\\\\ 0&e^\{\-j\\pi/2\}\\end\{pmatrix\}\}\_\{PS\}\\begin\{pmatrix\}x\_\{i\}\\\\ y\_\{i\}\\end\{pmatrix\}\\\\ &=\\frac\{1\}\{\\sqrt\{2\}\}\\begin\{pmatrix\}x\_\{i\}\+y\_\{i\}\\\\ j\(x\_\{i\}\-y\_\{i\}\)\\end\{pmatrix\}\\end\{split\}\(4\)The photodiode \(PD\) at each output port of DC, converts the signals into photocurrent, which is proportional to the accumulated optical intensities of the input signals\. Hence, the output current \(I0I\_\{0\}\) follows the relation ofI0∝x→⋅y→I\_\{0\}\\propto\\vec\{x\}\\cdot\\vec\{y\}\.
Figure 6:Data tiling mechanism in the accelerator\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\], which partitions data from matrixM1M\_\{1\}along theDvD\_\{v\}dimension and map them to different tiles\.
## IIIOur DxPTA Methodology
This methodology identifies the architecture parameters, analyzes the impact of parameters, and develops the constraint\-aware search algorithm; which are further described below \(an overview in Fig\.[4](https://arxiv.org/html/2606.06515#S1.F4)\)\.
Figure 7:\(a\)Power and area for different configurations of parameters\.\(b\)Energy consumption and latency for different configurations of parameters when running DeiT\-B\. We observe similar trends for different workloads \(DeiT\-T/S and BERT\-B/L\)\.### III\-AIdentifying the Architecture Parameters
Discussion in Section[I\-B](https://arxiv.org/html/2606.06515#S1.SS2)suggests that the configuration of architecture parameters is important for determining the performance and efficiency of the PTA\. Therefore,this step aims to identify parameters that should be customized when designing the PTA\. To achieve this, we first analyze the hierarchy of the PTA base architecture and its coherent optical dataflow, and make the following key observations\.
- •Combining the dynamically\-operated DDots and the coherent optical dataflow enables multi\-wavelength processing which maximizes spectral parallelism and throughput\. From such coherent dataflow and operations, we observe some characteristics below\. - –Multiple wavelengths can be processed in the same DDot unit without requiring prior data programming; seeAin Fig\.[5](https://arxiv.org/html/2606.06515#S2.F5)\. - –AnNhN\_\{h\}×\\timesNvN\_\{v\}DDot array is employed to perform multiplications, which defines the parallelism level in a core; seeBin Fig\.[5](https://arxiv.org/html/2606.06515#S2.F5)\. - –Multiple cores can be employed to process a chunk of data, which defines the parallelism level in a tile; seeCin Fig\.[6](https://arxiv.org/html/2606.06515#S2.F6)\. - –Multiple tiles can be employed to process multiple data chunks, which defines the parallelism level in a chip; seeDin Fig\.[6](https://arxiv.org/html/2606.06515#S2.F6)\.
- •On\-chip global SRAM should have the minimum required size for holding the largest activations in a layer due to layer\-by\-layer processing, and buffering a portion of off\-chip data based on the tiling approach \(see Fig\.[6](https://arxiv.org/html/2606.06515#S2.F6)\)\. Hence, global SRAM size should not be reduced below or increased significantly over this minimum required size, as this will increase the expensive off\-chip data access \(i\.e\., high access latency and access energy\) or aggravate the static power, respectively\[[19](https://arxiv.org/html/2606.06515#bib.bib143),[20](https://arxiv.org/html/2606.06515#bib.bib107),[21](https://arxiv.org/html/2606.06515#bib.bib2)\]\.
These observations expose the parameters that should be configured for developing an appropriate accelerator architecture, i\.e\.,number of tiles \(NtN\_\{t\}\),number of cores\-per\-tile \(NcN\_\{c\}\),number of input horizontal waveguides\-per\-core \(NhN\_\{h\}\),number of input vertical waveguides\-per\-core \(NvN\_\{v\}\), andnumber of wavelengths \(NλN\_\{\\lambda\}\)\.
### III\-BAnalyzing the Impact of Parameters
To ensure the accelerator meets the constraints, an appropriate configuration of parameters \(i\.e\.,NtN\_\{t\},NcN\_\{c\},NhN\_\{h\},NvN\_\{v\}, andNλN\_\{\\lambda\}\) is needed\. A promising solution is employing design space exploration \(DSE\)\. However, the design space is large due to a high number of possible configurations from different parameter sizes, indicating the need for an optimization\. Hence,this step aims to identify the significance of parameters, which will be used to efficiently guide the DSE process\. To achieve this, we propose to employ the following steps\.
- •We perform an experimental case study that varies the value of a specific parameter and analyze its impact on different metrics, i\.e\., area, power, energy, and latency\.
- •Afterward, we evaluate the significance score \(SS\) for each parameter on a specific metric using Eq\.[5](https://arxiv.org/html/2606.06515#S3.E5), whose mechanism is also presented as pseudocode in Alg\.[1](https://arxiv.org/html/2606.06515#alg1)\. Here,sis\_\{i\}is the ratio between the values of metric\-mm\(i\.e\., areaAAor powerPP\) from architecture withii\+1 units and architecture withiiunits, which represents the impact of unit addition; whileKKis the total number of ratios\. S=1K∑i=1Ksi=1K∑i=1Kmi\+1unitsmiunitswithm∈\{A,P\}\\begin\{split\}S=&\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}s\_\{i\}=\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\\frac\{m\_\{i\+1\\text\{ units\}\}\}\{m\_\{i\\text\{ units\}\}\}\\\\ &\\text\{with\}\\;\\;\\;m\\in\\\{A,P\\\}\\end\{split\}\(5\)
Experimental results are shown in Fig\.[7](https://arxiv.org/html/2606.06515#S3.F7), from which we make the following key observations\.
- •Increasing parameter size leads to higher power and area, but potentially reduces latency and energy due to higher parallelism\.
- •NtN\_\{t\}has the highest impact as each additional tile leads to the highest significanceSSwith 1\.26x higher power and 1\.24x larger area on average; seeE\. Introducing a new tile means addition of single/multiple cores and peripherals \(e\.g\., DAC and ADC\)\.
- •NcN\_\{c\}has relatively high impact as each additional core leads to high significanceSSwith 1\.23x higher power and 1\.20x larger area on average; seeF\. Introducing a new core means addition of a DDot array with increased size of peripherals \(e\.g\., accumulator\)\.
- •NvN\_\{v\},NhN\_\{h\}, orNλN\_\{\\lambda\}have comparable impact to each other, but they are lower thanNtN\_\{t\}andNcN\_\{c\}\. For each additional unit ofNvN\_\{v\},NhN\_\{h\}, orNλN\_\{\\lambda\}, power and area are increased by up to 1\.16x and 1\.06x, respectively; seeG\. Introducing new DDots within the core or new wavelengths only slightly increases circuit complexity and size \(e\.g\., broadcast unit in the array\)\.
Algorithm 1Observing the significance of architecture parameters0:Maximum number of observations \(
JJ=10\);
0:Significance score of each parameter on area \(
SAS\_\{A\}\) and power \(
SPS\_\{P\}\);BEGINProcess:
1:foreach investigated parameterdo
2:
NtN\_\{t\}= 4;
NcN\_\{c\}= 2;
NvN\_\{v\}= 12;
NhN\_\{h\}= 12;
NλN\_\{\\lambda\}= 12;// init default values
3:for\(
jj=1;
jj<<\(
JJ\+1\);
jj\+\+\)do
4:set the investigated parameter value with
jj;
5:
cfg\[j\]cfg\[j\]= construct\(
NtN\_\{t\},
NcN\_\{c\},
NvN\_\{v\},
NhN\_\{h\},
NλN\_\{\\lambda\}\);
6:
A\[j\]A\[j\],
P\[j\]P\[j\]= eval\_hw\(
cfg\[j\]cfg\[j\]\);
7:
ii=
jj\-1;
8:if\(
ii\>\>0\)then
9:
sA\[i\]s\_\{A\}\[i\]=
A\[iA\[i\+
1\]/A\[i\]1\]/A\[i\];// compute ratio for area
10:
sP\[i\]s\_\{P\}\[i\]=
P\[iP\[i\+
1\]/P\[i\]1\]/P\[i\];// compute ratio for power
11:
KK=
JJ\-1;
12:
SAS\_\{A\}=
1K∑i=1KsA\[i\]\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}s\_\{A\}\[i\];// compute significance score for area with Eq\.[5](https://arxiv.org/html/2606.06515#S3.E5)
13:
SPS\_\{P\}=
1K∑i=1KsP\[i\]\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}s\_\{P\}\[i\];// compute significance score for power with Eq\.[5](https://arxiv.org/html/2606.06515#S3.E5)
14:return
SAS\_\{A\},
SPS\_\{P\};END
### III\-CThe Constraint\-aware Search Algorithm
Based on observations in Section[III\-B](https://arxiv.org/html/2606.06515#S3.SS2), we develop the following optimization strategy for DSE process\.
- •ExploringNtN\_\{t\}andNcN\_\{c\}should be performed carefully since their slight changes may incur significant changes on area and power\.
- •ExploringNvN\_\{v\},NhN\_\{h\}, andNcN\_\{c\}may be performed more aggressively thanNtN\_\{t\}andNcN\_\{c\}, since their slight changes do not incur significant changes on area and power consumption\.
- •The data dimension is typically evenly sized across layers, which should be leveraged to maximize the resource utilization by selecting the evenly\-sized dimension for architecture parameters\.
We leverage this strategy for devisinga constraint\-aware search algorithm\. Its key ideas are shown in Alg\.[2](https://arxiv.org/html/2606.06515#alg2)and discussed below\.
- •We determine a set of values for each parameter, which defines its search space; see Alg\.[2](https://arxiv.org/html/2606.06515#alg2): lines 3\-10\. It is stored inTcndT\_\{cnd\},CcndC\_\{cnd\},VcndV\_\{cnd\},HcndH\_\{cnd\},GcndG\_\{cnd\}forNtN\_\{t\},NcN\_\{c\},NvN\_\{v\},NhN\_\{h\},NλN\_\{\\lambda\}, respectively\. The search spaces forTcndT\_\{cnd\}andCcndC\_\{cnd\}employ incremental values, while the ones forVcndV\_\{cnd\},HcndH\_\{cnd\}andGcndG\_\{cnd\}are optimized using progressive values with explorationstepstepbased on evenly\-sized data dimension\.
- •Then, each configuration candidate for the architecture is explored\. It is performed by investigating each combination of values from different parameters, and evaluating the area, power, energy, and latency profiles; see Alg\.[2](https://arxiv.org/html/2606.06515#alg2): lines 11\-14\.
- •We evaluate if these profiles meet all constraints and if the energy\-delay product \(EDP\) is lower than the recorded one\. If so, then the configuration is saved; see Alg\.[2](https://arxiv.org/html/2606.06515#alg2): lines 15\-22\. Here, EDP is the metric for reflecting both performance and energy efficiency\.
- •When the DSE process is finished, the last recorded configuration \(cfgsvdcfg\_\{svd\}\) is selected as the final solution; see Alg\.[2](https://arxiv.org/html/2606.06515#alg2): line 23\.
Algorithm 2Our constraint\-aware search algorithm0:\(1\)Maximum number of parameter sizes \(
NzN\_\{z\}=12\);\(2\)Progressive exploration step for non\-significant parameters \(
stepstep=2\);\(3\)Targeted pre\-trained network model \(
netnet\);\(4\)Constraints for area \(
constAconst\_\{A\}\), power \(
constPconst\_\{P\}\), energy \(
constEconst\_\{E\}\), and latency \(
constLconst\_\{L\}\);
0:\(1\)Final configuration of the architecture parameters \(
cfgsvdcfg\_\{svd\}\);BEGINInitialization:
1:
Z1Z\_\{1\}= \[\];
Z2Z\_\{2\}= \[\];
2:
EDPsvdEDP\_\{svd\}= 1000;Process:// Define the search space for each parameter
3:for\(
nzn\_\{z\}=1;
nzn\_\{z\}<<\(
NzN\_\{z\}\+1\);
nzn\_\{z\}\+\+\)do
4:
Z1Z\_\{1\}= append\(
Z1Z\_\{1\},
nzn\_\{z\}\);
5:if\(
nzn\_\{z\}mod
stepstep== 0\)then
6:
Z2Z\_\{2\}= append\(
Z2Z\_\{2\},
nzn\_\{z\}\);
7:
TcndT\_\{cnd\}=
Z1Z\_\{1\};
CcndC\_\{cnd\}=
Z1Z\_\{1\};
8:
VcndV\_\{cnd\}=
Z2Z\_\{2\};
HcndH\_\{cnd\}=
Z2Z\_\{2\};
GcndG\_\{cnd\}=
Z2Z\_\{2\};// Construct and evaluate the configurations
9:foreach combination from \(
∀\\forallntn\_\{t\}∈\\inTcndT\_\{cnd\}\), \(
∀\\forallncn\_\{c\}∈\\inCcndC\_\{cnd\}\), \(
∀\\forallnvn\_\{v\}∈\\inVcndV\_\{cnd\}\), \(
∀\\forallnhn\_\{h\}∈\\inHcndH\_\{cnd\}\), and \(
∀\\forallnλn\_\{\\lambda\}∈\\inGcndG\_\{cnd\}\)do
10:
cfgcndcfg\_\{cnd\}= construct\(
NtN\_\{t\},
NcN\_\{c\},
NvN\_\{v\},
NhN\_\{h\},
NλN\_\{\\lambda\}\);
11:
AcndA\_\{cnd\},
PcndP\_\{cnd\}= eval\_hw\(
cfgcndcfg\_\{cnd\}\);// evaluate areaAand powerP
12:
EcndE\_\{cnd\},
LcndL\_\{cnd\}= eval\_wload\(
cfgcndcfg\_\{cnd\},
netnet\);// evaluate energyEand latencyL
13:if\(
AcndA\_\{cnd\}<<constAconst\_\{A\}\) and \(
PcndP\_\{cnd\}<<constPconst\_\{P\}\) and \(
EcndE\_\{cnd\}<<constEconst\_\{E\}\) and \(
LcndL\_\{cnd\}<<constLconst\_\{L\}\)then
14:
EDPcndEDP\_\{cnd\}= calc\_EDP \(
EcndE\_\{cnd\},
LcndL\_\{cnd\}\);// evaluate EDP
15:if\(
EDPcndEDP\_\{cnd\}<<EDPsvdEDP\_\{svd\}\)then
16:
EDPsvdEDP\_\{svd\}=
EDPcndEDP\_\{cnd\};
17:
cfgsvdcfg\_\{svd\}=
cfgcndcfg\_\{cnd\};
18:return
cfgsvdcfg\_\{svd\};// finalNtN\_\{t\},NcN\_\{c\},NvN\_\{v\},NhN\_\{h\}, andNλN\_\{\\lambda\}END
## IVEvaluation Methodology
We implement the DxPTA methodology using PyTorch and run it on the Nvidia RTX 6000 Ada GPU machine; see Fig\.[8](https://arxiv.org/html/2606.06515#S4.F8)\. For hardware evaluation, we employ the state\-of\-the\-art PTA hardware simulator\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]with 4\-bit precision that has been evaluated using the Lumerical Interconnect tools, and incorporate it into the DxPTA\. For workloads, we use DeiT\-T, DeiT\-S, DeiT\-B, BERT\-B, and BERT\-L models\. For comparison partners, we consider the LT accelerators \(i\.e\., LT\-Base and LT\-Large\)\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]as the state\-of\-the\-art PTA designs, and an exhaustive approach as the search technique\. For constraints, we consider 50mm2area, 5W power, 50mJ energy, and 10ms latency, to show the applicability of DxPTA for providing solutions under any application requirements\. Evaluation metrics include area, power, energy, latency, and search time\.
Figure 8:Experimental setup used in this work\.
## VExperimental Results and Discussion
### V\-AEnsuring the PTA Architecture Design to Meet All Design Constraints
Fig\.[9](https://arxiv.org/html/2606.06515#S5.F9)provides experimental results of area, power, energy consumption, and latency for different samples of investigated architecture configurations during DSE process across different workloads\. These results show that, the state\-of\-the\-art accelerators \(i\.e\., LT\-Base and LT\-Large\) do not meet all constraints at once, as they incur significantly larger area and higher power than the respective constraints; see1\. Specifically, LT\-Base and LT\-Large incur about 60mm2and 112mm2, respectively \(\>\>50mm2constraint\)\. These sizes are dominated by memory, DAC, and cores; see Fig\.[10](https://arxiv.org/html/2606.06515#S5.F10)\(a\)\. In terms of power, LT\-Base and LT\-Large incur about 15W and 28W power, respectively \(\>\>5W constraint\)\. These power values are dominated by Mach\-Zender Modulator \(MZM\), DAC, photodetector, and ADC; see Fig\.[10](https://arxiv.org/html/2606.06515#S5.F10)\(b\)\. The reason is that, LT\-Base and LT\-Large designs have fixed configurations for accelerating diverse workloads, thus they may not be applicable for different requirements\.
In contrast, our DxPTA consistently finds the suitable configurations that meet all constraints and have the lowest EDP scores among the candidates across different workloads; see2for power and area,3for energy consumption,4for latency, and5for EDP\. The reason is that, DxPTA employs a search algorithm that incorporates all constraints in its exploration process, ensuring the selected configuration to fulfills the requirements\. Consequently, the DxPTA\-generated accelerators significantly reduce the area and power as compared to LT\-Base and LT\-Large, by achieving up to 76\.9% area saving and 82\.7% power saving; see6in Fig\.[10](https://arxiv.org/html/2606.06515#S5.F10)\. Furthermore, our DxPTA\-generated accelerators also achieve comparable area and power consumption to the accelerators whose configurations generated from exhaustive search \(i\.e\., Exh\-DeiT and Exh\-BERT\); see7in Fig\.[10](https://arxiv.org/html/2606.06515#S5.F10)\. The reason is that, DxPTA already considers the significance of parameters in its search strategy to ensure the coverage of potential configuration candidates in the DSE process\. Hence, DxPTA can find configurations that are close to the ones from the exhaustive search\.
Figure 9:Experimental results of\(a\)power vs\. area,\(b\)energy vs\. area,\(c\)latency vs\. area, and\(d\)EDP vs\. area for different samples of configurations during DSE process, across different workloads: DeiT\-T, DeiT\-S, DeiT\-B, BERT\-B, and BERT\-L\.Figure 10:Experimental results of\(a\)area and\(b\)power for LT\-Base, LT\-Large, exhaustive search\-based accelerators for DeiT/BERT models \(i\.e\., Exh\-DeiT/BERT\), and DxPTA\-based accelerators for DeiT/BERT models \(i\.e\., DxPTA\-DeiT/BERT\)\.
### V\-BEnabling High\-Performance and Energy\-Efficient Transformer Inference
Fig\.[11](https://arxiv.org/html/2606.06515#S5.F11)presents the performance \(i\.e\., FPS: frame\-per\-second\) and energy consumption of DeiT\-B processing using different platforms: CPU, GPU, electronic accelerators, state\-of\-the\-art PTAs, and our DxPTA\-based PTA\. These results show that, the DxPTA\-based PTA achieves comparable performance \(FPS\) to the LT\-Base and LT\-Large designs, and provides significant improvements from conventional platforms; see8\. Specifically, the DxPTA\-based PTA improves the performance by 189x from CPU, 4\.1x from GPU, 20\.1x from AutoVit\-4, and 17\.2x from HeatVIT\-8\. Such remarkable performance improvements come from the high\-speed nature of light propagation\. The DxPTA\-based PTA also achieves comparable energy\-efficiency to the LT\-Base and LT\-Large designs, and provides significant energy savings from conventional platforms; see9\. Specifically, the DxPTA\-based PTA saves energy consumption by 782\.1x from CPU, 15\.2x from GPU, 31\.6x from AutoVit\-4, and 27\.6x from HeatVIT\-8\. Such energy efficiency improvements come from its lightweight optical\-based processing and minimum energy consumption from electronic parts\. Note, all these improvements are achieved by DxPTA while meeting all given constraints at once, further highlighting the benefits of our DxPTA methodology\.
### V\-CArchitecture Searching Time Speedup
Fig\.[12](https://arxiv.org/html/2606.06515#S5.F12)presents the searching time of the exhaustive approach and the guided search in our DxPTA across different workloads\. These results highlight that DxPTA achieves a significant searching time speedup, i\.e\., by 15\.2x faster than the exhaustive one\. This speedup comes from the optimized search space in DxPTA through exploration steps, guided by data dimension, coherent optical dataflow, and parameter significance\. This speedup is beneficial to make the DxPTA methodology an efficient and scalable solution for developing appropriate PTA architectures under different possible design constraints\.
Figure 11:Experimental results of\(a\)performance and\(b\)energy consumption of DeiT\-B processing using CPU, GPU, electronic accelerators \(i\.e\., AutoViT\-4b\[[16](https://arxiv.org/html/2606.06515#bib.bib25)\]and HeatViT\-8b\[[5](https://arxiv.org/html/2606.06515#bib.bib31)\]\), state\-of\-the\-art PTAs \(i\.e\., LT\-Base and LT\-Large\[[37](https://arxiv.org/html/2606.06515#bib.bib32)\]\), and DxPTA\.Figure 12:Searching time profiles of the exhaustive search approach and the searching strategy in our DxPTA\.
## VIConclusion
We propose a novel DxPTA methodology to perform DSE for enabling efficient HW/SW co\-design of the PTA architecture that meets all given constraints\. DxPTA identifies the prominent architecture parameters based on the coherent optical dataflow, analyzes the significance of parameters, and then devises a constraint\-aware search algorithm\. Experiments show that, DxPTA successfully finds the appropriate PTA architectures for different transformer models, achieving up to 26mm2area, 4\.8W power, 39mJ energy, and 6ms latency, for constraints of 50mm2area, 5W power, 50mJ energy, and 10ms latency; with 15\.2x faster search time than the exhaustive approach\. These results demonstrate the potential of DxPTA methodology for enabling efficient PTA design automation for diverse AGI\-based applications\.
## References
- \[1\]\(2025\)A light\-speed large language model accelerator with optical stochastic computing\.InGreat Lakes Symposium on VLSI \(GLSVLSI\) 2025,pp\. 922–928\.Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[2\]S\. Afifi, O\. Alo, I\. Thakkar, and S\. Pasricha\(2025\)ASTRA: a stochastic transformer neural network accelerator with silicon photonics\.ACM Transactions on Embedded Computing Systems \(TECS\)\.Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[3\]S\. Afifi, F\. Sunny, M\. Nikdast, and S\. Pasricha\(2023\)Tron: transformer neural network acceleration with non\-coherent silicon photonics\.InGreat Lakes Symposium on VLSI \(GSVLSI\) 2023,pp\. 15–21\.Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[4\]W\. Chang, C\. Wu, and Y\. Lo\(2025\)P\-dac: power\-efficient photonic accelerators for llm inference\.In2025 62nd ACM/IEEE Design Automation Conference \(DAC\),Vol\.,pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.1109/DAC63849.2025.11132618)Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[5\]P\. Dong, M\. Sun, A\. Lu, Y\. Xie, K\. Liu, Z\. Kong, X\. Meng, Z\. Li, X\. Lin, Z\. Fang, and Y\. Wang\(2023\)HeatViT: hardware\-efficient adaptive token pruning for vision transformers\.In2023 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),Vol\.,pp\. 442–455\.External Links:[Document](https://dx.doi.org/10.1109/HPCA56546.2023.10071047)Cited by:[Figure 1](https://arxiv.org/html/2606.06515#S1.F1),[Figure 11](https://arxiv.org/html/2606.06515#S5.F11)\.
- \[6\]A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. Houlsby\(2021\)An image is worth 16x16 words: transformers for image recognition at scale\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[7\]J\. Feldmann, N\. Youngblood, M\. Karpov, H\. Gehring, X\. Li, M\. Stappers, M\. Le Gallo, X\. Fu, A\. Lukashchuk, A\. S\. Raja,et al\.\(2021\)Parallel convolutional processing using an integrated photonic tensor core\.Nature589\(7840\),pp\. 52–58\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[8\]J\. Gu, C\. Feng, H\. Zhu, R\. T\. Chen, and D\. Z\. Pan\(2022\)Light in ai: toward efficient neurocomputing with optical neural networks—a tutorial\.IEEE Transactions on Circuits and Systems II: Express Briefs69\(6\),pp\. 2581–2585\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[9\]K\. Han, Y\. Wang, H\. Chen, X\. Chen, J\. Guo, Z\. Liu, Y\. Tang, A\. Xiao, C\. Xu, Y\. Xu, Z\. Yang, Y\. Zhang, and D\. Tao\(2023\)A survey on vision transformer\.IEEE Transactions on Pattern Analysis and Machine Intelligence \(TPAMI\)45\(1\),pp\. 87–110\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2022.3152247)Cited by:[Figure 1](https://arxiv.org/html/2606.06515#S1.F1),[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[10\]N\. Jouppi, G\. Kurian, S\. Li, P\. Ma, R\. Nagarajan, L\. Nai, N\. Patil, S\. Subramanian, A\. Swing, B\. Towles, C\. Young, X\. Zhou, Z\. Zhou, and D\. A\. Patterson\(2023\)TPU v4: an optically reconfigurable supercomputer for machine learning with hardware support for embeddings\.InProceedings of the 50th Annual International Symposium on Computer Architecture \(ISCA\),External Links:[Document](https://dx.doi.org/10.1145/3579371.3589350)Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[11\]S\. Khan, M\. Naseer, M\. Hayat, S\. W\. Zamir, F\. S\. Khan, and M\. Shah\(2022\)Transformers in vision: a survey\.ACM Computing Surveys \(CSUR\)54\(10s\),pp\. 1–41\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[12\]H\. Li, D\. Chen, and T\. Mitra\(2025\)HyAtten: hybrid photonic\-digital architecture for accelerating attention mechanism\.In2025 Design, Automation & Test in Europe Conference \(DATE\),Vol\.,pp\. 1–7\.External Links:[Document](https://dx.doi.org/10.23919/DATE64628.2025.10993031)Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[13\]Y\. Li, A\. Louri, and A\. Karanth\(2022\)SPACX: silicon photonics\-based scalable chiplet accelerator for dnn inference\.In2022 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),Vol\.,pp\. 831–845\.External Links:[Document](https://dx.doi.org/10.1109/HPCA53966.2022.00066)Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[14\]Y\. Li, A\. Louri, and A\. Karanth\(2022\)SPRINT: a high\-performance, energy\-efficient, and scalable chiplet\-based accelerator with photonic interconnects for cnn inference\.IEEE Transactions on Parallel and Distributed Systems \(TPDS\)33\(10\),pp\. 2332–2345\.External Links:[Document](https://dx.doi.org/10.1109/TPDS.2021.3139015)Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[15\]Y\. Li, A\. Louri, and A\. Karanth\(2025\)MERIT: a sustainable dnn accelerator design with photonic phase\-change memory\.IEEE Transactions on Sustainable Computing \(TSUSC\)10\(4\),pp\. 705–716\.External Links:[Document](https://dx.doi.org/10.1109/TSUSC.2024.3521847)Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.
- \[16\]Z\. Li, M\. Sun, A\. Lu, H\. Ma, G\. Yuan, Y\. Xie, H\. Tang, Y\. Li, M\. Leeser, Z\. Wang,et al\.\(2022\)Auto\-vit\-acc: an fpga\-aware automatic acceleration framework for vision transformer with mixed\-scheme quantization\.In2022 32nd International Conference on Field\-Programmable Logic and Applications \(FPL\),pp\. 109–116\.Cited by:[Figure 1](https://arxiv.org/html/2606.06515#S1.F1),[Figure 11](https://arxiv.org/html/2606.06515#S5.F11)\.
- \[17\]S\. Lu, M\. Wang, S\. Liang, J\. Lin, and Z\. Wang\(2020\)Hardware accelerator for multi\-head attention and position\-wise feed\-forward in the transformer\.In2020 IEEE 33rd International System\-on\-Chip Conference \(SOCC\),pp\. 84–89\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[18\]A\. Mumuni and F\. Mumuni\(2025\)Large language models for artificial general intelligence \(agi\): a survey of foundational principles and approaches\.arXiv preprint arXiv:2501\.03151\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[19\]R\. V\. W\. Putra, M\. A\. Hanif, and M\. Shafique\(2020\)DRMap: a generic dram data mapping policy for energy\-efficient processing of convolutional neural networks\.In2020 57th ACM/IEEE Design Automation Conference \(DAC\),Vol\.,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/DAC18072.2020.9218672)Cited by:[2nd item](https://arxiv.org/html/2606.06515#S3.I1.i2.p1.1)\.
- \[20\]R\. V\. W\. Putra, M\. A\. Hanif, and M\. Shafique\(2021\)ROMANet: fine\-grained reuse\-driven off\-chip memory access management and data organization for deep neural network accelerators\.IEEE Transactions on Very Large Scale Integration Systems \(TVLSI\)29\(4\),pp\. 702–715\.External Links:[Document](https://dx.doi.org/10.1109/TVLSI.2021.3060509)Cited by:[2nd item](https://arxiv.org/html/2606.06515#S3.I1.i2.p1.1)\.
- \[21\]R\. V\. W\. Putra, M\. A\. Hanif, and M\. Shafique\(2024\)PENDRAM: enabling high\-performance and energy\-efficient processing of deep neural networks through a generalized dram data mapping policy\.arXiv preprint arXiv:2408\.02412\.Cited by:[2nd item](https://arxiv.org/html/2606.06515#S3.I1.i2.p1.1)\.
- \[22\]P\. Qi, E\. H\. Sha, Q\. Zhuge, H\. Peng, S\. Huang, Z\. Kong, Y\. Song, and B\. Li\(2021\)Accelerating framework of transformer by hardware design and model compression co\-optimization\.In2021 IEEE/ACM International Conference On Computer Aided Design \(ICCAD\),pp\. 1–9\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[23\]B\. J\. Shastri, A\. N\. Tait, T\. Ferreira de Lima, W\. H\. Pernice, H\. Bhaskaran, C\. D\. Wright, and P\. R\. Prucnal\(2021\)Photonics for artificial intelligence and neuromorphic computing\.Nature Photonics15\(2\),pp\. 102–114\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[24\]Y\. Shen, N\. C\. Harris, S\. Skirlo, M\. Prabhu, T\. Baehr\-Jones, M\. Hochberg, X\. Sun, S\. Zhao, H\. Larochelle, D\. Englund,et al\.\(2017\)Deep learning with coherent nanophotonic circuits\.Nature photonics11\(7\),pp\. 441–446\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[25\]K\. Shiflett, A\. Karanth, R\. Bunescu, and A\. Louri\(2021\)Albireo: energy\-efficient acceleration of convolutional neural networks via silicon photonics\.In2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture \(ISCA\),pp\. 860–873\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[26\]M\. Sun, H\. Ma, G\. Kang, Y\. Jiang, T\. Chen, X\. Ma, Z\. Wang, and Y\. Wang\(2022\)Vaqf: fully automatic software\-hardware co\-design framework for low\-bit vision transformer\.arXiv preprint arXiv:2201\.06618\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[27\]F\. Sunny, A\. Mirza, M\. Nikdast, and S\. Pasricha\(2021\)CrossLight: a cross\-layer optimized silicon photonic neural network accelerator\.In2021 58th ACM/IEEE design automation conference \(DAC\),pp\. 1069–1074\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[28\]F\. P\. Sunny, E\. Taheri, M\. Nikdast, and S\. Pasricha\(2021\)A survey on silicon photonics for deep learning\.ACM Journal of Emerging Technologies in Computing System \(JETC\)17\(4\),pp\. 1–57\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[29\]A\. N\. Tait, T\. F\. De Lima, E\. Zhou, A\. X\. Wu, M\. A\. Nahmias, B\. J\. Shastri, and P\. R\. Prucnal\(2017\)Neuromorphic photonic networks using silicon photonic weight banks\.Scientific Reports7\(1\),pp\. 7430\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[30\]H\. Touvron, M\. Cord, M\. Douze, F\. Massa, A\. Sablayrolles, and H\. Jégou\(2021\)Training data\-efficient image transformers & distillation through attention\.InInternational Conference on Machine Learning \(ICML\),pp\. 10347–10357\.Cited by:[Figure 1](https://arxiv.org/html/2606.06515#S1.F1),[Figure 2](https://arxiv.org/html/2606.06515#S1.F2),[§I\-B](https://arxiv.org/html/2606.06515#S1.SS2.p1.2),[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[31\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez,et al\.\(2017\)Attention is all you need\.Advances in Neural Information Processing Systems \(NIPS\)30\(1\),pp\. 261–272\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[32\]H\. Wang, Z\. Zhang, and S\. Han\(2021\)Spatten: efficient sparse attention architecture with cascade token and head pruning\.In2021 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),pp\. 97–110\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[33\]G\. Yenduri, R\. Murugan, P\. Kumar Reddy Maddikunta, S\. Bhattacharya, D\. Sudheer, and B\. Bhushan Savarala\(2025\)Artificial general intelligence: advancements, challenges, and future directions in agi research\.IEEE Access13\(\),pp\. 134325–134356\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2025.3592708)Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[34\]Z\. Yin, M\. Zhang, N\. Gangi, R\. Huang, J\. Zhang, and J\. Gu\(2025\)Simphony: a device\-circuit\-architecture cross\-layer modeling and simulation framework for heterogeneous electronic\-photonic ai system\.In2025 62nd ACM/IEEE Design Automation Conference \(DAC\),pp\. 1–7\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p2.1)\.
- \[35\]H\. You, Z\. Sun, H\. Shi, Z\. Yu, Y\. Zhao, Y\. Zhang, C\. Li, B\. Li, and Y\. Lin\(2023\)Vitcod: vision transformer acceleration via dedicated algorithm and accelerator co\-design\.In2023 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),pp\. 273–286\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[36\]M\. Zhou, W\. Xu, J\. Kang, and T\. Rosing\(2022\)Transpim: a memory\-based acceleration via software\-hardware co\-design for transformer\.In2022 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),pp\. 1071–1085\.Cited by:[§I](https://arxiv.org/html/2606.06515#S1.p1.1)\.
- \[37\]H\. Zhu, J\. Gu, H\. Wang, Z\. Jiang, Z\. Zhang, R\. Tang, C\. Feng, S\. Han, R\. T\. Chen, and D\. Z\. Pan\(2024\)Lightening\-transformer: a dynamically\-operated optically\-interconnected photonic transformer accelerator\.In2024 IEEE International Symposium on High\-Performance Computer Architecture \(HPCA\),Vol\.,pp\. 686–703\.External Links:[Document](https://dx.doi.org/10.1109/HPCA57654.2024.00059)Cited by:[Figure 1](https://arxiv.org/html/2606.06515#S1.F1),[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1),[§I\-B](https://arxiv.org/html/2606.06515#S1.SS2.p1.2),[§I](https://arxiv.org/html/2606.06515#S1.p2.1),[Figure 5](https://arxiv.org/html/2606.06515#S2.F5),[Figure 6](https://arxiv.org/html/2606.06515#S2.F6),[§II\-A](https://arxiv.org/html/2606.06515#S2.SS1.p1.2),[§II\-B](https://arxiv.org/html/2606.06515#S2.SS2.p1.7.1),[§IV](https://arxiv.org/html/2606.06515#S4.p1.1),[Figure 11](https://arxiv.org/html/2606.06515#S5.F11)\.
- \[38\]H\. Zhu, Z\. Zhou, S\. Ning, X\. Wu, R\. Chen, Y\. Wan, and D\. Pan\(2025\)ENLighten: lighten the transformer, enable efficient optical acceleration\.arXiv preprint arXiv:2510\.01673\.Cited by:[§I\-A](https://arxiv.org/html/2606.06515#S1.SS1.p1.1)\.Similar Articles
@Underfox3: In this paper is proposed a hardware-software co-design framework for N:M sparse vision Transformer inference, enabling…
This paper proposes a hardware-software co-design framework for N:M sparse vision Transformer inference, achieving over 2.2× latency speedup on GPUs while maintaining accuracy through a novel CUDA kernel (MD-SpMM) and a deployment-aware sparsity search.
Topology-Aware Data Movement for Disaggregated GPU Inference
This paper presents a topology-aware data movement orchestrator for disaggregated LLM inference, which dynamically selects optimal transport based on interconnect hierarchy and overlaps KV cache transfer with computation, achieving 3-18x transfer latency reduction over uniform RDMA.
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
SlideFormer introduces a heterogeneous co-design for full-parameter LLM fine-tuning on a single GPU, leveraging GPU/CPU/RAM/NVMe with a layer-sliding engine and optimized Triton kernels, enabling fine-tuning of 123B+ models on a single RTX 4090 with significant throughput improvements.
RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory
Proposes RED-PIM, an algorithm-architecture co-design that reduces inter-bank data movement from O(N^2) to O(N) and shrinks attention matrices, achieving significant inference time reductions (16% to 99.99%) for transformer models.
Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So
This paper investigates why transformer intermediate representations are off-axis relative to the readout direction, showing that this off-axis subspace functionally insulates composition from the vocabulary and proposing methods to impose this geometry via rotation.