@davsca1: Check out our #ECCV2026 paper "Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention", where …
Summary
This paper introduces Spatially-Sparse Linear Attention for low-latency object detection with event cameras, achieving state-of-the-art accuracy with 20x less computation than prior asynchronous methods.
View Cached Full Text
Cached at: 08/27/26, 07:43 AM
Check out our #ECCV2026 paper “Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention”, where we make linear attention sparse in space, recurrent in time, and parallel in training, enabling the first purely-linear-attention-based neural network for asynchronous object detection with #EventCameras, outperforming the previous best asynchronous method with 20x less computation with truly event-by-event inference on CPU! Code released!
Paper: https://arxiv.org/abs/2603.06228 Code: https://github.com/haohq19/ssla Video: https://youtube.com/watch?v=qaVeSqEt8IM…
Event cameras promise extremely low-latency vision, but to fully exploit them, the neural network must be low-latency too. We introduce #SpatiallySparseLinearAttention (#SSLA) for asynchronous object detection directly from raw events.
Linear attention is particularly appealing for event cameras: it can be trained efficiently in parallel on long event sequences, while at inference it operates recurrently, updating its prediction every time a new event arrives.
The problem is that conventional linear attention updates its entire state for every event. For object detection, where fine spatial resolution matters, this quickly becomes expensive.
Our key idea is simple: an event only carries information about a small spatial region, so why update the entire spatial state? SSLA updates only the relevant parts of the state, enabling fine-grained spatial representations while keeping per-event computation low. We achieve:
-
20× lower per-event computation than the strongest prior asynchronous baseline
- State-of-the-art accuracy among asynchronous object detection methods
- Truly event-by-event inference on CPU, designed to preserve the latency advantage of event cameras
Come to our poster on Friday September 11, 2026 from 4-6pm at ExHall #389
Reference: Haiqing Hao, Zhipeng Sui, Rong Zou, Zijia Dai, Nikola Zubić, Davide Scaramuzza, Wenhui Wang Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention ECCV, 2026
@Prophesee_ai @SynSenseNeuro @UZH_en @UZH_Science @ERC_Research @UZH_ai @uzh_ifi @Tesla @BYDCompany
#EventCameras #ComputerVision #Robotics #DeepLearning #NeuromorphicVision #AI
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
Source: https://arxiv.org/html/2603.06228 Haiqing Haohttps://orcid.org/0009-0002-0991-2009Affiliation:State Key Laboratory of Precision Measurement Technology and Instruments, Department of Precision Instrument, Tsinghua University, Beijing, ChinaZhipeng Suihttps://orcid.org/0000-0003-1913-1986Affiliation:State Key Laboratory of Precision Measurement Technology and Instruments, Department of Precision Instrument, Tsinghua University, Beijing, ChinaZijia Daihttps://orcid.org/0009-0007-0108-1921Affiliation:ShanghaiTech University, Shanghai, ChinaNikola Zubićhttps://orcid.org/0000-0001-9816-2718Affiliation:Robotics and Perception Group, University of Zurich, Zurich, SwitzerlandDavide Scaramuzzahttps://orcid.org/0000-0002-3831-6778Affiliation:Robotics and Perception Group, University of Zurich, Zurich, SwitzerlandWenhui Wanghttps://orcid.org/0000-0002-5884-6098††thanks:Corresponding author:E-mail[email protected]Affiliation:State Key Laboratory of Precision Measurement Technology and Instruments, Department of Precision Instrument, Tsinghua University, Beijing, China
Abstract
Event cameras provide sequential visual data with spatial sparsity and high temporal resolution, making them attractive for low-latency object detection. Existing asynchronous event-based neural networks exploit this low-latency advantage by updating predictions event by event, but still suffer from two bottlenecks: recurrent architectures are difficult to train efficiently on long sequences, and improving accuracy often increases per-event computation and latency. Linear attention is appealing because it enables parallel training and recurrent inference. However, its dense state updates make per-event computation scale with the state size, yielding a poor accuracy-efficiency trade-off for object detection, where accurate localization requires fine-grained spatial states. The key challenge is therefore to introducesparse state activationthat exploits the spatial sparsity of events while preservingefficient parallel training. We propose Spatially-Sparse Linear Attention (SSLA), which introduces a mixture-of-spaces state decomposition and a scatter-compute-gather training procedure, enabling state-level sparsity as well as training parallelism. Building on SSLA, we develop an end-to-end asynchronous linear attention model, SSLA-Det, for low-latency event-based object detection. On Gen1 and N-Caltech101, SSLA-Det achieves state-of-the-art accuracy among asynchronous methods, reaching 0.375 mAP and 0.515 mAP, respectively, while reducing per-event computation by over 20×\timescompared with the strongest prior asynchronous baseline, demonstrating the potential of linear attention for low-latency event-based vision. Code is available at:https://github.com/haohq19/ssla.
Keywords:
Event camera Linear attention Object detection
(a)Events
(b)SSLA
(c)Detection
(d)mAP vs. FLOPS
Figure 1:Our method processes (a) asynchronous event sequence with (b) a sparsely activated linear attention neural network for (c) low-latency event-based object detection. On the Gen1 dataset, our SSLA-Det models achieve SOTA asynchronous mAP and lower FLOPS compared with previous asynchronous baselines (d).∗refers to AP50.## 1Introduction
Event cameras provide sequential visual data with spatial sparsity and high temporal resolution, making them highly promising for low-latency perception[9,29,11]. Asynchronous event-based neural networks realize this potential by updating their predictions every time a new event arrives[36,37]. This event-driven processing paradigm is particularly appealing for object detection in latency-critical scenarios, such as autonomous driving[11], drone obstacle avoidance[7], and vision-based control[15].
Despite their much lower latency, existing asynchronous event-based neural networks still lag behind their synchronous counterparts in accuracy[49,33,50]. The gap stems from two coupled architectural bottlenecks. The first is the parallel-recurrent bottleneck: the event-by-event inference paradigm naturally relies on recurrent architectures, whereas efficient training on long event sequences requires parallelization along the sequence dimension[39,13]. The second is the accuracy-efficiency trade-off: improving accuracy typically requires larger and deeper models, while scaling the model increases per-event computation and consequently latency. A natural way to mitigate this trade-off is to exploit the spatial sparsity of event camera data through sparse neural network activation[37,36]. Nevertheless, as receptive fields expand layer by layer in deep networks, sparse inputs can still induce dense activations. To preserve sparsity in deep layers and reduce computation, prior work has designed specialized architectures[11,36], but this comes with additional architectural constraints that further limit accuracy.
Linear attention111For convenience, here we use linear attention as a shorthand for parallel-trainable linear recurrent models, including state space models (SSMs) and linear recurrent neural networks (linear RNNs).has emerged as a promising asynchronous event-based model architecture since it naturally addresses the parallel-recurrent bottleneck[47,46,31,20,12]. However, existing methods are limited to relatively simple global-level classification tasks[41,38], while more challenging local-level tasks, like object detection, remain unexplored. The main obstacle is the poor accuracy-efficiency trade-off due to the lack of state-level sparsity,i.e., linear attention updates all elements of its state, making per-event computation scale with the state size. This is problematic for object detection, where accurate localization requires fine-grained spatial representations and therefore a large state size[13]. Although sparsifying state activation is conceptually straightforward, the key challenge is to maintain parallel training while gaining sparsity.
We address this challenge by introducing Spatially-Sparse Linear Attention (SSLA), a linear attention module with state-level sparsity while preserving its parallel training advantages. To enable state-level sparsity, we introduce amixture-of-spaces(MOS) structure inspired by[6]that decomposes the global state into substates with spatially overlapping receptive fields, and each event activates only a few substates based on its location. To preserve the sparsity in deep networks, we aggregate the activations of each event from all its activated substates, preventing the expansion of the activated region. We further propose aposition-aware projection(PAP) that projects events based on their relative positions within every activated substate, injecting state-relative spatial priors. We derive ascatter-compute-gathertraining procedure that parallelizes this sparse activation structure by sequence-level reorganization. Specifically, events are scattered into state-specific subsequences, computed in parallel by linear attention, and then gathered back into the original event sequence. In this way, SSLA makes linear attention sparse in space, recurrent in time, and parallel in training.
Building on the SSLA module, we present SSLA-Det, to the best of our knowledge, the first end-to-end asynchronous linear attention model for low-latency event-based object detection (Fig.1(a)-(c)). Experiments on Gen1 and N-Caltech101 show that SSLA-Det achieves a substantially improved accuracy-efficiency trade-off over prior asynchronous methods (Fig.1(d)), including state-of-the-art (SOTA) asynchronous mAP (0.375 on Gen1 and 0.515 on N-Caltech101) at much lower computational cost (over20×20\timesreduction compared with previous SOTA[11]). Our contributions are as follows:
- •We propose an SSLA module for sequential event modeling, including a MOS structure for state-level sparsity, PAP for spatial prior encoding, and a scatter-compute-gather procedure for efficient parallel training.
- •We present SSLA-Det, to the best of our knowledge, the first end-to-end asynchronous linear attention model for event-based object detection.
- •SSLA-Det sets a new accuracy-efficiency frontier, reaching 0.375 mAP on Gen1 and 0.515 mAP on N-Caltech101, while reducing computation by over20×20\timescompared with the prior asynchronous SOTA method.
2Related Work
2.1Asynchronous Event-based Neural Networks
Asynchronous event-based neural networks treat event camera data as different geometric structures to leverage their spatial sparsity. This strategy includes graphs[22,37,3,4,11], submanifolds[25,36], point clouds[39,43], and sequences[19,13,38,41]. Graph-based methods[22,37,3,4,11,2]transform events into sparsely connected spatial-temporal graphs, and derive local update rules on the graph for recurrent inference. However, they have limitations in temporal accumulation[4], failing to handle long event sequences. Submanifold methods[25,36]assume that events lie on a spatial submanifold, and conduct convolution only on the submanifold to keep sparsity. However, this submanifold assumption does not strictly hold, and thus leads to suboptimal performance. Point cloud methods[43,39]represent events as spatial-temporal 3-D point clouds, and process with PointNet[35]. This approach is limited to shallow neural networks, and thus has difficulty in challenging tasks. Sequence-based methods use causal sequence-to-sequence models for event processing including softmax attention or linear recurrent models. For example,[19]uses attention to process batched events at each timestep. EventSSM[38]and S7[41]use SSMs, which enable parallel training and recurrent inference, but are limited to global-level classification tasks. EVA[13]explores linear attention for event-based object detection but still demands a dense backbone, while our method is fully end-to-end asynchronous.
2.2State-level Sparsity in Linear Attention
A recent direction to improve linear attention is to introduce state-level sparsity. Mixture-of-Memories[6]and Sparse State Expansion[28]maintain multiple independent states and use a learned router to send each token to only a few memories. Our SSLA follows the same sparse state activation idea but is fundamentally different in that the construction of substates is spatially structured, which enables geometric routing of event embeddings and admits the position-aware projection that encodes spatial inductive bias. Concurrent work[40]also introduces local states in linear attention for local-level event-based vision, but requires spatial contraction at discrete timestamps, which is not asynchronously, event-by-event trainable. Our work, instead, uses a scatter-compute-gather algorithm to reorganize events into subsequences and does not rely on spatial contraction, which is fully event-by-event asynchronous.
3Method
3.1Problem Formulation: Asynchronous Event Processing
Event camera data are represented as a sequence of eventsℰ={ei}i=1L\mathcal{E}=\{e_{i}\}_{i=1}^{L}, where each eventei=(𝐱i,ti,pi)e_{i}=(\mathbf{x}_{i},t_{i},p_{i})carries its spatial coordinates𝐱i∈ℝ2\mathbf{x}_{i}\in\mathbb{R}^{2}, a timestampti∈ℝt_{i}\in\mathbb{R}and a polaritypi∈{+1,−1}p_{i}\in\{+1,-1\}. The sequence is temporally ordered so thatti≤ti+1t_{i}\leq t_{i+1}. Asynchronous event processing learns a causal stateful neural networkℳ\mathcal{M}that takesℰ\mathcal{E}as input and makes predictions incrementally, which can be formulated as a causal sequence-to-sequence problem from{ei}i=1L\{e_{i}\}_{i=1}^{L}to{𝐲^i}i=1L\{\hat{\mathbf{y}}_{i}\}_{i=1}^{L}. Specifically, for each new incoming eventeie_{i}, the model updates its state𝐒i\mathbf{S}_{i}and produces a new prediction𝐲^i\hat{\mathbf{y}}_{i}, as(𝐲^i,𝐒i)=ℳ(ei,𝐒i−1)(\hat{\mathbf{y}}_{i},\mathbf{S}_{i})=\mathcal{M}(e_{i},\mathbf{S}_{i-1}).
3.2Preliminaries: Linear Attention
Linear attention models (linear RNNs, SSMs) are linear-time alternatives to softmax attention[44]for sequence-to-sequence modeling. Causal linear attention has an equivalent parallel and recurrent form, which enables both parallel training and recurrent inference[20]. Given an input sequence of embeddings{𝐳i}i=1L\{\mathbf{z}_{i}\}_{i=1}^{L}, a causal linear attention module updates its hidden state𝐒i\mathbf{S}_{i}and computes outputs{𝐨i}i=1L\{\mathbf{o}_{i}\}_{i=1}^{L}as
𝐒i\displaystyle\mathbf{S}_{i}=g(𝐳i)⊙𝐒i−1+ϕ(𝐳i),\displaystyle=g(\mathbf{z}_{i})\odot\mathbf{S}_{i-1}+\phi(\mathbf{z}_{i}),(1)𝐨i\displaystyle\mathbf{o}_{i}=ρ(𝐳i,𝐒i),\displaystyle=\rho(\mathbf{z}_{i},\mathbf{S}_{i}),whereg,ϕ,ρg,\phi,\rhoare learnable projections (gating, updating, and output) and⊙\odotis the Hadamard product. This recurrent linear attention is parallel trainable with parallel scan[12]or chunk-wise algorithms[46], which is training-efficient on long event sequences. For convenience, we denote the map of linear attention asLinearAttention:{𝐳i}i=1L↦{𝐨i}i=1L\textbf{LinearAttention}:\{\mathbf{z}_{i}\}_{i=1}^{L}\mapsto{\{\mathbf{o}_{i}\}_{i=1}^{L}}in this paper.
3.3Spatially-Sparse Linear Attention for Event Sequence Modeling
We introduce the Spatially-Sparse Linear Attention (SSLA) module for asynchronous event processing, which is illustrated inFig.2. The SSLA module takes an event sequenceℰ\mathcal{E}with embeddings𝐯i∈ℝDin\mathbf{v}_{i}\in\mathbb{R}^{D_{in}}as input, and outputsℰ\mathcal{E}with updated embeddings𝐨i∈ℝDout\mathbf{o}_{i}\in\mathbb{R}^{D_{out}}. To exploit the spatial sparsity of events, we first introduce amixture-of-spaces(MOS) decomposition of the hidden state, which enables state-level spatial sparsity. We incorporate aposition-aware projectionin the SSLA module to encode spatial priors into event embeddings. To preserve the training efficiency of linear attention, we derive ascatter-compute-gatherprocedure for parallel training by reorganizing events to independent subsequences.
Figure 2:Overview of the Spatially-Sparse Linear Attention (SSLA) module. (a) The global state is decomposed into independent substates, one per overlapping patch in the spatial domain. (1) and (2) are2×22\times 2patch examples that cover the eventeie_{i}(for brevity, we omit the other patches coveringeie_{i}). Their states are updated with the embedding ofeie_{i}, and the output embedding ofeie_{i}is the summation of the interim outputs from all the patches coveringeie_{i}. (b) In each patch, the embeddings are projected based on their relative position within the patch. (c) During training, with an event sequence as input, the events are scattered into patch-wise subsequences (duplicating each event across all the patches covering it), computed in parallel by a weight-shared linear attention, then gathered back to the original event order.#### Sparse State Activation with Mixture-of-Spaces.
We defineΩ\Omegaas the spatial domain of the event camera data, and decompose it into spatially local and overlapping patches, with each patch maintaining its own state independently. We construct these patches by applying a sliding window of sizeP×PP\times Pwith a stride of 1 overΩ\Omega. Let𝒫={1,…,K}\mathcal{P}=\{1,\dots,K\}denote the indices of these patches, where each patchk∈𝒫k\in\mathcal{P}covering a spatial regionℛk⊂Ω\mathcal{R}_{k}\subset\Omegamaintains state𝐒k\mathbf{S}_{k}.
Foreie_{i}at𝐱i\mathbf{x}_{i}, we activate only specific patches that contain𝐱i\mathbf{x}_{i}for sparse state activation. We define the set of active patches ofeie_{i}as
𝒦i={k∈𝒫∣𝐱i∈ℛk}.\mathcal{K}_{i}=\{k\in\mathcal{P}\mid\mathbf{x}_{i}\in\mathcal{R}_{k}\}.(2)We pad the image domain to ensure that each event activates a constant number of states,i.e.,|𝒦i|=P2≜A|\mathcal{K}_{i}|=P^{2}\triangleq A. TheseAAstates are updated byEq.1with event embeddings, and generate interim outputs{𝐨i,k∣k∈𝒦i}\{\mathbf{o}_{i,k}\mid k\in\mathcal{K}_{i}\}. All active patches share the same linear attention parameters but maintain independent states. We aggregate the interim outputs as the updated embedding from all activated patches:
𝐨i=∑k∈𝒦i𝐨i,k.\mathbf{o}_{i}=\sum_{k\in\mathcal{K}_{i}}\mathbf{o}_{i,k}.(3)After aggregation, the embedding sequence length is unchanged, which preserves sparsity in deep layers.
Position-Aware Projection.
Sharing identical embeddings𝐯i\mathbf{v}_{i}for one event in all its activated patches ignores the spatial prior of the event, since the event is located at different relative positions in different patches. We therefore introduce a position-aware projection (PAP) of the embedding, which projects the embedding based on the event’s relative positions inside the patches.
For an event at𝐱i\mathbf{x}_{i}in patchkkwith top-left global coordinates𝐜k∈Ω\mathbf{c}_{k}\in\Omega, we compute its relative position in the patch as
𝜹i,k=𝐱i−𝐜k,where𝜹i,k∈{0,…,P−1}2.\boldsymbol{\delta}_{i,k}=\mathbf{x}_{i}-\mathbf{c}_{k},\quad\text{where }\boldsymbol{\delta}_{i,k}\in\{0,...,P-1\}^{2}.(4)We define the PAP as a linear transform by𝐖in∈ℝP×P×Dout×Din\mathbf{W}^{\text{in}}\in\mathbb{R}^{P\times P\times D_{out}\times D_{in}}and𝐖out∈ℝP×P×Dout×Dout\mathbf{W}^{\text{out}}\in\mathbb{R}^{P\times P\times D_{out}\times D_{out}}. The input embedding𝐯i\mathbf{v}_{i}is projected by
𝐮i,k∈ℝDout←𝐖in[𝜹i,k]𝐯i.\mathbf{u}_{i,k}\in\mathbb{R}^{D_{out}}\leftarrow\mathbf{W}^{\text{in}}[\boldsymbol{\delta}_{i,k}]\mathbf{v}_{i}.(5)Then we use𝐮i,k\mathbf{u}_{i,k}in patchkkas input to linear attention. Similarly, before aggregating the interim outputs, we also conduct a PAP, andEq.3becomes
𝐨i=∑k∈𝒦i𝐖out[𝜹i,k]𝐨i,k.\mathbf{o}_{i}=\sum_{k\in\mathcal{K}_{i}}\mathbf{W}^{\text{out}}[\boldsymbol{\delta}_{i,k}]\mathbf{o}_{i,k}.(6)This projection encodes spatial priors with learnable parameters, expanding the capability of the model to capture spatial patterns and information.
Algorithm 1Spatially-Sparse Linear Attention (SSLA) Module Training0:Events
ℰ\mathcal{E}, embeddings
{𝐯i}i=1L\{\mathbf{v}_{i}\}_{i=1}^{L}, coordinates
{𝐱i}i=1L\{\mathbf{x}_{i}\}_{i=1}^{L}, Patches
𝒫\mathcal{P}, lookup table
T:𝐱∈Ω↦{(k,δ)}T:\mathbf{x}\in\Omega\mapsto\{\left(k,\delta\right)\}.
0:Updated embeddings
{𝐨i}i=1L\{\mathbf{o}_{i}\}_{i=1}^{L}in the same order as the input.
1:Initialize an empty embedding sequence
𝒰\mathcal{U}of length
ALAL.
2:foreach event
eie_{i}in paralleldo
3:Lookup
𝐱i\mathbf{x}_{i}: active patch indices
{k∣k∈𝒦i}\{k\mid k\in\mathcal{K}_{i}\}, relative positions
{𝜹i,k∣k∈𝒦i}\{\boldsymbol{\delta}_{i,k}\mid k\in\mathcal{K}_{i}\};
4:Position-aware projection:
𝐮i,k←𝐖in[𝜹i,k]𝐯i\mathbf{u}_{i,k}\leftarrow\mathbf{W}^{\text{in}}[\boldsymbol{\delta}_{i,k}]\mathbf{v}_{i};
5:Assign
𝒰[A(i−1):Ai]←{𝐮i,k∣k∈𝒦i}\mathcal{U}[A(i-1):Ai]\leftarrow\{\mathbf{u}_{i,k}\mid k\in\mathcal{K}_{i}\} 6:endfor
7:Scatter: stable-sort
𝒰\mathcal{U}based on
kk, and cache the permutation
π\pi;
8:Split
𝒰\mathcal{U}to patch-specific subsequences
{𝒰k∣k∈𝒫}\{\mathcal{U}_{k}\mid k\in\mathcal{P}\};
9:foreach patch
k∈𝒫k\in\mathcal{P}in paralleldo
10:
𝒪k←LinearAttention(𝒰k)\mathcal{O}_{k}\leftarrow\textbf{LinearAttention}(\mathcal{U}_{k})usingEq.1;
11:endfor
12:Concat
𝒪k\mathcal{O}_{k}into interim output sequence
𝒪\mathcal{O};
13:Gather:
𝒪←𝒪[π−1]\mathcal{O}\leftarrow\mathcal{O}[\pi^{-1}]. We get
𝒪[A(i−1):Ai]={𝐨i,k∣k∈𝒦i}\mathcal{O}[A(i-1):Ai]=\{\mathbf{o}_{i,k}\mid k\in\mathcal{K}_{i}\};
14:foreach event
eie_{i}in paralleldo
15:Position-aware projection:
𝐨i,k←𝐖out[𝜹i,k]𝐨i,k\mathbf{o}_{i,k}\leftarrow\mathbf{W}^{\text{out}}[\boldsymbol{\delta}_{i,k}]\mathbf{o}_{i,k};
16:endfor
17:Aggregate:
𝐨i=∑𝒪[A(i−1):Ai]\mathbf{o}_{i}=\sum\mathcal{O}[A(i-1):Ai];
18:return
{𝐨i}i=1L\{\mathbf{o}_{i}\}_{i=1}^{L}
Parallelizable Training with Scatter-Compute-Gather.
While the MOS structure enables state-level sparsity, it breaks the single state form of linear attention, and thus poses challenges in efficient parallel training on GPUs. To address this, we derive ascatter-compute-gatheralgorithm, which allows both intra-patch (temporal) and inter-patch (spatial) parallelism to speed up training.
Scatter.
Let𝐮i,k\mathbf{u}_{i,k}denote the projected embeddings of eventeie_{i}for the active patchk∈𝒦ik\in\mathcal{K}_{i}. We first reorganize the input event sequenceℰ\mathcal{E}toKKpatch-specific subsequences, by constructing a subsequence for each patchkkas
𝒰k={𝐮i,k∣isuch thatk∈𝒦i},∀k∈𝒫,\mathcal{U}_{k}=\{\mathbf{u}_{i,k}\mid i\text{ such that }k\in\mathcal{K}_{i}\},\quad\forall k\in\mathcal{P},(7)where𝒰k\mathcal{U}_{k}is ordered bytit_{i}. We implement this with a precomputed lookup table that maps each coordinate inΩ\Omegato the indices of the patches covering it, together with its relative position inside each patch. Using the table, we expandℰ\mathcal{E}to a projected sequence𝒰\mathcal{U}of lengthALAL, where each event contributesAAconsecutive projected embeddings, one for each of its activated patches. We apply stable sorting on𝒰\mathcal{U}based on the patch indices of each element, forming a reorganized sequence with embeddings in the same patch grouped together as the subsequence, while preserving the temporal order within each𝒰k\mathcal{U}_{k}. The resulting permutation is cached and later reused in the gather step.
Compute.
All patches share the same parameters but maintain independent states, so that theKKsubsequences can be computed in parallel, achieving inter-patch parallelism. We apply linear attention (Eq.1) to each subsequence𝒰k\mathcal{U}_{k}
𝒪k=LinearAttention(𝒰k),\mathcal{O}_{k}=\textbf{LinearAttention}(\mathcal{U}_{k}),(8)where𝒪k\mathcal{O}_{k}contains the interim outputs for events contained in patchkk. In each patch, the linear attention is standard, which achieves intra-patch parallelism.
Gather.
Finally, we restore the outputs to the original expanded event sequence order using the cached permutation and aggregate the interim outputs from all active patches of each event. Specifically, we first apply the PAP to each interim output𝐨i,k\mathbf{o}_{i,k}, and then sum them overk∈𝒦ik\in\mathcal{K}_{i}according toEq.6. This gather step is efficient because it only involves indexed reordering and reduction. The training procedure of the SSLA module is summarized inAlgorithm1.
3.4SSLA-Det Architecture
We propose the SSLA-Det model for asynchronous event-based object detection. An overview of our neural network is shown inFig.3, which consists of an asynchronous backbone and a YOLOX detection head[10]. The backbone has 4 stages, each containing 2 SSLA module layers. In the SSLA layers, we use residual connections[14]and layer normalization[1]to stabilize training. In the first 3 stages, we use one sparse pooling and one temporal dropout layer introduced in[36], which compresses event sequence to reduce computation while preserving the high temporal resolution of events. All of the above layers are asynchronous, which makes the backbone fully asynchronous.
SSLA module is agnostic to the specific design of linear attention mechanism, allowing for the integration of any variant, including linear RNNs and SSMs. We use a real-valued Linear Recurrent Unit[27]implemented by Triton[42]in our model for hardware efficiency. The embedding dimensionDoutD_{out}has an expansion of 2×\timesin each stage.
For each step, the input to the model is a raw event, with polarity and time difference𝐯i=[pi,Δti]∈ℝ2\mathbf{v}_{i}=\left[p_{i},\mathrm{\Delta}t_{i}\right]\in\mathbb{R}^{2}as the embedding, and the backbone generates𝐨i\mathbf{o}_{i}. We form a spatially fine-grained representation𝐑∈ℝHout×Wout×Dout\mathbf{R}\in\mathbb{R}^{H_{out}\times W_{out}\times D_{out}}from the asynchronous output, by updating𝐨i\mathbf{o}_{i}to𝐑[𝐱i]\mathbf{R}[\mathbf{x}_{i}]. To achieve an end-to-end asynchronicity, we modify the YOLOX head by changing all the convolutions to1×11\times 1. Each backbone output only updates the head predictions at position𝐱i\mathbf{x}_{i}, making the YOLOX head also asynchronous. The whole SSLA-Det model is therefore end-to-end fully asynchronous, leading to minimal latency.
Figure 3:Overview of the SSLA-Det model.Top:Thefully asynchronousevent-based object detector. Events are processed by a 4-stage asynchronous backbone, and each stage doubles the output embedding dimension. Each output embedding updates the representation of its position, and an asynchronous YOLOX head gives the detections. The red region marks the area sparsely activated by an event.Bottom:Stage layout. Each stage has 2 SSLA layers followed by sparse pooling and temporal dropout (only SSLA layers at stage 4). In one SSLA layer, event embeddings are processed by the SSLA module, including two position-aware projections and a patch-wise linear attention. A residual connection and a layer normalization are used for training stability.
4Experiments
4.1Experimental Setup
Datasets.
Following previous work[11], we evaluate on the N-Caltech101 Detection[26]and the Gen1 Detection[5]datasets. N-Caltech101 consists of recordings captured by a DAVIS240 event camera with a resolution of240×180240\times 180pixels, undergoing saccadic motion in front of a projector displaying Caltech101 images. Bounding box annotations were manually added in post-processing, with 101 classes. Gen1 is a more challenging, large-scale benchmark for automotive scenarios, recorded by an ATIS event camera with a resolution of304×240304\times 240pixels. The dataset has two categories of 228,123 cars and 27,658 pedestrians. Following previous works[11,34], we filter out bounding boxes with a diagonal below 30 pixels or width below 20 pixels in Gen1.
Training Details.
All experiments were conducted with PyTorch 2.6.0[30]on NVIDIA Ampere GPUs (A800/RTX 3090). We design four variants of our SSLA-Det model: small (SSLA-S), base (SSLA-B), medium (SSLA-M), and large (SSLA-L), by scaling the embedding dimensionDoutD_{out}of the first stage to 12, 16, 24 and 32. We use AdamW[24]optimizer. For Gen1, we train 40 epochs using a batch size of 32 and a base learning rate of1×10−31\times 10^{-3}with a cosine decay. We use random flipping with probability 0.5 and random dropout of input events with the keep ratio sampled from𝒰(0.8,1.0)\mathcal{U}(0.8,1.0). For N-Caltech101, we train 200 epochs with a batch size of 64. In addition to random dropout, we apply random cropping to 75%\%of the full resolution with probability 0.2 and random translation by up to 10%\%of the full resolution following the implementations of[11]. We also use exponential model averaging[17].
4.2Results
Gen1 Automotive.
We compare our SSLA-Det models with asynchronous event-based detection baselines[22,25,37,36,11], and use synchronous methods as reference[13,49,8,33,45,32]. The performance is evaluated in accuracy with mean average precision (mAP)[23]and efficiency with average floating point operations (FLOPS) for every new event. We observed that some prior works report AP50whereas others report mAP, and these values are sometimes presented together in the literature. To avoid potentially misleading comparisons, we separate them inTab.1.
Table 1:Object detection results on the Gen1 Detection dataset.Async. refers to asynchronous methods.Our SSLA-Det models consistently improve the accuracy-efficiency trade-off over existing asynchronous baselines. Notably, compared to the strongest prior asynchronous baseline DAGr-L[11], our smallest model SSLA-S achieves higher mAP (0.334 vs. 0.321) while reducing the computational cost by about 171×\times(0.102 vs. 17.4 MFLOPS/ev). Our largest model, SSLA-L, achieves an mAP of 0.375, setting a new SOTA for asynchronous detection on Gen1, with>20×>20\timesreduction of FLOPS to previous best model DAGr-L (0.724 M/ev vs. 17.4 M/ev).
Fig.4visualizes the detection results on Gen1.Fig.4(a)-(d) andFig.4(e)-(h) show the detected cars and pedestrians, respectively.Fig.4(i)-(l) provide qualitative understandings of the typical failure cases. Specifically,Fig.4(i) and (j) show the false negatives, mainly caused by the lack of relative motion between the event camera and the target, which leads to missing event data.Fig.4(k) and (l) show the false positives caused by missing annotations.
(a)
(b)
(c)
(d)
(e)
(f)
(g)
(h)
(i)
(j)
(k)
(l)
Figure 4:Visualization of the detection results on the Gen1 dataset.Green boxes denote ground truth and orange boxes denote predictions with confidence scores. Predicted boxes with confidence scores below 0.5 are removed. (a)-(d): Cars. (e)-(h): Pedestrians. (i)-(l): Failure cases.Table 2:Object detection results on the N-Caltech101 Detection dataset.Async. refers to asynchronous methods.
N-Caltech101.
Tab.2presents the performance of our models on the N-Caltech101 dataset. Consistent with the Gen1 results, our models achieve a superior accuracy-efficiency trade-off among asynchronous methods. In particular, SSLA-L reaches an AP50of 0.743 with a computational cost of only 0.926 MFLOPS/ev. Compared with the previous best asynchronous baseline DAGr-L, SSLA-L improves AP50by 1.1 points (0.743 vs. 0.732) while using20×20\timesfewer MFLOPS per event (0.926 vs. 18.9).
4.3Timing Experiments
Training Efficiency.
We show the training efficiency benefit of linear attention with sequential parallelism. We replace linear attention with a long short-term memory (LSTM) baseline[16]with the same hidden dimension as SSLA-S.Tab.3compares SSLA with an LSTM baseline from the official PyTorch implementation under the same training setup on Gen1. We compare the training time per epoch, which is measured on 4 NVIDIA A800 GPUs. SSLA-S reduces the epoch time from 1.05 to 0.25 hours (4.2×\times), but at the cost of a drop in mAP (from 0.353 to 0.334), mainly caused by a lower FLOPS. At a similar FLOPS level, SSLA-B achieves a comparable mAP to LSTM (0.351 vs. 0.353) and higher AP50(0.655 vs. 0.631), while reducing train time from 1.05 to 0.28 hours (3.8×\times).
Table 3:Comparison of training efficiency between SSLA and an LSTM baseline on Gen1.Training time per epoch is measured on 4 NVIDIA A800 GPUs.
Inference Latency.
We measure the latency of SSLA-Det as the time required to process one newly arrived event in a recurrent, event-by-event setting. We implement a recurrent C++ version of SSLA-Det and benchmark the latency on a single core of an AMD Ryzen 9 9950X3D CPU. As shown inTab.4, our models achieve a low latency of less than10μs10\,\mu s, which is lower than the sensor transmission latency of approximately200μs200\,\mu s[11]. Interestingly, a smaller model does not yield lower latency in our setting (e.g.on Gen1,3.43μs3.43\,\mu sfor SSLA-S and2.44μs2.44\,\mu sfor SSLA-B), because the actual runtime also depends on hardware factors such as vectorization efficiency and memory access patterns. The SSLA module has a constant per-event inference FLOPS of𝒪(P2Dout2)\mathcal{O}(P^{2}D^{2}_{out}), making the latency independent of resolution. While CPU does not fully translate our FLOPS efficiency into latency gains[11], further runtime latency reduction could be achieved on specific hardware, such as FPGAs[18]or neuromorphic accelerators[48].
Table 4:Inference latency.Latency refers to the time used for the recurrent model to process a new event.
4.4Ablation Study
Efficiency Attribution.
To isolate the sources of efficiency in SSLA-Det, we compare SSLA-S with three variants: (i) removing temporal dropout (TD), (ii) further replacing the SSLA module with a dense-activation counterpart that retains the MOS decomposition and PAP but activates all patches per event, and (iii) removing sparse pooling (SP). As shown inTab.5, the SSLA module contributes the dominant computational cost reduction (380×\times) at no accuracy cost, while TD provides an additional 10×\timesreduction as an accuracy-efficiency trade-off. SP does not change the per-event FLOPS since it only downscales the coordinates of events without reducing the event count, but is essential for detection.
Table 5:Efficiency Attribution.We isolate the contribution of SSLA-Det components on the validation set of Gen1. TD refers to temporal dropout, SP refers to sparse pooling, and we remove SSLA by changing it to its dense counterpart and keeping the MOS and PAP designs.
Effect of Spatial Sparsity.
We replace the SSLA module with a standard linear attention (LRU). We useDoutD_{out}of the first stage 12 and 36, to keep sameDoutD_{out}and similar FLOPS as SSLA-S.Tab.6shows that using a standard linear attention fails in both cases, which is considered mainly due to the lack of fine-grained state. In particular, SSLA-S maintains a state 380×\timeslarger than LA (Dout=36D_{out}=36) with similar FLOPS (0.102 M/ev vs. 0.093 M/ev). This demonstrates the importance of state-level sparsity.
Table 6:Effect of Spatial Sparsity.We use linear attention (LA) with same embedding dimension (first stageDout=12D_{out}=12) and similar FLOPS (first stageDout=36D_{out}=36) compared to SSLA-S. We report accuracy on the validation set of Gen1. State refers to the size of hidden state in the last layer, which reflects the capability to model fine-grained spatial representations.
Effect of Position-Aware Projection.
To ablate PAP, we replace it with a position-irrelevant learnable linear projection. The results are summarized inTab.7. Removing either input or output PAP causes a significant accuracy drop, and removing both results in catastrophic failure, with the mAP collapsing to only 0.014, which highlights its importance for encoding spatial priors in the SSLA module.
Table 7:Effect of Position-Aware Projection.Input and Output refer to the PAP with𝐖in\mathbf{W}^{\text{in}}and𝐖out\mathbf{W}^{\text{out}}, respectively. We report accuracy on the validation set of Gen1.
Effect of Patch Size.
PPcontrols the receptive field of event interaction. As shown inTab.8, a smaller patch size (P=2P=2) reduces FLOPS (0.047 M/ev) but leads to a mAP drop (0.200). Conversely,P=4P=4boosts the mAP to 0.371 but increases the FLOPS to 0.179 M/ev, showing an accuracy-efficiency trade-off. Besides, for training, increasingPPresults in higher GPU memory consumption and longer training time. Therefore, we selectP=3P=3as our default configuration as it yields a reasonable trade-off between efficiency and accuracy.
Table 8:Effect of Patch Size.We report accuracy on the validation set of Gen1.
5Limitation and Discussion
This work focuses on event-based low-latency object detection. While hybrid event-image models have become a recent research trend[11,21], they also introduce additional challenges, including event-image alignment, extra sensor requirements, and high system complexity. Besides, in principle, our method is also compatible with the hybrid framework, since image features from dense models can be injected into the intermediate layers of our model. Exploring this event-image fusion in SSLA is an interesting direction for our future work.
Although SSLA-Det achieves SOTA performance among asynchronous event-based object detection methods, a gap remains compared with synchronous methods. For example, as shown inTab.1on Gen1, SSLA-L has an mAP of 0.375, whereas synchronous SOTA methods exceed 0.5 mAP. This gap is expected, as asynchronous and synchronous methods target fundamentally different objectives along the accuracy-efficiency trade-off and are not directly comparable. Synchronous methods accumulate events into image-like representations and perform dense image-level inference, which allows for more information aggregation, the use of image-based neural network architectures and pretrained weights[49,13], and models with larger parameter count[49,8], but at the cost of larger computational cost (Tab.1) and millisecond-level latency. Asynchronous models, in contrast, aim to realize the low-latency advantage of event cameras at the neural network level, giving predictions event-by-event at minimal latency. This structurally constrains parameter count, information aggregation, and architectural choices, naturally limiting accuracy. Therefore, the remaining accuracy gap should be understood as part of the accuracy-latency trade-off in low-latency event-based perception, rather than a methodological shortcoming. Further improving this trade-off while preservingμ\mus-level per-event latency remains an important direction for future work.
6Conclusion
In this paper, we propose SSLA, a novel linear attention module with spatial sparsity and efficient parallel training capability for event sequence modeling. We develop SSLA-Det, the first end-to-end asynchronous linear attention-based model for event-based object detection. Experimental results on Gen1 and N-Caltech101 show that SSLA-Det achieves SOTA asynchronous accuracy with significantly lower FLOPS than previous asynchronous baselines. We believe that SSLA provides a promising direction for low-latency, high-performance event-based perception.
Acknowledgements
This work was supported by the State Key Laboratory of Precision Measurement Technology and Instruments (2025PMTI03), and STI 2030-Major Projects (2021ZD0200300).
References
- [1]J. L. Ba, J. R. Kiros, and G. E. Hinton(2016)Layer normalization.arXiv preprint arXiv:1607.06450.Cited by:§3.4.
- [2]H. Chen, L. Luo, M. Mo, Z. Wu, G. Xiao, J. Gan, J. Leng, and X. Gao(2025)EHGCN: hierarchical euclidean-hyperbolic fusion via motion-aware gcn for hybrid event stream perception.arXiv preprint arXiv:2504.16616.Cited by:§2.1,Table 2.
- [3]T. Dalgaty, T. Mesquida, D. Joubert, A. Sironi, P. Vivet, and C. Posch(2023)Hugnet: hemi-spherical update graph neural network applied to low-latency event-based optical flow.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp. 3953–3962.Cited by:§2.1.
- [4]M. Dampfhoffer, T. Mesquida, D. Joubert, T. Dalgaty, P. Vivet, and C. Posch(2025)Graph neural network combining event stream and periodic aggregation for low-latency event-based vision.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp. 6909–6918.Cited by:§2.1.
- [5]P. De Tournemire, D. Nitti, E. Perot, D. Migliore, and A. Sironi(2020)A large scale event-based detection dataset for automotive.arXiv preprint arXiv:2001.08499.Cited by:§4.1.
- [6]J. Du, W. Sun, D. Lan, J. Hu, T. Zhang, and Y. Cheng(2026)MoM: linear sequence modeling with mixture-of-memories.InThe Fourteenth International Conference on Learning Representations,Cited by:§1,§2.2.
- [7]D. Falanga, K. Kleber, and D. Scaramuzza(2020)Dynamic obstacle avoidance for quadrotors with event cameras.Science Robotics5(40),pp. eaaz9712.Cited by:§1.
- [8]R. Fan, W. Hao, J. Guan, L. Rui, L. Gu, T. Wu, F. Zeng, and Z. Zhu(2025)Eventpillars: pillar-based efficient representations for event data.InProceedings of the AAAI Conference on Artificial Intelligence,Vol.39,pp. 2861–2869.Cited by:§4.2,Table 1,§5.
- [9]G. Gallego, T. Delbrück, G. Orchard, C. Bartolozzi, B. Taba, A. Censi, S. Leutenegger, A. J. Davison, J. Conradt, K. Daniilidis,et al.(2020)Event-based vision: a survey.IEEE transactions on pattern analysis and machine intelligence44(1),pp. 154–180.Cited by:§1.
- [10]Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun(2021)Yolox: exceeding yolo series in 2021.arXiv preprint arXiv:2107.08430.Cited by:§3.4.
- [11]D. Gehrig and D. Scaramuzza(2024)Low-latency automotive vision with event cameras.Nature629(8014),pp. 1034–1040.Cited by:§1,§1,§1,§2.1,§4.1,§4.1,§4.2,§4.2,§4.3,Table 1,Table 1,Table 1,Table 1,Table 2,Table 2,Table 2,Table 2,§5.
- [12]A. Gu and T. Dao(2024)Mamba: linear-time sequence modeling with selective state spaces.InFirst conference on language modeling,Cited by:§1,§3.2.
- [13]H. Hao, N. Zubic, W. He, Z. Sui, D. Scaramuzza, and W. Wang(2026)Maximizing asynchronicity in event-based neural networks.InThe Fourteenth International Conference on Learning Representations,Cited by:§1,§1,§2.1,§4.2,Table 1,§5.
- [14]K. He, X. Zhang, S. Ren, and J. Sun(2016)Deep residual learning for image recognition.InProceedings of the IEEE conference on computer vision and pattern recognition,pp. 770–778.Cited by:§3.4.
- [15]W. He, J. Zhu, Y. Feng, F. Liang, K. You, H. Chai, Z. Sui, H. Hao, G. Li, J. Zhao,et al.(2024)Neuromorphic-enabled video-activated cell sorting.Nature communications15(1),pp. 10792.Cited by:§1.
- [16]S. Hochreiter and J. Schmidhuber(1997)Long short-term memory.Neural computation9(8),pp. 1735–1780.Cited by:§4.3.
- [17]P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson(2018)Averaging weights leads to wider optima and better generalization.arXiv preprint arXiv:1803.05407.Cited by:§4.1.
- [18]K. Jeziorek, P. Wzorek, K. Blachut, H. Nakano, M. Dampfhoffer, T. Mesquida, H. Nishi, T. Dalgaty, and T. Kryjak(2026)Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on soc fpga.arXiv preprint arXiv:2602.16442.Cited by:§4.3.
- [19]U. Kamal, S. Dash, and S. Mukhopadhyay(2023)Associative memory augmented asynchronous spatiotemporal representation learning for event-based perception.InThe Eleventh International Conference on Learning Representations,Cited by:§2.1.
- [20]A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret(2020)Transformers are rnns: fast autoregressive transformers with linear attention.InInternational conference on machine learning,pp. 5156–5165.Cited by:§1,§3.2.
- [21]D. Li, J. Li, X. Liu, X. Fan, and Y. Tian(2025)Asynchronous collaborative graph representation for frames and events.InProceedings of the Computer Vision and Pattern Recognition Conference,pp. 1655–1666.Cited by:§5.
- [22]Y. Li, H. Zhou, B. Yang, Y. Zhang, Z. Cui, H. Bao, and G. Zhang(2021)Graph-based asynchronous event processing for rapid object recognition.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp. 934–943.Cited by:§2.1,§4.2,Table 1,Table 2.
- [23]T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick(2014)Microsoft coco: common objects in context.InEuropean conference on computer vision,pp. 740–755.Cited by:§4.2.
- [24]I. Loshchilov and F. Hutter(2019)Decoupled weight decay regularization.InInternational Conference on Learning Representations,Cited by:§4.1.
- [25]N. Messikommer, D. Gehrig, A. Loquercio, and D. Scaramuzza(2020)Event-based asynchronous sparse convolutional networks.InEuropean Conference on Computer Vision,pp. 415–431.Cited by:§2.1,§4.2,Table 1,Table 2.
- [26]G. Orchard, A. Jayawant, G. K. Cohen, and N. Thakor(2015)Converting static image datasets to spiking neuromorphic datasets using saccades.Frontiers in neuroscience9,pp. 437.Cited by:§4.1.
- [27]A. Orvieto, S. L. Smith, A. Gu, A. Fernando, C. Gulcehre, R. Pascanu, and S. De(2023)Resurrecting recurrent neural networks for long sequences.InInternational conference on machine learning,pp. 26670–26698.Cited by:§3.4.
- [28]Y. Pan, Y. An, Z. Li, Y. Chou, R. Zhu, X. Wang, M. Wang, J. Wang, and G. Li(2026)Scaling linear attention with sparse state expansion.InThe Fourteenth International Conference on Learning Representations,Cited by:§2.2.
- [29]F. Paredes-Vallés, J. J. Hagenaars, J. Dupeyroux, S. Stroobants, Y. Xu, and G. C. de Croon(2024)Fully neuromorphic vision and control for autonomous drone flight.Science Robotics9(90),pp. eadi0591.Cited by:§1.
- [30]A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga,et al.(2019)Pytorch: an imperative style, high-performance deep learning library.Advances in neural information processing systems32.Cited by:§4.1.
- [31]B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski,et al.(2023)Rwkv: reinventing rnns for the transformer era.InFindings of the association for computational linguistics: EMNLP 2023,pp. 14048–14077.Cited by:§1.
- [32]Y. Peng, H. Li, Y. Zhang, X. Sun, and F. Wu(2024)Scene adaptive sparse transformer for event-based object detection.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp. 16794–16804.Cited by:§4.2,Table 1.
- [33]Y. Peng, Y. Zhang, Z. Xiong, X. Sun, and F. Wu(2023)Get: group event transformer for event-based vision.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp. 6038–6048.Cited by:§1,§4.2,Table 1.
- [34]E. Perot, P. De Tournemire, D. Nitti, J. Masci, and A. Sironi(2020)Learning to detect objects with a 1 megapixel event camera.Advances in Neural Information Processing Systems33,pp. 16639–16652.Cited by:§4.1.
- [35]C. R. Qi, H. Su, K. Mo, and L. J. Guibas(2017)Pointnet: deep learning on point sets for 3d classification and segmentation.InProceedings of the IEEE conference on computer vision and pattern recognition,pp. 652–660.Cited by:§2.1.
- [36]R. Santambrogio, M. Cannici, and M. Matteucci(2024)Farse-cnn: fully asynchronous, recurrent and sparse event-based cnn.InEuropean conference on computer vision,pp. 1–18.Cited by:§1,§1,§2.1,§3.4,§4.2,Table 1.
- [37]S. Schaefer, D. Gehrig, and D. Scaramuzza(2022)Aegnn: asynchronous event-based graph neural networks.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp. 12371–12381.Cited by:§1,§1,§2.1,§4.2,Table 1,Table 2.
- [38]M. Schöne, N. M. Sushma, J. Zhuge, C. Mayr, A. Subramoney, and D. Kappel(2024)Scalable event-by-event processing of neuromorphic sensory signals with deep state-space models.In2024 International Conference on Neuromorphic Systems (ICONS),pp. 124–131.Cited by:§1,§2.1.
- [39]Y. Sekikawa, K. Hara, and H. Saito(2019)Eventnet: asynchronous recursive event processing.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp. 3887–3896.Cited by:§1,§2.1.
- [40]Y. Sekikawa, J. Nagata, I. Araki, and A. Girbau(2026)CoL2A: convolution-free local linear attention for spatiotemporal event processing.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp. 4869–4880.Cited by:§2.2.
- [41]T. Soydan, N. Zubić, N. Messikommer, S. Mishra, and D. Scaramuzza(2024)S7: selective and simplified state space layers for sequence modeling.arXiv preprint arXiv:2410.03464.Cited by:§1,§2.1.
- [42]P. Tillet, H. Kung, and D. Cox(2019)Triton: an intermediate language and compiler for tiled neural network computations.InProceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages,pp. 10–19.Cited by:§3.4.
- [43]C. M. Turrero, M. Bouvier, M. Breitenstein, P. Zanuttigh, and V. Parret(2024)ALERT-transformer: bridging asynchronous and synchronous machine learning for real-time event-based spatio-temporal data.InForty-first International Conference on Machine Learning,Cited by:§2.1.
- [44]A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin(2017)Attention is all you need.Advances in neural information processing systems30.Cited by:§3.2.
- [45]N. Yang, Y. Wang, Z. Liu, M. Li, Y. An, and X. Zhao(2025)Smamba: sparse mamba for event-based object detection.InProceedings of the AAAI Conference on Artificial Intelligence,Vol.39,pp. 9229–9237.Cited by:§4.2,Table 1.
- [46]S. Yang, B. Wang, Y. Shen, R. Panda, and Y. Kim(2024)Gated linear attention transformers with hardware-efficient training.InInternational Conference on Machine Learning,pp. 56501–56523.Cited by:§1,§3.2.
- [47]S. Yang, B. Wang, Y. Zhang, Y. Shen, and Y. Kim(2024)Parallelizing linear transformers with the delta rule over sequence length.Advances in neural information processing systems37,pp. 115491–115522.Cited by:§1.
- [48]X. Zhang, M. Hu, S. Lu, S. Kim, E. Y. Lee, Y. Liu, and W. D. Lu(2026)Compute-in-memory implementation of state space models for event sequence processing.Nature Communications17(1),pp. 1513.Cited by:§4.3.
- [49]N. Zubić, D. Gehrig, M. Gehrig, and D. Scaramuzza(2023)From chaos comes order: ordering event representations for object recognition and detection.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp. 12846–12856.Cited by:§1,§4.2,Table 1,§5.
- [50]N. Zubić, M. Gehrig, and D. Scaramuzza(2024)State space models for event cameras.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),pp. 5819–5828.Cited by:§1.
Similar Articles
@songhan_mit: Speed-of-light block sparse attention :
Sol-Engine weekly update announces integration of Sol Attention, a training-free sparse attention method for video diffusion, with full paper coming next week.
SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference
SparDA proposes a decoupled sparse attention architecture that adds a lightweight 'Forecast' projection to predict future KV cache needs, enabling lookahead prefetching from CPU to GPU and reducing selection overhead. On 8B sparse-pretrained models, it achieves up to 1.25× prefill and 1.7× decode speedup, with up to 5.3× higher decode throughput over non-offload baselines.
LVSA: Training-Free Sparse Attention for Long Video Diffusion
LVSA introduces a training-free sparse attention mechanism for video diffusion models, reducing compute up to 3.17x while enabling generation beyond training horizons without quality loss.
@eliebakouch: the new sparse attention method introduced with this model is basically a combination of components from existing ones.…
Meituan introduces LongCat-2.0, a 1.6T parameter MoE model with 48B active parameters and 1M context length, featuring a new LongCat Sparse Attention (LSA) method that combines components from existing sparse attention techniques.
Elastic Attention Cores for Scalable Vision Transformers [R]
This article presents a new paper on Elastic Attention Cores for Vision Transformers, proposing a core-periphery block-sparse attention structure that improves scalability and accuracy compared to dense self-attention methods like DINOv3.