@Phoenixyin13: I think this is a top-notch work in ICML 2026. The attention mechanism of traditional Transformers is essentially point-to-point matching: it cuts input into a bunch of tokens (discrete points), computes similarity between Query and Key, and then weights the Value. In NLP...
Summary
Introduces the ICML 2026 paper Functional Attention, which treats functions as first-class citizens and replaces softmax point-to-point similarity with structured linear operators. It addresses issues of discretization, resolution sensitivity, and high computational complexity in traditional Transformers when handling continuous functions. Achieves or surpasses SOTA in tasks like PDE solving and 3D segmentation, and exhibits strong OOD generalization.
View Cached Full Text
Cached at: 06/25/26, 07:13 AM
I believe this is top-tier work at ICML 2026.
The attention mechanism in traditional Transformers is essentially a point-to-point matching: it cuts the input into a bunch of tokens (discrete points), computes the similarity between Query and Key, and then weights the Value.
This works well for NLP and images, but for continuous functions and fields — such as PDEs (partial differential equations) in physical simulation, 3D geometry, fluids, and material mechanics — it has serious issues:
- Continuity is forcibly discretized into tokens, ignoring the global functional structure.
- It is sensitive to resolution; changing the sampling rate easily breaks it.
- Computation is O(n²), which becomes expensive when n is large.
Thus, Functional Attention was born.
This time, FuncAttn treats the function itself as a first-class citizen. Instead of using softmax to compute point-to-point similarity, it employs a structured linear operator to model the functional correspondence between two function spaces.
The inspiration comes from Functional Maps in geometry processing, which cleverly transforms complex 3D shape correspondence problems into simple linear operators.
The result is more compact (k×k operators, where k << n), resolution-invariant, better at capturing global dependencies, achieves or surpasses SOTA on tasks such as PDE solving, 3D segmentation, and regression, and performs exceptionally well in OOD generalization.
In the future, Functional Attention will be highly practical for AI applications, especially in scientific computing + AI and scenarios that require handling continuous, physical-world data.
It will become a powerful upgrade component in the Neural Operator and Scientific AI domains.
This paper will take a big step toward making AI excel in the continuous physical world. It is extremely valuable for accelerating scientific research, engineering simulation, embodied AI, BCI, and other cutting-edge applications.
This is precisely one of the core infrastructures of my cognitive neuroscience vision.
Similar Articles
@tetsuoai: Attention is a lookup. Each token builds a query, compares it against every key in the sequence, and pulls value vector…
Explains attention in transformers as a lookup operation where each token builds a query, compares against keys, and retrieves weighted value vectors, with a video covering the full pipeline.
Functional Attention: From Pairwise Affinities to Functional Correspondences
Functional Attention is a novel attention mechanism that reinterprets attention as a functional correspondence between adaptive bases, replacing softmax affinities with structured linear operators inspired by geometric functional maps. The method achieves state-of-the-art performance on operator learning tasks including PDE solving and 3D segmentation while remaining resolution-invariant.
MassAlloc Attention: Let Attention Allocate Its Own Compute
The article introduces two attention mechanisms, CoWindow Attention and MassAlloc Attention, which optimize compute allocation in transformers, achieving significant speedups and reduced training FLOPs in benchmarks.
@waterloo_intern: After reading up a bit on ML research post transformer era, I was upset that it seems to have converged on hyper-optimi…
This tweet discusses the convergence of ML research on attention-based, matmul-optimized algorithms due to hardware constraints, drawing on the 'hardware lottery' concept and noting OpenAI's 9-month chip tape-out as a potential sign of hardware-research co-design.
@snowboat84: Continuing the discussion on applying physical models in AI. Today's Transformer uses attention to allow information at different positions in a sequence to interact, but this mixing is very likely lossy and irreversible. Independent pieces of information get blurred and lost as they mix, and it's impossible to precisely reverse-engineer the input from the output. What we want is a different kind of interaction...
The author discusses the problem of information loss and irreversibility caused by the Transformer's attention mechanism, and proposes drawing inspiration from the physical model of soliton propagation to design a reversible, zero-loss interaction layer as a direction for improving existing AI model architectures.