@Phoenixyin13: I think this is a top-notch work in ICML 2026. The attention mechanism of traditional Transformers is essentially point-to-point matching: it cuts input into a bunch of tokens (discrete points), computes similarity between Query and Key, and then weights the Value. In NLP...

X AI KOLs Timeline Papers

Summary

Introduces the ICML 2026 paper Functional Attention, which treats functions as first-class citizens and replaces softmax point-to-point similarity with structured linear operators. It addresses issues of discretization, resolution sensitivity, and high computational complexity in traditional Transformers when handling continuous functions. Achieves or surpasses SOTA in tasks like PDE solving and 3D segmentation, and exhibits strong OOD generalization.

I think this is a top-notch work in ICML 2026. The attention mechanism of traditional Transformers is essentially point-to-point matching: it slices input into a bunch of tokens (discrete points), computes the similarity between Query and Key, and then weights the Value. This works well for NLP and images, but for continuous functions and fields—such as PDEs in physics simulations, 3D geometry, fluids, and material mechanics—it has serious issues: 1. Continuous data is forcibly discretized into tokens, ignoring the global function structure. 2. It is sensitive to resolution; changing the sampling rate easily breaks it. 3. Computation is O(n²), expensive when n is large. Hence, **Functional Attention** emerged. This time, FuncAttn treats the function itself as a first-class citizen, and instead of using softmax to compute point-to-point similarity, it uses **structured linear operators** to model functional correspondences between two function spaces. Inspired by **Functional Maps** from geometry processing, which turns complex 3D shape correspondence problems into simple linear operators, this approach is quite clever. The result is more compact (k×k operator, k << n), resolution-invariant, captures global dependencies better, achieves or surpasses SOTA in tasks like PDE solving, 3D segmentation, and regression, and even excels in OOD generalization. In the future, Functional Attention will be highly useful for AI applications, particularly in scientific computing + AI and scenarios that require processing continuous physical-world data. It will become a powerful upgrade component for Neural Operators and scientific AI. This paper will move AI a big step forward toward mastering the continuous physical world. It is extremely valuable for accelerating research, engineering simulations, embodied AI, BCI, and other cutting-edge applications. This is one of the core infrastructures of my cognitive neuroscience vision.
Original Article
View Cached Full Text

Cached at: 06/25/26, 07:13 AM

I believe this is top-tier work at ICML 2026.

The attention mechanism in traditional Transformers is essentially a point-to-point matching: it cuts the input into a bunch of tokens (discrete points), computes the similarity between Query and Key, and then weights the Value.

This works well for NLP and images, but for continuous functions and fields — such as PDEs (partial differential equations) in physical simulation, 3D geometry, fluids, and material mechanics — it has serious issues:

  1. Continuity is forcibly discretized into tokens, ignoring the global functional structure.
  2. It is sensitive to resolution; changing the sampling rate easily breaks it.
  3. Computation is O(n²), which becomes expensive when n is large.

Thus, Functional Attention was born.

This time, FuncAttn treats the function itself as a first-class citizen. Instead of using softmax to compute point-to-point similarity, it employs a structured linear operator to model the functional correspondence between two function spaces.

The inspiration comes from Functional Maps in geometry processing, which cleverly transforms complex 3D shape correspondence problems into simple linear operators.

The result is more compact (k×k operators, where k << n), resolution-invariant, better at capturing global dependencies, achieves or surpasses SOTA on tasks such as PDE solving, 3D segmentation, and regression, and performs exceptionally well in OOD generalization.

In the future, Functional Attention will be highly practical for AI applications, especially in scientific computing + AI and scenarios that require handling continuous, physical-world data.

It will become a powerful upgrade component in the Neural Operator and Scientific AI domains.

This paper will take a big step toward making AI excel in the continuous physical world. It is extremely valuable for accelerating scientific research, engineering simulation, embodied AI, BCI, and other cutting-edge applications.

This is precisely one of the core infrastructures of my cognitive neuroscience vision.

Similar Articles

Functional Attention: From Pairwise Affinities to Functional Correspondences

Hugging Face Daily Papers

Functional Attention is a novel attention mechanism that reinterprets attention as a functional correspondence between adaptive bases, replacing softmax affinities with structured linear operators inspired by geometric functional maps. The method achieves state-of-the-art performance on operator learning tasks including PDE solving and 3D segmentation while remaining resolution-invariant.

MassAlloc Attention: Let Attention Allocate Its Own Compute

Hugging Face Daily Papers

The article introduces two attention mechanisms, CoWindow Attention and MassAlloc Attention, which optimize compute allocation in transformers, achieving significant speedups and reduced training FLOPs in benchmarks.

@snowboat84: Continuing the discussion on applying physical models in AI. Today's Transformer uses attention to allow information at different positions in a sequence to interact, but this mixing is very likely lossy and irreversible. Independent pieces of information get blurred and lost as they mix, and it's impossible to precisely reverse-engineer the input from the output. What we want is a different kind of interaction...

X AI KOLs Timeline

The author discusses the problem of information loss and irreversibility caused by the Transformer's attention mechanism, and proposes drawing inspiration from the physical model of soliton propagation to design a reversible, zero-loss interaction layer as a direction for improving existing AI model architectures.