Tag
This paper proposes learnable wavelet activations to combat plasticity loss in continual learning, decomposing activations into low- and high-frequency components with dynamic injection and regularization, achieving state-of-the-art results on benchmarks.
Proposes a deep wireless physical neural network where multi-hop MIMO relays realize trainable linear transforms and power amplifier nonlinearities serve as activation functions, enabling over-the-air inference for image classification.
An Arxiv AI paper on polynomial approximation of activation functions ends with the author's wish to marry his high school sweetheart, highlighting a personal touch in Chinese tech academia.
DECO is a sparse MoE architecture that matches dense Transformer performance with only 20% activated experts and a 3x acceleration kernel, utilizing ReLU-based routing, learnable scaling, and the NormSiLU activation function.
Blog post surveys fast hyperbolic tangent approximations—Taylor, Padé, splines, and bit-level tricks—for neural-network and real-time audio use.