Tag
该论文提出 SW-KAN,一种基于 Stieltjes-Wigert q-正交多项式的 Kolmogorov-Arnold 网络,通过指数-tanh 域映射和 O(N) 三递推计算,在图像分类与函数逼近任务中实现更优的精度-效率权衡。
UCLA researchers unveiled a hybrid optical-neural processor that detects deepfake videos with ~97.79% accuracy while analyzing 15-18 video streams in parallel, using light propagation to replace energy-hungry digital decoding. The system achieves 99.86% sensitivity and can be deepened with passive diffractive layers for higher accuracy at minimal energy cost.
The paper proposes CS-MoE, a novel Transformer architecture that shares neural experts across layers to improve parameter utilization, achieving lower perplexity with only 55% of parameters activated.
GAPS introduces dimension-level gating for conditional activation steering in language models, combining static and dynamic gates to selectively intervene and improve behavior-capability trade-off, with significant gains in toxicity mitigation and concept removal tasks.
This paper proposes 4MAS, a novel neural architecture inspired by biological bilaterality and memory consolidation, to address catastrophic forgetting in lifelong learning, achieving competitive results on benchmark datasets.
This paper introduces MixerLoop, a method that allocates recurrent compute by selectively looping the mixer component in language models while applying the feed-forward network once, achieving performance improvements with reduced computational costs.
SAMPAT is a three-layer neural architecture that can learn smooth, interpretable functions via closed-form algebraic expressions, offering competitive performance while enabling full interpretability.
This paper presents an empirical study comparing how different neural architectures (MLPs, CNNs, RNNs, pretrained transformers) degrade under temporal distribution shift across image and text domains, finding that models exploiting localized features degrade fastest while pretrained encoders drift more gradually.
GoldWorm is a native Rust cognitive engine that uses the biologically mapped 302-neuron connectome of C. elegans to process language through a transparent, zero-trust architecture with dual-stream processing and Hebbian associative memory.
Introduces ChainzRule, a neural architecture using Polynomial Engine and Differential Regularization to balance accuracy, hardware efficiency, and functional stability, outperforming standard models with 15.5x fewer parameters and smoother gradients.
WorldString is a neural architecture that models object state manifolds from point clouds or RGB-D video streams, serving as a foundational component for physical world models with differentiable structure for policy learning integration.
Researchers introduce Raven, a novel sequence model that merges state space model efficiency with a selective slot-updating mechanism inspired by sliding window attention to improve long-context retrieval. The approach offers a more principled alternative to existing linear-time models.
CTNet introduces a novel neural architecture where computation is framed as the evolution of a persistent state rather than successive rewrites, incorporating re-entrant memory, multi-scale coherence, and projective output.
OpenAI introduces the Sparse Transformer, a deep neural network that improves the attention mechanism from O(N²) to O(N√N) complexity, enabling modeling of sequences 30x longer than previously possible across text, images, and audio. The model uses sparse attention patterns and checkpoint-based memory optimization to train networks up to 128 layers deep, achieving state-of-the-art performance across multiple domains.