neural-architecture

Tag

Cards List
#neural-architecture

SW-KAN: Kolmogorov-Arnold Networks with Stieltjes-Wigert q-Orthogonal Polynomials

arXiv cs.LG ↗ · yesterday Cached

该论文提出 SW-KAN,一种基于 Stieltjes-Wigert q-正交多项式的 Kolmogorov-Arnold 网络,通过指数-tanh 域映射和 O(N) 三递推计算,在图像分类与函数逼近任务中实现更优的精度-效率权衡。

0 favorites 0 likes
#neural-architecture

This light-powered AI can spot deepfakes with nearly 98% accuracy

Reddit r/artificial ↗ · yesterday Cached

UCLA researchers unveiled a hybrid optical-neural processor that detects deepfake videos with ~97.79% accuracy while analyzing 15-18 video streams in parallel, using light propagation to replace energy-hungry digital decoding. The system achieves 99.86% sensitivity and can be deepened with passive diffractive layers for higher accuracy at minimal energy cost.

0 favorites 0 likes
#neural-architecture

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

arXiv cs.LG ↗ · 2026-09-22 Cached

The paper proposes CS-MoE, a novel Transformer architecture that shares neural experts across layers to improve parameter utilization, achieving lower perplexity with only 55% of parameters activated.

0 favorites 0 likes
#neural-architecture

GAPS: Dimension-Level Gates for Conditional Activation Steering

arXiv cs.CL ↗ · 2026-09-03 Cached

GAPS introduces dimension-level gating for conditional activation steering in language models, combining static and dynamic gates to selectively intervene and improve behavior-capability trade-off, with significant gains in toxicity mitigation and concept removal tasks.

0 favorites 0 likes
#neural-architecture

In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models

arXiv cs.LG ↗ · 2026-08-21 Cached

This paper proposes 4MAS, a novel neural architecture inspired by biological bilaterality and memory consolidation, to address catastrophic forgetting in lifelong learning, achieving competitive results on benchmark datasets.

0 favorites 0 likes
#neural-architecture

Allocating Recurrent Compute in Looped Language Models

arXiv cs.LG ↗ · 2026-08-20 Cached

This paper introduces MixerLoop, a method that allocates recurrent compute by selectively looping the mixer component in language models while applying the feed-forward network once, achieving performance improvements with reduced computational costs.

0 favorites 0 likes
#neural-architecture

All you need is SAMPAT

arXiv cs.LG ↗ · 2026-07-13 Cached

SAMPAT is a three-layer neural architecture that can learn smooth, interpretable functions via closed-form algebraic expressions, offering competitive performance while enabling full interpretability.

0 favorites 0 likes
#neural-architecture

Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift

arXiv cs.LG ↗ · 2026-07-08 Cached

This paper presents an empirical study comparing how different neural architectures (MLPs, CNNs, RNNs, pretrained transformers) degrade under temporal distribution shift across image and text domains, finding that models exploiting localized features degrade fastest while pretrained encoders drift more gradually.

0 favorites 0 likes
#neural-architecture

A native Rust cognitive engine that routes language through a biologically faithful neural substrate

Reddit r/artificial ↗ · 2026-06-30

GoldWorm is a native Rust cognitive engine that uses the biologically mapped 302-neuron connectome of C. elegans to process language through a transparent, zero-trust architecture with dual-stream processing and Hebbian associative memory.

0 favorites 0 likes
#neural-architecture

Layer-wise Derivative Controlled Networks

arXiv cs.LG ↗ · 2026-05-18 Cached

Introduces ChainzRule, a neural architecture using Polynomial Engine and Differential Regularization to balance accuracy, hardware efficiency, and functional stability, outperforming standard models with 15.5x fewer parameters and smoother gradients.

0 favorites 0 likes
#neural-architecture

Actionable World Representation

Hugging Face Daily Papers ↗ · 2026-05-18 Cached

WorldString is a neural architecture that models object state manifolds from point clouds or RGB-D video streams, serving as a foundational component for physical world models with differentiable structure for policy learning integration.

0 favorites 0 likes
#neural-architecture

@_albertgu: Introducing a new sequence model Raven which pushes the boundary of fixed-state-size sequence models! Raven bridges pop…

X AI KOLs Timeline ↗ · 2026-05-07

Researchers introduce Raven, a novel sequence model that merges state space model efficiency with a selective slot-updating mechanism inspired by sliding window attention to improve long-context retrieval. The approach offers a more principled alternative to existing linear-time models.

0 favorites 0 likes
#neural-architecture

He presentado CTNet: una arquitectura donde el cómputo ocurre como evolución de un estado persistente [D]

Reddit r/MachineLearning ↗ · 2026-04-23

CTNet introduces a novel neural architecture where computation is framed as the evolution of a persistent state rather than successive rewrites, incorporating re-entrant memory, multi-scale coherence, and projective output.

0 favorites 0 likes
#neural-architecture

Generative modeling with sparse transformers

OpenAI Blog ↗ · 2019-04-23 Cached

OpenAI introduces the Sparse Transformer, a deep neural network that improves the attention mechanism from O(N²) to O(N√N) complexity, enabling modeling of sequences 30x longer than previously possible across text, images, and audio. The model uses sparse attention patterns and checkpoint-based memory optimization to train networks up to 128 layers deep, achieving state-of-the-art performance across multiple domains.

0 favorites 0 likes
← Back to home

Submit Feedback