feed-forward-networks

Tag

Cards List
#feed-forward-networks

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

arXiv cs.CL · 2026-08-24 Cached

TriPLU introduces direct trilinear product feed-forward networks for tiny decoder-only language models, showing improved validation loss over SwiGLU in low-compute regimes.

0 favorites 0 likes
#feed-forward-networks

A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers

arXiv cs.LG · 2026-08-18 Cached

A reproducibility study reveals asymmetric effects when removing residual connections in Pre-LN transformers: attention-skip removal leads to collapse, while FFN-skip removal allows partial recovery at smaller scales.

0 favorites 0 likes
#feed-forward-networks

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

arXiv cs.CL · 2026-08-05 Cached

The paper introduces DUD (Decoupled Update Dynamics), a framework that separates Feed-Forward Network and Attention contributions via causal interventions to improve uncertainty quantification and calibration in large language models, outperforming state-of-the-art baselines.

0 favorites 0 likes
#feed-forward-networks

Searching the Space of Feed-Forward Neural-Network Weight-Update Rules with Fixed Depth Symbolic Regression

arXiv cs.LG · 2026-07-27 Cached

This paper investigates using symbolic regression to discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks, achieving an aggregate MSE reduction of 44.47% in 25 out of 30 benchmark/network combinations.

0 favorites 0 likes
#feed-forward-networks

More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval

arXiv cs.LG · 2026-07-21 Cached

This paper investigates the signed nature of FFN residual writes in long-context retrieval, finding that FFN writes act as suppressors or amplifiers depending on layer and task, and proposes a gradient-based diagnostic to distinguish these roles.

0 favorites 0 likes
#feed-forward-networks

Kolmogorov--Arnold Networks for Small Language Models

arXiv cs.AI · 2026-07-20 Cached

This paper evaluates Kolmogorov-Arnold Networks (KANs) as interpretable components and replacements for transformer feed-forward networks in small language models, finding that while KANs provide a practical audit interface, they show no consistent benchmark advantage over MLP baselines.

0 favorites 0 likes
← Back to home

Submit Feedback