Tag
This paper investigates using symbolic regression to discover explicit neural network weight-update rules that outperform standard hand-designed optimizers on small symbolic regression benchmarks, achieving an aggregate MSE reduction of 44.47% in 25 out of 30 benchmark/network combinations.
This paper investigates the signed nature of FFN residual writes in long-context retrieval, finding that FFN writes act as suppressors or amplifiers depending on layer and task, and proposes a gradient-based diagnostic to distinguish these roles.
This paper evaluates Kolmogorov-Arnold Networks (KANs) as interpretable components and replacements for transformer feed-forward networks in small language models, finding that while KANs provide a practical audit interface, they show no consistent benchmark advantage over MLP baselines.