Tag
OceanMoE introduces a Mixture-of-Experts framework with structured conditional sparse computation to balance shared ocean context and adaptive specialization for multivariate ocean forecasting, demonstrating improved accuracy in long-horizon predictions.
This paper introduces Accessibility Plasticity, a principle where neural systems adapt by reorganizing which existing computations can interact, rather than solely modifying parameters. A proof-of-concept on sequential learning tasks shows that accessibility adaptation reduces capability modification while maintaining comparable performance.
TriRoute introduces a single lightweight controller that jointly decides attention mode, expert selection, and KV-cache bit-width for each token, achieving superior efficiency and robustness compared to independently tuned combinations of MoD, MoE, and KV-quantization.
This paper proposes a dense-to-sparse continual training method for LLMs, using a predictor-gated bank-wise sparsity to achieve 4x FFN sparsity, and demonstrates it on Qwen2.5-8B with long-context training.
OpenAI proposes a practical L₀ regularization method for neural networks that encourages weights to become exactly zero during training, enabling network pruning for improved speed and generalization. The method uses stochastic gates and introduces the hard concrete distribution to make the non-differentiable L₀ norm optimization tractable via gradient descent.