conditional-computation

Tag

Cards List
#conditional-computation

OceanMoE: Structured Conditional Sparse Computation for Long-Horizon Multivariate Ocean Forecasting

arXiv cs.LG ↗ · 2026-09-18 Cached

OceanMoE introduces a Mixture-of-Experts framework with structured conditional sparse computation to balance shared ocean context and adaptive specialization for multivariate ocean forecasting, demonstrating improved accuracy in long-horizon predictions.

0 favorites 0 likes
#conditional-computation

Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence

arXiv cs.LG ↗ · 2026-07-28 Cached

This paper introduces Accessibility Plasticity, a principle where neural systems adapt by reorganizing which existing computations can interact, rather than solely modifying parameters. A proof-of-concept on sequential learning tasks shows that accessibility adaptation reduces capability modification while maintaining comparable performance.

0 favorites 0 likes
#conditional-computation

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

arXiv cs.LG ↗ · 2026-07-09 Cached

TriRoute introduces a single lightweight controller that jointly decides attention mode, expert selection, and KV-cache bit-width for each token, achieving superior efficiency and robustness compared to independently tuned combinations of MoD, MoE, and KV-quantization.

0 favorites 0 likes
#conditional-computation

Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs

arXiv cs.CL ↗ · 2026-06-10 Cached

This paper proposes a dense-to-sparse continual training method for LLMs, using a predictor-gated bank-wise sparsity to achieve 4x FFN sparsity, and demonstrates it on Qwen2.5-8B with long-context training.

0 favorites 0 likes
#conditional-computation

Learning sparse neural networks through L₀ regularization

OpenAI Blog ↗ · 2017-12-04 Cached

OpenAI proposes a practical L₀ regularization method for neural networks that encourages weights to become exactly zero during training, enabling network pruning for improved speed and generalization. The method uses stochastic gates and introduces the hard concrete distribution to make the non-differentiable L₀ norm optimization tractable via gradient descent.

0 favorites 0 likes
← Back to home

Submit Feedback