Tag
This paper proposes a method for improving frozen forecasters by combining static and online correctors to control downside risk, demonstrating gains up to 11.5% in benchmarks and reduced errors in electricity load forecasting.
This paper introduces randomized-pass replay (RPR) to bound rehearsal gaps in online continual learning, showing improved accuracy over independent class-balanced retrieval in experience replay methods like ER-ACE.
This paper introduces the Lightweight Ranking Heads framework to accelerate multi-task experimentation in production recommender systems by enabling dynamic task injection without retraining backbone models, reducing iteration cycles from weeks to days at YouTube scale.
A promotional post for TensorTonic, an online platform for learning machine learning through coding practice, starting with an explanation of the RBF kernel.
This paper improves the optimal expected mistake bound for randomized proper online learning, showing it is O(L(H) log T), which is optimal up to a universal constant for worst-case classes.
The paper introduces RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification that optimizes question selection based on retrieval feedback to improve performance across multiple interaction rounds.
Keet is an educational platform launched on Product Hunt, offering interactive video courses on any topic with AI integration.
This paper introduces HACK GPs, a method that uses online learning with expert advice for kernel selection in Gaussian Processes, enhancing robustness in sequential decision-making tasks such as Bayesian optimization and active learning.
The paper introduces the Free Inference dimension, a combinatorial complexity measure for zero-collision navigation in meta-reinforcement learning under hypothesis mixtures, proving its relationship with VC-dimension and generalization bounds.
The paper introduces MASA, a method that uses frozen multimodal large language models to break the self-referential loop in wild test-time adaptation by providing structured semantic descriptions for more reliable adaptation.
This paper proposes an online method for warped Gaussian processes that jointly updates latent GP moments and optimizes warping parameters using exact recursive gradient computation.
This paper proposes a variational Bayesian last layer method for graph neural networks to perform online node classification under distribution shift, achieving superior accuracy and uncertainty calibration compared to baselines.
ReCAST proposes a method for per-reward, timestep-dependent credit assignment in diffusion model fine-tuning, separating user preferences from temporal allocation to improve alignment and informativeness.
A curated list of 10 free online courses for content creation, covering topics such as content marketing, blogging, and monetization, offered by platforms like HubSpot and Alison.
Dictoterix is a free language learning service that uses the dichotic method, playing native and target languages in different ears simultaneously to improve learning efficiency.
Microsoft Research presents FrogNano, a 4B coding agent trained exclusively via reinforcement learning with online task synthesis, achieving competitive performance without distillation from larger models.
This paper proposes algorithms for adaptively routing prompts to LLM experts in an online setting with limited feedback, formulated as a bandit problem to minimize regret and maximize response quality.
HyperTrace is a training-free framework for online LLM personalization that uses interpretable natural-language hypotheses to trace latent user preferences, improving response alignment and consistency.
This paper formulates the adaptive routing of prompts to large language model experts as a contextual bandit problem with limited feedback, proposing algorithms that achieve sublinear regret and demonstrate efficient learning of high-quality routing strategies.
A tweet recommends LabEx, an online platform for hands-on learning in skills like Linux, Docker, and cybersecurity, featuring interactive environments and Chinese support.