Tag
The user questions whether Mixture of Experts (MoE) routers can be designed to predict future expert needs for token sequences to enable faster caching between RAM and VRAM, or if a separate neural network could be trained for this purpose.
The article criticizes current AI agents as overhyped and fundamentally limited, arguing that they are essentially LLMs with scaffolding and that their reliance on next-token prediction makes them unreliable for long-horizon planning and accountability.
A study comparing Olmo Hybrid and Olmo 3 transformers at the token level shows hybrid models better predict meaningful tokens like nouns/verbs, while transformers excel at copying tokens from input.
This paper proves that learning by predicting latent representations (as in world models like JEPA and data2vec) requires exponentially less data than predicting tokens (as in LLMs) for hierarchical data with hidden structure.
This paper proposes STOP (SuperTOken for Pruning), a systematic framework for pruning inefficient reasoning paths early in parallel reasoning with Large Reasoning Models. The method achieves superior efficiency and effectiveness across models from 1.5B to 20B parameters, boosting GPT-OSS-20B accuracy on AIME25 from 84% to 90% under fixed compute budgets.