token-prediction

Tag

Cards List
#token-prediction

Why can't we make MoE routers predict experts needed in the next 5-10 tokens?

Reddit r/LocalLLaMA ↗ · 2026-08-27

The user questions whether Mixture of Experts (MoE) routers can be designed to predict future expert needs for token sequences to enable faster caching between RAM and VRAM, or if a separate neural network could be trained for this purpose.

0 favorites 0 likes
#token-prediction

Current AI Agents Are Overhyped and Fundamentally Limited

Reddit r/AI_Agents ↗ · 2026-08-23

The article criticizes current AI agents as overhyped and fundamentally limited, arguing that they are essentially LLMs with scaffolding and that their reliance on next-token prediction makes them unreliable for long-horizon planning and accountability.

0 favorites 0 likes
#token-prediction

Which tokens does a hybrid model predict better?

Hugging Face Blog ↗ · 2026-06-25 Cached

A study comparing Olmo Hybrid and Olmo 3 transformers at the token level shows hybrid models better predict meaningful tokens like nouns/verbs, while transformers excel at copying tokens from input.

0 favorites 0 likes
#token-prediction

@MatthieuWyart: LLMs learn by predicting tokens. World models (JEPA, data2vec) learn by predicting their own abstractions. Which needs …

X AI KOLs Timeline ↗ · 2026-06-01 Cached

This paper proves that learning by predicting latent representations (as in world models like JEPA and data2vec) requires exponentially less data than predicting tokens (as in LLMs) for hierarchical data with hidden structure.

0 favorites 0 likes
#token-prediction

Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper proposes STOP (SuperTOken for Pruning), a systematic framework for pruning inefficient reasoning paths early in parallel reasoning with Large Reasoning Models. The method achieves superior efficiency and effectiveness across models from 1.5B to 20B parameters, boosting GPT-OSS-20B accuracy on AIME25 from 84% to 90% under fixed compute budgets.

0 favorites 0 likes
← Back to home

Submit Feedback