Tag
This paper conducts a scaling study for fMRI foundation models, revealing that performance depends on the combination of pretraining data size, model size, and training duration, not just compute.
This paper presents a controlled scaling study comparing lexical, dense, graph-based, and agentic RAG paradigms across corpus sizes from 1,000 to 512,000 documents, finding that BM25 provides the best accuracy-cost tradeoff, while graph-based RAG faces high construction costs that limit scalability.
A controlled scaling study of retrieval-augmented generation paradigms finds that BM25 lexical retrieval outperforms agentic and graph-based retrieval at scale, while agentic search only leads on small corpora.
The paper presents the largest controlled scaling study for Earth-observation foundation models, showing that pretraining loss poorly predicts downstream performance and providing an optimal compute allocation rule. It trains scaled pixel-wise models (0.5B and 1B parameters) and distills them into compact student models that outperform larger open and proprietary models.
This paper introduces the Very Big Video Reasoning (VBVR) dataset and benchmark, a large-scale resource with over one million video clips across 200 reasoning tasks, enabling systematic study of spatiotemporal reasoning and showing early signs of emergent generalization.