Tag
A Statista chart estimates that approximately a quarter of the training data for Meta's Llama AI model came from books, highlighting the composition of sources used in large language model development.
This paper organizes embodied data sources into a five-level pyramid (real-robot, UMI, egocentric/exocentric, simulation, general vision-language), analyzing their trade-offs between scalability and robot alignment, and reviews recent embodied foundation models in terms of data recipes. It also discusses open challenges for building next-generation embodied systems.