Tag
Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.
Call for papers for the Scaling H2R workshop at CoRL 2026, focusing on scaling laws and diversity in human-to-robot learning. Submissions due Oct 7, 2026.
This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.
This paper studies the gap between synthetic and human data for evaluating LLM personalization across three stages: attribute extraction, relevance matching, and response generation. Results show models perform worse on real human data, and the authors introduce lightweight training interventions to improve alignment.
Microsoft releases technical details of MAI-Thinking-1 training: uses purely human data to train a base model, then trains three domain expert models, merges capabilities back into the base model via distillation, and then applies reinforcement learning to enable the model to flexibly utilize different capabilities.