human-data

Tag

Cards List
#human-data

@seclink: 视频里是加速了的,实际上动作慢很多 ...

X AI KOLs Timeline · 3d ago Cached

Lightwheel AI and Hugging Face have open-sourced EgoSuite-Open100K, the largest fully annotated egocentric human dataset with 100,000 hours of data, emphasizing scaling laws for physical AI.

0 favorites 0 likes
#human-data

@KimD0ing: Call for Papers: Scaling H2R @ CoRL 2026 Human data became the most important data for robot learning. But what can we …

X AI KOLs Timeline · 2026-08-09 Cached

Call for papers for the Scaling H2R workshop at CoRL 2026, focusing on scaling laws and diversity in human-to-robot learning. Submissions due Oct 7, 2026.

0 favorites 0 likes
#human-data

Should Reddit users care how their posts are being used to train AI?

Reddit r/artificial · 2026-07-03 Cached

This article argues that Reddit's messy, authentic human conversations are becoming increasingly valuable for training AI as the web fills with synthetic content, highlighting the economic shift toward scarce human behavioral data.

0 favorites 0 likes
#human-data

Re-Centering Humans in LLM Personalization

arXiv cs.CL · 2026-06-08 Cached

This paper studies the gap between synthetic and human data for evaluating LLM personalization across three stages: attribute extraction, relevance matching, and response generation. Results show models perform worse on real human data, and the authors introduce lightweight training interventions to improve alignment.

0 favorites 0 likes
#human-data

@Datou: Microsoft values its reputation, deliberately avoiding synthetic data. They trained a base model using only human data, then split it into three expert models for different domains. They then distilled these three capabilities back into the base model (weight ratio allocation requires experience), followed by a round of reinforcement learning to enable the distilled model to flexibly apply the right capability based on the problem.

X AI KOLs Timeline · 2026-06-02 Cached

Microsoft releases technical details of MAI-Thinking-1 training: uses purely human data to train a base model, then trains three domain expert models, merges capabilities back into the base model via distillation, and then applies reinforcement learning to enable the model to flexibly utilize different capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback