Tag
This paper presents a platform based on FitLayout for creating visual-aware representations of web pages to support machine learning applications, demonstrating its use with graph neural networks for recognizing key content elements.
BrainCause framework uses generative and brain models to identify causal neural representations in the human brain, demonstrating that activation alone is insufficient for confirming concept representation.
This paper introduces RLA-WM, a visual feature-based world model that leverages residual latent actions and flow matching to efficiently predict future visual states. The method outperforms existing video-diffusion and feature-based approaches while enabling novel robot learning techniques from offline, actionless demonstration videos.