Tag
SafeAtlas-VL introduces a large-scale multimodal safety dataset with graded risk labels and trains guard models that predict continuous safety scores, achieving state-of-the-art generalization in multimodal safety assessment.
The paper introduces LibriBrain100, a large-scale MEG dataset with over 100 hours of data for neural speech decoding, achieving state-of-the-art performance on word classification benchmarks.
TouchThinker introduces a million-scale tactile reasoning dataset and benchmark to scale tactile commonsense reasoning to open-world settings, using action-aware representation for efficient reasoning.
Introduces MMIOC-1M, a large-scale multi-modal benchmark for industrial defect detection, and proposes RTVPNet, a refined text-visual prompt network achieving state-of-the-art performance.
LocateAnything proposes Parallel Box Decoding for unified visual grounding and object detection, decoding geometric elements as atomic units to improve throughput and localization accuracy, supported by a large-scale dataset of 138M samples.