Tag
This paper introduces a label-free method to find a valence axis from nine emotion examples that transfers across text, vision, audio, and brain modalities, achieving competitive sentiment classification with minimal labels.
This paper investigates the scaling properties of native multimodal pre-training, deriving compute-optimal model sizes and token counts under a fixed budget, and revealing distinct scaling behaviors for language and multimodal objectives.
Researchers at MIT Lincoln Laboratory propose 'principle-driven foundation models' that encode signal-theoretic physical principles (Fourier decomposition, energy conservation, symmetry) instead of learning statistical correlations from large paired datasets. Trained exclusively on RF data, their 1.99M parameter frozen encoder achieves 77.7% average accuracy across 15 diverse tasks spanning audio, images, text, and video without any fine-tuning on target domains.