Tag
SMart is a new time series representation learning framework that uses multi-phase recurrence plots recovery and a source dataset selector to enhance representation transfer from multiple datasets, showing improved performance in classification and regression tasks.
This paper presents V2TATC, a joint voice–trajectory embedding framework for air traffic controller situational awareness, and introduces a novel dataset for cross-modal retrieval experiments in congested airspaces.
The paper proposes SciJEPA, a citation-free framework for scientific document representation using asymmetric within-document predictive learning, and demonstrates that regularization improves performance.
Google Research introduces GlucoFM, a lightweight self-supervised foundation model for continuous glucose monitoring that improves metabolic prediction tasks.
LingBot-Vision is a family of self-supervised vision transformer backbones for dense spatial perception, using masked boundary modeling to capture semantic and geometric structures for tasks like depth estimation and segmentation.
PRISM is a training-free test-time adaptation method that reverses low-rank affine noise distortions in audio-text models using frozen text prototypes, showing significant improvements under severe acoustic noise.
Introduces SBCO, a self-supervised verifier-grounded harness optimizer for planning agents that improves agent outputs via approximate block coordinate ascent, matching or exceeding self-modifying baselines with far less compute.
This paper introduces a bidirectional latent diffusion model that steps dynamical systems forward or backward in time, using round-trip consistency as a self-supervised test-time error signal to predict rollout errors without ground truth or ensembles.
This paper introduces Self-Supervised Skill Optimization (SSO), a framework that learns and optimizes reusable agent skills from unlabeled task instances using LLM-judged pairwise comparisons, without requiring ground-truth labels or rewards. SSO outperforms existing ground-truth-free prompt optimizers and approaches ground-truth-based methods on closed-ended benchmarks.
The paper proposes an unsupervised self-evolving agent framework inspired by diffusion models, using self-supervised semantic diffusion to train external skill libraries for LLM agents in specialized domains like creative screenwriting, without requiring weight access or external supervision.
PhiZero is a physical world model that learns a compact discrete representation called 'physical language' from videos and uses it to reason about world state transitions before rendering future videos, improving physical coherence in generation and understanding tasks.
Proposes Self Gradient Forcing (SGF), a two-pass training strategy for autoregressive video diffusion models that provides missing supervision for writing useful context memory, enabling strong long-video extrapolation even from short training windows.
Presents a multimodal voice activity projection framework extending audio-only VAP to audio-visual inputs for turn-taking prediction in social robots, using pretrained backbones and low-rank adaptation. Achieves improvements on NoXi and Haru EDR corpora.
A researcher reports a surprising 50-point accuracy gap between frozen SigLIP2 (92%) and DINOv2 (41%) embeddings on a fine-grained car classification task using k-NN, seeking insight on whether a linear probe would close the gap or if DINOv2 is unsuited for retrieval.
This paper introduces Lorentz Encoding (LE), a physics-informed framework that uses implicit neural representations and physical constraints to reconstruct high-resolution CEST MRI from sparsely sampled data, achieving superior performance over existing methods.
LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.
Robbyant, an embodied AI company under Ant Group, released LingBot-Vision, a self-supervised vision backbone family ranging from 21M to 1.1B parameters, under Apache-2.0. It matches or beats DINOv3 on several depth and segmentation benchmarks despite using less than one third of the training data, highlighting a push for open perception models.
AdaJEPA introduces an adaptive latent world model that continuously updates during test-time via closed-loop model predictive control, significantly improving planning success under distribution shift.
Proposes SAOT, a structure-aware optimal transport framework for self-supervised continual graph learning that preserves relational structure across tasks. Achieves significant performance gains over state-of-the-art methods on multiple benchmarks, including up to 15% improvement on Products-CL.
This paper proposes a self-supervised theorem-discovery algorithm that starts from axioms and inference rules alone, building a theorem library without human priors. Experiments show the discovered theorems are meaningful and improve LLM proof performance.