self-supervised

Tag

Cards List
#self-supervised

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

arXiv cs.CL · 2026-07-09 Cached

Presents a multimodal voice activity projection framework extending audio-only VAP to audio-visual inputs for turn-taking prediction in social robots, using pretrained backbones and low-rank adaptation. Achieves improvements on NoXi and Haru EDR corpora.

0 favorites 0 likes
#self-supervised

DINOv2 way worse than SigLIP in k-NN. Is this expected? [R]

Reddit r/MachineLearning · 2026-07-08

A researcher reports a surprising 50-point accuracy gap between frozen SigLIP2 (92%) and DINOv2 (41%) embeddings on a fine-grained car classification task using k-NN, seeking insight on whether a linear probe would close the gap or if DINOv2 is unsuited for retrieval.

0 favorites 0 likes
#self-supervised

Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding

arXiv cs.LG · 2026-07-08 Cached

This paper introduces Lorentz Encoding (LE), a physics-informed framework that uses implicit neural representations and physical constraints to reconstruct high-resolution CEST MRI from sparsely sampled data, achieving superior performance over existing methods.

0 favorites 0 likes
#self-supervised

@AdinaYakup: LingBot Vision A self-supervised vision backbone family for dense spatial perception from Ant Group @robbyant_brain - A…

X AI KOLs Timeline · 2026-07-07 Cached

LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.

0 favorites 0 likes
#self-supervised

Meta ships DINOv3 behind an access gate under its own license. Ant's Robbyant just shipped a full vision backbone family under Apache-2.0. What happens when perception goes free and small?

Reddit r/ArtificialInteligence · 2026-07-06

Robbyant, an embodied AI company under Ant Group, released LingBot-Vision, a self-supervised vision backbone family ranging from 21M to 1.1B parameters, under Apache-2.0. It matches or beats DINOv3 on several depth and segmentation benchmarks despite using less than one third of the training data, highlighting a push for open perception models.

0 favorites 0 likes
#self-supervised

@yingwww_: Warm take: Your world model should never stop learning Introducing AdaJEPA, an adaptive WM that plans, acts, and adapts…

X AI KOLs Following · 2026-07-05 Cached

AdaJEPA introduces an adaptive latent world model that continuously updates during test-time via closed-loop model predictive control, significantly improving planning success under distribution shift.

0 favorites 0 likes
#self-supervised

SAOT: Self-Supervised Continual Graph Learning with Structure-Aware Optimal Transport

arXiv cs.LG · 2026-07-02 Cached

Proposes SAOT, a structure-aware optimal transport framework for self-supervised continual graph learning that preserves relational structure across tasks. Achieves significant performance gains over state-of-the-art methods on multiple benchmarks, including up to 15% improvement on Products-CL.

0 favorites 0 likes
#self-supervised

Self-Supervised Theorem Discovery in a Formal Axiomatic System

arXiv cs.AI · 2026-06-30 Cached

This paper proposes a self-supervised theorem-discovery algorithm that starts from axioms and inference rules alone, building a theorem library without human priors. Experiments show the discovered theorems are meaningful and improve LLM proof performance.

0 favorites 0 likes
#self-supervised

DREAM: Dense Retrieval Embeddings via Autoregressive Modeling

Hugging Face Daily Papers · 2026-06-23 Cached

DREAM trains dense retrieval embeddings by using autoregressive language model attention to supervise query-document similarity, eliminating the need for labeled data. It consistently outperforms baselines on BEIR and RTEB benchmarks across model scales.

0 favorites 0 likes
#self-supervised

BadWorld: Adversarial Attacks on World Models

Hugging Face Daily Papers · 2026-06-15 Cached

BadWorld is a label-free adversarial framework that reveals structural vulnerabilities in visual world models by generating imperceptible perturbations that cause catastrophic failures in future rollouts.

0 favorites 0 likes
#self-supervised

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

arXiv cs.AI · 2026-06-11 Cached

This paper proposes a self-supervised reinforcement learning framework that uses consistency verifiers—reward functions checking geometric and semantic consistency under transformations—to improve spatial reasoning in large reasoning models without requiring ground-truth annotations. The method approaches the accuracy of supervised fine-tuning and generalizes across diverse tasks.

0 favorites 0 likes
#self-supervised

Pretrained self-supervised speech models can recognize unseen consonants

arXiv cs.CL · 2026-06-11 Cached

This paper investigates whether pretrained self-supervised speech models like Wav2Vec2 and HuBERT can accurately recognize click consonants, which are rare in training data, by fine-tuning on Khoisan languages. Results show the models recognize clicks more accurately than non-clicks, indicating generalization to uncommon phonemes.

0 favorites 0 likes
#self-supervised

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

arXiv cs.CL · 2026-06-10 Cached

A novel method for multilingual word-level forced alignment combines self-supervised representations from MMS and a phoneme boundary detector with a learned dynamic programming decoder, outperforming existing aligners on English and unseen languages without further training.

0 favorites 0 likes
#self-supervised

MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training

Hugging Face Daily Papers · 2026-06-07 Cached

MaskAlign proposes a token-subset representation alignment method that improves diffusion transformer training by reducing reliance on complete token sets and maintaining stable alignment under perturbations.

0 favorites 0 likes
#self-supervised

Self-supervised User Profile Generation for Personalization

arXiv cs.CL · 2026-06-05 Cached

Introduces BUMP, a self-supervised framework for training a profile generator for LLM personalization without task labels, using bidirectional in-batch ranking and GRPO. It matches or outperforms supervised methods on the LaMP benchmark.

0 favorites 0 likes
#self-supervised

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

Hugging Face Daily Papers · 2026-06-04 Cached

Retrospective Harness Optimization (RHO) is a self-supervised method that improves LLM agent performance using only past trajectories, achieving a 78% pass rate on SWE-Bench Pro without external grading.

0 favorites 0 likes
#self-supervised

Unsupervised Skill Discovery for Agentic Data Analysis

Hugging Face Daily Papers · 2026-06-04

DataCOPE is an unsupervised verifier-guided skill discovery framework for data-analytic agents that derives verifier signals from exploration trajectories without labeled supervision. It improves performance by 9.71% and 32.30% on report-style and reasoning-style data analysis tasks respectively.

0 favorites 0 likes
#self-supervised

MemTrain: Self-Supervised Context Memory Training

arXiv cs.CL · 2026-06-03 Cached

MemTrain proposes a self-supervised training framework that uses masked reconstruction and intermediate memory recall proxy tasks on Wikipedia corpora to enhance LLM agents' context memory, achieving up to 17.67 point gains on downstream memory-intensive QA benchmarks.

0 favorites 0 likes
#self-supervised

MindZero: Learning Online Mental Reasoning With Zero Annotations

arXiv cs.AI · 2026-06-02 Cached

MindZero introduces a self-supervised reinforcement learning framework that trains multimodal large language models for efficient and robust online mental reasoning without requiring mental state annotations, outperforming model-based methods in accuracy and efficiency.

0 favorites 0 likes
#self-supervised

RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

Hugging Face Daily Papers · 2026-05-29 Cached

RayDer is a unified feed-forward transformer that consolidates camera estimation, scene reconstruction, and rendering for self-supervised novel view synthesis from real-world video, achieving clean power-law scaling and strong zero-shot performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback