Tag
This paper proposes a neural ODE-based regularization method that enforces latent embeddings in reinforcement learning agents to follow consistent ODE flows, aligning representation learning with environment dynamics and yielding performance gains on Atari and gridworld benchmarks.
This paper introduces MiGHT-EHR, a multi-task graph transformer for heterogeneous temporal EHR data, jointly modeling clinical entities, temporal trajectories, and task dependencies. It outperforms state-of-the-art methods on MIMIC-III and MIMIC-IV across drug recommendation, length-of-stay, mortality, and readmission prediction.
CellWorld introduces a latent-space predictive pretraining approach for spatial transcriptomics foundation models, predicting latent representations of masked cells instead of reconstructing gene measurements. Across held-out datasets, even small variants outperform existing baselines on all benchmarks, showing that scaling and broad biological diversity improve transferability.
This paper proposes MSB-GFM, a multi-semantic basis graph foundation model for cross-domain multi-label node classification, addressing semantic entanglement by representing nodes as adaptive compositions of semantic bases.
This paper proposes OG-SPR, a model-free visual RL algorithm that combines latent self-prediction with observation prediction to learn dynamics-aware representations, achieving improved sample efficiency on DeepMind Control Suite tasks.
BioM-JEPA introduces a joint-embedding predictive architecture that learns single-cell representations by predicting graph-connected gene blocks instead of individual genes, showing improved efficiency and downstream performance in perturbation-response tasks.
NeuroPB is a framework that scales neural decoding by pretraining a motor encoder on large-scale behavioral data (including robotic trajectories) and aligning neural activity to that representation space, improving trajectory decoding and generalization with limited neural data.
This paper introduces EddyFlow, a deep learning framework for kilometer-scale sea surface temperature downscaling that balances predictive accuracy, scale-dependent structure, and regional generalization. It achieves strong zero-shot performance and near-ideal spectral fidelity across multiple ocean regions.
SJEPA introduces a reconstruction-free JEPA framework that learns hybrid symbolic-neural latent dynamics, aiming for the simplest adequate predictive representation. Experiments show it discovers simpler symbolic dynamics with lower rollout error than post-hoc fitting, while controlling symbolic-neural allocation under grammar misspecification.
This paper presents Tactus, an open-vocabulary tactile recognition model that maps low-cost pressure-array data to text embeddings, matching or exceeding a supervised closed-set CNN baseline on the STAG benchmark with only 187 training recordings and no classifier head.
This paper presents an AI-driven multimodal mediation framework using variational autoencoders to discover latent pathways linking socioeconomic disadvantage, psychosocial factors, and cardiometabolic multimorbidity in the All of Us Research Program cohort.
This paper proposes PLAN, a lightweight parallel liquid-inspired approximation network for efficient representation learning in flexible job shop scheduling, achieving better makespan and lower inference latency with fewer parameters than state-of-the-art baselines.
The paper studies the topology of learned representations in predictive coding networks using persistent homology, finding that smaller models simplify topology earlier than larger ones and that earlier simplification correlates with worse reconstruction performance.
This preprint introduces hierarchical self-supervised world models for music co-creation agents, with fast CPU-friendly models and a live demo for MIDI inpainting and generation.
This paper introduces CoCoS, a contrastive pretraining framework that learns whole-cell representations from complementary transcriptomic views, addressing limitations of masked gene reconstruction in single-cell foundation models. Experiments on cell-type annotation and gene regulatory network inference show competitive transfer performance.
This paper proposes Frontier Learning, a framework that combines representations and predictions from multiple black-box and white-box pretrained models to construct a unified target-domain representation, guaranteeing performance no worse than any individual reuse baseline under distribution shift. Evaluations on visual domain adaptation and clinical mortality prediction show consistent gains over strong baselines.
This paper proposes RHEA, a reliability-aware framework for multimodal-attributed graph clustering that estimates node-specific modality reliability from neighborhood consensus, reconstructs unreliable modalities, and uses reliability-aware fusion and optimal transport clustering. Experiments on four benchmarks show consistent gains, especially under noisy or missing attributes.
This paper introduces HP-JEPA, a hierarchical partitioning framework for multi-resolution graph joint-embedding predictive learning, which outperforms the fixed-resolution Graph-JEPA baseline on most graph classification and regression benchmarks.
This arXiv paper introduces ProGFM, a Propagation-aware Graph Foundation Model that treats propagation relationships between edges and feature dimensions as transferable knowledge units, enabling adaptive aggregation and improved cross-domain generalization.
Proposes FATE, a frame-level audio-visual temporal embedding method that aligns frame sequences on a physical timeline, enabling joint semantic and temporal understanding. It outperforms baselines on temporal retrieval, event localization, and generation evaluation metrics.