Tag
This paper introduces MAG, a manifold-guided framework for semi-supervised multi-modal in-context demonstration selection, leveraging unlabeled data to improve few-shot ICL for MLLMs. Experiments on eight benchmarks show consistent gains in label-scarce regimes.
RedNote has released dots3-note preview, the first open-weight model in the dots3 family, featuring 280B parameters with 16B active, multi-modal understanding (text, image, video, audio), 512K context length, and Apache 2.0 license, with strong agent capabilities.
A survey paper systematically analyzing the evolving safety landscape of multi-modal large language models, covering emerging threats such as adversarial attacks, data poisoning, jailbreaks, and hallucinations, and reviewing updated safety strategies.
Introduces GEOID-Flood, a large-scale multi-modal benchmark dataset for flood segmentation with over 14,000 tiles from 219 events across 65 countries, evaluating foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols.
VideoAgent is an all-in-one open-source framework for comprehensive video intelligence, combining understanding, editing, and creative generation through a unified agentic workflow.
mere.run is a local-first inference runtime for Apple Silicon and headless Linux that provides a single CLI for text, image, video, music, 3D, and more without requiring Python.
This paper proposes Expert-Guided Mutual Distillation (EGMD) to address domain bias and semantic misalignment in multimodal fake news detection, achieving state-of-the-art accuracy and reducing domain bias by up to 57.3% across four datasets.
This paper introduces Explorative Modeling, a new generative modeling paradigm that factors the training loop by exploring candidate matches between model generations and data. It establishes a third pretraining axis beyond parameters and data, improves scaling efficiency across images, video, and language, and enables end-to-end generative modeling with far fewer inference steps.
StatePlay proposes a state-aware game world model that jointly predicts visual content and game states to generate mechanics-consistent game rollouts, achieving 18.6% improvement in mechanics fidelity.
Introduces VlogReward, a reward model for evaluating vlog editing plans across six dimensions, along with a large-scale dataset and benchmark, achieving state-of-the-art results against GPT-5 and Gemini-3-Pro.
GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.
The paper proposes a Regime-Aware Multi-Modal Learning (RAML) method for Bitcoin price direction prediction that adaptively fuses social sentiment and technical features based on market volatility. Evaluated on hourly data from July 2024 to September 2025, RAML achieves moderate improvements over static fusion baselines.
Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.
BFL has introduced FLUX 3, a multi-modal AI model capable of generating images, videos, and audio.
Fusion Embedding introduces a family of models that add audio to a frozen vision-language embedding backbone, enabling a unified space for text, image, video, and audio retrieval. The models train only lightweight adapters and achieve audio-image retrieval without paired audio-visual data.
MindLab Research released Macaron-V1-Venti, a new multi-modal AI model available on HuggingFace, likely for text-to-image or image generation tasks.
This paper proposes UMMT, a token-level cross-modal transformer with contrastive multi-task learning for breast cancer subtype classification and survival prediction, achieving state-of-the-art results on METABRIC and TCGA-BRCA datasets.
RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.
Thinking Machines releases Inkling, an open-source multi-modal reasoning model with innovations in post-training RL, achieving stable scaling to 30M+ rollouts and controllable thinking effort. The model exhibits compressed reasoning and will be available soon for fine-tuning.
Hallo4D is a model-agnostic framework that leverages large multimodal language models to detect and correct spatial and temporal hallucinations in 3D and 4D generation, improving consistency across viewpoints and time without requiring retraining.