multi-modal

Tag

Cards List
#multi-modal

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

arXiv cs.LG · yesterday Cached

This paper introduces MAG, a manifold-guided framework for semi-supervised multi-modal in-context demonstration selection, leveraging unlabeled data to improve few-shot ICL for MLLMs. Experiments on eight benchmarks show consistent gains in label-scarce regimes.

0 favorites 0 likes
#multi-modal

@AdinaYakup: More players are joining the open source summer RedNote just released dots3-note preview The first open weight model in…

X AI KOLs Following · yesterday Cached

RedNote has released dots3-note preview, the first open-weight model in the dots3 family, featuring 280B parameters with 16B active, multi-modal understanding (text, image, video, audio), 512K context length, and Apache 2.0 license, with strong agent capabilities.

0 favorites 0 likes
#multi-modal

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards

arXiv cs.LG · 4d ago Cached

A survey paper systematically analyzing the evolving safety landscape of multi-modal large language models, covering emerging threats such as adversarial attacks, data poisoning, jailbreaks, and hallucinations, and reviewing updated safety strategies.

0 favorites 0 likes
#multi-modal

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation

Hugging Face Daily Papers · 2026-08-03 Cached

Introduces GEOID-Flood, a large-scale multi-modal benchmark dataset for flood segmentation with over 14,000 tiles from 219 events across 65 countries, evaluating foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols.

0 favorites 0 likes
#multi-modal

@tom_doerr: VideoAgent is an all-in-one open-source framework for comprehensive video intelligence, combining understanding, editin…

X AI KOLs Timeline · 2026-08-02 Cached

VideoAgent is an all-in-one open-source framework for comprehensive video intelligence, combining understanding, editing, and creative generation through a unified agentic workflow.

0 favorites 0 likes
#multi-modal

Show HN: Local text, image, video, music and 3D from one CLI, no Python

Hacker News Top · 2026-07-30 Cached

mere.run is a local-first inference runtime for Apple Silicon and headless Linux that provides a single CLI for text, image, video, music, 3D, and more without requiring Python.

0 favorites 0 likes
#multi-modal

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

arXiv cs.CL · 2026-07-30 Cached

This paper proposes Expert-Guided Mutual Distillation (EGMD) to address domain bias and semantic misalignment in multimodal fake news detection, achieving state-of-the-art accuracy and reducing domain bias by up to 57.3% across four datasets.

0 favorites 0 likes
#multi-modal

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

Hugging Face Daily Papers · 2026-07-29 Cached

This paper introduces Explorative Modeling, a new generative modeling paradigm that factors the training loop by exploring candidate matches between model generations and data. It establishes a third pretraining axis beyond parameters and data, improves scaling efficiency across images, video, and language, and enables end-to-end generative modeling with far fewer inference steps.

0 favorites 0 likes
#multi-modal

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

Hugging Face Daily Papers · 2026-07-29 Cached

StatePlay proposes a state-aware game world model that jointly predicts visual content and game states to generate mechanics-consistent game rollouts, achieving 18.6% improvement in mechanics fidelity.

0 favorites 0 likes
#multi-modal

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

arXiv cs.AI · 2026-07-28 Cached

Introduces VlogReward, a reward model for evaluating vlog editing plans across six dimensions, along with a large-scale dataset and benchmark, achieving state-of-the-art results against GPT-5 and Gemini-3-Pro.

0 favorites 0 likes
#multi-modal

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

Hugging Face Daily Papers · 2026-07-27 Cached

GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.

0 favorites 0 likes
#multi-modal

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features

Hugging Face Daily Papers · 2026-07-25 Cached

The paper proposes a Regime-Aware Multi-Modal Learning (RAML) method for Bitcoin price direction prediction that adaptively fuses social sentiment and technical features based on market volatility. Evaluated on hourly data from July 2024 to September 2025, RAML achieves moderate improvements over static fusion baselines.

0 favorites 0 likes
#multi-modal

Black Forest Lab's Flux 3: Omni-modality for image, video, audio & action prediction

Reddit r/singularity · 2026-07-23

Black Forest Lab's Flux 3 is a new omni-modal AI model capable of generating and predicting images, video, audio, and actions.

0 favorites 0 likes
#multi-modal

BFL Introduces FLUX 3 - multi-modal model for Image, Video and Audio

Reddit r/ArtificialInteligence · 2026-07-23

BFL has introduced FLUX 3, a multi-modal AI model capable of generating images, videos, and audio.

0 favorites 0 likes
#multi-modal

Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio

arXiv cs.CL · 2026-07-22 Cached

Fusion Embedding introduces a family of models that add audio to a frozen vision-language embedding backbone, enabling a unified space for text, image, video, and audio retrieval. The models train only lightweight adapters and achieve audio-image retrieval without paired audio-visual data.

0 favorites 0 likes
#multi-modal

mindlab-research/Macaron-V1-Venti • HuggingFace

Reddit r/LocalLLaMA · 2026-07-21

MindLab Research released Macaron-V1-Venti, a new multi-modal AI model available on HuggingFace, likely for text-to-image or image generation tasks.

0 favorites 0 likes
#multi-modal

Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction

arXiv cs.LG · 2026-07-21 Cached

This paper proposes UMMT, a token-level cross-modal transformer with contrastive multi-task learning for breast cancer subtype classification and survival prediction, achieving state-of-the-art results on METABRIC and TCGA-BRCA datasets.

0 favorites 0 likes
#multi-modal

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Hugging Face Daily Papers · 2026-07-20 Cached

RynnBrain 1.1 is a family of embodied foundation models (2B, 9B, 122B-A10B) that improve perception, spatial reasoning, and manipulation, achieving state-of-the-art results on VSI-Bench, MMSI, and RefSpatial-Bench, and outperforming baselines in real-robot experiments.

0 favorites 0 likes
#multi-modal

@LiTianleli: Incredibly proud of the team. After countless late nights, Inkling is out, and I especially want to highlight the post-…

X AI KOLs Timeline · 2026-07-15 Cached

Thinking Machines releases Inkling, an open-source multi-modal reasoning model with innovations in post-training RL, achieving stable scaling to 30M+ rollouts and controllable thinking effort. The model exhibits compressed reasoning and will be available soon for fine-tuning.

0 favorites 0 likes
#multi-modal

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Hugging Face Daily Papers · 2026-07-15 Cached

Hallo4D is a model-agnostic framework that leverages large multimodal language models to detect and correct spatial and temporal hallucinations in 3D and 4D generation, improving consistency across viewpoints and time without requiring retraining.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback