multi-modal

Tag

Cards List
#multi-modal

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

arXiv cs.CL ↗ · 5d ago Cached

PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.

0 favorites 0 likes
#multi-modal

Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper introduces Ovis-Embedding, a state-of-the-art omni-modal embedding model that uses a shared backbone to encode text, image, video, and audio in a common representation space, achieving top performance on benchmarks like MMEB-v3 and MVEB.

0 favorites 0 likes
#multi-modal

Steer LLMs and Agents at the Token Level: An interactive tool for token visualization & control, model inspection and data annotation.

Reddit r/LocalLLaMA ↗ · 2026-09-19

onPanda is an interactive tool for token-level visualization and control of LLMs and agents, featuring data annotation, model inspection, and support for multiple modalities including browser-based execution.

0 favorites 0 likes
#multi-modal

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

UFO is a unified framework for simultaneous evaluation of omni-condition alignment in multi-modal image generation. It introduces an Atomized Chain-of-Evaluation paradigm and UFO-Bench benchmark.

0 favorites 0 likes
#multi-modal

AlexWortega/openjev

Hugging Face Models Trending ↗ · 2026-09-16 Cached

openjev is a model based on Qwen3.5 trained for entailment tasks, enabling applications in reranking, grading, and real-time game playing, with the v2 version adding multi-modal capabilities and improved zero-shot performance.

0 favorites 0 likes
#multi-modal

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Hugging Face Daily Papers ↗ · 2026-09-16 Cached

This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.

0 favorites 0 likes
#multi-modal

@svpino: I’ve been using ElevenLabs since forever for audio. Now, they added images. Everyone now supports everything. AI is blu…

X AI KOLs Following ↗ · 2026-09-14 Cached

ElevenLabs has added image and video generation to its MCP platform, expanding beyond audio to support multi-modal content creation.

0 favorites 0 likes
#multi-modal

Convergent Emergence of In-Context Learning Across Modalities

Hugging Face Daily Papers ↗ · 2026-09-12 Cached

The paper proposes the Convergent Emergence Hypothesis, stating that few-shot in-context learning emerges with a common cross-modality difficulty profile, and provides empirical support through experiments on six modalities, showing correlated effects in five.

0 favorites 0 likes
#multi-modal

Omni Interaction Agent Technical Report

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

This paper presents Gander, an end-to-end framework for omni interaction and agentic tasks, enabling real-time full-duplex interaction across multiple modalities with a Cerebellum-Brain architecture.

0 favorites 0 likes
#multi-modal

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

arXiv cs.LG ↗ · 2026-09-04 Cached

SimpleDesign introduces a joint model for protein sequence and structure codesign, trained end-to-end in data space using a single-stage objective, achieving competitive performance on co-design and generation benchmarks.

0 favorites 0 likes
#multi-modal

MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

The paper introduces Multi3IR, a benchmark for multi-perspective, multi-domain, multi-modal information retrieval, and proposes SPIN, a method to improve perspective coverage in retrieval systems.

0 favorites 0 likes
#multi-modal

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

Hugging Face Daily Papers ↗ · 2026-08-31 Cached

DICS introduces a self-scoring metric for visual instruction data selection that enhances vision-language model performance with significantly less data by ensuring intra-sample consistency, outperforming state-of-the-art methods.

0 favorites 0 likes
#multi-modal

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

arXiv cs.CL ↗ · 2026-08-24 Cached

Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.

0 favorites 0 likes
#multi-modal

@DevAdventur3s: Ox Alpha = Gemini 4.0 has anyone tried it on @opencode or @OpenRouter ?

X AI KOLs Timeline ↗ · 2026-08-22 Cached

The tweet announces that Ox Alpha, equivalent to Gemini 4.0, is available for free on OpenCode for a week, featuring 1M context, multi-modal capabilities, and zero data retention.

0 favorites 0 likes
#multi-modal

A stealth model called Ox-Alpha has been released, outperforming Fable on SWE.

Reddit r/singularity ↗ · 2026-08-21 Cached

A stealth AI model named Ox-Alpha has been released, reportedly outperforming Fable on SWE benchmarks, and is available for free with features like multi-modal support and zero data retention.

0 favorites 0 likes
#multi-modal

@thegenioo: Why the fuck is this stealth model so good at frontend?

X AI KOLs Timeline ↗ · 2026-08-20 Cached

OpenRouter has released a new stealth AI model named Ox Alpha, optimized for efficient coding and agentic tasks with a 1M token context window and support for text, image, and video inputs.

0 favorites 0 likes
#multi-modal

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

arXiv cs.LG ↗ · 2026-08-18 Cached

This paper diagnoses perceptual-decision misalignment in Omni-LLMs and proposes a training-free inference-time framework called Modality Subspace Activation to mitigate it by dynamically balancing modality strengths.

0 favorites 0 likes
#multi-modal

Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper proposes a Multi-Modal Generative Fuzzy System (MMGFS) to enhance multimodal question answering by addressing modality bias and uncertainty through fuzzy inference and multi-hop reasoning, demonstrating improved performance on multiple benchmarks.

0 favorites 0 likes
#multi-modal

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper introduces MAG, a manifold-guided framework for semi-supervised multi-modal in-context demonstration selection, leveraging unlabeled data to improve few-shot ICL for MLLMs. Experiments on eight benchmarks show consistent gains in label-scarce regimes.

0 favorites 0 likes
#multi-modal

@AdinaYakup: More players are joining the open source summer RedNote just released dots3-note preview The first open weight model in…

X AI KOLs Following ↗ · 2026-08-14 Cached

RedNote has released dots3-note preview, the first open-weight model in the dots3 family, featuring 280B parameters with 16B active, multi-modal understanding (text, image, video, audio), 512K context length, and Apache 2.0 license, with strong agent capabilities.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback