Tag
PhysioBench introduces a unified benchmark for physiological signal question answering, harmonizing 22 datasets into 61.4 million questions across 30 tasks to evaluate the performance of various AI models.
The paper introduces Ovis-Embedding, a state-of-the-art omni-modal embedding model that uses a shared backbone to encode text, image, video, and audio in a common representation space, achieving top performance on benchmarks like MMEB-v3 and MVEB.
onPanda is an interactive tool for token-level visualization and control of LLMs and agents, featuring data annotation, model inspection, and support for multiple modalities including browser-based execution.
UFO is a unified framework for simultaneous evaluation of omni-condition alignment in multi-modal image generation. It introduces an Atomized Chain-of-Evaluation paradigm and UFO-Bench benchmark.
openjev is a model based on Qwen3.5 trained for entailment tasks, enabling applications in reranking, grading, and real-time game playing, with the v2 version adding multi-modal capabilities and improved zero-shot performance.
This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.
ElevenLabs has added image and video generation to its MCP platform, expanding beyond audio to support multi-modal content creation.
The paper proposes the Convergent Emergence Hypothesis, stating that few-shot in-context learning emerges with a common cross-modality difficulty profile, and provides empirical support through experiments on six modalities, showing correlated effects in five.
This paper presents Gander, an end-to-end framework for omni interaction and agentic tasks, enabling real-time full-duplex interaction across multiple modalities with a Cerebellum-Brain architecture.
SimpleDesign introduces a joint model for protein sequence and structure codesign, trained end-to-end in data space using a single-stage objective, achieving competitive performance on co-design and generation benchmarks.
The paper introduces Multi3IR, a benchmark for multi-perspective, multi-domain, multi-modal information retrieval, and proposes SPIN, a method to improve perspective coverage in retrieval systems.
DICS introduces a self-scoring metric for visual instruction data selection that enhances vision-language model performance with significantly less data by ensuring intra-sample consistency, outperforming state-of-the-art methods.
Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.
The tweet announces that Ox Alpha, equivalent to Gemini 4.0, is available for free on OpenCode for a week, featuring 1M context, multi-modal capabilities, and zero data retention.
A stealth AI model named Ox-Alpha has been released, reportedly outperforming Fable on SWE benchmarks, and is available for free with features like multi-modal support and zero data retention.
OpenRouter has released a new stealth AI model named Ox Alpha, optimized for efficient coding and agentic tasks with a 1M token context window and support for text, image, and video inputs.
This paper diagnoses perceptual-decision misalignment in Omni-LLMs and proposes a training-free inference-time framework called Modality Subspace Activation to mitigate it by dynamically balancing modality strengths.
This paper proposes a Multi-Modal Generative Fuzzy System (MMGFS) to enhance multimodal question answering by addressing modality bias and uncertainty through fuzzy inference and multi-hop reasoning, demonstrating improved performance on multiple benchmarks.
This paper introduces MAG, a manifold-guided framework for semi-supervised multi-modal in-context demonstration selection, leveraging unlabeled data to improve few-shot ICL for MLLMs. Experiments on eight benchmarks show consistent gains in label-scarce regimes.
RedNote has released dots3-note preview, the first open-weight model in the dots3 family, featuring 280B parameters with 16B active, multi-modal understanding (text, image, video, audio), 512K context length, and Apache 2.0 license, with strong agent capabilities.