multimodal

Tag

Cards List
#multimodal

ArGuard Shared Task: Harmful Content Detection in Arabic Memes and LLM Prompts

arXiv cs.CL ↗ · yesterday Cached

ArGuard is a shared task focused on detecting harmful content in Arabic memes and LLM prompts, highlighting challenges in fine-grained classification and releasing datasets for further research.

0 favorites 0 likes
#multimodal

Gemini 3.8 Live with Live Avatar (1 minute read)

TLDR AI ↗ · yesterday Cached

Gemini 3.8 Live with Live Avatar brings real-time visual presence to conversational AI, allowing enterprises to create more natural and interactive digital experiences with dynamic avatars and multilingual support.

0 favorites 0 likes
#multimodal

@_philschmid: Gemini 3.8 Flash. Just use Gemini for multimodal understanding.

X AI KOLs Following ↗ · yesterday Cached

A tweet suggests using Gemini 3.8 Flash for multimodal understanding and references a visual test comparing AI models' ability to name people from a drawing.

0 favorites 0 likes
#multimodal

Flux 3 Action: open weights 7B World Action Model

Reddit r/singularity ↗ · 2d ago

FLUX 3 Action is an open-weight 7B AI model that sets a new standard on the RoboLab benchmark, outperforming previous open models while being faster and more efficient, with applications in robotics and other environments requiring visual action planning.

0 favorites 0 likes
#multimodal

Space Bunny is the new stealth model in OpenCode. Free to try, multimodal

Reddit r/LocalLLaMA ↗ · 2d ago

A new stealth multimodal AI model named Space Bunny has been released in OpenCode, available for free trial, with the releasing company unknown.

0 favorites 0 likes
#multimodal

@heyshrutimishra: Xiaomi’s MiMo V2.6 Pro is now the #1 global open-source model on Artificial Analysis. Also #1 among Chinese-developed m…

X AI KOLs Timeline ↗ · 2d ago Cached

Xiaomi's MiMo V2.6 Pro is now the top-ranked global open-source AI model on Artificial Analysis, nearly matching GPT-5.6 Sol in intelligence while offering cost-effective pricing for real-world applications.

0 favorites 0 likes
#multimodal

@yibie: https://x.com/yibie/status/2102565806466912408

X AI KOLs Timeline ↗ · 3d ago Cached

Jev-Omni is the first open-weight model to extend typed decisions to multimodal, supporting text, images, audio, and video, and directly returning the probability distribution of options without generating explanations.

0 favorites 0 likes
#multimodal

@omarsar0: Impressive paper showing how much the harness changes a coding agent's results. Harnesses do play a huge role in what y…

X AI KOLs Following ↗ · 3d ago Cached

The paper introduces ReFigBench, a benchmark for evaluating coding agents in reconstructing scientific figures into editable PowerPoint slides, showing that harnesses significantly affect performance across model families.

0 favorites 0 likes
#multimodal

@yifanzhang_: Very interesting model architecture that uses block-sparse attention. https://pixverse.ai/en/blog/pixverse-r2-scaling-r…

X AI KOLs Timeline ↗ · 3d ago Cached

PixVerse R2 introduces a unified scaling architecture for real-time audiovisual world models, leveraging block-sparse attention and continuous pretraining to enhance video generation and interactive control.

0 favorites 0 likes
#multimodal

@AdinaYakup: RedNote don’t release that often, but each one is solid https://huggingface.co/dots-studio/dots3-note-prev…

X AI KOLs Timeline ↗ · 4d ago Cached

dots3-note Preview is a multimodal Mixture-of-Experts AI model with 280B total parameters and 16B activated, supporting up to 512K token context, released as an open-weight model on Hugging Face.

0 favorites 0 likes
#multimodal

Xiaomi open-sources MiMo-V2.6 Pro and Flash models (2 minute read)

TLDR AI ↗ · 4d ago Cached

Xiaomi has open-sourced the MiMo-V2.6 series, introducing two omnimodal AI models, Pro and Flash, that perform competitively with top proprietary models and are available for various applications.

0 favorites 0 likes
#multimodal

XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

Reddit r/LocalLLaMA ↗ · 4d ago Cached

MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.

0 favorites 0 likes
#multimodal

Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving

arXiv cs.AI ↗ · 5d ago Cached

This paper proposes RRDrive, a risk-aware occupancy representation for register-based end-to-end autonomous driving, which improves trajectory prediction and selection in complex scenes, achieving notable performance gains.

0 favorites 0 likes
#multimodal

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

arXiv cs.CL ↗ · 5d ago Cached

This paper introduces Omni Demand Understanding (ODU), a benchmark to evaluate how well multimodal AI models infer user demands from complex audio-visual interactions, revealing significant performance gaps in current models.

0 favorites 0 likes
#multimodal

VISPATH: Visual-Intent-Guided Path Reasoning for Multimodal Knowledge Graph Question Answering

arXiv cs.CL ↗ · 5d ago Cached

VisPath introduces a visual-intent-guided path reasoning framework for multimodal knowledge graph question answering, achieving significant improvements over baselines on a new benchmark, VisPath-Bench, and existing datasets.

0 favorites 0 likes
#multimodal

VideoGen-Agent: Reinforcing Video Generation Agents

Hugging Face Daily Papers ↗ · 5d ago Cached

The paper presents VideoGen-Agent, a reinforcement learning-based multimodal agent that coordinates tools for video generation, significantly improving performance on the new VABench benchmark.

0 favorites 0 likes
#multimodal

akhilaaa3/Jev-Omni

Hugging Face Models Trending ↗ · 5d ago Cached

Jev-Omni is a multimodal decision classifier built on Gemma 4 12B IT, fine-tuned to handle text, images, audio, and video, providing probability scores for options based on questions.

0 favorites 0 likes
#multimodal

@HuggingModels: OCR just got a major upgrade. jina-ocr-v1 is a multimodal vision language model built for document intelligence. It rea…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

Jina-ocr-v1 is a multimodal vision language model designed for advanced document intelligence, reading text from images in multiple languages.

0 favorites 0 likes
#multimodal

@Saccc_c: To be honest, actually, most people don't really need Astra for the majority of their work. I've been using ds v4.1 fla…

X AI KOLs Following ↗ · 2026-09-18 Cached

A user shares their positive experience switching to ds v4.1 flash from Astra, highlighting its multimodal capabilities and fast speed.

0 favorites 0 likes
#multimodal

@lillian_ma_: We threw the biggest summer multimodal event in the Bay Area, at my favorite museum Always feel a little extra hosting …

X AI KOLs Following ↗ · 2026-09-18 Cached

Lillian Ma hosted a large multimodal event in the Bay Area at a museum, showcasing products like Inference, MCP, and AgentBox with over 200 attendees from partner companies.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback