multimodal

Tag

Cards List
#multimodal

@lillian_ma_: We threw the biggest summer multimodal event in the Bay Area, at my favorite museum Always feel a little extra hosting …

X AI KOLs Following ↗ · 2026-09-18 Cached

Lillian Ma hosted a large multimodal event in the Bay Area at a museum, showcasing products like Inference, MCP, and AgentBox with over 200 attendees from partner companies.

0 favorites 0 likes
#multimodal

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

arXiv cs.CL ↗ · 2026-09-18 Cached

The paper proposes a discriminative adaptation of SpeechLLMs for emotion recognition, improving performance and interpretability by using a linear classification head on the hidden state of the final prompt token, which removes hallucinations and enhances analysis of emotion directions.

0 favorites 0 likes
#multimodal

Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery (18 minute read)

TLDR AI ↗ · 2026-09-18

Qwen3.8-Omni-Flash is an omnimodal AI model with a 1M-token context window, supporting text, image, audio, and video inputs, and achieving performance comparable to or better than Gemini 3.8 Flash, now available on the Qianwen AI Platform.

0 favorites 0 likes
#multimodal

Alibaba releases Qwen 3.8 Omni Flash

Hacker News Top ↗ · 2026-09-17

Alibaba has released the Qwen 3.8 Omni Flash AI model, which likely features multimodal capabilities and is optimized for speed.

0 favorites 0 likes
#multimodal

Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint

Hacker News Top ↗ · 2026-09-17 Cached

Introducing Ternary Bonsai 2 27B, a highly compressed AI model that retains 98.2% of performance while being 9x smaller in footprint, enabling efficient local deployment for tasks like reasoning, coding, and multimodal processing.

0 favorites 0 likes
#multimodal

WeVisDoc from tencent

Reddit r/LocalLLaMA ↗ · 2026-09-17

WeVisDoc is an end-to-end document parser fine-tuned from Qwen3-VL models that converts page images to structured Markdown with LaTeX and HTML, achieving top performance on benchmarks like OmniDocBench v1.6.

0 favorites 0 likes
#multimodal

Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper introduces Re2A, a framework for situated conversational recommendation that uses rubric-based preference reasoning and preference-conditioned alignment to enhance user preference satisfaction and situation consistency.

0 favorites 0 likes
#multimodal

Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper proposes a multimodal generative framework integrating fine-tuned Stable Diffusion XL with LLaMA for controllable and culturally faithful Ulos motif generation, validated through ablation studies and qualitative evaluations.

0 favorites 0 likes
#multimodal

NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

arXiv cs.AI ↗ · 2026-09-17 Cached

NeMo Data Designer (NDD) is an open-source, extensible framework for generating multimodal synthetic data, designed for intuitive use and reproducibility in AI model development.

0 favorites 0 likes
#multimodal

Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts

arXiv cs.CL ↗ · 2026-09-17 Cached

This paper introduces the Mult2EMo dataset for studying emotion expression and perception in multimodal social media posts, finding that reconstruction is challenging, particularly when posts rely heavily on images.

0 favorites 0 likes
#multimodal

TabPFN-3.5: Technical Report

arXiv cs.LG ↗ · 2026-09-17 Cached

TabPFN-3.5 is a new tabular foundation model that sets state-of-the-art performance across multiple benchmarks, with improvements in inference speed and multimodal capabilities.

0 favorites 0 likes
#multimodal

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.

0 favorites 0 likes
#multimodal

New stealth model: Union Alpha

Reddit r/singularity ↗ · 2026-09-16

A new stealth AI model named Union Alpha has been released for free on OpenRouter and OpenCode, featuring multimodal capabilities with a 256K context and claiming frontier-level general-purpose performance, sparking speculation about its origin.

0 favorites 0 likes
#multimodal

@seclink: Doubao is seriously badass this time, leaving other domestic large models in the dust by a huge margin.

X AI KOLs Following ↗ · 2026-09-16 Cached

Doubao large model 2.1 Pro released the 0915 version update, focusing on improvements in Agent task delivery, multimodal code writing, multimodal understanding, and inference costs.

0 favorites 0 likes
#multimodal

Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction

arXiv cs.LG ↗ · 2026-09-16 Cached

This paper proposes a multimodal anomaly detection framework for fault detection in mechanical systems using self-supervised cross-modal reconstruction and adaptive thresholding to improve robustness under distribution shifts.

0 favorites 0 likes
#multimodal

Negation Beyond the Verbal Channel: Temporal Multimodal Correlates in Dialogue

arXiv cs.CL ↗ · 2026-09-16 Cached

This paper investigates how spoken negation cues in human dialogue are reflected in multimodal nonverbal behavior, using time-series classification models to distinguish negation contexts from control contexts without lexical or acoustic input.

0 favorites 0 likes
#multimodal

GraMRAG: Orchestrating Multi-Agent Multi-Step Reasoning via Graph Memory with Reinforcement Learning

arXiv cs.CL ↗ · 2026-09-15 Cached

GraMRAG is a new multi-agent RAG framework that uses graph memory and reinforcement learning to improve reasoning depth and memory structure in complex multimodal tasks, achieving state-of-the-art performance.

0 favorites 0 likes
#multimodal

With iOS 27, I’m actually using Siri again

TechCrunch AI ↗ · 2026-09-14 Cached

Apple's iOS 27 update significantly improves Siri with Google's Gemini models, enabling complex requests and on-screen context understanding.

0 favorites 0 likes
#multimodal

@svpino: Building an agent that works with real-life documents is still crazy hard. All of you are talking about AGI, and this i…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The post highlights the difficulty of processing complex real-life documents and promotes LlamaParse as a tool that uses multiple AI agents for efficient, scalable, and multimodal document parsing.

0 favorites 0 likes
#multimodal

Intern-S2-397B (multimodal, reasoning, coding, and scientific agent capabilities)

Reddit r/LocalLLaMA ↗ · 2026-09-14

Intern-S2-397B is a large multimodal AI model with capabilities in reasoning, coding, and scientific agent tasks, now available on Hugging Face with day-0 support from vLLM.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback