Tag
Lillian Ma hosted a large multimodal event in the Bay Area at a museum, showcasing products like Inference, MCP, and AgentBox with over 200 attendees from partner companies.
The paper proposes a discriminative adaptation of SpeechLLMs for emotion recognition, improving performance and interpretability by using a linear classification head on the hidden state of the final prompt token, which removes hallucinations and enhances analysis of emotion directions.
Qwen3.8-Omni-Flash is an omnimodal AI model with a 1M-token context window, supporting text, image, audio, and video inputs, and achieving performance comparable to or better than Gemini 3.8 Flash, now available on the Qianwen AI Platform.
Alibaba has released the Qwen 3.8 Omni Flash AI model, which likely features multimodal capabilities and is optimized for speed.
Introducing Ternary Bonsai 2 27B, a highly compressed AI model that retains 98.2% of performance while being 9x smaller in footprint, enabling efficient local deployment for tasks like reasoning, coding, and multimodal processing.
WeVisDoc is an end-to-end document parser fine-tuned from Qwen3-VL models that converts page images to structured Markdown with LaTeX and HTML, achieving top performance on benchmarks like OmniDocBench v1.6.
The paper introduces Re2A, a framework for situated conversational recommendation that uses rubric-based preference reasoning and preference-conditioned alignment to enhance user preference satisfaction and situation consistency.
This paper proposes a multimodal generative framework integrating fine-tuned Stable Diffusion XL with LLaMA for controllable and culturally faithful Ulos motif generation, validated through ablation studies and qualitative evaluations.
NeMo Data Designer (NDD) is an open-source, extensible framework for generating multimodal synthetic data, designed for intuitive use and reproducibility in AI model development.
This paper introduces the Mult2EMo dataset for studying emotion expression and perception in multimodal social media posts, finding that reconstruction is challenging, particularly when posts rely heavily on images.
TabPFN-3.5 is a new tabular foundation model that sets state-of-the-art performance across multiple benchmarks, with improvements in inference speed and multimodal capabilities.
Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.
A new stealth AI model named Union Alpha has been released for free on OpenRouter and OpenCode, featuring multimodal capabilities with a 256K context and claiming frontier-level general-purpose performance, sparking speculation about its origin.
Doubao large model 2.1 Pro released the 0915 version update, focusing on improvements in Agent task delivery, multimodal code writing, multimodal understanding, and inference costs.
This paper proposes a multimodal anomaly detection framework for fault detection in mechanical systems using self-supervised cross-modal reconstruction and adaptive thresholding to improve robustness under distribution shifts.
This paper investigates how spoken negation cues in human dialogue are reflected in multimodal nonverbal behavior, using time-series classification models to distinguish negation contexts from control contexts without lexical or acoustic input.
GraMRAG is a new multi-agent RAG framework that uses graph memory and reinforcement learning to improve reasoning depth and memory structure in complex multimodal tasks, achieving state-of-the-art performance.
Apple's iOS 27 update significantly improves Siri with Google's Gemini models, enabling complex requests and on-screen context understanding.
The post highlights the difficulty of processing complex real-life documents and promotes LlamaParse as a tool that uses multiple AI agents for efficient, scalable, and multimodal document parsing.
Intern-S2-397B is a large multimodal AI model with capabilities in reasoning, coding, and scientific agent tasks, now available on Hugging Face with day-0 support from vLLM.