Tag
The paper introduces Libra, a decoupled vision-language architecture for multimodal large language models that enables both image-to-text understanding and text-to-image generation, demonstrating strong performance on benchmarks.
This paper presents MoE-ViE, a mixture-of-experts vision encoder that efficiently scales for image and video understanding, outperforming larger dense models with lower inference latency.
ClinFusion is a new open medical multimodal LLM from Alibaba DAMO Academy, available in 8B and 32B sizes (Apache 2.0), with unified 2D + 3D image understanding and state-of-the-art results on medical benchmarks.
A tweet claims Qwen3.8-Max is the best object detection VLM, excelling across satellite, infrared, document, and hand-drawn images, with examples shared.
Inkling is a large open-weights multimodal model (975B total, 41B active parameters) using a sparse MoE architecture, accepting text, image, and audio inputs and generating text outputs, intended for agentic systems, coding assistants, and chatbots.
Snippai is a fully open source and free screenshot recognition AI tool that supports formula conversion to LaTeX, text extraction, table conversion to Markdown, image understanding, screenshot problem solving, code explanation, color analysis, and screenshot translation.
The user shows two outfits in real-time via GPT-Live, and the AI gives specific clothing suggestions for the scenario of meeting parents, demonstrating multi-round image understanding and contextual recommendation capabilities.