image-understanding

Tag

Cards List
#image-understanding

Decoupled Vision-Language System for Multimodal Understanding and Generation

arXiv cs.CL · yesterday Cached

The paper introduces Libra, a decoupled vision-language architecture for multimodal large language models that enables both image-to-text understanding and text-to-image generation, demonstrating strong performance on benchmarks.

0 favorites 0 likes
#image-understanding

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

Hugging Face Daily Papers · 2026-08-18 Cached

This paper presents MoE-ViE, a mixture-of-experts vision encoder that efficiently scales for image and video understanding, outperforming larger dense models with lower inference latency.

0 favorites 0 likes
#image-understanding

@AdinaYakup: ClinFusion new open medical multimodal LLM from Alibaba DAMO Academy - 8B/32B (Apache2.0) - Unified 2D + 3D image under…

X AI KOLs Following · 2026-08-05 Cached

ClinFusion is a new open medical multimodal LLM from Alibaba DAMO Academy, available in 8B and 32B sizes (Apache 2.0), with unified 2D + 3D image understanding and state-of-the-art results on medical benchmarks.

0 favorites 0 likes
#image-understanding

@skalskip92: Qwen3.8-Max is the best object detection VLM - satellite images - infrared images - documents - techical drawings - han…

X AI KOLs Timeline · 2026-08-03 Cached

A tweet claims Qwen3.8-Max is the best object detection VLM, excelling across satellite, infrared, document, and hand-drawn images, with examples shared.

0 favorites 0 likes
#image-understanding

thinkingmachines/Inkling

Hugging Face Models Trending · 2026-07-14 Cached

Inkling is a large open-weights multimodal model (975B total, 41B active parameters) using a sparse MoE architecture, accepting text, image, and audio inputs and generating text outputs, intended for agentic systems, coding assistants, and chatbots.

0 favorites 0 likes
#image-understanding

@Jolyne_AI: Let me recommend a great screenshot recognition AI tool: Snippai. It's fully open source and free, with stable, accurate, and fast recognition. It can not only extract text and formulas from images, but also understand image content, convert tables into usable formats, and even provide ideas and answers for problems in screenshots. GitHub: http://githu…

X AI KOLs Timeline · 2026-07-03 Cached

Snippai is a fully open source and free screenshot recognition AI tool that supports formula conversion to LaTeX, text extraction, table conversion to Markdown, image understanding, screenshot problem solving, code explanation, color analysis, and screenshot translation.

0 favorites 0 likes
#image-understanding

Image interaction with GPT-Live

YouTube AI Channels · 2026-07-09 Cached

The user shows two outfits in real-time via GPT-Live, and the AI gives specific clothing suggestions for the scenario of meeting parents, demonstrating multi-round image understanding and contextual recommendation capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback