multi-modal

Tag

Cards List
#multi-modal

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv cs.AI · 2026-07-13 Cached

TheBioCollection is a 52.6B-token pre-training corpus for biology that consolidates heterogeneous resources into a unified format for training large language models, and includes a matched evaluation suite. Training on this corpus more than doubles overall scores on biology benchmarks while preserving general language ability.

0 favorites 0 likes
#multi-modal

@AdinaYakup: SenseNova-Vision SenseTime's new model treats all of computer vision as generation - 7B - CC BY-NC 4.0 ( non commercial…

X AI KOLs Following · 2026-07-08 Cached

SenseTime released SenseNova-Vision, a 7B parameter model that unifies computer vision tasks as generation, with open weights, instruction corpus, benchmark, paper, and demo.

0 favorites 0 likes
#multi-modal

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

arXiv cs.AI · 2026-07-07 Cached

BRAID is a framework that formulates interleaved text-image-text reasoning as a unified Markov decision process, enabling joint optimization of textual and visual generation via reinforcement learning with a VLM judge providing dense turn-level feedback.

0 favorites 0 likes
#multi-modal

Multi-modal Rail Crossing Safety Analysis

arXiv cs.LG · 2026-07-03 Cached

This paper proposes a proof-of-concept AI pipeline that uses multi-modal data (images and accident reports) to assess railway crossing safety, achieving a macro F1 score of 0.757 for risk classification and an RMSE of 0.078 for safety score estimation using a fine-tuned compact VLM.

0 favorites 0 likes
#multi-modal

nvidia/Cosmos3-Edge

Hugging Face Models Trending · 2026-07-01 Cached

NVIDIA releases Cosmos3-Edge, an omnimodal world foundation model that generates video, image, audio, and action commands from text, image, video, and action trajectory inputs, targeting Physical AI applications in robotics, autonomous driving, and smart spaces.

0 favorites 0 likes
#multi-modal

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

arXiv cs.AI · 2026-07-01 Cached

This paper introduces HealthAgentBench, a suite of 54 realistic healthcare tasks for evaluating frontier AI agents. It finds that even the best agent (Codex GPT-5.5) achieves only ~42% success, highlighting substantial room for improvement.

0 favorites 0 likes
#multi-modal

Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

arXiv cs.AI · 2026-06-29 Cached

This paper introduces MMG-Pop, a unified benchmark for multi-modal graph-based social media popularity prediction, and proposes MMG-PopNet, a model that jointly models multimodal content and temporal social interactions. Experiments on Bluesky and Reddit datasets demonstrate superior performance and provide insights into cross-platform generalization and multi-task prediction.

0 favorites 0 likes
#multi-modal

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

arXiv cs.LG · 2026-06-29 Cached

This paper systematically compares seven state-of-the-art sensor fusion methods for multi-modal human activity recognition on the HARMES dataset, showing that Gated Multi-modal Fusion achieves the highest macro F1-score of 0.82, outperforming the baseline by 6 percentage points.

0 favorites 0 likes
#multi-modal

@exploraX_: google's notebookLM charges $19.99/mo for pro. this open-source alternative gives you unlimited access for $0. zero sub…

X AI KOLs Timeline · 2026-06-27 Cached

An open-source alternative to Google's NotebookLM offers free unlimited access, runs locally under MIT license, and includes features like podcast generation and multi-modal ingestion.

0 favorites 0 likes
#multi-modal

@oliviscusAI: Microsoft open-sourced a system that lets one AI control hundreds of other AI models. It's called JARVIS. • Handles tex…

X AI KOLs Timeline · 2026-06-26 Cached

Microsoft open-sourced JARVIS, a system that uses a GPT controller to orchestrate hundreds of AI models from HuggingFace for multi-modal tasks.

0 favorites 0 likes
#multi-modal

OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility

arXiv cs.AI · 2026-06-24 Cached

OmniPath is a multi-modal agentic framework that combines OpenStreetMap network topology with aerial LiDAR data to audit wheelchair accessibility by analyzing physical barriers like slope and surface discontinuities at high resolution, validated against field surveys.

0 favorites 0 likes
#multi-modal

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Hugging Face Daily Papers · 2026-06-23 Cached

IV-CoT decomposes visual conditioning into structural and semantic cascades for improved structure-aware image generation, using training-only sketch supervision to guide structural queries. It achieves state-of-the-art results on GenEval and T2I-CompBench.

0 favorites 0 likes
#multi-modal

ChartWalker: Benchmarking the Cross-Chart RAG Task

Hugging Face Daily Papers · 2026-06-22 Cached

ChartWalker introduces a novel framework for cross-chart retrieval-augmented generation (RAG) using hierarchical knowledge graph construction and structure-aware sampling. It releases a challenging benchmark (ChartWalker-Bench) and an agentic baseline (ChartWalker-Agent), revealing significant performance gaps in current RAG paradigms.

0 favorites 0 likes
#multi-modal

@mervenoyann: day 2 findings on this pipeline > it works, got map@50=0.8028 on road sign detection against human annotations, with on…

X AI KOLs Timeline · 2026-06-17 Cached

Merve (@mervenoyann) shares day two findings of a pipeline using multiple small VLMs as judges for road sign detection, achieving map@50=0.8028 with only 1.3k examples. The thread compares model rejection rates and discusses dataset shrinking, super-specific prompts, and plans to generalize the library.

0 favorites 0 likes
#multi-modal

QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks

arXiv cs.AI · 2026-06-16 Cached

This research paper proposes a framework for fair token allocation and private data valuation in decentralized multi-modal agentic systems, using differentially private prototypes to balance privacy and utility while scheduling limited edge AI resources.

0 favorites 0 likes
#multi-modal

@mishig25: Open source is so back http://hf.co/mistralai/Mistral-Medium-3.5-128B…

X AI KOLs Following · 2026-06-15 Cached

Mistral AI releases Mistral Medium 3.5, an open-source 128B dense model with 256k context, multimodal input, configurable reasoning, and agentic capabilities.

0 favorites 0 likes
#multi-modal

LLM Gateway Chat

Product Hunt · 2026-06-15

LLM Gateway Chat is a platform that provides access to multiple AI models for chat, image, video, and audio generation.

0 favorites 0 likes
#multi-modal

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

Hugging Face Daily Papers · 2026-06-15 Cached

PermaVid introduces a multi-modal context memory that disentangles appearance and geometric structure to maintain long-term video consistency after editing operations, outperforming prior methods.

0 favorites 0 likes
#multi-modal

@PyTorch: In this clip from his PyTorch Conference Europe 2026 keynote, Patrick von Platen (@MistralAI) discusses why real-world …

X AI KOLs Following · 2026-06-12 Cached

At PyTorch Conference Europe 2026, Mistral AI's Patrick von Platen explains why real-world AI interaction requires streaming architectures that process continuous input and produce continuous output, using Vox Real Time as a live transcription example.

0 favorites 0 likes
#multi-modal

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models

arXiv cs.AI · 2026-06-12 Cached

This paper introduces a Multi-Modal Agent framework for power distribution defect detection, evaluating foundation models on perception, reasoning, and tool usage capabilities, with a new domain-specific dataset and benchmark.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback