temporal-grounding

Tag

Cards List
#temporal-grounding

GigaChat Audio: Time-aware Large Audio Language Model

Hugging Face Daily Papers · 2026-07-11 Cached

This paper introduces GigaChat Audio, a time-aware large audio language model that answers questions with explicit timestamps for up to 120 minutes of audio, using interleaved periodic time markers and synthetic supervision. The model achieves strong temporal grounding accuracy on benchmarks and the authors release model weights and datasets.

0 favorites 0 likes
#temporal-grounding

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

Hugging Face Daily Papers · 2026-07-01 Cached

VideoSearch-R1 introduces an agentic framework that iteratively retrieves videos and refines search queries using continuous latent space refinement and policy optimization, achieving state-of-the-art performance on video corpus moment retrieval and temporal grounding tasks.

0 favorites 0 likes
#temporal-grounding

Towards One-to-Many Temporal Grounding

Hugging Face Daily Papers · 2026-06-04 Cached

This paper introduces One-to-Many Temporal Grounding (OMTG), a new task for localizing multiple disjoint video segments from a single text query, along with a benchmark, evaluation metrics, a 56k-sample dataset, and novel reward functions that achieve state-of-the-art results, outperforming Gemini 2.5 Pro and Seed-1.8.

0 favorites 0 likes
#temporal-grounding

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

arXiv cs.CL · 2026-05-29 Cached

MusTBench is a benchmark for evaluating temporal grounding in Large Audio-Language Models (LALMs) for music understanding. The authors propose MusT, a four-stage training recipe that significantly improves temporal grounding performance over existing models.

0 favorites 0 likes
#temporal-grounding

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

Hugging Face Daily Papers · 2026-05-26 Cached

OmniInteract introduces a streaming benchmark for real-time omnimodal LLMs, evaluating online audio-visual processing with temporal grounding and interactive response requirements. Experiments show that current models perform poorly, with the best overall IA-QTF1 score reaching only 0.368.

0 favorites 0 likes
#temporal-grounding

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Hugging Face Daily Papers · 2026-05-25 Cached

LLaVA-OneVision-2 introduces codec-stream tokenization and windowed attention for efficient video understanding, achieving state-of-the-art performance across multiple multimodal benchmarks including video, spatial, and tracking tasks.

0 favorites 0 likes
#temporal-grounding

NemoStation/Marlin-2B

Hugging Face Models Trending · 2026-05-13

NemoStation/Marlin-2B is a fine-tuned model based on Qwen3.5-2B for video-text-to-text tasks, supporting video captioning and temporal grounding.

0 favorites 0 likes
← Back to home

Submit Feedback