Tag
This paper proposes a YOLO- and CLIP-based vision-language framework to classify mosquito flight frames for Dengue virus detection, achieving 98.54% accuracy and 99.91% sensitivity at frame level, with complete video-level performance after temporal aggregation.
This paper introduces V2N, the first complete visual piano transcription system that jointly predicts onset, offset, keyhold, and velocity from video, achieving state-of-the-art results on PianoVAM and R3 benchmarks.
NeuroVidz is a tool that shows how a brain reacts to your video clip, likely using AI or neuroscience for audience insights.
An analysis of 10 AI startup launch videos over six months reveals a consistent pattern distinguishing those that achieved over 1 million views from those that failed.
This paper evaluates the effect of frame sampling rate on sequence-based classification of autism-related self-stimulatory hand behaviors using LSTM and GRU models, achieving up to 98.75% accuracy at a 15-frame interval, and analyzes data augmentation strategies for small behavioral video datasets.
Lapian-Notes is an open-source AI-assisted film analysis tool that automatically extracts frames, generates scene timelines and emotion curves, runs entirely locally, and requires no data uploads.
The article announces the open-source release of small_vlm_video_analysis, a tool that uses small vision-language models locally to verify if procedural videos comply with predefined SOPs, including frame description, rule-based judging, and a visualization viewer.
Claude Code gains the ability to watch videos by using FFmpeg scene detection to capture frames at meaningful moments and integrating caption extraction from YouTube or Whisper for Zoom/Loom, enabling users to analyze recordings and save insights to a knowledge base.
Introduces the claude-video tool on GitHub, which enables AI agents like Claude/Codex to truly 'understand' video content, applicable for deconstructing viral videos, analyzing ad creatives, diagnosing bugs, and more.
Introducing the open-source tool claude-real-video, which uses scene change detection and audio transcription to allow any LLM to analyze videos locally, solving the pain points of fixed-frame sampling and cloud uploads.
Veridive is a tool that lets you find key moments in videos via chat, enabling quick discovery of important segments.
Interhuman.ai has launched a Streaming API for its Inter-1 model, enabling real-time detection of 12 social signals from live video streams via WebSocket, along with engagement tracking and conversation quality scoring.
Artifact-Bench is a comprehensive benchmark that evaluates multimodal large language models on detecting and analyzing artifacts in AI-generated videos, revealing significant limitations and misalignment with human perception.
Introduces Knowly AI tool, capable of interpreting YouTube videos and arXiv papers with impressive results. Interaction and interpretation quality rival NotebookLM. Comes with a Chrome extension already featured by Google. Drawbacks: limited free quota and slightly slow vector processing.
Perceptron Inc. released its flagship video analysis model Mk1, claiming 80-90% lower cost than competitors while achieving strong performance on spatial and video reasoning benchmarks.
This paper introduces three parameter-efficient methods for multi-view proficiency estimation on the Ego-Exo4D dataset, shifting from discriminative classification to generative feedback. The proposed models achieve state-of-the-art accuracy with significantly fewer parameters and training epochs than video-transformer baselines.
A tool that gives Claude the ability to watch and analyze videos by extracting captions and frames, enabling video-based queries and summaries.