Tag
The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.
This paper introduces Segment-Snap, a method that combines geometric and semantic cues to improve interaction understanding in 3D scenes, achieving significant gains in motion-gated AP and handle detection metrics.
SAM 3.1 can detect, segment, and track objects in images and video using text prompts on the Meta Model API, with specified inference costs.
A user showcased a project using OpenAI's Codex to code a game controlled by hand gestures with computer vision and projection mapping, with a timelapse and process to be shared in replies.
A test showcasing the Qwen-Image-2.1 model's ability to handle multiple reference images.
UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.
A Stanford deep learning course on computer vision taught by Professor Fei-Fei Li is available on YouTube, discussing the evolutionary role of vision in developing intelligence.
Fei-Fei Li's AI model for 3D reconstruction faces significant challenges in shipping as a reliable product, highlighting the gap between research and real-world application. The article compares it to LiDAR, discussing trade-offs in accuracy and accessibility for industries like construction and design.
Hackers hacked Flock Safety cameras, copying data to reveal that the system tracks both vehicles and people in detail, raising privacy concerns and exposing surveillance capabilities.
The author tested an AI system to turn raw inventory photos into eBay listings, finding that ensuring correct reasoning and handling marketplace quirks are more challenging than initial item identification.
This paper proposes a training-adaptive convolutional sparse coding framework that leverages information bottleneck principles for robust visual representation, achieving improved performance on CIFAR and ImageNet under input perturbations.
FAMOS is a feed-forward model that predicts movable-part segmentation and joint parameters from sparse point clouds using a Multi-state Articulation Transformer and a procedural data generator, showing consistent improvements over baselines in experiments.
Stability AI's research team presented new work on color consistency for AI-generated images at the 19th European Conference on Computer Vision, addressing production challenges in ensuring color matching across shots.
NVIDIA released FoundationPose on Hugging Face, a unified foundation model for 6-DoF object pose estimation and tracking that works on novel objects without fine-tuning.
This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.
The paper introduces TAPe+MLv3, a compact computer vision system using structured representation for multi-task tasks, achieving competitive performance on benchmarks like COCO with fewer than 100,000 parameters.
Adversarial fashion is emerging as a creative response to AI-powered surveillance, using colorful patterns and designs to evade object detection systems, with projects and products like noRecognition and Cap_able already available.
A browser extension called ChessInsights AI has been developed for real-time chessboard detection and analysis using 100% client-side computer vision, ensuring privacy by processing all data locally without server uploads.
Sentradel is hiring and describes their autonomous counter-drone systems that detect, track, and engage small drones cost-effectively using vision and thermal sensing.
rtmlib is a lightweight, open-source pose estimation library that supports full-body, hand, face, and animal pose tracking, built on rtmpose and vitpose models, with a built-in Gradio web UI.