Tag
Hackers hacked Flock Safety cameras, copying data to reveal that the system tracks both vehicles and people in detail, raising privacy concerns and exposing surveillance capabilities.
The author tested an AI system to turn raw inventory photos into eBay listings, finding that ensuring correct reasoning and handling marketplace quirks are more challenging than initial item identification.
This paper proposes a training-adaptive convolutional sparse coding framework that leverages information bottleneck principles for robust visual representation, achieving improved performance on CIFAR and ImageNet under input perturbations.
FAMOS is a feed-forward model that predicts movable-part segmentation and joint parameters from sparse point clouds using a Multi-state Articulation Transformer and a procedural data generator, showing consistent improvements over baselines in experiments.
Stability AI's research team presented new work on color consistency for AI-generated images at the 19th European Conference on Computer Vision, addressing production challenges in ensuring color matching across shots.
NVIDIA released FoundationPose on Hugging Face, a unified foundation model for 6-DoF object pose estimation and tracking that works on novel objects without fine-tuning.
This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.
The paper introduces TAPe+MLv3, a compact computer vision system using structured representation for multi-task tasks, achieving competitive performance on benchmarks like COCO with fewer than 100,000 parameters.
Adversarial fashion is emerging as a creative response to AI-powered surveillance, using colorful patterns and designs to evade object detection systems, with projects and products like noRecognition and Cap_able already available.
A browser extension called ChessInsights AI has been developed for real-time chessboard detection and analysis using 100% client-side computer vision, ensuring privacy by processing all data locally without server uploads.
Sentradel is hiring and describes their autonomous counter-drone systems that detect, track, and engage small drones cost-effectively using vision and thermal sensing.
rtmlib is a lightweight, open-source pose estimation library that supports full-body, hand, face, and animal pose tracking, built on rtmpose and vitpose models, with a built-in Gradio web UI.
RelateAnything is a lightweight, real-time open-vocabulary relation prediction model that accepts arbitrary predicate vocabularies and region sources, trained on a large geometrically verified dataset and evaluated on new cross-dataset benchmarks, showing significant performance gains over comparable methods.
Tesla's Full Self-Driving (FSD) system effectively handles sunlight glare scenarios, as demonstrated by a user's positive experience shared in a tweet, highlighting its advanced photon count reconstruction capabilities.
MultiMatte is a promptable image background removal model fine-tuned from SAM 3 using LoRA, achieving higher accuracy on benchmarks by outputting alpha mattes for better handling of fuzzy boundaries.
A basketball AI system that combines local AI with GPT-6 Astra to detect and track players, perform OCR for player numbers, and map trajectories for sports analytics.
This paper presents EPIC-Contact, an in-the-wild dataset for 3D hand-object pose estimation, and HOPformer, a transformer model that jointly predicts hand and object poses from a single RGB image.
This paper introduces TRACE, a benchmark for post-fire object understanding, and proposes a Feature Recovery Module (FRM) to restore degraded features and improve detection and vision-language tasks under severe physical damage.
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates task-specific inductive biases and achieves state-of-the-art accuracy on benchmarks like Sintel and KITTI-2015.
The paper presents World in World, a training-free interface that enables flexible camera and time control in frozen autoregressive video world models by using correspondence-guided queries and evidence-wise attention guidance.