computer-vision

Tag

Cards List
#computer-vision

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.

0 favorites 0 likes
#computer-vision

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

This paper introduces Segment-Snap, a method that combines geometric and semantic cues to improve interaction understanding in 3D scenes, achieving significant gains in motion-gated AP and handle detection metrics.

0 favorites 0 likes
#computer-vision

Segment Anything Model (SAM) 3.1 (2 minute read)

TLDR AI ↗ · 2026-09-21

SAM 3.1 can detect, segment, and track objects in images and video using text prompts on the Meta Model API, with specified inference costs.

0 favorites 0 likes
#computer-vision

@h4nkdog: I used Codex to vibe code a game on my living room wall! I control it with my hand using computer vision + projection m…

X AI KOLs Following ↗ · 2026-09-20 Cached

A user showcased a project using OpenAI's Codex to code a game controlled by hand gestures with computer vision and projection mapping, with a timelapse and process to be shared in replies.

0 favorites 0 likes
#computer-vision

@DrTBehrens: Qwen-Image-2.1 Test: Multiple references.

X AI KOLs Timeline ↗ · 2026-09-20 Cached

A test showcasing the Qwen-Image-2.1 model's ability to handle multiple reference images.

0 favorites 0 likes
#computer-vision

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Hugging Face Daily Papers ↗ · 2026-09-19 Cached

UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.

0 favorites 0 likes
#computer-vision

@rohanpaul_ai: Stanford deep learning for computer Vision taught by Professor Fei-Fei Li ( @drfeifei ) "Evolutionary forces drives int…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

A Stanford deep learning course on computer vision taught by Professor Fei-Fei Li is available on YouTube, discussing the evolutionary role of vision in developing intelligence.

0 favorites 0 likes
#computer-vision

Fei-Fei Li just admitted the hard part of her own product isn't the AI. It's shipping it.

Reddit r/artificial ↗ · 2026-09-17

Fei-Fei Li's AI model for 3D reconstruction faces significant challenges in shipping as a reliable product, highlighting the gap between research and real-world application. The article compares it to LiDAR, discussing trade-offs in accuracy and accessibility for industries like construction and design.

0 favorites 0 likes
#computer-vision

Hackers reveal how Flock cameras really track cars and people

Ars Technica ↗ · 2026-09-17 Cached

Hackers hacked Flock Safety cameras, copying data to reveal that the system tracks both vehicles and people in detail, raising privacy concerns and exposing surveillance capabilities.

0 favorites 0 likes
#computer-vision

102 raw inventory photos → 22 live eBay listings. Now I’m trying to use less AI

Reddit r/artificial ↗ · 2026-09-17

The author tested an AI system to turn raw inventory photos into eBay listings, finding that ensuring correct reasoning and handling marketplace quirks are more challenging than initial item identification.

0 favorites 0 likes
#computer-vision

Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

This paper proposes a training-adaptive convolutional sparse coding framework that leverages information bottleneck principles for robust visual representation, achieving improved performance on CIFAR and ImageNet under input perturbations.

0 favorites 0 likes
#computer-vision

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Hugging Face Daily Papers ↗ · 2026-09-17 Cached

FAMOS is a feed-forward model that predicts movable-part segmentation and joint parameters from sparse point clouds using a Multi-state Articulation Transformer and a procedural data generator, showing consistent improvements over baselines in experiments.

0 favorites 0 likes
#computer-vision

@StabilityAI: Color consistency is a persistent friction point in production. Our interactive research team just returned from The 19…

X AI KOLs Timeline ↗ · 2026-09-15

Stability AI's research team presented new work on color consistency for AI-generated images at the 19th European Conference on Computer Vision, addressing production challenges in ensuring color matching across shots.

0 favorites 0 likes
#computer-vision

@HuggingPapers: NVIDIA just released FoundationPose on Hugging Face A unified foundation model for 6-DoF object pose estimation and tra…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

NVIDIA released FoundationPose on Hugging Face, a unified foundation model for 6-DoF object pose estimation and tracking that works on novel objects without fine-tuning.

0 favorites 0 likes
#computer-vision

Discovering and Preserving Category Correlation Knowledge via Adaptive Reciprocal Knowledge Distillation

arXiv cs.LG ↗ · 2026-09-15 Cached

This paper proposes adaptive reciprocal knowledge distillation (AR-KD), a novel method that improves knowledge transfer from teacher to student models by simplifying the teacher's output distribution through relational alignment, achieving up to 7.13% accuracy gain on CIFAR-100 and ImageNet-1k datasets.

0 favorites 0 likes
#computer-vision

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

Hugging Face Daily Papers ↗ · 2026-09-15 Cached

The paper introduces TAPe+MLv3, a compact computer vision system using structured representation for multi-task tasks, achieving competitive performance on benchmarks like COCO with fewer than 100,000 parameters.

0 favorites 0 likes
#computer-vision

Adversarial Fashion Makes a Statement on AI Panopticon

Hacker News Top ↗ · 2026-09-14 Cached

Adversarial fashion is emerging as a creative response to AI-powered surveillance, using colorful patterns and designs to evade object detection systems, with projects and products like noRecognition and Cap_able already available.

0 favorites 0 likes
#computer-vision

[P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P]

Reddit r/MachineLearning ↗ · 2026-09-14

A browser extension called ChessInsights AI has been developed for real-time chessboard detection and analysis using 100% client-side computer vision, ensuring privacy by processing all data locally without server uploads.

0 favorites 0 likes
#computer-vision

@sts_3d: Sentradel is hiring! See https://sentradel.com for more. If you DM me, please include details on relevant projects / ex…

X AI KOLs Following ↗ · 2026-09-13 Cached

Sentradel is hiring and describes their autonomous counter-drone systems that detect, track, and engage small drones cost-effectively using vision and thermal sensing.

0 favorites 0 likes
#computer-vision

@oliviscusAI: this tool can track perfect 3D motion. rtmlib is a lightweight pose estimation library covering full body, hands, face,…

X AI KOLs Following ↗ · 2026-09-13 Cached

rtmlib is a lightweight, open-source pose estimation library that supports full-body, hand, face, and animal pose tracking, built on rtmpose and vitpose models, with a built-in Gradio web UI.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback