computer-vision

Tag

Cards List
#computer-vision

Desert Ant Labs

Product Hunt ↗ · 2026-09-09 Cached

Desert Ant Labs offers small, specialized AI models for speech, text, and vision that run offline on devices, with an SDK for easy integration and no per-use costs.

0 favorites 0 likes
#computer-vision

@kwangmoo_yi: Leroy et al., "BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors" We are back t…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

A research paper on bundle adjustment for any image set using multi-view matching and monocular priors, achieving state-of-the-art results in computer vision.

0 favorites 0 likes
#computer-vision

@NVIDIAAI: Visual AI needs to recognize a scene, understand what happened and predict what could happen next. For the AI City Chal…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

The AI City Challenge at ECCV 2026 showcased winners developing robust visual AI systems, with top solutions in video forecasting using NVIDIA Cosmos world foundation models for tasks like scene understanding and prediction.

0 favorites 0 likes
#computer-vision

@svpino: I used to work on computer vision models for robots and drones, and nothing is slower than having to run the robot in o…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

Antioch Robotics announced a $32 million Series A funding round led by GreylockVC to develop digital twins for simulating physical world testing in robotics and drones, aiming to accelerate physical autonomy.

0 favorites 0 likes
#computer-vision

@FinanceYF5: Data labeling is "disappearing"! GPT-6 Astra can effortlessly distinguish between Celtics and Knicks players, and deter…

X AI KOLs Timeline ↗ · 2026-09-08

GPT-6 Astra can automatically identify and label various elements in basketball footage, such as players and referees, potentially making manual data labeling obsolete.

0 favorites 0 likes
#computer-vision

@LearnOpenCV: 1/20 I trained two segmentation models using GPT-6 Astra Ultra as the orchestrator with no human labels supplied. My pr…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

An experiment where two segmentation models were trained using GPT-6 Astra Ultra as an orchestrator without human labels, with lazy prompts due to time constraints, showing predictions on held-out videos.

0 favorites 0 likes
#computer-vision

SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

SynthGait-19K is a large synthetic video dataset for gait analysis that enables benchmarking and demonstrates synthetic supervision transfers to real data. It includes methods like Gait2Vid for video generation and GaitXFormer for estimation.

0 favorites 0 likes
#computer-vision

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Hugging Face Daily Papers ↗ · 2026-09-08 Cached

Marigold V2 repurposes diffusion transformers for monocular depth estimation via single-step inference and a novel fine-tuning protocol, achieving sharper depth maps and significant improvements on benchmarks like KITTI and ETH3D.

0 favorites 0 likes
#computer-vision

@ErenChenAI: WuJi just open-sourced its MINT model and EgoPipeline for reconstructing 3D hand + camera motion from first-person vide…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

WuJi has open-sourced its MINT model and EgoPipeline for reconstructing 3D hand and camera motion from first-person video, along with 1,021 hours of egocentric data to aid robot learning.

0 favorites 0 likes
#computer-vision

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

RelightFormer introduces a feed-forward generative Transformer for direct single- and multi-view image relighting, using cross-attention for illumination injection and permutation-invariant encodings for unordered views, trained on a massive synthetic dataset to achieve state-of-the-art visual quality.

0 favorites 0 likes
#computer-vision

Fei Fei Li: The Race to Build World Models For AI (45 minute podcast)

TLDR AI ↗ · 2026-09-07

World Labs' Atlas unifies AI generation and 3D reconstruction through new-view prediction, using sparse images to infer scenes. The podcast features Fei Fei Li discussing the race to build world models for AI.

0 favorites 0 likes
#computer-vision

@IntuitMachine: Count the number of people in a crowd instantly

X AI KOLs Timeline ↗ · 2026-09-06 Cached

A demonstration highlights how drone footage of crowded streets can be analyzed with just 10 lines of Python to detect and count hundreds of people, illustrating how accessible computer vision has become.

0 favorites 0 likes
#computer-vision

TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation

Hugging Face Daily Papers ↗ · 2026-09-06 Cached

TransNormal-2 improves monocular normal estimation by correcting VAE reconstruction errors with geometry-aware training losses and a lightweight RGB-guided refinement module, achieving strong results with minimal annotations.

0 favorites 0 likes
#computer-vision

@zhiwen_fan_: paper from dust3r’s team

X AI KOLs Timeline ↗ · 2026-09-05 Cached

A paper from the DUSt3R team proposes sparse auto-regressive modeling for 3D scene generation from multi-view images, using a voxel-aligned 3D latent space and an occupancy-aware masked autoregressive transformer.

0 favorites 0 likes
#computer-vision

Why AI food looks like that

The Verge ↗ · 2026-09-04 Cached

This article explores why AI-generated food images often appear unappetizing, explaining the technical limitations of diffusion models and human psychological responses to these visuals.

0 favorites 0 likes
#computer-vision

Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields

arXiv cs.LG ↗ · 2026-09-04 Cached

This paper develops NTK-KIP, MetaQuill, and MetaQuill-KIP algorithms to improve neural field reconstruction from sparse observations, making NTK-driven neural fields non-linear and meta-learnable for efficient few-shot adaptation.

0 favorites 0 likes
#computer-vision

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

LatentStream introduces a progressive latent working memory framework that internalizes streaming visual evidence for continuous reasoning, achieving state-of-the-art results on video benchmarks.

0 favorites 0 likes
#computer-vision

@LinusEkenstam: We’re not sprinting towards a generated world, we are on a rocket ship entering hyperspace. This is extremely powerful.…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

World Labs introduces Atlas, the first multimodal world model that generates images and videos with pixel-perfect camera control and reconstructs them in 3D.

0 favorites 0 likes
#computer-vision

YOLO26-RGB: repurposing YOLO26's depth-trained backbone for image deraining [P]

Reddit r/MachineLearning ↗ · 2026-09-01

This research investigates repurposing the backbone from YOLO26's depth estimation model for image deraining, demonstrating that depth-trained initialization provides a consistent improvement over random initialization in controlled experiments.

0 favorites 0 likes
#computer-vision

Temperature-Adaptive Transformed Teacher Matching

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper proposes a sample-wise adaptive temperature scaling method for Transformed Teacher Matching in knowledge distillation, improving performance on image classification benchmarks by locally minimizing KL divergence between teacher and student distributions.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback