computer-vision

Tag

Cards List
#computer-vision

@lillyguisnet: A quality metric is a little difficult here since I haven't had the patience to make ground truth labels, but a 3-way a…

X AI KOLs Following ↗ · 13h ago Cached

The tweet discusses the difficulty of setting a quality metric for polygon output without ground truth labels, but proposes a 3-way agreement method as reliable. It highlights impressive results for a non-specialized Vision-Language Model despite false positives/negatives, with improvements noted using a grid prompt.

0 favorites 0 likes
#computer-vision

@yibie: https://x.com/yibie/status/2104024011692753228

X AI KOLs Timeline ↗ · yesterday Cached

This article explains how to use the logprobs parameter in LLMs to implement Jev-style encapsulation for supporting structured queries, and extend to visual models via the attachments field for image processing, offering a flexible and customizable computer vision approach.

0 favorites 0 likes
#computer-vision

Video models are getting good

Reddit r/singularity ↗ · 2d ago

The article discusses the increasing capabilities and improvements in AI video models, highlighting their growing effectiveness in generating or processing video content.

0 favorites 0 likes
#computer-vision

@airesearch12: Introducing ImageJevBench Jev-Omni decider-2b-vision Reflex 4B The video shows 89 public synthetic images. Running the …

X AI KOLs Timeline ↗ · 2d ago Cached

ImageJevBench v0.1 is introduced as a benchmark for evaluating AI models on image decision tasks, ranking systems like Jev-Omni and decider-2b-vision based on performance, cost, and calibration metrics.

0 favorites 0 likes
#computer-vision

@AdinaYakup: Gen-HumanEgo New dataset for robot learning from @GenrobotAI - 1,800+ hours of egocentric human data - 10K+ unique task…

X AI KOLs Timeline ↗ · 3d ago Cached

Gen-HumanEgo is a new dataset released by GenrobotAI for robot learning, featuring over 1,800 hours of egocentric human data with 10K+ unique tasks and synchronized camera views including 3D hand keypoints and structured annotations.

0 favorites 0 likes
#computer-vision

@LinusEkenstam: Now this is bonkers

X AI KOLs Timeline ↗ · 3d ago Cached

AI capabilities for expression recognition, object detection, and finger counting have achieved less than 1 second latency, highlighting notable real-time performance gains.

0 favorites 0 likes
#computer-vision

The promise and peril of using visual AI to study cities

MIT News — Artificial Intelligence ↗ · 4d ago Cached

Researchers from MIT Senseable City Lab discuss the use of visual AI to analyze urban environments, highlighting its potential for urban planning while addressing concerns about privacy and fairness in a new book.

0 favorites 0 likes
#computer-vision

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper introduces RGBD20K, a large-scale benchmark dataset for RGB-D semantic segmentation with 160 fine-grained categories and 20,000 image pairs, featuring high-quality annotations and a novel score-purified fusion method.

0 favorites 0 likes
#computer-vision

@paul_cal: These capabilities are insane. This isn't perfect but you'd be talking about multiple talented people working for weeks…

X AI KOLs Timeline ↗ · 4d ago Cached

Ryan Sael demonstrates Opus 5.5's ability to create an interactive lens lab for camera focus education in under two hours at a low cost, highlighting advanced AI prototyping capabilities.

0 favorites 0 likes
#computer-vision

@yoheinakajima: got object detection down to below 0.35 sec latency locally

X AI KOLs Timeline ↗ · 4d ago Cached

A developer shares their achievement of reducing object detection latency to below 0.35 seconds on local hardware, highlighting progress in AI performance optimization.

0 favorites 0 likes
#computer-vision

SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts

arXiv cs.CL ↗ · 6d ago Cached

This paper introduces SCoR, a hierarchical framework for forecasting relations between scientific concepts, with a benchmark and model that improve research-direction discovery by predicting typed relations.

0 favorites 0 likes
#computer-vision

ImIR: Image-Instruction Tuning for All-in-One Image Restoration

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

ImIR adapts a pretrained image-editing model for six image restoration tasks using image-derived instructions, enabling efficient and task-agnostic restoration.

0 favorites 0 likes
#computer-vision

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.

0 favorites 0 likes
#computer-vision

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

This paper introduces Segment-Snap, a method that combines geometric and semantic cues to improve interaction understanding in 3D scenes, achieving significant gains in motion-gated AP and handle detection metrics.

0 favorites 0 likes
#computer-vision

Segment Anything Model (SAM) 3.1 (2 minute read)

TLDR AI ↗ · 2026-09-21

SAM 3.1 can detect, segment, and track objects in images and video using text prompts on the Meta Model API, with specified inference costs.

0 favorites 0 likes
#computer-vision

@h4nkdog: I used Codex to vibe code a game on my living room wall! I control it with my hand using computer vision + projection m…

X AI KOLs Following ↗ · 2026-09-20 Cached

A user showcased a project using OpenAI's Codex to code a game controlled by hand gestures with computer vision and projection mapping, with a timelapse and process to be shared in replies.

0 favorites 0 likes
#computer-vision

@DrTBehrens: Qwen-Image-2.1 Test: Multiple references.

X AI KOLs Timeline ↗ · 2026-09-20 Cached

A test showcasing the Qwen-Image-2.1 model's ability to handle multiple reference images.

0 favorites 0 likes
#computer-vision

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Hugging Face Daily Papers ↗ · 2026-09-19 Cached

UltraTex is an efficient framework for high-resolution multi-view diffusion-based 3D texturing, introducing techniques to reduce redundancy and achieve significant speedups in training and inference.

0 favorites 0 likes
#computer-vision

@rohanpaul_ai: Stanford deep learning for computer Vision taught by Professor Fei-Fei Li ( @drfeifei ) "Evolutionary forces drives int…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

A Stanford deep learning course on computer vision taught by Professor Fei-Fei Li is available on YouTube, discussing the evolutionary role of vision in developing intelligence.

0 favorites 0 likes
#computer-vision

Fei-Fei Li just admitted the hard part of her own product isn't the AI. It's shipping it.

Reddit r/artificial ↗ · 2026-09-17

Fei-Fei Li's AI model for 3D reconstruction faces significant challenges in shipping as a reliable product, highlighting the gap between research and real-world application. The article compares it to LiDAR, discussing trade-offs in accuracy and accessibility for industries like construction and design.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback