Tag
The paper presents PANORAMA, a vision-language model for panoptic grounded captioning that uses mask proposal selection to ground captions with pixel-level masks, and introduces the PanoCaps benchmark for evaluation.
MultiMatte is a promptable image background removal model fine-tuned from SAM 3 using LoRA, achieving higher accuracy on benchmarks by outputting alpha mattes for better handling of fuzzy boundaries.
A technique is shared for using the Segment Anything Model (SAM) to segment images without prompting by exploiting similarities in microscopy images, enabling zero-shot mask propagation similar to video.
A user shares enthusiastic feedback about SAM 3.1's ability to accurately segment images using simple text prompts like 'worm', highlighting significant improvements over SAM 1.
SAM 3 introduces a unified model for promptable concept segmentation and tracking, achieving state-of-the-art performance with a decoupled recognition and localization architecture and a scalable data engine.
BiRefNet is an AI model for image segmentation available on Replicate, offering cost-effective inference on Nvidia A100 GPUs and is open-source.