Tag
提出了MEGA(Mesh Extraction from Gaussians)这一“先分割后网格化”框架,通过Spatial Visual Distillation(SVD)和掩码引导的神经曲面重建,从3DGS场景中提取物体级、水密的网格,实现高质量的物体级几何占用和物理交互。
Anthropic and OpenAI are winning market share in 2026 through business model segmentation rather than technical innovation: Anthropic's enterprise metered billing doubled its revenue in a quarter, while OpenAI's 80% price cut on its cheapest Luna model pushed it toward $70B in run rate. Both companies could near $100B in revenue by year-end, but margins and customer lock-in remain uncertain.
A PhD student shares the release of their paper on AI programs for bioimaging analysis, utilizing SAM and SAM2 models to automate segmentation and data extraction in scientific applications.
This paper introduces Segment-Snap, a method that combines geometric and semantic cues to improve interaction understanding in 3D scenes, achieving significant gains in motion-gated AP and handle detection metrics.
This paper assesses the generalization of nnU-Net for brain tumor segmentation in the BraTS-GoAT 2026 challenge, reporting performance metrics and analyzing failure cases.
An experiment where two segmentation models were trained using GPT-6 Astra Ultra as an orchestrator without human labels, with lazy prompts due to time constraints, showing predictions on held-out videos.
The article analyzes the segmentation of the frontier AI market through exclusive partnerships, access restrictions, and government regulations, highlighting Nvidia's investments in open ecosystems.
This paper introduces SAPO, a segment-level automatic prompt optimization method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on weak and strong examples. Evaluated across several benchmarks, SAPO outperforms strong APO baselines including APE, OPRO, EvoPrompt, GEPA, and StraGO on GPT-3.5-Turbo and GPT-4o-mini.
This paper introduces ConCor-1, a grounding model that treats vision-language grounding as bidirectional concept correspondence, jointly recovering text spans, image segments, and cross-modal matches without prespecified phrases. It unifies phrase grounding, referring expression grounding, and open-vocabulary detection, achieving significant F1 improvements on long-caption and zero-shot LVIS benchmarks.
iFAN is a training framework that improves mask transformers for segmentation by aligning query ranking with mask quality and distilling intermediate predictions to the final layer, yielding consistent gains across benchmarks.
This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.
GLI-AL is a new label resource for glioma MRI that unifies anatomy and lesion labels, expanding supervision to include healthy tissues and previously unlabeled abnormalities. It provides 1,251 label sets aligned with BraTS-GLI cases.
This paper addresses the challenge of maintaining boundary faithfulness in backbone models when processing subsequent video frames beyond the first.
This paper presents a systematic comparison of two geospatial foundation models, TerraMind and THOR, developed under ESA's φ-lab, analyzing how architectural choices like patch size and decoder type affect performance across ten use cases in Earth observation tasks.
The tweet praises SAM3.1 for improved segmentation with faint borders and better tail coverage, leading to more accurate length measurements.
This paper introduces VIP-SAM for instance-level garment segmentation and CtrlVTON, a controllable virtual try-on framework that treats try-on as an image editing problem, allowing precise control over garment layout, style, and placement. Both methods achieve state-of-the-art results on their respective tasks.
REBASE is a training-free framework that suppresses spurious contextual correspondences in in-context segmentation by projecting features onto the orthogonal complement of a low-rank background subspace, achieving state-of-the-art results among training-free methods on several datasets.
This paper introduces SAM-MT, an extension of SAM2 for real-time interactive multi-target video segmentation, achieving high FPS independent of target count.
This paper describes FBK's submission to the IWSLT 2026 Instruction Following shared task, developing SpeechLLMs for short-form and long-form speech instruction following, exploring segmentation methods and achieving robust long-form performance with fixed 30-second segmentation.
This paper presents CALHippo, a framework for 3D mapping of neurons and glial cells in the human hippocampus using state-of-the-art segmentation and density estimation models.