Tag
LiveGrid is a free browser-based tool that lets users watch multiple live streams from platforms like Twitch, YouTube, and Kick in a customizable grid layout for events such as esports and sports.
A research paper on bundle adjustment for any image set using multi-view matching and monocular priors, achieving state-of-the-art results in computer vision.
CoVeR is a training-free spatial token selector that improves 3D reasoning in Vision-Language Models by enforcing exact token budgets and full scene coverage, outperforming prior state-of-the-art methods.
A paper from the DUSt3R team proposes sparse auto-regressive modeling for 3D scene generation from multi-view images, using a voxel-aligned 3D latent space and an occupancy-aware masked autoregressive transformer.
Presents MBDiff, a multi-view behavior-aware diffusion model for probabilistic utility data imputation that learns user behavior from global, local, and instance-level views and uses a conditional attentional denoising network. Evaluated on real utility data from Florida, it outperforms state-of-the-art baselines.
DSTFView is a dual-input spatio-temporal-frequency multi-view framework for cloud-edge workload forecasting, jointly modeling closeness and period dependencies with an adaptive fusion mechanism to capture abrupt changes.
CodeNib is a multi-view data system that serves repository context to coding agents by building reusable lexical, dense, and structural views per commit, enabling faster updates and efficient context serving.
G-MAD is an open-source framework using Arma 3 to generate synchronized multi-view RGB-T data for aerial object detection, addressing limitations of real-world datasets. It also introduces the AMOD benchmark.
MultiView-Bench is a diagnostic benchmark for evaluating vision-language models on their ability to integrate multiple viewpoints into a coherent 3D mental model, revealing systematic failures in 3D spatial reasoning, and introducing ViewNavigator to mitigate these issues.
Deform360 is a large-scale visuotactile dataset with 198 objects and over 215 hours of observations for studying deformable object dynamics, enabling comparison between 2D video and 3D particle world models for robotic manipulation.
Proposes a method to embed graphs in high-dimensional space and search for informative 2D viewpoints that optimize aesthetic and readability metrics, enabled by a novel differentiable surrogate for edge crossings. Introduces an interactive system, DataFly, for exploring multiple candidate viewpoints.
This paper proposes a feed-forward framework that decomposes 3D scenes into instance-structured token groups from unposed multi-view images, enabling direct object-level reconstruction, segmentation, and manipulation without 3D annotations.
MVTrack4Gen introduces a training framework that uses multi-view point tracking as geometric supervision to enhance motion-aware diffusion models, achieving state-of-the-art geometric consistency and motion fidelity in novel-view video generation from monocular video.
DR-MV3D presents a map-grounded learning framework with dense rewards to improve multi-view 3D visual question answering through global map construction, view-trajectory planning, and egocentric grounding.
PAIWorld enhances diffusion-transformer world models with geometric awareness and cross-view attention to improve multi-view 3D consistency for robotic manipulation tasks, achieving state-of-the-art results on benchmarks.
This paper introduces a non-parametric multi-view Gaussian process framework for detecting machine-generated text that is robust to adversarial manipulations like paraphrasing. By combining complementary features and providing calibrated uncertainty, it outperforms existing detectors on held-out attacks.
WEAVER is a multi-view world model for robotic manipulation that achieves high fidelity, consistency, and efficiency using flow-matching loss, demonstrating superior performance in policy evaluation, improvement, and test-time planning with significant real-world improvements.
X-Stream introduces the first benchmark for multi-stream video understanding, evaluating MLLMs as multiplexers across multiple concurrent streams. The study reveals that current MLLMs achieve only about 50% accuracy, exposing significant limitations in handling multiple streams.
PaGeR adapts the multi-view perspective foundation model Depth Anything 3 to predict scale-invariant and metric depth, surface normals, and sky segmentation from a single equirectangular image, using a fixed cubemap representation that keeps VRAM and runtime constant. The paper also releases the ZüriPano and PanoInfinigen datasets.
Introduces GARD, a diffusion-based framework that operates in the feature space of a feed-forward 3D reconstructor to jointly recover scene geometry and high-quality imagery from degraded inputs.