Tag
UniH3 proposes a unified framework for all-in-one medical image restoration by leveraging hierarchical homogeneity and heterogeneity, achieving state-of-the-art performance on benchmarks.
EditVid is a unified training-free video editing framework that supports instruction-guided and subject-guided edits using sparse causal memory, token injection, and soft latent blending, achieving high fidelity and outperforming baseline methods in benchmarks.
RouteTS is a unified forecasting framework that routes time series components between frequency and time domains based on spectral characteristics, improving accuracy and efficiency in handling periodicity and transience.
GenRouter is a unified routing framework for agentic image generation that adaptively directs prompts to optimal workflows, significantly reducing costs and latency while improving visual alignment through demand profiling and self-evolution.
MUGEN introduces a unified motion-language framework that avoids discrete codebooks and iterative decoding, using a single adaptive-length autoencoder with continuous latent slots and one-shot generation to achieve efficient, high-quality text-to-motion and motion-to-text performance across HumanML3D and SnapMoGen benchmarks.
This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.
Introduces Flex-Forcing, a unified training and inference framework that allows video diffusion models to operate under both bidirectional and autoregressive regimes via a flexible chunking mechanism over temporal and denoising steps, achieving better video quality, long-video stability, and faster inference.
Presents Qwen-RobotManip, a Vision-Language-Action foundation model for robotic manipulation that achieves generalization through unified alignment across representation, motion, and behavior dimensions, enabling large-scale training on diverse data sources. It outperforms prior state-of-the-art models across multiple out-of-distribution benchmarks and demonstrates emergent capabilities like zero-shot instruction following and cross-embodiment transfer.
This paper presents a unified multi-modal framework integrating reinforcement learning, high-frequency trading, game-theoretic approaches, and cross-modal sentiment analysis for intelligent financial systems, claiming significant improvements over single-domain systems.
This paper introduces a unified framework for test-time diverse generation in large language models, categorizing methods by where diversity is injected (surface-level vs. specification-level). It proposes specification-level methods that generate diverse intermediate specifications, achieving better output diversity across five open-ended tasks and four backbone models while maintaining quality.
ThinkBooster is a unified framework for test-time compute scaling of LLM reasoning, providing a modular Python library, a performance-efficiency benchmark, an OpenAI-compatible proxy service, and a visual debugger. Empirical results on math and coding tasks demonstrate practical gains with quality-cost trade-offs.
CIPER is a unified transformer framework that jointly performs city-scale retrieval and precise 3-DoF pose estimation from cross-view images, overcoming limitations of cascade pipelines.
OmniRetrieval is a framework that unifies retrieval across heterogeneous knowledge sources (text, tables, graphs) by dispatching native queries to appropriate execution engines, outperforming single-source baselines on a benchmark of 13 datasets and 309 knowledge bases.
FashionLens proposes a unified fashion image retrieval framework using multimodal large language models with adaptive calibration and sampling, achieving state-of-the-art performance across diverse retrieval scenarios.
Aurora is an agentic video editing framework that pairs a tool-augmented vision-language model agent with a diffusion transformer to automatically resolve textual and visual underspecification in user requests, enabling unified video editing tasks like replacement, removal, style transfer, and reference-driven insertion.
Skill1 is a unified framework that trains a single policy to co-evolve skill selection, utilization, and distillation using a shared task-outcome objective. Experiments on ALFWorld and WebShop show it outperforms existing baselines in complex task environments.
The article discusses the UniVidX paper, which introduces a unified multimodal framework for video generation using diffusion priors and discusses its cross-modal coherence mechanisms.
UniMesh introduces a single model that jointly handles 3D mesh generation and understanding via a Mesh Head, Chain-of-Mesh iterative editing, and a self-reflection error-correction mechanism.