Tag
Vivix.AI launched A1, a real-time video model for interactive characters and worlds that generates audio and video together in small consecutive segments, enabling seamless interruption and live, long-form conversation with persistent state.
The paper proposes AV-GRPO, a modality-anchored reinforcement learning framework for joint audio-video generation that improves generation quality, semantic alignment, and cross-modal synchronization over existing methods.
Temporal Context Routing improves script-aligned timing in joint audio-video generation by mapping script timing onto shared video-audio temporal axes, reducing errors in shot boundaries and dialogue.
DreamX-Creator 1.0 is a compact 7B native joint audio-video generator that uses cross-modal attention, progressive training, and reinforcement learning to produce synchronized 2K resolution outputs, aiming to democratize high-quality audio-video generation.
Multi2AV-Safety is the first benchmark for evaluating safety in multimodal-to-audio-Video generation, covering all 11 non-singleton conditioning configurations with 11,024 attack instances, and revealing compositional risks where harmful semantics emerge from benign inputs.
LightX2V releases an open, local LoRA prompt rewriter for MiniMax-H3 text-to-audio-video generation, fine-tuned on Qwen3.6-27B to expand short prompts into structured audio-video descriptions.
This paper introduces MultiRef-Compass, a comprehensive benchmark for multi-reference-to-audio-video generation, comprising 350 curated samples and an evaluation protocol with four dimensions including Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.
AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.
MSAVBench is the first comprehensive benchmark and adaptive evaluation framework for multi-shot audio-video generation, assessing 19 models across diverse tasks and achieving high alignment with human judgment.
LTX-2 is the first DiT-based audio-video foundation model from Lightricks, offering synchronized audio and video generation, high fidelity, and production-ready outputs, with open-source code and open model weights.