Tag
MovieGrid is a multi-grid post-training paradigm that decomposes long videos into spatially arranged chunks to improve multi-shot coherence and efficiency, achieving state-of-the-art intra-shot and inter-shot consistency in video generation.
LiveAnimate presents a 14B-parameter video diffusion transformer enabling real-time, long-form pose-driven human animation with stable streaming and constant memory usage via PR-Sink attention. Achieves 19.63 FPS inference on two H100 GPUs, maintaining quality over three-minute streams.
This paper proposes a novel online reinforcement learning method to improve factuality in reasoning LLMs by designing a reward function that balances factual precision, detail, and relevance, achieving a 23.1 percentage point reduction in hallucination rate on six benchmarks.
Marc Andreessen and Joe Rogan had a conversation lasting over 3 hours. This article summarizes 17 noteworthy points.
Introduces Magnet, a multi-agent goal-driven narrative engine for long-form story generation with persona-grounded characters, and Atlas, a graph-based pipeline for detecting hallucinations in generated narratives. The framework improves coherence and reduces hallucinations compared to single-model baselines and IBSEN.
SwanVoice is a zero-shot text-to-speech model designed for expressive long-form monologue and dialogue synthesis, combining VAE, flow-matching DiT, and diffusion post-training to achieve higher richness and hierarchy scores than existing baselines.
Swanbench-Speech is a comprehensive benchmark for evaluating long-form speech generation across diverse scenarios, using multi-dimensional metrics covering acoustics, semantics, and expressiveness, revealing limitations of current models.
This paper constructs a large dataset of 263,911 long-form stories annotated with TTCW-based creativity metrics and fine-tunes Qwen3 models to generate structured review reports. It finds that non-reasoning fine-tuning outperforms reasoning-supervised fine-tuning, which suffers from parse failures and irrelevant repetition.
YuE is a family of open foundation models that generates long-form music with aligned lyrics and coherent structure, using innovative techniques like track-decoupled next-token prediction and structural progressive conditioning.