Tag
UnityShots is a memory-driven multi-shot audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. It outperforms open-source baselines on cross-shot coherence metrics and matches closed-source systems.
UniDDT proposes a decoupled diffusion transformer framework that unifies multimodal understanding and generation by leveraging a Noisy ViT encoder and LLM for semantic encoding, achieving strong performance on both tasks.
Anthropic introduced Claude Fable 5, a powerful Mythos-class model, and a developer shared a repository of 50+ landing pages generated with it, preserving prompts for community use.
The comment acknowledges that the model is state-of-the-art for editing but not for generation.
Introduces UniTok, a universal tokenizer that transforms continuous time series into discrete tokens, and UniTok-FM, a foundation model pretrained via next-token prediction that enables zero-shot and prompt-boosted forecasting as well as few-shot generation and classification through training-free in-context inference.
SenseNova U1 is a unified model that handles understanding, reasoning, and generation of text and images in the same architecture, enabling tasks like planning infographics end-to-end.
Introduces six forms of AI agent workflows: classify and execute, fan-out and consolidate, adversarial validation, generate and filter, tournament, loop until done.
The paper proposes SF-Re2G, a method that improves document-grounded dialogue systems by leveraging document structure to enhance retrieval, reranking, and generation. It validates on Chinese and English datasets.
This paper presents a comprehensive taxonomy of 3D vision research, covering geometric representations, datasets, learning paradigms, and applications in reconstruction, generation, and video modeling.
NAVA proposes a native audio-visual alignment framework for joint audio-video generation using an Align-then-Fuse MMDiT architecture, achieving improved synchronization and controllability with 6.3B parameters.
A tweet pondering whether the first generation raised with AI will become superhumans or emotionally disconnected, sparking debate on AI's societal impact.
A tweet shares a curated list of 50 LLM interview questions covering fundamentals, fine-tuning, generation, advanced concepts, and math, compiled by Hao Hoang.
RankE introduces an end-to-end post-training framework for discrete text-to-image generation that jointly optimizes both the generator and decoder to address the latent covariate shift problem, improving alignment and fidelity simultaneously.
A new AI model generates impressively realistic video and audio, with many observers noting the high quality of the output.
LatentUMM introduces dual latent alignment to improve cross-modal consistency in unified multimodal models by aligning transformations and stabilizing latent dynamics.
Introduces Token Time Continuous Diffusion (TTCD), a new diffusion language model that operates in continuous space with per-token times, outperforming discrete models at high speedups in conditional generation and Sudoku solving.