Tag
Videoclaw is a newly launched AI tool that simplifies video creation by allowing users to generate and edit videos through prompts, blending human and generated footage.
Proposes ReCoGen, a two-stage framework for multimodal-conditioned time-series generation under irregular missingness, achieving state-of-the-art downstream utility on physiological benchmarks.
This paper introduces LingT2I, a 10-language, 33K-prompt benchmark for evaluating cross-lingual consistency in text-to-image generation, revealing linguistic inequality and language-dependent trade-offs across content generation and text rendering.
RA-CAD presents a state-aware agent for text-to-CAD generation that uses a Generate–Execute–Critique–Rewrite loop, with feedback-driven agent optimization via Group Relative Policy Optimization. It achieves state-of-the-art execution validity and geometric quality on CADFusion and Text2CAD benchmarks.
TraceCAD is a persistent recovery layer for LLM-based CAD agents, diagnosing faulty operations and performing localized, reusable repairs to improve final geometric quality and repair reliability.
MotifRole-Diff proposes a role-aware corruption schedule for masked discrete diffusion on molecular graphs, allocating masking rates based on denoising difficulty and graph-level perturbation impact, demonstrating improved validity and reduced FCD on QM9 and MOSES benchmarks.
Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.
Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.
A tweet discusses Claude Fable 5 Max generating a Pokémon crossword on a Klein bottle, but expresses skepticism that the model may have memorized the design from training data.
China has developed the ability to produce highly realistic AI-generated videos up to one minute long, marking a significant advancement in AI video synthesis.
This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.
Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.
This paper introduces SpectraReward, a training-free reward function that leverages pretrained multimodal large language models (MLLMs) as zero-shot reward models for reinforcement learning in text-to-image generation, demonstrating consistent improvements over prior methods.
GPT-5.6-Sol autonomously generated a precise voxel-based Manhattan over nearly a week.
User praises Grok 4.5's speed and quality. It generated the result in 3 minutes after uploading their personal website's PRD, with good results.
This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.
Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.
A2e is an AI tool for video and image generation or processing.