Tag
Proposes ReCoGen, a two-stage framework for multimodal-conditioned time-series generation under irregular missingness, achieving state-of-the-art downstream utility on physiological benchmarks.
This paper introduces LingT2I, a 10-language, 33K-prompt benchmark for evaluating cross-lingual consistency in text-to-image generation, revealing linguistic inequality and language-dependent trade-offs across content generation and text rendering.
RA-CAD presents a state-aware agent for text-to-CAD generation that uses a Generate–Execute–Critique–Rewrite loop, with feedback-driven agent optimization via Group Relative Policy Optimization. It achieves state-of-the-art execution validity and geometric quality on CADFusion and Text2CAD benchmarks.
TraceCAD is a persistent recovery layer for LLM-based CAD agents, diagnosing faulty operations and performing localized, reusable repairs to improve final geometric quality and repair reliability.
MotifRole-Diff proposes a role-aware corruption schedule for masked discrete diffusion on molecular graphs, allocating masking rates based on denoising difficulty and graph-level perturbation impact, demonstrating improved validity and reduced FCD on QM9 and MOSES benchmarks.
Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.
Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.
A tweet discusses Claude Fable 5 Max generating a Pokémon crossword on a Klein bottle, but expresses skepticism that the model may have memorized the design from training data.
China has developed the ability to produce highly realistic AI-generated videos up to one minute long, marking a significant advancement in AI video synthesis.
This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.
Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.
This paper introduces SpectraReward, a training-free reward function that leverages pretrained multimodal large language models (MLLMs) as zero-shot reward models for reinforcement learning in text-to-image generation, demonstrating consistent improvements over prior methods.
GPT-5.6-Sol autonomously generated a precise voxel-based Manhattan over nearly a week.
User praises Grok 4.5's speed and quality. It generated the result in 3 minutes after uploading their personal website's PRD, with good results.
This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.
Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.
A2e is an AI tool for video and image generation or processing.
Continuous batching has been added to TRL for GRPO, improving speed and VRAM usage without needing vLLM. The tweet explains how it works and when to use it.