Tag
Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.
Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.
A tweet discusses Claude Fable 5 Max generating a Pokémon crossword on a Klein bottle, but expresses skepticism that the model may have memorized the design from training data.
China has developed the ability to produce highly realistic AI-generated videos up to one minute long, marking a significant advancement in AI video synthesis.
This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.
Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.
This paper introduces SpectraReward, a training-free reward function that leverages pretrained multimodal large language models (MLLMs) as zero-shot reward models for reinforcement learning in text-to-image generation, demonstrating consistent improvements over prior methods.
GPT-5.6-Sol autonomously generated a precise voxel-based Manhattan over nearly a week.
User praises Grok 4.5's speed and quality. It generated the result in 3 minutes after uploading their personal website's PRD, with good results.
This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.
Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.
A2e is an AI tool for video and image generation or processing.
Continuous batching has been added to TRL for GRPO, improving speed and VRAM usage without needing vLLM. The tweet explains how it works and when to use it.
UnityShots is a memory-driven multi-shot audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. It outperforms open-source baselines on cross-shot coherence metrics and matches closed-source systems.
UniDDT proposes a decoupled diffusion transformer framework that unifies multimodal understanding and generation by leveraging a Noisy ViT encoder and LLM for semantic encoding, achieving strong performance on both tasks.
Anthropic introduced Claude Fable 5, a powerful Mythos-class model, and a developer shared a repository of 50+ landing pages generated with it, preserving prompts for community use.
The comment acknowledges that the model is state-of-the-art for editing but not for generation.
Introduces UniTok, a universal tokenizer that transforms continuous time series into discrete tokens, and UniTok-FM, a foundation model pretrained via next-token prediction that enables zero-shot and prompt-boosted forecasting as well as few-shot generation and classification through training-free in-context inference.