generation

Tag

Cards List
#generation

Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII

Reddit r/ArtificialInteligence · 8h ago

Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.

0 favorites 0 likes
#generation

SciForma: Structure-Faithful Generation of Scientific Diagrams

Hugging Face Daily Papers · 2d ago Cached

Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.

0 favorites 0 likes
#generation

@charliermarsh: Impressive on first glance, but I'm pretty sure this Klein bottle was in the training data

X AI KOLs Following · 2d ago Cached

A tweet discusses Claude Fable 5 Max generating a Pokémon crossword on a Klein bottle, but expresses skepticism that the model may have memorized the design from training data.

0 favorites 0 likes
#generation

China already has the capability of making extremely realistic 1-min long AI videos

Reddit r/ArtificialInteligence · 6d ago

China has developed the ability to produce highly realistic AI-generated videos up to one minute long, marking a significant advancement in AI video synthesis.

0 favorites 0 likes
#generation

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

arXiv cs.LG · 6d ago Cached

This paper introduces a unified conceptual framework for discrete diffusion models, analyzing their design space through tokenization, state space construction, and highlighting trade-offs in training, inference, and scaling.

0 favorites 0 likes
#generation

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

Reddit r/LocalLLaMA · 2026-07-13

Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.

0 favorites 0 likes
#generation

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

Hugging Face Daily Papers · 2026-07-13 Cached

This paper introduces SpectraReward, a training-free reward function that leverages pretrained multimodal large language models (MLLMs) as zero-shot reward models for reinforcement learning in text-to-image generation, demonstrating consistent improvements over prior methods.

0 favorites 0 likes
#generation

@mattshumer_: GPT-5.6-Sol one-shotted this voxel-based Manhattan. Just look at the precision... it's insane. It ran for almost a week…

X AI KOLs Timeline · 2026-07-09 Cached

GPT-5.6-Sol autonomously generated a precise voxel-based Manhattan over nearly a week.

0 favorites 0 likes
#generation

@Saccc_c: How can Grok 4.5 be this fast? And the quality is excellent. After uploading my personal website's PRD, it was generated in 3 minutes. The overall outcome is very impressive.

X AI KOLs Following · 2026-07-09 Cached

User praises Grok 4.5's speed and quality. It generated the result in 3 minutes after uploading their personal website's PRD, with good results.

0 favorites 0 likes
#generation

Attending to Multimodal Generation One Token at a Time

Hugging Face Daily Papers · 2026-07-04 Cached

This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.

0 favorites 0 likes
#generation

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

arXiv cs.AI · 2026-06-30 Cached

This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.

0 favorites 0 likes
#generation

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

Hugging Face Daily Papers · 2026-06-29 Cached

AVTok proposes a unified 1D tokenizer for audio-video generation using a dual-stream transformer with shared encoder-decoder and modal-specific queries, achieving compact latent representations and excelling in reconstruction and downstream generation tasks.

0 favorites 0 likes
#generation

I'm eager for a 15x speedup on my strix halo

Reddit r/LocalLLaMA · 2026-06-23

Nvidia claims a 15x speedup in text generation using a diffusion model, generating entire blocks at once.

0 favorites 0 likes
#generation

A2e ai video and image

Reddit r/AI_Agents · 2026-06-21

A2e is an AI tool for video and image generation or processing.

0 favorites 0 likes
#generation

@SergioPaniego: continuous batching just landed in TRL for GRPO at 64 generations it runs faster and uses less VRAM than plain generate…

X AI KOLs Following · 2026-06-19 Cached

Continuous batching has been added to TRL for GRPO, improving speed and VRAM usage without needing vLLM. The tweet explains how it works and when to use it.

0 favorites 0 likes
#generation

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

Hugging Face Daily Papers · 2026-06-19 Cached

UnityShots is a memory-driven multi-shot audio-video generation system that maintains consistent subject appearance and audio across video cuts using fixed-size long-term and short-term memory slots with boundary-conditioned gates and discrete cut-type priors. It outperforms open-source baselines on cross-shot coherence metrics and matches closed-source systems.

0 favorites 0 likes
#generation

UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer

Hugging Face Daily Papers · 2026-06-15 Cached

UniDDT proposes a decoupled diffusion transformer framework that unifies multimodal understanding and generation by leveraging a Noisy ViT encoder and LLM for semantic encoding, achieving strong performance on both tasks.

0 favorites 0 likes
#generation

@_pulkitxm: I spent thousands of dollars building this repo… so you don't have to 50+ landing pages, hero sections & interactive pr…

X AI KOLs Timeline · 2026-06-12 Cached

Anthropic introduced Claude Fable 5, a powerful Mythos-class model, and a developer shared a repository of 50+ landing pages generated with it, preserving prompts for community use.

0 favorites 0 likes
#generation

A bit weird, but okay. (Don't get me wrong it's SOTA for editing, but definitely not generation) Thoughts?

Reddit r/singularity · 2026-06-11

The comment acknowledges that the model is state-of-the-art for editing but not for generation.

0 favorites 0 likes
#generation

Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models

arXiv cs.LG · 2026-06-10 Cached

Introduces UniTok, a universal tokenizer that transforms continuous time series into discrete tokens, and UniTok-FM, a foundation model pretrained via next-token prediction that enables zero-shot and prompt-boosted forecasting as well as few-shot generation and classification through training-free in-context inference.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback