image-to-video

Tag

Cards List
#image-to-video

@_akhaliq: SCoPE Sightline-Coordinate Positional Encoding for Video Diffusion Transformers model: https://huggingface.co/TencentAR…

X AI KOLs Timeline · 2026-08-13 Cached

SCoPE is a model from TencentARC that adds camera sightlines as positional coordinates to a pretrained video diffusion transformer, enabling camera trajectory control while preserving the image-to-video prior. The release includes a self-contained checkpoint for Wan2.2-I2V-A14B inference.

0 favorites 0 likes
#image-to-video

@yoheinakajima: one man’s blur is another man’s motion data to decode

X AI KOLs Following · 2026-08-11 Cached

Introduces Blur2Vid, a method that generates video from motion-blurred images, published at SIGGRAPH Asia 2025.

0 favorites 0 likes
#image-to-video

lightx2v/Minimax-h3-Turbo

Hugging Face Models Trending · 2026-08-07 Cached

Hugging Face page for the Minimax-h3-Turbo video generation model, with instructions for using it via Diffusers and Colab/Kaggle notebooks.

0 favorites 0 likes
#image-to-video

China's MiniMax H3 is the first open model to top an AI video ranking (2 minute read)

TLDR AI · 2026-08-04 Cached

MiniMax released the open-weights H3 video model, the first open model to top an AI video ranking, ranking first in video editing and second in text-to-video. The 33B parameter model handles text, images, video, and audio, with some components like 2K resolution and H3-Context-IR remaining closed.

0 favorites 0 likes
#image-to-video

@heyshrutimishra: MiniMax H3 just became the best video editor in AI. #1 on Artificial Analysis's Video Editing Leaderboard, #2 in Text t…

X AI KOLs Following · 2026-07-31 Cached

MiniMax H3 tops the Artificial Analysis video editing leaderboard, offering instruction-based editing on existing clips with multimodal input and competitive pricing, while planning open weights for commercial use.

0 favorites 0 likes
#image-to-video

Comfy-Org/MiniMax-H3

Hugging Face Models Trending · 2026-07-30 Cached

Comfy-Org repackaged MiniMax-H3 model files for ComfyUI, including diffusion models, text encoders, and VAEs, with workflow templates for text-to-video, image-to-video, and reference-to-video generation.

0 favorites 0 likes
#image-to-video

Flux 3

Hacker News Top · 2026-07-24 Cached

Black Forest Labs announces FLUX 3, a multimodal foundation model that jointly learns from images, videos, and audio, enabling unified generation and understanding across modalities with early access now available.

0 favorites 0 likes
#image-to-video

GraphVid: Interactive Graph-Controllable Video Generation

Hugging Face Daily Papers · 2026-07-23 Cached

GraphVid introduces a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs, outperforming prior methods with significant reductions in FID and FVD.

0 favorites 0 likes
#image-to-video

Japan accelerates video generation with new series of anime generation models.

Reddit r/singularity · 2026-07-20 Cached

Japan's AIdea Labs released AnimeGen, a free AI model for anime-style video generation, capable of text-to-video and image-to-video, with commercial use allowed.

0 favorites 0 likes
#image-to-video

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Hugging Face Daily Papers · 2026-07-04 Cached

This paper presents CineMobile, a method for efficient on-device image-to-video generation that achieves a 40x speedup over the teacher model through distillation-guided pruning, diffusion distillation, and hybrid quantization, enabling cinematic camera motion effects on mobile devices.

0 favorites 0 likes
#image-to-video

Pixlie

Product Hunt · 2026-06-19

Pixlie is an AI video studio that enables text and image to video conversion with real control.

0 favorites 0 likes
#image-to-video

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

Hugging Face Daily Papers · 2026-06-18 Cached

Go-with-the-Track unifies motion control and reference image compositing in video generation using point-track embeddings with spatial-aware encoding and video diffusion transformers, achieving superior motion and reference control in a single model.

0 favorites 0 likes
#image-to-video

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

Hugging Face Daily Papers · 2026-06-04 Cached

PhaseLock is a training-free framework that preserves motion priors from early-step inference to improve physical consistency in image-to-video diffusion models, achieving 6.2 point improvement with minimal overhead.

0 favorites 0 likes
#image-to-video

AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation

Hugging Face Daily Papers · 2026-06-02 Cached

AAD-1 introduces asymmetric adversarial distillation with phased training to achieve one-step autoregressive video generation, outperforming prior methods on VBench.

0 favorites 0 likes
#image-to-video

@HuggingPapers: ByteDance just dropped Bernini on Hugging Face Generate or edit videos from text, images, or references Rivals the best…

X AI KOLs Following · 2026-06-01 Cached

ByteDance released Bernini, an open-source video generation and editing model on Hugging Face that rivals top closed-source models.

0 favorites 0 likes
#image-to-video

@FinanceYF5: Same company, two generations of products, a 52-point gap. Grok-Imagine-Video-1.5-Preview (720p) has taken first place in the Image-to-Video Arena! Compared to Grok-Imagine-Video (720p), it has significantly improved…

X AI KOLs Timeline · 2026-05-31 Cached

Grok-Imagine-Video-1.5-Preview (720p) has secured first place in the Image-to-Video Arena, improving by 52 points over its predecessor and surpassing top video models like Seedance-2.0 and HappyHorse.

0 favorites 0 likes
#image-to-video

nvidia/Cosmos3-Super-Image2Video

Hugging Face Models Trending · 2026-05-21 Cached

NVIDIA releases Cosmos3-Super-Image2Video, a model that generates temporally coherent video sequences from an input image and text instructions, part of the Cosmos 3 omnimodal world model platform for Physical AI applications.

0 favorites 0 likes
#image-to-video

Efficient-Large-Model/SANA-WM_bidirectional

Hugging Face Models Trending · 2026-05-18 Cached

SANA-WM is an efficient 2.6B-parameter open-source world model for minute-scale video generation with precise camera control. It uses a hybrid linear diffusion transformer and a two-stage pipeline to produce 720p videos from images and text prompts.

0 favorites 0 likes
#image-to-video

I've been building something for the AI community and would like some early feedback.

Reddit r/AI_Agents · 2026-05-17

Dhee is a new agentic video generation AI that creates videos from a single description, handling prompting, image generation, and assembly, with shot-by-shot editing for refinement.

0 favorites 0 likes
#image-to-video

@Jaaneek: Not only chat Generate images, search x, image to video All with just normal grok subscription!

X AI KOLs Following · 2026-05-15 Cached

Grok AI now offers image generation, X search, and image-to-video capabilities to all subscribers with a normal subscription.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback