@AdinaYakup: JD just released JoyAI-Echo An interesting long video generation model 5 minute multi shot video generation Cross modal…

X AI KOLs Following Models

Summary

JD released JoyAI-Echo, a long video generation model capable of 5-minute multi-shot video with cross-modal memory for character and voice consistency, native audio+video generation, and 7.5x speed improvement via DMD distillation.

JD just released JoyAI-Echo 📹 An interesting long video generation model ✨ 5 minute multi shot video generation ✨ Cross modal memory for character & voice consistency ✨ Native audio + video generation ✨ 7.5× faster via DMD distillation (without quality loss) https://t.co/qIel5Gc8qX
Original Article
View Cached Full Text

Cached at: 06/03/26, 01:51 PM

JD just released JoyAI-Echo 📹 An interesting long video generation model

✨ 5 minute multi shot video generation ✨ Cross modal memory for character & voice consistency ✨ Native audio + video generation ✨ 7.5× faster via DMD distillation (without quality loss) https://t.co/qIel5Gc8qX

Similar Articles

jdopensource/JoyAI-Echo

Hugging Face Models Trending

JD Open Source releases JoyAI-Echo (Echo-LongVideo), a text-to-audio-video diffusion model capable of generating minute-level multi-shot videos with consistent character identity and voice, using DMD distillation for 7.5x speedup.

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

Hugging Face Daily Papers

This paper presents JoyAI-VL-Interaction, an open-source 8B-scale vision-language model that operates continuously in real-time, deciding autonomously when to respond or delegate. It includes a complete deployable system and a training recipe, outperforming Doubao and Gemini in human evaluations.