@TONGYI_SpeechAI: Most audio VAEs treat every latent channel equally — same Gaussian prior everywhere. But audio isn't uniform: low frequ…

X AI KOLs Timeline Papers

Summary

The tweet critiques standard audio VAEs for using a uniform Gaussian prior on all latent channels, ignoring the structured nature of low frequencies and chaotic high frequencies, a problem termed 'disordered information packing.'

Most audio VAEs treat every latent channel equally — same Gaussian prior everywhere. But audio isn't uniform: low frequencies are structured, high frequencies are chaotic. Treating them the same has a name: disordered information packing. https://t.co/rBmmgNii9x
Original Article
View Cached Full Text

Cached at: 07/16/26, 06:22 PM

Most audio VAEs treat every latent channel equally — same Gaussian prior everywhere. But audio isn’t uniform: low frequencies are structured, high frequencies are chaotic. Treating them the same has a name: disordered information packing. https://t.co/rBmmgNii9x

Similar Articles

When Vision Speaks for Sound

Hugging Face Daily Papers

This paper identifies that video-capable multimodal LLMs often appear to understand audio but actually rely on visual cues, a failure mode termed the audio-visual Clever Hans effect. It introduces Thud, an intervention-driven probing framework to diagnose this issue, and proposes an alignment recipe that improves audio-visual consistency by 28 percentage points.

stabilityai/stable-audio-3-medium

Hugging Face Models Trending

Stability AI releases Stable Audio 3, a family of latent diffusion models for variable-length audio generation and editing, with weights for small and medium models available on Hugging Face.