@TONGYI_SpeechAI: Most audio VAEs treat every latent channel equally — same Gaussian prior everywhere. But audio isn't uniform: low frequ…
Summary
The tweet critiques standard audio VAEs for using a uniform Gaussian prior on all latent channels, ignoring the structured nature of low frequencies and chaotic high frequencies, a problem termed 'disordered information packing.'
View Cached Full Text
Cached at: 07/16/26, 06:22 PM
Most audio VAEs treat every latent channel equally — same Gaussian prior everywhere. But audio isn’t uniform: low frequencies are structured, high frequencies are chaotic. Treating them the same has a name: disordered information packing. https://t.co/rBmmgNii9x
Similar Articles
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
VIBE is a framework that evaluates generative bias in Large Audio-Language Models using open-ended tasks with human-recorded speech, revealing systematic biases triggered by gender and accent cues.
Voice feels like the underrated output layer for AI agents
The article discusses the underutilized potential of voice as an output layer for AI agents, highlighting practical use cases and workflow challenges beyond simple text-to-speech.
$\mathbf{\lambda}$-VAE: Variance Equalization for Posterior Collapse
Identifies two coupled causes of posterior collapse in VAEs and introduces λ-VAE, a modification to the reparameterization step that equalizes variance across latent dimensions, reducing collapse and improving information capacity.
When Vision Speaks for Sound
This paper identifies that video-capable multimodal LLMs often appear to understand audio but actually rely on visual cues, a failure mode termed the audio-visual Clever Hans effect. It introduces Thud, an intervention-driven probing framework to diagnose this issue, and proposes an alignment recipe that improves audio-visual consistency by 28 percentage points.
stabilityai/stable-audio-3-medium
Stability AI releases Stable Audio 3, a family of latent diffusion models for variable-length audio generation and editing, with weights for small and medium models available on Hugging Face.