Tag
This Twitter thread showcases Astra AI model's ability to identify sounds from spectrogram images, illustrating ongoing discoveries of its capabilities.
NAPE is a self-supervised audio learning framework that uses causal Transformers to predict next spectrogram patch embeddings, achieving state-of-the-art performance on multiple audio and speech benchmarks with a minimalist design.