Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Summary
This preprint introduces hierarchical self-supervised world models for music co-creation agents, with fast CPU-friendly models and a live demo for MIDI inpainting and generation.
View Cached Full Text
Cached at: 08/07/26, 01:54 AM
Paper page - Helping Music Co-Creation Agents ‘Listen’ Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Source: https://huggingface.co/papers/2608.04378 Excited to share a new Preprint & LIVE DEMO! This is >2 years of my life building fast, lightweight CPU-friendly models to facilitate songwriters’ iterative workflows, with Representations usable for Understanding and Generation, trained in a Self-Supervised way. First, the demo: https://drscotthawley-midi-rae-jepa-son.hf.space
This turned outso funthat I had to pause paper-writing to revamp it into a user-friendly app to send to friends! You draw an inpainting mask that controls where spatial dropout occurs, and the “EQ”-looking sliders scale the abstraction level.
The preprint is here:https://arxiv.org/abs/2608.04378The first 6 pages + refs are submitted to the NeurIPS Creative AI Track (“single-blind, preprints ok” 👍) Preprint adds 10 pages of Supplemental Materials: ablation studies, hyperparameter surveys, etc.
You may have seen another preprint from me a couple weeks ago (https://arxiv.org/abs/2607.14537), which was the project state mid-April 2026 submitted to ISMIR (double-blind, gag order) but the new preprint is the more mature one (despite being the “-SON”!)
The story started with the “graphical prompts” of “Pictures of MIDI” ca. Jan. 2024 (https://picturesofmidi.github.io/PicturesOfMIDI/) but that was big & slow. The quest to streamline it led me to Flow Models (fast sampling), Representation AutoEncoders (reusable rep’s), and LeJEPA (to resist collapse).
You could wade through my messy “midi-rae” work repo but maybe best to wait til I push a clean distro-repo in the coming weeks. Stay tuned via the links on the project website:https://drscotthawley.github.io/midi-rae-jepa-son/Weights are closed for now; I’ll probably open them closer to NeurIPS time.
Similar Articles
@iScienceLuvr: Music-JEPA: Learning a World Model of Sound from Action "we propose to learn a world model of piano sound using JEPA by…
This paper proposes Music-JEPA, a world model that learns piano sound representations by framing audio as a state and piano roll as an action. It captures action-sound relationships and enables downstream tasks like beat tracking and piano transcription via planning.
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
This paper introduces Live Music Diffusion Models (LMDMs), which modify the diffusion process to enable efficient block-wise processing and novel training paradigms for real-time interactive music generation on consumer hardware, outperforming discrete autoregressive models in inference complexity and enabling stable post-training alignment.
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
Wan-Dancer introduces a hierarchical framework for generating minute-scale coherent dances from music, addressing long-duration choreography generation.
Learning Implicit Causal World Models from Multi-Agent Demonstrations
This paper presents a method for learning implicit causal world models from multi-agent demonstrations, enabling agents to infer causal structures from observed behavior.
new AI music model dropped, demos sound surprisingly real
A new AI music model has been released, with demos that sound surprisingly realistic.