MusiChat: Vibe Composing for Music Creation

arXiv cs.AI Papers

Summary

MusiChat presents a conversational system for human-AI music co-creation that enables iterative refinement through natural language interaction, achieving high accuracy in multi-turn editing.

arXiv:2607.24873v1 Announce Type: new Abstract: Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas. We present MusiChat, a conversational vibe composing system that enables collaborative human-AI music creation through natural-language interaction and iterative refinement. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure-preserving edits. The system integrates a large language model with a hybrid symbolic music engine through a memory-augmented architecture that maintains the active composition state and user history across interactions. A hybrid intent-routing mechanism further enables efficient interpretation of both precise musical edits and open-ended creative requests. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent. We evaluate MusiChat through objective analysis and human studies, achieving 95.31% and 100% accuracy for single- and multi-turn interactions, respectively, and obtaining like-to-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality. Our results demonstrate that MusiChat supports coherent multi-turn music authoring and interactive human-AI co-creation through a conversational interface.
Original Article
View Cached Full Text

Cached at: 07/29/26, 09:52 AM

# MusiChat: Vibe Composing for Music Creation
Source: [https://arxiv.org/html/2607.24873](https://arxiv.org/html/2607.24873)
Callie C\. Liao∗ Stanford University Stanford, CA ccliao@cs\.stanford\.edu &Duoduo Liao∗ George Mason University Fairfax, VA dliao2@gmu\.edu &Ellie L\. Zhang IntelliSky McLean, VA elzhang@intellisky\.org

###### Abstract

Recent advances in AI music generation have enabled users to create complete musical pieces from natural\-language prompts\. However, most existing systems follow a prompt\-and\-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas\. We present MusiChat, a conversational vibe composing system that enables collaborative human–AI music creation through natural\-language interaction and iterative refinement\. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric\-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure\-preserving edits\. The system integrates a large language model with a hybrid symbolic music engine through a memory\-augmented architecture that maintains the active composition state and user history across interactions\. A hybrid intent\-routing mechanism further enables efficient interpretation of both precise musical edits and open\-ended creative requests\. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent\. We evaluate MusiChat through objective analysis and human studies, achieving 95\.31% and 100% accuracy for single\- and multi\-turn interactions, respectively, and obtaining like\-to\-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality\. Our results demonstrate that MusiChat supports coherent multi\-turn music authoring and interactive human–AI co\-creation through a conversational interface\.

MusiChat: Vibe Composing for Music Creation

Callie C\. Liao∗Stanford UniversityStanford, CAccliao@cs\.stanford\.eduDuoduo Liao∗George Mason UniversityFairfax, VAdliao2@gmu\.eduEllie L\. ZhangIntelliSkyMcLean, VAelzhang@intellisky\.org

\*\*footnotetext:Equal contribution\.11footnotetext:MusiChat:[demo video](https://www.intellisky.org/share/musichat_demo.mp4)and[AWS\-based web tool](https://musichat.intellisky.ai/)\.## 1Introduction

![Refer to caption](https://arxiv.org/html/2607.24873v1/latex/figures/musichat_ui.png)Figure 1:The MusiChat workspace\. Left: engraved sheet music, MIDI player \(instrument \+ chords controls\), and piano roll\. Right: the conversation and style selector\.![Refer to caption](https://arxiv.org/html/2607.24873v1/x1.png)Figure 2:MusiChat framework for vibe composing for music co\-creation\. Multimodal inputs are first abstracted into an explicit lyrical representation via LLM\-based reasoning\. Lyrics serve as a persistent intermediate reasoning representation that guides hybrid symbolic music generation\. A conversational interface enables iterative refinement by operating directly over these representations, transforming chat into an interactive reasoning environment\.With the proliferation of existing, well\-known music generation models such as SunoSuno \([2026](https://arxiv.org/html/2607.24873#bib.bib125)\), Google’s LyriaCaillonet al\.\([2025](https://arxiv.org/html/2607.24873#bib.bib126)\), Stable Audio 3\.0Evanset al\.\([2026](https://arxiv.org/html/2607.24873#bib.bib127)\), and MusicGenCopetet al\.\([2024](https://arxiv.org/html/2607.24873#bib.bib128)\), music generation can now be accomplished with a single text, image, or audio prompt, making music creation possible for creators with little to no music experience\. Some have also provided a wide variety of features to edit and refine the generated music, including allowing the artist to edit the music manually themselves within the product, adjusting the parameters as the music generates, and regenerating selected segments of the song through lyric editing or by just normal regeneration\.

While they may offer segment regeneration, changes in music style as the music generates, and manual editing, a wish to regenerate only a portion of a musical feature \(such as a melody or instrumental part\) thatmaintainsthe musical continuity of the original song remains out of scope, as users frequently receive a drastically different regeneration of that portion instead\. Asking music generation models to revise a portion of the song is similar to asking state\-of\-the\-art Large Language Models \(LLMs\) such as ChatGPT and Claude to modify only a paragraph of an essay while keeping others the same—in\-turn refinement is currently not possible, as they only provide one\-shot generation\(Huanget al\.,[2018](https://arxiv.org/html/2607.24873#bib.bib13); Payne,[2019](https://arxiv.org/html/2607.24873#bib.bib123); Dhariwalet al\.,[2020](https://arxiv.org/html/2607.24873#bib.bib12); Yuet al\.,[2025](https://arxiv.org/html/2607.24873#bib.bib7)\), where semantic content, structure, and surface realization are entangled in opaque representations, limiting interpretability, controllability, and collaborative use\(Agostinelliet al\.,[2023](https://arxiv.org/html/2607.24873#bib.bib23); Wuet al\.,[2023](https://arxiv.org/html/2607.24873#bib.bib8); Zhouet al\.,[2024](https://arxiv.org/html/2607.24873#bib.bib18)\)\. Symbolic models like MelodyGLM\(Wuet al\.,[2023](https://arxiv.org/html/2607.24873#bib.bib8)\)and MeloTrans\(Wanget al\.,[2024](https://arxiv.org/html/2607.24873#bib.bib6)\)offer more transparency but use single\-pass pipelines, while SongComposer\(Dinget al\.,[2025](https://arxiv.org/html/2607.24873#bib.bib19)\)advances lyric\-aware generation, yet all remain largely data\-driven and end\-to\-end\.

Some alternatives to this end\-to\-end method would be to ask for a complete regeneration of the entire song but specify a slight change to only a portion of the song; however, that would likely result in a song with little to no similarity compared to the original, and due to a focus on expressing high\-level intent, many music generation models currently do not have the ability to exactly pinpoint the area to make slight modifications to anyway\.

Under such limitations, artists are limited to either many multi\-turn regenerations using conversation as a medium\(Zhaoet al\.,[2025](https://arxiv.org/html/2607.24873#bib.bib11); Labanet al\.,[2025](https://arxiv.org/html/2607.24873#bib.bib2); Zhanget al\.,[2024](https://arxiv.org/html/2607.24873#bib.bib10); Yiet al\.,[2025](https://arxiv.org/html/2607.24873#bib.bib1)\)until they reach a somewhat coherent, satisfying result or editing the music in another software themselves\. The former does not address the direct problem, and the latter requires a higher music production experience, confining many aspiring musicians\.

Therefore, we introduce MusiChat, avibe composingsymbolic natural\-language\-to\-music generation tool that provides reliable melodic, rhythmic, and lyrical in\-turn refinement preserving all other parts of the song while producing musically coherent, intent\-aware segments and retaining the musical structure\. Similar to the termvibe coding111https://x\.com/karpathy/status/1886192184808149383, the idea that a person can build software by conversing with an LLM rather than writing code directly, we introduce the termvibe composingfor music, where any melodic, rhythmic, and lyrical modifications to the music at any scope can be made purely using natural language\. Contrary to current music generation methods, we incorporate a mix of hybrid symbolic music generation models to generate the desired music\. By particularly including a deterministic design, in\-turn refinement becomes feasible\.

In our demonstration, users not only have the ability to make basic modifications to features such as the author and title, but also to selected measures and lyrical segments through specific regeneration requests communicated with natural language\. In addition to the listed features, users can alter the song style while preserving the song skeleton, hence allowing general refinement as well\. With MusiChat, users are able to prompt, create, and refine musically, rhythmically, and lyrically without needing extensive musical experience, extending beyond the previous limitation placed upon precise musical and rhythmic modification requests for existing music generation models\.

Our contributions are multifold:

- •*MusiChat: A conversational vibe composing system*and an LLM\-powered system that illustrates how conversational AI can facilitate collaborative human–AI music creation through continuous dialogue and iterative refinement\.
- •*A hierarchical controllable music generation framework*that forms the foundation of vibe composition by separating lyric\-aligned musical structure generation from expressive surface realization, enabling flexible, structure\-preserving musical transformations\.
- •*A comprehensive system evaluation*that evaluates interaction, musical structures, expressive variation, and perceptual quality through objective analysis and human studies\.

We demonstrate MusiChat through representative composition scenarios, showing how conversational interaction, persistent memory, and incremental editing support a natural and iterative workflow for vibe composing with AI\-assisted music generation\.

## 2System Architecture

### 2\.1System Overview

MusiChat is a two\-panel web\-based workspace \(Figure[1](https://arxiv.org/html/2607.24873#S1.F1)\) designed for vibe composing through conversational music creation and iterative refinement\. The left panel presents the generated composition through engraved sheet music, a MIDI player with instrument and chord controls, and a collapsible piano roll for inspecting musical structure\. The right panel provides the conversational interface for creating, modifying, and refining compositions through natural\-language requests, exploring musical ideas, and applying stylistic transformations\. A set of buttons on the far right side provides access to additional functions, including sign\-in, composition history, input examples, and other utilities that support managing and revisiting creative work\.

Figure[2](https://arxiv.org/html/2607.24873#S1.F2)illustrates the end\-to\-end system architecture\. User requests are sent through an interface to a serverless backend, which performs intent routing, invokes foundation models for tasks such as lyric generation, and calls an in\-process neuro\-symbolic hybrid music engine for composition generation and transformation\. Generated MusicXML and MIDI artifacts are then delivered to the client for notation rendering and audio playback\.

The system maintains two complementary forms of state to support iterative vibe composing\.*Short\-term composition state*captures the active piece across conversational turns, while*persistent composition history*stores previous compositions, prompts, and user ratings for authenticated users in a document database, enabling users to revisit and continue prior creative explorations\.

![Refer to caption](https://arxiv.org/html/2607.24873v1/x2.png)Figure 3:*Left:*mean opinion score with 95% Confidence Intervals \(CIs\) per realizer\.*Right:*percentile\-bootstrap 95% CIs for every pairwise per\-participant Mean Opinion Score \(MOS\) difference\.
### 2\.2Natural Language Interaction

The conversational loop forms the core interaction paradigm of vibe composing\. Users create and refine compositions through natural\-language instructions or image inputs, which are interpreted and mapped to structured musical operations\. Requests are categorized into generation, parameter modification \(e\.g\., key, tempo, and time signature\), symbolic editing \(e\.g\., title or pitch changes\), style transformation, musical analysis, or general conversation\. This intent\-driven workflow enables iterative refinement while maintaining user control over both structural and expressive aspects of the composition\.

##### Hybrid intent routing\.

MusiChat employs a hybrid intent\-routing architecture that combines deterministic rules with LLM\-based reasoning\. Lightweight*deterministic detectors*using regular expressions and keyword matching handle frequent, unambiguous edits \(e\.g\. “change to C major” or “make it faster”\) without LLM calls\. For open\-ended or ambiguous requests, an*LLM\-based router*classifies intent and delegates tasks to the appropriate agent, including music generation, key, time signature, tempo, analysis, or general conversation\. This hybrid design reduces latency and cost for structured edits while retaining the flexibility needed for natural\-language vibe composing\.

##### Multimodal lyric generation\.

The system supports multimodal lyric generation by allowing users to initiate composition from either textual or visual inputs\. When provided with text, an LLM generates short lyrics that serve as the semantic foundation for music generation\. When users provide an*image*, a multimodal model interprets the visual content and generates image\-inspired lyrics, supporting visual experiences such as photographs to be transformed into musical compositions\. The resulting lyrics are hierarchically aligned with the generated melody at the levels of sections, phrases, words, and syllables, providing the foundation for subsequent lyric\-aware composition and editing\.

##### Persistent composition state\.

Edits operate on the*current*composition rather than as independent generation requests\. The backend maintains the working song state, while the client re\-sends relevant context, including the previous prompt, composition metadata, and pending suggestions across interactions\. This design ensures iterative refinement despite stateless serverless execution, allowing each conversational turn to transform the same musical artifact\. A “regenerate” request rebuilds the composition from the initial generation process, whereas an in\-place edit updates the existing piece and replaces the current composition card rather than creating a new version\.

### 2\.3Hybrid Symbolic Music Engine

Music generation is performed by a self\-contained Python engine integrated into the backend, without relying on an external model server\. The engine adopts a*hierarchical representation*that forms the foundation of vibe composition\. It first derives a stable*musical skeleton*from the input lyrics and then applies a pluggable*surface realizer*to produce style\-specific musical realizations while keeping the underlying skeleton unchanged\.

##### Lyric\-to\-music skeleton generation\.

Lyrics provide the semantic and prosodic foundation for melody generation\. The engine analyzes linguistic*stress*patterns to infer an appropriate time signature and hierarchically segments lyrics into sections, phrases, words, and syllablesLiaoet al\.\([2023](https://arxiv.org/html/2607.24873#bib.bib129)\)\. Melodic notes are assigned to syllables, with rests inserted at phrase boundaries to reinforce musical structure\. Pitch selection follows a buffer\-constrained, predominantly stepwise traversal across measures, with smoothing at phrase transitions to promote melodic continuity\. The resulting melody, rhythm, phrase structure, and lyric alignment form the compositional backbone preserved throughout subsequent editing operationsLiaoet al\.\([2024](https://arxiv.org/html/2607.24873#bib.bib75),[2025](https://arxiv.org/html/2607.24873#bib.bib65),[2022](https://arxiv.org/html/2607.24873#bib.bib130)\)\.

##### Surface realization \(three realizers\)\.

A post\-generation styling subsystem transforms the neutral compositional scaffold into style\-specific renditions through a fixed pipeline of rhythm, harmony, melody, dynamics, and phrasing transformations\. Parameterized by specifications for eight genres \(*Classical, Pop, Jazz, Rock, Waltz, Lo\-fi, Cinematic*, and*R&B*\), the subsystem decouples style realization from the underlying skelton, supporting flexible vibe composing while preserving melodic continuity, rhythmic structure, and lyric alignment\.

The framework provides three interchangeable realizers over the same scaffold: \(i\) a*deterministic*realizer applying rule\-based transformations; \(ii\) an*LLM\-refined*realizer enriching the deterministic output with ornaments, chromatic embellishments, and articulation; and \(iii\) an*LLM\-driven style\-replacement*realizer generating surface\-level style realization through the language model\. All realizers preserve the melodic scaffold and note onsets while modifying surface attributes such as pitch realization, harmony, articulation, dynamics, and expression\. Each styling operation starts from a fresh copy of the base composition to prevent accumulated transformations\.

##### Chord harmonization\.

During generation, the system assigns chord symbols to the notation and produces a harmonized melody with chord accompaniment\. Harmony is modeled as an independent layer over the musical skeleton, enabling flexible vibe composing while preserving the melody\. Users can control chord realization through interface controls or natural\-language commands, with the*Chords*toggle managing playback and chord\-symbol visibility\.

##### Multi\-instrument playback and robustness\.

The playback system supports ten instruments \(e\.g\.,*Grand Piano, Electric Piano, Guitar, Flute, Violin, Saxophone, Cello, Clarinet,*and*Drums*\) through client\-side soundfont rendering, enabling diverse sonic exploration during vibe composing\. To maintain a robust creative workflow, the generation pipeline uses graceful degradation through a fallback hierarchy: styled composition with chords, chords\-only output, and finally, melody\-only output\. This ensures that transformation failures reduce expressive richness rather than preventing playable results, preserving iterative music creation\.

## 3Demonstration and Experiments

We conducted a mix of quantitative feature editing and human evaluation to evaluate this paradigm mainly on its consistency, ability to retain the correct features over multiple turns, the naturalness or melodic consistency of the generated music, and the overall musical quality\. The former two qualities are evaluated through feature editing, and the latter two are evaluated through our conducted human study\. More detailed evaluation for objective, music\-theory\-based metrics have been conducted inLiaoet al\.\([2025](https://arxiv.org/html/2607.24873#bib.bib65)\)\.

##### Experimental Setup\.

To evaluate the proposed MusiChat paradigm, we developed a prototype web tool that facilitates both image\-to\-lyrics and text\-to\-lyrics generation\. The MusiChat prototype is deployed both locally and on AWS, and uses the Amazon Nova 2 multimodal AI model for LLM\-based lyric generation from images and text description as well as for general agentic reasoning\.

Table 1:Interactive Feature Editing Performance\. \(a\) Single\-Turn\. \(b\) Multi\-Turn\.\(a\) Single\-Turn

FeatureHLFeatureHLTitle80Title\+KS80KS80KS\+Tempo80Tempo80KS\+Tempo\+Pitch71Pitch62Tempo\+Pitch\+Title80TOTAL302311Accuracy: 95\.31%
\(b\) Multi\-Turn

FeatureHLFeatureHLTitle,KS80KS\+Pitch,Tempo50KS,Tempo80KS,Tempo\+Pitch50KS,Tempo,Pitch50Pitch,KS50TOTAL260100Accuracy: 100%
Overall Accuracy: 97% \(H=97, L=3\)

H: Hit, L: Loss, KS: Key Signature\.

### 3\.1Feature Editing Evaluation

We evaluated feature editing performance across single\-turn and multi\-turn settings, each further divided into single\- and multi\-feature edits, where each turn contains one or more editing requests\. In total, we collected 64 single\-turn and 36 multi\-turn edits involving musical attributes such as pitch, tempo, and key signature, either individually or in combination\. For instance, the "KS\+Pitch, Tempo” category featured two turns: one turn for a key signature and pitch change, and a separate turn for a tempo change\. For each feature, the number of hits and losses was recorded, with hits defined as having all correct editing and losses defined as having one or more incorrect changes\. Detailed results are reported in Table[1](https://arxiv.org/html/2607.24873#S3.T1)\. Across these settings, the system consistently demonstrates strong performance\. In the single\-turn case, it achieves an accuracy of 95\.31%, with minor errors primarily in pitch\-related configurations\. In the multi\-turn setting, it attains 100% accuracy, indicating that iterative editing does not introduce error accumulation and remains stable under sequential or compositional edits\. Overall, the system achieves 97% accuracy, highlighting its robustness and scalability across both single\- and multi\-feature editing scenarios\.

### 3\.2Human Evaluation

#### 3\.2\.1Study Design

The survey contained 35 participants with self\-reported musical experience: 18 reported some musical training, 9 no formal training, and 8 were professional musicians or had extensive musical experience\. Participants were not informed about how the excerpts were produced to avoid algorithm\-aversion bias\.

Each participant was asked to evaluate the musical coherence and quality of each short excerpt from 1 to 5, with 1 being incoherent or low quality and 5 being very coherent and high quality\. In each section, participants listened to three melodies presented in randomized order from the same set of lyrics\. Each melody was generated with a randomized musical genre, and it was a distinct realization from the three categories: deterministic, LLM\-refined, and LLM\-replaced\. The specific information regarding each melody was not disclosed to the participants, each melody being labeled as A, B, and C to maintain objectivity\.

#### 3\.2\.2Results

From the left figure of Figure[3](https://arxiv.org/html/2607.24873#S2.F3), the Mean Opinion Scores \(MOS\) with 95% Confidence Intervals \(CI\) for all three realizers are around average, but slightly on the higher end\. The highest mean opinion score among the three realizers is the deterministic realizer\. Looking at the figure on the right, the pairwise differences between the realizers are all around zero, so there is no significant difference between the three realizers for their mean opinion scores; however, the deterministic realizer overall has higher pairwise differences compared to the rest of the combinations\.

Figure[4](https://arxiv.org/html/2607.24873#Ax1.F4)sorts the naturalness ratings by listener experience and the realizer\. From a distribution perspective, the group with zero musical experience and the group with extensive musical experience had distributions that skewed left, whereas the group with some musical experience did not particularly concentrate at any point\. Notably, the mean is the highest for the group with zero musical experience and the lowest for the group with some musical experience, and the mean for all three groups concentrated in the region between 3 and 4, indicating that at best, the melodies were somewhat natural, and at worst, the listeners felt indifferent to the naturalness\. Based on this figure and values from Table[2](https://arxiv.org/html/2607.24873#Ax1.T2), within each group, the deterministic realizer \(3\.26\) had the highest average ratings for listeners with some experience; LLM\-refined \(3\.90\) had the highest average for listeners with no experience; and LLM\-replace \(3\.46\) had the highest average for listeners with extensive experience\.

Figure[5](https://arxiv.org/html/2607.24873#Ax1.F5)displays a pooling of ratings across all realizers, songs, and participants, with the only categorization being based on the survey questions\. For naturalness ratings, approximately 50% of the ratings were 4 or higher compared to the 25% that were 2 or lower, giving an overall 2:1 ratio on naturalness\. Similarly, the overall quality ratings observe a 55% and 18% like and dislike rate, giving an approximately 3:1 ratio on the category\. With both questions heavily favoring the ratings of 4 or higher, it indicates that the listeners generally liked the generated music\.

### 3\.3Conversational Evaluation

Figure[6](https://arxiv.org/html/2607.24873#Ax1.F6)evaluates the consistency and the preservation of multiple features across multiple instances of regenerating and editing songs\. On the top left, the figure demonstrates that composition consistency is 100%, while intent fulfillment is the same with the exception of mood consistency\. On the top right, the melody preservation for the MusiChat editor is at minimum more than two times the melody fidelity of a basic regenerating baseline, and at most more than three times the baseline\. At the bottom left, while both methods achieved full lyric alignment, the MusiChat editor achieved 2\-4 times better preservation than the baseline\. At the bottom right, the figure reveals that despite having 100% preservation for aspects that did not require change, the baseline tended to change the melody over time and experienced error accumulation, whereas MusiChat largely preserved it compared to the baseline\.

Overall, results demonstrate coherent, perceptually realistic, and structurally consistent music generation under interactive use\. As a result, the autonomous generation can provide immediate inspiration and support iterative refinement while preserving flexibility in aligning with user intent\. Additionally, the combined approach of being unified, language\-driven, and multimodal helps improve coherence and creative control\.

## 4Conclusions

We present MusiChat, a conversational vibe composing text\-to\-music generation system that utilizes a hybrid architecture of deterministic and music generation models\. Due to this architecture, MusiChat supports in\-turn iterative refinement for each desired revision that still semantically, melodically, and rhythmically preserves the untouched features, positioning natural language as the main driver for co\-creation in creative AI\. The workflow that MusiChat supports is what we would callvibe composing\. Our 35\-person study on the naturalness and overall quality of the melody showed that the ratio of like to dislike was 2:1 for naturalness and 3:1 for overall quality, and MusiChat achieved 95\.31% and 100% accuracies in single\-turn and multi\-turn refinements, respectively\. By supporting a full vibe composing workflow, MusiChat can further foster human creativity by expanding the music creative process to users with little to no musical experience\. In future work, we plan to scale the extent of music generation to multi\-instrument or polyphonic compositions and expand to more modalities\.

## Broader Impact

This work advances interpretable and collaborative creative AI by reframing music generation as a language\-centric reasoning process rather than opaque synthesis, enabling inspection, editing, and human oversight that are difficult to achieve in end\-to\-end neural systems\. The approach supports responsible creative AI by reducing reliance on black\-box generative models for musical output while maintaining real\-time interactivity\.

Potential risks include over\-reliance on AI\-generated lyrical interpretations, which may narrow creative exploration\. We mitigate these risks by exposing all intermediate representations, supporting conversational revision, and allowing users to override or refine both lyrical interpretations and algorithmic composition decisions\. Overall, the framework demonstrates how language\-based reasoning combined with algorithmic realization enables transparent, scalable, and human\-centered creative systems\.

## Limitations

Our framework centers natural language as an explicit reasoning representation, using lyrics and conversation to guide music generation\. While this improves interpretability and control, it depends on the clarity and expressiveness of linguistic input, as ambiguous or underspecified lyrics may lead to under\-constrained musical structure\. Moreover, the multimodal\-to\-lyrics stage relies on LLMs to translate multimodal inputs into descriptive lyrical representations\. Errors, biases, or omissions in multimodal–semantic interpretation may propagate into downstream reasoning\. Furthermore, the melody generation in the lyrics\-to\-music pipeline is implemented using a purely algorithmic composition engine, ensuring determinism, interpretability, and real\-time performance, but it may result in less flexibility for certain musical genres\. Additionally, while the symbolic stage is computationally lightweight, upstream LLM\-based reasoning may still reflect model biases\.

## Ethics Statement

The lyrics\-to\-melody algorithm powering the music generation does not contain any copyright concerns since it is purely algorithmic\. However, the chord generation, LLM\-Replace and LLM\-Refine and multimodal feature extraction do contain the usage of LLMs\. All photo examples presented in this paper and incorporated into the MusiChat web tool were taken by the authors and are free from copyright restrictions\.

## References

- A\. Agostinelli, T\. I\. Denk, Z\. Borsos, J\. Engel, M\. Verzetti, A\. Caillon, Q\. Huang, A\. Jansen, A\. Roberts, M\. Tagliasacchi, M\. Sharifi, N\. Zeghidour, and C\. Frank \(2023\)MusicLM: generating music from text\.External Links:2301\.11325,[Link](https://arxiv.org/abs/2301.11325)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- A\. Caillon, B\. McWilliams, C\. Tarakajian, I\. Simon, I\. Manco, J\. Engel, N\. Constant, Y\. Li, T\. I\. Denk, A\. Lalama, A\. Agostinelli, C\. A\. Huang, E\. Manilow, G\. Brower, H\. Erdogan, H\. Lei, I\. Rolnick, I\. Grishchenko, M\. Orsini, M\. Kastelic, M\. Zuluaga, M\. Verzetti, M\. Dooley, O\. Skopek, R\. Ferrer, S\. Petridis, Z\. Borsos, Ä\. van den Oord, D\. Eck, E\. Collins, J\. Baldridge, T\. Hume, C\. Donahue, K\. Han, and A\. Roberts \(2025\)Live music models\.External Links:2508\.04651,[Link](https://arxiv.org/abs/2508.04651)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p1.1)\.
- Simple and controllable music generation\.External Links:2306\.05284,[Link](https://arxiv.org/abs/2306.05284)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p1.1)\.
- P\. Dhariwal, H\. Jun, C\. Payne, J\. W\. Kim, A\. Radford, and I\. Sutskever \(2020\)Jukebox: a generative model for music\.External Links:2005\.00341,[Link](https://arxiv.org/abs/2005.00341)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- S\. Ding, Z\. Liu, X\. Dong, P\. Zhang, R\. Qian, J\. Huang, C\. He, D\. Lin, and J\. Wang \(2025\)SongComposer: a large language model for lyric and melody generation in song composition\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 7108–7127\.External Links:[Link](https://aclanthology.org/2025.acl-long.352/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.352),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- Z\. Evans, J\. D\. Parker, M\. Rice, C\. Carr, Z\. Zukowski, J\. Taylor, and J\. Pons \(2026\)Stable audio 3\.External Links:2605\.17991,[Link](https://arxiv.org/abs/2605.17991)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p1.1)\.
- C\. A\. Huang, A\. Vaswani, J\. Uszkoreit, N\. Shazeer, I\. Simon, C\. Hawthorne, A\. M\. Dai, M\. D\. Hoffman, M\. Dinculescu, and D\. Eck \(2018\)Music transformer\.External Links:1809\.04281,[Link](https://arxiv.org/abs/1809.04281)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- P\. Laban, H\. Hayashi, Y\. Zhou, and J\. Neville \(2025\)LLMs get lost in multi\-turn conversation\.External Links:2505\.06120,[Link](https://arxiv.org/abs/2505.06120)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p4.1)\.
- C\. C\. Liao, D\. Liao, and J\. Guessford \(2022\)Multimodal lyrics\-rhythm matching\.In2022 IEEE International Conference on Big Data \(Big Data\),Vol\.,pp\. 3622–3630\.External Links:[Document](https://dx.doi.org/10.1109/BigData55660.2022.10021009)Cited by:[§2\.3](https://arxiv.org/html/2607.24873#S2.SS3.SSS0.Px1.p1.1)\.
- C\. C\. Liao, D\. Liao, and J\. Guessford \(2023\)Automatic time signature determination for new scores using lyrics for latent rhythmic structure\.In2023 IEEE International Conference on Big Data \(BigData\),pp\. 4485–4494\.External Links:[Link](http://dx.doi.org/10.1109/BigData59044.2023.10386875),[Document](https://dx.doi.org/10.1109/bigdata59044.2023.10386875)Cited by:[§2\.3](https://arxiv.org/html/2607.24873#S2.SS3.SSS0.Px1.p1.1)\.
- C\. C\. Liao, D\. Liao, and E\. L\. Zhang \(2024\)Relationships between Keywords and Strong Beats in Lyrical Music\.In2024 IEEE International Conference on Big Data \(BigData\),Vol\.,Los Alamitos, CA, USA,pp\. 3191–3199\.External Links:ISSN,[Document](https://dx.doi.org/10.1109/BigData62323.2024.10825973),[Link](https://doi.ieeecomputersociety.org/10.1109/BigData62323.2024.10825973)Cited by:[§2\.3](https://arxiv.org/html/2607.24873#S2.SS3.SSS0.Px1.p1.1)\.
- C\. C\. Liao, D\. Liao, and E\. L\. Zhang \(2025\)MusicAIR: a multimodal ai music generation framework powered by an algorithm\-driven core\.External Links:2511\.17323,[Link](https://arxiv.org/abs/2511.17323)Cited by:[§2\.3](https://arxiv.org/html/2607.24873#S2.SS3.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2607.24873#S3.p1.1)\.
- C\. Payne \(2019\)MuseNet\.OpenAI Blog\.Note:https://openai\.com/research/musenetCited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- Suno \(2026\)Suno\.Note:Accessed: July 11, 2026External Links:[Link](https://suno.com/)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p1.1)\.
- Y\. Wang, W\. Yang, Z\. Dai, Y\. Zhang, K\. Zhao, and H\. Wang \(2024\)MeloTrans: a text to symbolic music generation model following human composition habit\.External Links:2410\.13419,[Link](https://arxiv.org/abs/2410.13419)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- X\. Wu, Z\. Huang, K\. Zhang, J\. Yu, X\. Tan, T\. Zhang, Z\. Wang, and L\. Sun \(2023\)MelodyGLM: multi\-task pre\-training for symbolic melody generation\.External Links:2309\.10738,[Link](https://arxiv.org/abs/2309.10738)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- Z\. Yi, J\. Ouyang, Z\. Xu, Y\. Liu, T\. Liao, H\. Luo, and Y\. Shen \(2025\)A survey on recent advances in llm\-based multi\-turn dialogue systems\.ACM Comput\. Surv\.58\(6\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3771090),[Document](https://dx.doi.org/10.1145/3771090)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p4.1)\.
- J\. Yu, X\. Wu, Y\. Xu, T\. Zhang, S\. Wu, L\. Ma, and K\. Zhang \(2025\)SongGLM: lyric\-to\-melody generation with 2d alignment encoding and multi\-task pre\-training\.InProceedings of the Thirty\-Ninth AAAI Conference on Artificial Intelligence and Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’25/IAAI’25/EAAI’25\.External Links:ISBN 978\-1\-57735\-897\-8,[Link](https://doi.org/10.1609/aaai.v39i24.34766),[Document](https://dx.doi.org/10.1609/aaai.v39i24.34766)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.
- Z\. Zhang, A\. Zhang, M\. Li, H\. Zhao, G\. Karypis, and A\. Smola \(2024\)Multimodal chain\-of\-thought reasoning in language models\.External Links:2302\.00923,[Link](https://arxiv.org/abs/2302.00923)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p4.1)\.
- Y\. Zhao, M\. Yang, Y\. Lin, X\. Zhang, F\. Shi, Z\. Wang, J\. Ding, and H\. Ning \(2025\)AI\-enabled text\-to\-music generation: a comprehensive review of methods, frameworks, and future directions\.Electronics14\(6\)\.External Links:[Link](https://www.mdpi.com/2079-9292/14/6/1197),ISSN 2079\-9292,[Document](https://dx.doi.org/10.3390/electronics14061197)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p4.1)\.
- Z\. Zhou, Y\. Wu, Z\. Wu, X\. Zhang, R\. Yuan, Y\. Ma, L\. Wang, E\. Benetos, W\. Xue, and Y\. Guo \(2024\)Can llms "reason" in music? an evaluation of llms’ capability of music understanding and generation\.External Links:2407\.21531,[Link](https://arxiv.org/abs/2407.21531)Cited by:[§1](https://arxiv.org/html/2607.24873#S1.p2.1)\.

## Appendix

### Demo Video and Live System

- •Explore[the AWS\-based web tool](https://musichat.intellisky.ai/)to interact with MusiChat to generate music online, analyze and edit generated sheet music, and play the music in the conversational interface in real time\.
- •Watch the[MusiChat demo video](https://www.intellisky.org/share/musichat_demo.mp4)featuring interactive demonstrations of human\-AI music co\-creation, including simple or complex queries, editing, analysis, playback, etc\.

Naturalness MOSDet\. \(P1\)Refine \(P2b\)Replace \(P2a\)No training \(n=9\)3\.863\.903\.83Some training \(n=18\)3\.263\.173\.20Professional \(n=8\)3\.433\.223\.46Table 2:Naturalness MOS by musical experience\. Non\-experts \(target users\) give the highest ratings, while realizers remain comparable across groups\.![Refer to caption](https://arxiv.org/html/2607.24873v1/x3.png)Figure 4:Raincloud of naturalness ratings by listener experience×\\timesrealizer\.![Refer to caption](https://arxiv.org/html/2607.24873v1/x4.png)Figure 5:Pooling of ratings across all realizers, songs, and participants\.![Refer to caption](https://arxiv.org/html/2607.24873v1/x5.png)Figure 6:Conversational evaluation\.\(A\)The two headline metrics with Wilson 95% confidence intervals \(CIs\)\.\(B\)MusiChat melody preservation vs\. the regenerate\-from\-scratch baseline, per edit type\.\(C\)Preservation by feature\.\(D\)MusiChat and baseline fidelities versus the rate of contract preserved\.

Similar Articles

Composer

Product Hunt

Composer is a multiplayer markdown editor designed for team collaboration and AI agents.

HumToBeats

Product Hunt

HumToBeats turns humming into AI-generated beats, enabling users to create music from simple vocal input.

VibeAround

Product Hunt

VibeAround is a local AI coding agent that users can chat with from any instant messaging app or web browser.

Mistral AI to get Code and Apps features on Vibe (2 minute read)

TLDR AI

Mistral AI is adding dedicated Code and Apps sections to its Vibe (Le Chat) web platform, turning it from a conversational interface into a development and app-building environment. A new large, sparse mixture-of-experts model is also confirmed for summer release as open weights.

Spotify expands its AI push with a ChatGPT-like music assistant

TechCrunch AI

Spotify has introduced a beta feature for Premium users that enables interactive conversations with the app to choose music, using a mix of its own AI and models from multiple providers, initially available in the U.S., Ireland, and Sweden.