- Main
Diffusion Models for Musician-Centered Music AI: Generation, Separation, Interpretation, and Performance
- Karchkhadze, Tornike
- Advisor(s): Dubnov, Shlomo
Abstract
Recent breakthroughs in generative AI have produced a new generation of powerful music synthesis models; the dominant paradigm to emerge is text-to-music generation, in which complete musical pieces are synthesized from natural language descriptions. While technically impressive, this paradigm does not always reflect how real-life musicians actually work. Music-making rarely begins with a sentence — it is a practice that spans orchestrating individual parts, separating and analyzing musical textures, interpreting notation, accompanying other performers, and ultimately engaging in real-time musical interaction. This dissertation presents four works that collectively explore these dimensions of music-making through the lens of diffusion models. The works presented in this dissertation trace a coherent arc through the musical process as a practitioner may experiences it. The first addresses the structural understanding and creation of music — how individual tracks relate and interact — and demonstrates that any track or subset of tracks can be simultaneously generated and separated within a shared musical context. The second introduces a diffusion-based refinement on top of established state-of-the-art music source separation methods, making them more powerful and, crucially, fast enough for practical use through consistency distillation — a consideration of fundamental importance from the musician’s perspective. The third turns to interpretation, exploring how AI can engage with non-standard notation and give voice to visual scores that resist conventional musical language. The fourth brings these capabilities into live performance, where the efficiency gained through consistency distillation is no longer a convenience — it is the condition for real-time musical interaction. What began as a question of musical structure ends as a question of musical presence: can a generative model become a genuine co-performer, listening and responding with accompaniment in the moment? Each chapter answers its own specific question, and together they propose diffusion models not only as a tool for straightforward music generation, but as a foundation for AI that participates in music-making the way musicians do.