🤖 AI Summary
This study addresses the limitation of existing music generation systems, which predominantly produce one-shot outputs and lack the capacity for continuous, incremental editing of symbolic scores. To overcome this, we propose an operation-aware state transition framework that repurposes general-purpose instruction-tuned large language models (LLMs) into reusable symbolic score editing operators. Leveraging Llama 3.1 Instruct with LoRA fine-tuning on a corpus of 490,000 dialogue samples, our approach enables precise incremental composition in ABC notation by explicitly constraining modification and preservation logic during editing. The proposed method significantly improves compliance rates in persistent symbolic editing tasks, demonstrating the technical feasibility of employing LLMs as controllable and reusable tools for music editing.
📝 Abstract
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: not as a one-shot music generator, but as a reusable operator over an evolving symbolic score. We formulate incremental composition as a sequence of operation-aware state transitions over persistent ABC notation, with explicit requirements on what each operation may change and what it must preserve. The interaction includes two artifact-initialization variants and three editing operations -- chord addition, inpainting, and transposition. We instantiate the formulation by adapting Llama 3.1 8B Instruct with Low-Rank Adaptation (LoRA) on 496,038 operation-aware dialogue records derived from Irish traditional music. The comparison with the unadapted model is used to test the feasibility of learning this interaction contract, not to claim novelty for fine-tuning itself. Across 500 dialogues per model (1,750 attempted output states), checker admission rises from 29.37% to 99.37%, while compliance conditional on admission rises from 0.7205 to 0.9798. Strict eligibility for reference-relative musical-feature analysis increases from 14 to 1,548 outputs, and Longest Common Subsequence analysis does not show a systematic increase in high-overlap sequences relative to held-out baselines under the specified protocol. The results support the technical feasibility of persistent, operation-aware symbolic editing with a general-purpose instruction LLM. They do not establish superior musical quality or human-AI co-creativity, which remain questions for musician-centered evaluation.