Score
Designs and implements autoregressive vision–language transcription systems that take images of musical notation and generate sequences of symbolic music tokens and embedded textual content, jointly transcribing musical symbols and text. Builds and evaluates vision-LM OMR models and pipelines that capture local notation details and larger structural relationships by predicting token sequences conditioned on visual context.
This work addresses the challenge of jointly extracting symbolic notation and semantic knowledge from complex, long-form musical scores containing embedded textual annotations. We propose the first neural optical music recognition (OMR) pipeline capable of system-level sequential processing, integrating system-level score segmentation with an autoregressive vision-language model to process staves in reading order. This approach enables unified modeling of local note-level details and global structural context, producing complete symbolic transcriptions that preserve embedded text such as titles and annotations. Experimental results demonstrate that our method achieves new state-of-the-art performance across multiple benchmarks, significantly outperforming existing approaches. Furthermore, we show that high-quality symbolic transcriptions generated by our pipeline substantially enhance large language models’ ability to semantically interpret music documents.
Visual autoregressive modeling faces challenges in scalability, long-range dependency capture, computational efficiency, and geometric/physical consistency—particularly for image, video, 3D, and multimodal generation. Method: We systematically survey ~250 works, unifying pixel-level, token-level, and scale-level representations; integrating sequence modeling, discrete representation learning (e.g., VQ-VAE, DALL·E tokenizer), causal attention, and hierarchical autoregressive decoding; and establishing connections to diffusion models and GANs. Contribution/Results: We propose the first comprehensive taxonomy spanning representation granularities and cross-cutting dimensions (hierarchical, multimodal, task-agnostic), construct an open-source knowledge base, identify core bottlenecks—including 3D structural priors and inference latency—and outline future directions in scalable architectures, efficient sampling, and physics-aware generation.
Existing end-to-end, full-page optical music recognition (OMR) methods rely on multi-stage pipelines and dedicated layout analysis, limiting generalization. This work introduces the first truly end-to-end full-page piano score OMR system, directly mapping raw page images to structured MusicXML representations. Our approach features: (1) the first full-page, end-to-end OMR architecture integrating convolutional feature extraction with an autoregressive Transformer decoder; (2) a curriculum learning–driven progressive synthetic data generation and training paradigm; and (3) zero-shot performance on real piano scores surpassing leading commercial OMR software. Experiments demonstrate state-of-the-art (SOTA) accuracy on both synthetic data and two real-world benchmarks—achieving significant improvements over existing tools under both zero-shot and fine-tuned settings. The proposed method eliminates hand-crafted heuristics and stage-wise dependencies, enabling robust, unified transcription of entire musical pages without intermediate structural assumptions.
Existing large vision-language models (LVLMs) apply autoregressive supervision only to text, limiting their ability to leverage caption-free images, causing visual detail omission, and preventing modeling of purely visual content. Method: We propose Autoregressive Semantic Visual Reconstruction (ASVR), a unified autoregressive framework jointly modeling vision and language by reconstructing discrete semantic tokens—derived from images via a learned tokenizer—rather than raw pixels, enabling fine-grained visual understanding. Contribution/Results: We empirically demonstrate, for the first time, that semantic-level autoregressive visual reconstruction consistently improves VLM performance, whereas pixel-level reconstruction is ineffective or even detrimental. ASVR enables efficient mapping from continuous visual features to discrete semantic tokens. Compatible with mainstream architectures (e.g., LLaVA), it boosts LLaVA-1.5 by +5% on average across 14 multimodal benchmarks, exhibiting robustness across data scales (556K–2M samples) and diverse LLM backbones. Code is publicly available.
Existing symbolic music generation models treat note attributes as a unidirectional sequential dependency, yet empirical evidence reveals no strict temporal or hierarchical constraints among them—attributes are inherently unordered and concurrent. Method: We propose Amadeus—a hybrid framework that decouples sequence and attribute generation. It employs an autoregressive model to generate the note sequence backbone, while a bidirectional discrete diffusion model concurrently models all note attributes. To enhance representation learning, we introduce MLSDES (Music Latent Space Discriminability Enhancement via Contrastive Learning) and CIEM (Attention-Driven Conditional Information Enhancement Module). Contribution/Results: Evaluated on AMD—the largest open-source MIDI dataset—we demonstrate that Amadeus surpasses state-of-the-art methods in both unconditional and text-conditioned generation across multiple metrics. It achieves ≥4× faster inference and supports training-free, fine-grained attribute control.
This study addresses the scarcity of manual annotations and the sequence inconsistencies caused by local predictions in automatic music transcription. To overcome these challenges, we propose a unified framework integrating synthetic data supervision, structured decoding, and autoregressive distillation. Specifically, MIDI-rendered audio is leveraged to generate synthetic supervision signals, alleviating the data bottleneck. We introduce the first task-specific structured decoding approach, which employs dynamic programming to ensure global coherence of the transcribed musical scores. Furthermore, knowledge distillation is utilized to transfer the capabilities of autoregressive models into more efficient architectures, balancing accuracy with inference speed. Experimental results demonstrate that the proposed method surpasses most existing systems across eight benchmarks, significantly improving both transcription accuracy and sequential consistency.
This work addresses the limited performance of existing optical music recognition (OMR) systems on real-world handwritten piano scores, which stems primarily from the scarcity of diverse, realistically annotated training data—most datasets rely on digitally generated notation that fails to capture the visual variability of handwritten manuscripts, while manual annotation remains prohibitively expensive. To tackle this challenge under resource-constrained conditions, the authors propose a domain-adaptive approach that leverages Music Notation Graphs (MuNGs) and the Smashcima synthesis framework to generate photorealistic handwritten scores using out-of-domain symbols, thereby substantially reducing reliance on finely annotated real data. The study establishes the first end-to-end OMR baseline for complex handwritten piano manuscripts and demonstrates through experiments that the proposed method significantly improves recognition accuracy on authentic historical music documents, advancing the practical applicability of OMR in music heritage preservation.
This study systematically compares three major classes of generative models—autoregressive LSTMs with attention, latent-variable models (including recurrent VAEs and VQ-VAEs), and GANs—in the task of generating polyphonic symbolic music in the style of Bach. Evaluated on a unified MIDI benchmark within a common framework, the models are assessed for musical coherence, quality of learned latent representations, and stylistic consistency. Results indicate that autoregressive LSTMs produce the most coherent sequences; VQ-VAEs effectively mitigate posterior collapse through vector quantization, yielding clearer structural outputs; and GANs, while capable of capturing local pitch patterns, suffer from training instability and poor generalization. The findings elucidate the respective strengths and failure modes of each approach, offering empirical guidance for polyphonic music generation.
This work proposes BandTok, a generative two-dimensional tokenizer for Mel-spectrogram representation that addresses the limitations of existing high-fidelity music codecs relying on residual vector quantization. Such approaches suffer from strong sequential dependencies after flattening, which hinder autoregressive modeling and exacerbate error accumulation. In contrast, BandTok assigns discrete tokens to individual Mel frequency bands per time frame using a single shared codebook, yielding a physically interpretable and structurally disentangled time–frequency token grid. Coupled with a 2D Rotary Position Embedding–enhanced autoregressive language model, this method formulates music generation as image-like time–frequency modeling. Evaluated under data-constrained conditions, BandTok significantly outperforms residual codebook–based baselines in both reconstruction fidelity and controllable generation. Code and audio samples are publicly released.
This work proposes a unified autoregressive multimodal large model framework designed to jointly advance image understanding, generation, and editing while aligning with human preferences. Leveraging a discrete semantic visual tokenizer, the approach converts visual content into a unified discrete representation, enabling multi-task modeling under a single next-token prediction paradigm. The framework integrates multi-objective supervised training with reinforcement learning–driven preference optimization, substantially enhancing both generation quality and instruction-following capabilities. Experimental results demonstrate state-of-the-art performance on text-to-image generation and instruction-guided editing tasks—evidenced by improvements in WISE from 0.50 to 0.56 and GEdit-Bench-EN G_O from 5.75 to 6.68—and reveal notable cross-task synergistic gains.