vision-language omr

Designs and implements autoregressive vision–language transcription systems that take images of musical notation and generate sequences of symbolic music tokens and embedded textual content, jointly transcribing musical symbols and text. Builds and evaluates vision-LM OMR models and pipelines that capture local notation details and larger structural relationships by predicting token sequences conditioned on visual context.

vision-languageomr

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of jointly extracting symbolic notation and semantic knowledge from complex, long-form musical scores containing embedded textual annotations. We propose the first neural optical music recognition (OMR) pipeline capable of system-level sequential processing, integrating system-level score segmentation with an autoregressive vision-language model to process staves in reading order. This approach enables unified modeling of local note-level details and global structural context, producing complete symbolic transcriptions that preserve embedded text such as titles and annotations. Experimental results demonstrate that our method achieves new state-of-the-art performance across multiple benchmarks, significantly outperforming existing approaches. Furthermore, we show that high-quality symbolic transcriptions generated by our pipeline substantially enhance large language models’ ability to semantically interpret music documents.

multimodal recognitionmusic notationoptical music recognition

Autoregressive Models in Vision: A Survey

Nov 08, 2024
JX
Jing Xiong
🏛️ The University of Hong Kong | Tsinghua University | Duke University | University of Rochester | The Ohio State University | Bytedance | The University of North Carolina at Chapel Hill | The Hong Kong Polytechnic University | Apple | Princeton University

Visual autoregressive modeling faces challenges in scalability, long-range dependency capture, computational efficiency, and geometric/physical consistency—particularly for image, video, 3D, and multimodal generation. Method: We systematically survey ~250 works, unifying pixel-level, token-level, and scale-level representations; integrating sequence modeling, discrete representation learning (e.g., VQ-VAE, DALL·E tokenizer), causal attention, and hierarchical autoregressive decoding; and establishing connections to diffusion models and GANs. Contribution/Results: We propose the first comprehensive taxonomy spanning representation granularities and cross-cutting dimensions (hierarchical, multimodal, task-agnostic), construct an open-source knowledge base, identify core bottlenecks—including 3D structural priors and inference latency—and outline future directions in scalable architectures, efficient sampling, and physics-aware generation.

Comparing representation strategies: pixel, token, and scale levelsExploring challenges and future directions for vision autoregressive modelsSurveying autoregressive models in computer vision applications

End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

May 20, 2024
AR
Antonio R'ios-Vila
🏛️ University of Alicante | University of Rouen

Existing end-to-end, full-page optical music recognition (OMR) methods rely on multi-stage pipelines and dedicated layout analysis, limiting generalization. This work introduces the first truly end-to-end full-page piano score OMR system, directly mapping raw page images to structured MusicXML representations. Our approach features: (1) the first full-page, end-to-end OMR architecture integrating convolutional feature extraction with an autoregressive Transformer decoder; (2) a curriculum learning–driven progressive synthetic data generation and training paradigm; and (3) zero-shot performance on real piano scores surpassing leading commercial OMR software. Experiments demonstrate state-of-the-art (SOTA) accuracy on both synthetic data and two real-world benchmarks—achieving significant improvements over existing tools under both zero-shot and fine-tuned settings. The proposed method eliminates hand-crafted heuristics and stage-wise dependencies, enabling robust, unified transcription of entire musical pages without intermediate structural assumptions.

Eliminating multi-stage pipelines for music score recognitionHandling complex layouts without dedicated layout analysisTranscribing full-page pianoform sheet music end-to-end

Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better

Jun 10, 2025
DW
Dianyi Wang
🏛️ Fudan University | Westlake University | Zhejiang University | University of Southern California | Shanghai AI Lab

Existing large vision-language models (LVLMs) apply autoregressive supervision only to text, limiting their ability to leverage caption-free images, causing visual detail omission, and preventing modeling of purely visual content. Method: We propose Autoregressive Semantic Visual Reconstruction (ASVR), a unified autoregressive framework jointly modeling vision and language by reconstructing discrete semantic tokens—derived from images via a learned tokenizer—rather than raw pixels, enabling fine-grained visual understanding. Contribution/Results: We empirically demonstrate, for the first time, that semantic-level autoregressive visual reconstruction consistently improves VLM performance, whereas pixel-level reconstruction is ineffective or even detrimental. ASVR enables efficient mapping from continuous visual features to discrete semantic tokens. Compatible with mainstream architectures (e.g., LLaVA), it boosts LLaVA-1.5 by +5% on average across 14 multimodal benchmarks, exhibiting robustness across data scales (556K–2M samples) and diverse LLM backbones. Code is publicly available.

Captions may omit critical visual detailsLVLMs lack visual modality in autoregressive learningVision-centric content inadequately conveyed through text

Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music

Aug 28, 2025
HS
Hongju Su
🏛️ Beijing University of Posts and Telecommunications | University of Surrey

Existing symbolic music generation models treat note attributes as a unidirectional sequential dependency, yet empirical evidence reveals no strict temporal or hierarchical constraints among them—attributes are inherently unordered and concurrent. Method: We propose Amadeus—a hybrid framework that decouples sequence and attribute generation. It employs an autoregressive model to generate the note sequence backbone, while a bidirectional discrete diffusion model concurrently models all note attributes. To enhance representation learning, we introduce MLSDES (Music Latent Space Discriminability Enhancement via Contrastive Learning) and CIEM (Attention-Driven Conditional Information Enhancement Module). Contribution/Results: Evaluated on AMD—the largest open-source MIDI dataset—we demonstrate that Amadeus surpasses state-of-the-art methods in both unconditional and text-conditioned generation across multiple metrics. It achieves ≥4× faster inference and supports training-free, fine-grained attribute control.

Enabling training-free fine-grained control over musical note attributesEnhancing symbolic music generation performance and generation speedModeling concurrent unordered note attributes instead of fixed sequential dependencies

Latest Papers

What's happening recently
View more

This study addresses the scarcity of manual annotations and the sequence inconsistencies caused by local predictions in automatic music transcription. To overcome these challenges, we propose a unified framework integrating synthetic data supervision, structured decoding, and autoregressive distillation. Specifically, MIDI-rendered audio is leveraged to generate synthetic supervision signals, alleviating the data bottleneck. We introduce the first task-specific structured decoding approach, which employs dynamic programming to ensure global coherence of the transcribed musical scores. Furthermore, knowledge distillation is utilized to transfer the capabilities of autoregressive models into more efficient architectures, balancing accuracy with inference speed. Experimental results demonstrate that the proposed method surpasses most existing systems across eight benchmarks, significantly improving both transcription accuracy and sequential consistency.

data scarcitylead-sheetmusic transcription

This work addresses the limited performance of existing optical music recognition (OMR) systems on real-world handwritten piano scores, which stems primarily from the scarcity of diverse, realistically annotated training data—most datasets rely on digitally generated notation that fails to capture the visual variability of handwritten manuscripts, while manual annotation remains prohibitively expensive. To tackle this challenge under resource-constrained conditions, the authors propose a domain-adaptive approach that leverages Music Notation Graphs (MuNGs) and the Smashcima synthesis framework to generate photorealistic handwritten scores using out-of-domain symbols, thereby substantially reducing reliance on finely annotated real data. The study establishes the first end-to-end OMR baseline for complex handwritten piano manuscripts and demonstrates through experiments that the proposed method significantly improves recognition accuracy on authentic historical music documents, advancing the practical applicability of OMR in music heritage preservation.

domain adaptationmusical cultural heritageOptical Music Recognition

This study systematically compares three major classes of generative models—autoregressive LSTMs with attention, latent-variable models (including recurrent VAEs and VQ-VAEs), and GANs—in the task of generating polyphonic symbolic music in the style of Bach. Evaluated on a unified MIDI benchmark within a common framework, the models are assessed for musical coherence, quality of learned latent representations, and stylistic consistency. Results indicate that autoregressive LSTMs produce the most coherent sequences; VQ-VAEs effectively mitigate posterior collapse through vector quantization, yielding clearer structural outputs; and GANs, while capable of capturing local pitch patterns, suffer from training instability and poor generalization. The findings elucidate the respective strengths and failure modes of each approach, offering empirical guidance for polyphonic music generation.

Bach-style musicgenerative modelingpolyphonic sequences

This work proposes BandTok, a generative two-dimensional tokenizer for Mel-spectrogram representation that addresses the limitations of existing high-fidelity music codecs relying on residual vector quantization. Such approaches suffer from strong sequential dependencies after flattening, which hinder autoregressive modeling and exacerbate error accumulation. In contrast, BandTok assigns discrete tokens to individual Mel frequency bands per time frame using a single shared codebook, yielding a physically interpretable and structurally disentangled time–frequency token grid. Coupled with a 2D Rotary Position Embedding–enhanced autoregressive language model, this method formulates music generation as image-like time–frequency modeling. Evaluated under data-constrained conditions, BandTok significantly outperforms residual codebook–based baselines in both reconstruction fidelity and controllable generation. Code and audio samples are publicly released.

audio tokenizerautoregressive music generationerror accumulation

This work proposes a unified autoregressive multimodal large model framework designed to jointly advance image understanding, generation, and editing while aligning with human preferences. Leveraging a discrete semantic visual tokenizer, the approach converts visual content into a unified discrete representation, enabling multi-task modeling under a single next-token prediction paradigm. The framework integrates multi-objective supervised training with reinforcement learning–driven preference optimization, substantially enhancing both generation quality and instruction-following capabilities. Experimental results demonstrate state-of-the-art performance on text-to-image generation and instruction-guided editing tasks—evidenced by improvements in WISE from 0.50 to 0.56 and GEdit-Bench-EN G_O from 5.75 to 6.68—and reveal notable cross-task synergistic gains.

autoregressive modeldiscrete representationimage editing

Hot Scholars

LC

Lele Cao

Senior Principal AI/ML Researcher and Research Lead, Microsoft (ABK); ACM & IEEE Member
Machine LearningGraph LearningTime Series ModelingFinance
ZS

Zineb Senane

Machine Learning Engineer, Fever Energy
Machine LearningSelf-supervised LearningTime SeriesDiffusion Models
KL

Keqin Li

AMA University
RoboticMachine learningArtificial intelligenceComputer vision
TM

Tao Meng

Central South University of Forestry and Technology
Graph Neural NetworkMultimodal Emotion RecognitionText ClassificationEntity Alignment