Score
Designs and implements methods that produce timestamped correspondences between an audio recording and its lyric text and/or melody sequence, mapping words, syllables, or notes to time instants or intervals in the audio. This competence covers detecting onsets and segment boundaries and generating event- or frame-level alignments that link symbolic lyric/melody representations with recorded audio.
This work addresses end-to-end audio-to-time-aligned musical score transcription—simultaneously predicting note pitch, onset/offset timestamps, and precise note durations. To tackle the scarcity of duration annotations, we introduce, for the first time, explicit note duration modeling within an end-to-end note-level transcription framework. Our approach features a duration-aware tokenization scheme and a pseudo-labeling-based data augmentation strategy. The model is trained via multi-objective joint optimization, and we propose novel evaluation metrics that jointly assess temporal precision and duration consistency. Experiments demonstrate state-of-the-art performance across multiple benchmarks. Visual analysis further confirms the method’s high accuracy and robustness in modeling note durations, particularly in challenging polyphonic vocal recordings.
This work addresses the limitations of existing symbolic music tokenization approaches, which typically rely on event sequences with irregular time steps and struggle to explicitly model rhythmic regularities. The authors propose a novel tokenization method that uses fixed temporal units—such as beats—as the fundamental building blocks, merging all events sharing the same pitch within a single time step into one token, thereby yielding a sparse piano-roll-like representation. This approach enables explicit alignment of temporal structure for the first time. Integrated with a Transformer architecture, it significantly enhances generation quality, structural coherence, and long-range dependency modeling in tasks such as music continuation and accompaniment generation, while also achieving superior efficiency and rhythmic consistency compared to prevailing event-based tokenization methods.
Existing song generation methods struggle to model the temporal evolution of musical structure and dynamics, resulting in insufficient fine-grained controllability. To address this, we propose a non-autoregressive song generation framework. Our method introduces three key innovations: (1) a paragraph-level temporal alignment control mechanism that explicitly coordinates lyric paragraphs with musical structure; (2) an LLM-driven duration prediction approach for LRC-formatted lyrics, enabling construction of high-quality temporally aligned training data; and (3) a multi-component prompting strategy integrating local-global text prompts, time-broadcasted conditioning injection, and multi-granularity prompt fusion. Experiments demonstrate significant improvements over baselines in paragraph alignment accuracy, vocal attribute consistency, and musical coherence. The framework enables high-fidelity, structurally controllable song synthesis. Furthermore, we introduce novel evaluation metrics to quantitatively assess controllability and cross-modal consistency.
This paper addresses cross-modal melody-to-lyrics matching (MLM), a task of retrieving singable lyrics given a symbolic melody. We propose a contrastive alignment learning framework that requires no aligned melody–lyrics annotations. Methodologically, we introduce syllable-level *sylphone* encoding to explicitly model phonemic, vocalic, and stress features; integrate self-supervised representation learning with a contrastive loss to achieve semantic alignment between melodies and lyrics in a shared latent space. Unlike prior generative approaches, we are the first to formally formulate MLM as a retrieval problem and learn cross-modal associations without explicit temporal alignment supervision. Experiments demonstrate significant improvements in lyrical singability and semantic matching accuracy. Open-sourced code and visualization examples further validate the method’s effectiveness and practicality in real-world music applications.
Machine-generated lyrics often suffer from poor singability due to misaligned rhythmic phrasing, inaccurate line counts, and inconsistent syllable numbers per line. To address this, we propose a melody-to-lyrics joint generation framework that introduces, for the first time, a format-aware training objective—explicitly modeling musicological constraints (e.g., meter, phrase structure) as fine-grained melody–lyric alignment penalties. Our method employs a two-stage pretraining strategy: initially infusing length and rhythm awareness into large language models using pure lyric corpora, then optimizing with a music-driven format alignment loss. Evaluated on standard benchmarks, our approach achieves absolute improvements of 3.75% and 21.44% in line-count and syllable-per-line accuracy, respectively. In both objective metrics and human evaluations, it outperforms state-of-the-art methods by 63.92% and 74.18% on melody–lyric compatibility and overall quality, significantly narrowing the singability gap between human-composed and AI-generated lyrics.
This study investigates the internal mechanisms by which large audio-language models (LALMs) process the temporal structure of sound events, which remain poorly understood. Employing mechanistic interpretability, representation analysis, and activation steering techniques, we examine how temporal information is bound to sound events within LALMs. Our findings reveal that event positional information is encoded along low-dimensional curved trajectories within intermediate-layer entity name representations. Intervening in these representations alters temporal reasoning beliefs without affecting precise timestamp predictions. This work demonstrates that coarse-grained temporal reasoning and fine-grained localization rely on independent neural pathways, thereby elucidating the internal temporal binding mechanisms of LALMs.
为解决音乐标注数据稀缺问题,TimeCues Studio提供了一个开源工作空间,支持团队进行音乐集合的标注与算法开发、比较及原型设计。
This study addresses the challenge of pattern matching in unpitched, polyphonic music by proposing a point-set geometric framework for musical representation and transformation. Methodologically, it replaces traditional sequential representations with point-set geometric structures, modeling musical operations through affine transformations such as translation and scaling. The work innovatively defines eight pitch-time representations and transformation classes, integrating morphetic and chroma pitch representations with graph-theoretic approaches to construct transformation graphs among patterns. Applied to the analysis of Ravel’s compositions, the framework precisely reveals intricate musical relationships involving Haydn themes. Ultimately, this research provides a rigorous geometric theoretical foundation for cross-work pattern recognition in computational musicology.
为解决大型音频语言模型无法标注时间戳的问题,提出TEMPO方法,通过原子时间戳标记、时间感知投影器和距离感知高斯损失进行监督微调,并使用强化学习进一步优化。
研究通过引入时间音乐定位任务及MusicGroundingBench基准测试,评估音频-文本大模型在音乐理解上的准确性,发现特定任务训练能显著提升性能。