pitch tracking

Estimating fundamental frequency and related prosodic cues and coupling those estimates with rhythmic and lyrical alignment metrics to produce block‑aligned multimodal scores and capture conversational prosody and rhythm.

pitchtracking

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Current spoken dialogue systems lack interpretable prosodic evaluation methods that adapt to speaker characteristics and interaction states. This work addresses this gap by proposing a condition-matched human reference benchmark and a percentile-based evaluation protocol, leveraging over 4,000 hours of dyadic English conversational data to enable fine-grained analysis of acoustic prosodic features such as fundamental frequency, speech rate, and pause duration. By incorporating hierarchical matching and out-of-bounds flagging mechanisms, the method significantly enhances behavioral plausibility and interpretability of prosodic assessments. Validation on held-out human data demonstrates an anomaly flagging rate close to the nominal 10%, outperforming conventional aggregate statistics while clearly indicating the direction of deviation, thereby providing an effective behavioral plausibility check for synthetic speech.

conversational speechprosody evaluationreference-based metrics

Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens

Sep 24, 2025
IR
Ismail Rasim Ulgen
🏛️ Johns Hopkins University | National University of Singapore

Existing TTS evaluation metrics (e.g., WER, F0-RMSE) suffer from limited dimensionality, weak correlation with human perceptual judgments, and heavy reliance on reference audio. To address these limitations, we propose TTScore—the first reference-free, dual-path evaluation framework based on discrete speech tokens. TTScore employs a text-conditioned sequence-to-sequence model to jointly predict two distinct token sequences: content tokens (capturing intelligibility) and prosody tokens (encoding intonation, rhythm, and other prosodic attributes), thereby enabling decoupled, fine-grained, and interpretable modeling of intelligibility and prosody. By eliminating dependence on reference speech, TTScore introduces, for the first time, a collaborative dual-sequence predictor architecture to model intrinsic speech properties. Extensive experiments on three major benchmarks—SOMOS, VoiceMOS, and TTSArena—demonstrate that TTScore significantly outperforms existing metrics, achieving 12–28% improvements in Pearson and Spearman correlations with human subjective ratings.

Developing reference-free evaluation framework using discrete token predictionEvaluating synthesized speech intelligibility and prosody objectivelyOvercoming limitations of existing metrics like WER and F0-RMSE

RhythmTA: A Visual-Aided Interactive System for ESL Rhythm Training via Dubbing Practice

Jul 25, 2025
CC
Chang Chen
🏛️ Hong Kong University of Science and Technology

ESL learners exhibit overreliance on instructor feedback in prosodic training and demonstrate limited autonomy in self-directed practice. To address this, we propose a dubbing-based, interactive visualization system that supports independent prosody perception and pronunciation training through a three-stage workflow: synchronized listening, guided shadowing, and comparative reflection. The system integrates an automated prosodic feature extraction algorithm with a multi-view visual design, and its interaction logic is refined based on pedagogical expert input. A controlled user study demonstrates statistically significant improvement in learners’ prosodic perception (p < 0.01) and robust gains in rhythmic accuracy of pronunciation. These findings provide empirical validation and a scalable technical framework for autonomous ESL speech learning.

Helps ESL learners practice English speech rhythm independentlyProvides visual aids for dubbing-based rhythm trainingReduces reliance on instructor feedback for rhythm improvement

LOAF-M2L: Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation

Jul 05, 2023
LO
Longshen Ou
🏛️ National University of Singapore

Machine-generated lyrics often suffer from poor singability due to misaligned rhythmic phrasing, inaccurate line counts, and inconsistent syllable numbers per line. To address this, we propose a melody-to-lyrics joint generation framework that introduces, for the first time, a format-aware training objective—explicitly modeling musicological constraints (e.g., meter, phrase structure) as fine-grained melody–lyric alignment penalties. Our method employs a two-stage pretraining strategy: initially infusing length and rhythm awareness into large language models using pure lyric corpora, then optimizing with a music-driven format alignment loss. Evaluated on standard benchmarks, our approach achieves absolute improvements of 3.75% and 21.44% in line-count and syllable-per-line accuracy, respectively. In both objective metrics and human evaluations, it outperforms state-of-the-art methods by 63.92% and 74.18% on melody–lyric compatibility and overall quality, significantly narrowing the singability gap between human-composed and AI-generated lyrics.

Improves adherence to line and syllable counts without degrading text qualityJointly learns wording and formatting for singable melody-to-lyric generationReduces singability gap by capturing prosodic and structural patterns

Melody-Guided Music Generation

Sep 30, 2024
SW
Shaopeng Wei
🏛️ Guangxi University | Southwestern University of Finance and Economics

This work addresses melody-guided text-to-music generation by proposing a controllable diffusion framework that jointly leverages implicit semantic alignment and explicit melody retrieval. Methodologically: (1) it introduces Contrastive Language–Music Pretraining (CLMP), the first approach to jointly align textual, audio, and melodic representations; (2) it designs a retrieval-augmented melody-conditioned diffusion mechanism, integrating melody-aligned embeddings with a lightweight text–audio joint encoder to achieve high-fidelity generation within a minimal architectural footprint. Experiments demonstrate that the model outperforms all existing open-source methods on MusicCaps and MusicBench—despite using fewer than one-third the parameters and less than 0.5% of the training data. Human evaluation confirms significant superiority across five dimensions—including realism and melody consistency—validating its dual strengths in content fidelity and harmonic coherence.

Melody GenerationQuality and RelevanceText-to-Music

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing automatic singing quality assessment methods, which often rely on a single modality and struggle to jointly evaluate lyrical accuracy and musical expressiveness. To overcome this, the authors propose MusicJudge, a framework that leverages block-level multimodal alignment to simultaneously assess lyric correctness and pitch–rhythm fidelity. Its key innovations include integrating semantic embeddings, lexical similarity, and phoneme alignment to accurately identify semantically coherent lyric blocks, as well as introducing a modality-guided LoRA fine-tuning strategy to enhance the robustness of automatic speech recognition (ASR) for sung transcription. Experimental results demonstrate that MusicJudge achieves high agreement with human expert ratings across multiple datasets and significantly outperforms current state-of-the-art approaches, confirming its effectiveness and generalization capability.

automatic singing quality assessmentexpressive variationslyric correctness

This work addresses the challenge that large language models often generate musically incoherent melodies from lyrics due to neglecting fundamental musical constraints, resulting in issues such as erratic rhythm and inappropriate pitch ranges. To mitigate this, the authors propose a novel two-stage alignment framework that requires no human annotations: first, preference data are automatically generated using unsupervised music-theoretic rules; then, domain knowledge is injected into the model by integrating Direct Preference Optimization (DPO) with Kahneman-Tversky Optimization (KTO). This approach uniquely unifies rule-based constraints with preference-based learning, significantly enhancing both the musical validity and perceptual quality of generated melodies. Experimental results demonstrate that the aligned model consistently outperforms existing baselines across objective metrics and subjective evaluations, producing outputs that are markedly more musical and coherent.

constraint violationlyric-to-melody generationmusical plausibility

It remains unclear whether current neural text-to-speech (TTS) systems can accurately model fine-grained segmental prosodic phenomena—specifically, consonant-induced fundamental frequency (F0) perturbations—and how well they generalize to low-frequency words. This work proposes a linguistically motivated, segment-level prosody probing framework, training Tacotron 2 and FastSpeech 2 models on the LJ Speech corpus and evaluating their F0 modeling capabilities through large-scale multi-system comparisons and lexical frequency-stratified analyses. Results indicate that while both systems perform well on high-frequency words, they exhibit limited generalization on low-frequency items, suggesting reliance on lexical memorization rather than abstract segment-to-prosody rules. These findings highlight a critical gap in the systematic modeling of fine-grained prosody within current TTS architectures and offer a novel perspective for assessing the naturalness of synthetic speech.

consonant-induced effectsF0 perturbationneural TTS

This study addresses the limitations of existing second-language pronunciation assessment approaches, which predominantly focus on segmental features while underrepresenting suprasegmental dimensions such as rhythm and intonation, and often rely heavily on annotated data, limiting their applicability in low-resource settings. To overcome these challenges, this work proposes a text- and annotation-free multidimensional pronunciation evaluation framework. Leveraging self-supervised WavLM representations and dynamic time warping (DTW), the method quantifies rhythmic proficiency through the warping path distortion of DTW alignments and evaluates intonation by integrating prosodic residuals, fundamental frequency, and intensity features. Experimental results demonstrate that the proposed approach surpasses human inter-rater consistency in phoneme-level scoring, achieves near-human performance in rhythm assessment, and, despite modest effectiveness in intonation evaluation, validates the feasibility of an unsupervised pathway for comprehensive pronunciation assessment.

intonationL2 speech assessmentlow-resource settings

Existing automatic evaluation methods for text-to-music generation systems struggle to optimize ranking metrics and exhibit weak cross-modal consistency. This work proposes DeRA-MOS, a novel framework that decouples listwise ranking from modality alignment objectives for the first time. It employs a batch-aware listwise ranking loss to optimize the ranking performance of musical impressions and integrates a score-anchored modality alignment loss to enhance semantic consistency between text and music. By explicitly addressing pointwise training bias and modality drift, the proposed approach significantly improves Spearman rank correlation on the MusicEval benchmark, establishing a new paradigm for large-scale evaluation of text-to-music generation systems.

evaluationmean opinion scoremodality alignment

Hot Scholars

AR

Alain Riou

PhD student, Sony CSL × Télécom Paris
self-supervised learningmusic
GP

Geoffroy Peeters

Télécom Paris (previously IRCAM - STMS)
audio signal processingmachine learningmusic information retrieval
JB

Jerrin Bright

University of Waterloo
3D Human ModelingComputer VisionAutonomous Navigation
GH

Gaëtan Hadjeres

Research Scientist at Sony AI
Automated Music CompositionGenerative ModelsArtificial IntelligenceMusic
SL

Stefan Lattner

Sony CSL Paris (Music Team)
Deep LearningAudio GenerationAI-assisted Music ProductionMusic Information Retrieval