Score
Designs annotation schemas, tools, and scalable pipelines to label individual words with fine-grained prosodic and acoustic measurements (e.g., pitch, duration, intensity, stress, boundary markers). Produces and curates word-level prosody datasets and performs quality control and analysis of those labels for use in modeling, synthesis, or phonetic investigation.
This study addresses the challenge of automatic prosodic labeling to support training prosody-controllable text-to-speech (TTS) systems. We propose a multimodal feature fusion framework that jointly integrates representations from speech self-supervised learning (SSL) models, the Whisper encoder, and phoneme-level language models (PnG BERT and PL-BERT)—the first approach to synergistically model acoustic and linguistic features at the phoneme level. Through feature concatenation and end-to-end joint optimization, our method achieves state-of-the-art prosody prediction performance on Japanese: 89.8% accuracy for accent nucleus detection, 93.2% for pitch contour (high/low tone) classification, and 94.3% for phrase boundary (break index) prediction. The approach provides a high-accuracy, scalable, and fully automated solution for prosodic annotation—particularly valuable for low-resource languages—and advances fine-grained prosody modeling for controllable TTS.
This paper addresses the grapheme–phoneme and grapheme–prosody inconsistency problem in speech phoneme and prosody annotation. We propose an end-to-end grapheme-consistency modeling framework that jointly integrates implicit grapheme modeling—via a BERT-based prompt encoder—and explicit grapheme constraints—implemented through grapheme-consistency pruning—to construct speech–annotation–text triple parallel data. To our knowledge, this is the first approach to achieve fully automatic speech annotation with strict grapheme consistency *without* manual alignment. We validate its effectiveness on downstream tasks including text-to-speech (TTS) and accent estimation: the generated parallel data significantly improves accent recognition accuracy. Our work establishes a reliable weakly supervised annotation paradigm for speech representation learning, offering both methodological novelty—through unified implicit/explicit grapheme modeling—and practical utility in low-resource annotation scenarios.
为解决多维度语音标注成本高、依赖外部服务等问题,提出SpeechAnnotator框架,利用开源工具和多智能体协作实现高效自动标注。
This study addresses the challenge of precisely distinguishing turns, feedback, and pauses in spontaneous dialogue by proposing a semi-automatic annotation pipeline. The method integrates voice activity detection, energy filtering, automatic speech recognition, and contextual post-processing to automatically extract turns and feedback while generating consistent initial annotations to facilitate manual review. Experimental results demonstrate that the pipeline achieves an overall F1 score of 0.621 with a boundary error of approximately 0.15 seconds, exhibiting robustness to variations in listening conditions. By standardizing the dialogue annotation workflow, this work significantly enhances the reproducibility of conversational dynamics analysis.
This study addresses the prevalent issue in existing AI speech datasets where disfluent speech—such as that associated with stuttering—is often annotated by crowdworkers lacking lived or clinical experience, leading to inconsistent or distorted labels. To counter this, the authors engaged individuals who stutter and domain experts through semi-structured interviews and co-design workshops, integrating embodied disability experiences into the annotation process. They propose a participatory annotation paradigm centered on “disability-first” and “diversity-aware” principles. The resulting inclusive annotation framework challenges conventional static labeling systems and yields a detailed annotation guideline grounded in the lived realities of people who stutter. This approach not only markedly improves dataset quality but also offers a scalable, inclusive data practice for integrating disability perspectives throughout the AI development lifecycle.
Behavioral profiling (BP) annotation is challenging to automate due to its multidimensional, multilingual nature, and conventional task-level evaluation obscures underlying skill heterogeneity. This work proposes a novel “skill feasibility” paradigm, decomposing BP annotation into 14 operationalizable annotation skills and implementing a schema-guided, skill-document-driven pipeline. Evaluation over a 300-instance validation set—through two rounds of testing involving human annotators and large language models (GPT-5.4 and three open-source models)—reveals a “shared categorization, independent execution” pattern: humans and GPT exhibit high agreement at the skill level but diverge in instance-level execution. The study identifies five directly feasible skills, four recoverable via relabeling, and five structurally undefined. GPT-5.4 demonstrates reliable performance on feasible skills (accuracy = 0.678, κ = 0.665, weighted F1 = 0.695), whereas open-source models primarily fail in translating schemas into executable skills.
Existing evaluation methods for audio captioning struggle to accurately assess the fidelity of multimodal semantics and acoustic attributes in structured audio descriptions. This work proposes the first multi-axis evaluation framework tailored for structured audio captioning, integrating large language model (LLM)-based semantic judgments with deterministic acoustic metrics across five orthogonal dimensions: label sets, descriptive content, logical reasoning, numerical measurements, and spectral contours. The framework incorporates a controlled perturbation protocol to validate its ability to distinguish between semantic preservation and acoustic distortion. Experiments on the AudioCards dataset demonstrate that the proposed approach effectively differentiates semantically consistent paraphrases from genuine errors, significantly outperforming existing methods in both reliability and sensitivity.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
This work addresses the high cost of speaker diarization annotation and the lack of quantifiable evaluation metrics by introducing an open-source annotation tool that guides human annotators through automatically generated initial hypotheses. For the first time, annotation cost—measured in edit operations and time—is treated as a primary output metric. The system features a React-based frontend integrated with the pyannote ecosystem and a stride-accelerated logging engine, supporting automatic initialization, uncertainty-aware highlighting, and a novel “phantom” attention-check mechanism to ensure annotation quality. Experiments on the AMI dataset demonstrate that automatic initialization substantially reduces annotation effort while improving accuracy, with the uncertainty-highlighting strategy yielding the best performance among the evaluated approaches.
This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.