perform manual transcription

Design and produce accurate, convention-compliant written transcripts from speech recordings, including time-alignment, speaker labeling, and capture of consent and metadata; implement and carry out quality-control and verification procedures to ensure consistency and correctness.

performmanualtranscription

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical gap in the evaluation of controllable speech generation systems, which typically focus on target prompt alignment while neglecting the stability of non-target attributes. The study systematically reveals, for the first time, that adjusting a target attribute often induces unintended shifts in other acoustic or speaker characteristics. To investigate this phenomenon, the authors construct an audit dataset comprising 5,940 samples and perform controlled pairwise evaluations using multidimensional metrics spanning acoustics, prosody, content, and speaker traits. They further propose VoDER-Cal, a training-free, inference-time candidate re-ranking method that significantly enhances attribute preservation. In a three-candidate setting, VoDER-Cal improves joint success rate from 4.8% to 14% and reduces non-target deviation from 0.344 to 0.276, demonstrating its effectiveness in mitigating unintended perturbations during targeted voice editing.

attribute controloff-target deviationprompt adherence

This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.

ambiguitycontent creationedit instruction

Automated Quality Control for Language Documentation: Detecting Phonotactic Inconsistencies in a Kokborok Wordlist

Oct 24, 2025
KP
Kellen Parker van Dam
🏛️ University of Passau | Charles University

Lexical data in linguistic documentation frequently contain transcription errors and unannotated loanwords, introducing bias into phonological analysis. This paper addresses these challenges for the low-resource language Kokborok by proposing an unsupervised anomaly detection method that innovatively integrates character-level and syllable-aware phonological features to identify both transcription errors and covert loanwords in lexical inventories. Evaluated on a Kokborok–Bengali multilingual dataset, the method significantly outperforms a character-only baseline, achieving high recall while maintaining practical applicability and systematic rigor. Although precision is constrained by the subtle, linguistically embedded nature of certain anomalies, the method’s strong recall enables field linguists to perform actionable data quality diagnostics. It thus establishes a novel paradigm for automated cleaning and annotation of lexical data in low-resource language documentation.

Detecting phonotactic inconsistencies in Kokborok wordlistsIdentifying transcription errors and undocumented borrowingsImproving data quality through automated anomaly detection

Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?

Nov 13, 2023
CA
Cristina Aggazzotti
🏛️ Johns Hopkins University | Université du Québec à Montréal

This work investigates the feasibility of authorship attribution models for speaker identification in speech transcription texts. Unlike written text, transcriptions lack punctuation and capitalization but contain speech-specific patterns such as fillers and backchannels. To address this, we introduce the first benchmark for speaker attribution in manually transcribed dialogues and propose a topic-controlled verification paradigm to mitigate topic confounding bias. Methodologically, we integrate contextual language models (BERT/RoBERTa) with n-gram and stylometric features, incorporating transcription-style analysis and domain-adaptive fine-tuning on speech-transcribed text. Experiments show that general-purpose models achieve moderate speaker discrimination under relaxed settings, but performance drops substantially under topic control. Crucially, fine-tuning on speech-transcribed corpora significantly improves speaker identification accuracy, demonstrating that speech-style features—e.g., disfluencies and interactional cues—are discriminative and recoverable from transcriptions.

Assessing performance of text attribution models on controlled speech dataDeveloping a benchmark for speaker attribution in conversational transcriptsExploring authorship verification in transcribed speech with unique challenges

Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data

Apr 05, 2024
J(
Jingyu (Jack) Zhang
🏛️ Johns Hopkins University

To address the insufficient verifiability of large language model (LLM) outputs, this paper proposes Quote-Tuning—a paradigm that compels models to verbatim quote original statements from trusted pretraining corpora, enabling zero-cost human verification. Methodologically: (1) we design an efficient membership inference function for millisecond-level quote detection; (2) we formulate a quantitative quoting reward and curate a specialized preference dataset emphasizing faithful citation; (3) we integrate these components into a quoting-aware alignment pipeline via Reinforcement Learning from Human Feedback (RLHF). Experiments demonstrate that Quote-Tuning increases high-quality document quoting frequency by 130% across diverse tasks without degrading response quality, generalizes robustly across domains and model families, and concurrently improves factual accuracy—marking the first work to natively embed verifiability as an intrinsic capability of LLMs.

Aligning language models to quoteEnhancing verifiability of model outputsImproving trustworthiness through verbatim quoting

Latest Papers

What's happening recently
View more

Existing text-to-speech (TTS) models struggle to directly modify reference timbres and achieve segment-level local expressiveness control. To address these limitations, this work proposes EDICT, a framework that pioneers the construction of shared timbre anchors within the codec token space to unify global timbre editing with local expressiveness control. Specifically, the method generates edited acoustic references to anchor target timbres and integrates segment-wise instructions with dynamic KV cache reconstruction techniques, effectively balancing instruction following, speaker consistency, and transition quality. Experimental results demonstrate that EDICT significantly improves timbre editing performance and overall quality on benchmarks such as TimbreEdit-Bench.

expressive speech synthesislocal instruction controltext-to-speech

This study addresses the challenge of manually verifying residual target-speaker speech following the de-identification of psychiatric audio recordings. To this end, we propose a symmetric dual-pipeline automated verification system built upon Gemma and Nemotron series open-source audio and large language models. The framework introduces a novel symmetric dual-pipeline architecture combined with a multi-model view disjunctive ensemble strategy, significantly enhancing detection robustness through model complementarity. Experimental evaluation on a corpus of 48 dialogue recordings demonstrates that the proposed method achieves an overall F1 score of 0.478 and a recall of 0.870, substantially outperforming single-model baselines. These results indicate that the approach offers an efficient, automated quality-control solution for clinical audio privacy protection.

Audio Language ModelsClinical PsychiatryDyadic Dialogue

This study addresses the difficulty of attributing model behaviors to their origins due to the lack of generative provenance in synthetic speech data. We propose a compact provenance contract and auditing protocol, formally establishing provenance as a necessary but insufficient condition for behavioral attribution. Methodologically, we construct synthetic research objects that bind source specifications to content within a Japanese nursing care scenario, implementing audits through immutable manifests, disjoint versioning of scenario seeds, and multimodal asset linkage. Experimentally, we audit 1.55 hours of speech, revealing impediments to precise upstream attribution and establishing a candidate causal graph framework. This work provides a novel paradigm for enhancing the traceability and causal analysis of synthetic data.

behavior attributioncausal graphdata auditing

This study addresses the limited accuracy of existing forced alignment methods under long audio, complex acoustic conditions, and ASR transcription errors by introducing AlignBench, a dedicated evaluation benchmark, and FuseAlign, a Transformer-based model. FuseAlign performs joint audio-visual contextual modeling with millisecond-level boundary refinement. It incorporates an online label correction mechanism via exponential moving average (EMA) snapshots to detect missing words without relying on lexicons or Viterbi decoding. Furthermore, convolutional upsampling and large-scale pseudo-labeled training are employed to enhance robustness. Experimental results demonstrate that FuseAlign significantly outperforms baselines on AlignBench and maintains robust performance in real-world ASR transcription scenarios, validating the critical contribution of each proposed module.

ASR errorsevaluation benchmarkforced alignment

Hot Scholars

SM

Samar M. Magdy

The University of British Columbia
LinguisticsComputational LinguisticsNLP
EB

Emmanouil Benetos

Queen Mary University of London
Machine listeningAudio signal processingMusic information retrievalMachine learning
MP

Madhurananda Pahar

Research Fellow at University of Sheffield
Machine LearningSignal ProcessingDigital HealthSmart Sensors
HC

Heidi Christensen

University of Sheffield
pathological speech processingdisordered speech recognitionassistive technologyhuman computer interaction
AA

Aisha Alansari

Graduate Assistant, Information and Computer Science Department, KFUPM
Machine LearningNatural Language ProcessingDeep LearningLLMs