transcription

Designs, implements, or evaluates systems and processes that convert an input signal or representation into an accurate, structured textual or symbolic rendering — for example producing written transcripts, time‑aligned captions, or symbolic notations. Work includes segmentation and alignment, speaker/label annotation, normalization and punctuation, handling noise and ambiguous input, and measuring transcription accuracy and reliability.

transcription

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inconsistency and poor reproducibility of character error rate (CER) and word error rate (WER) metrics in evaluating automatic transcription models, which stem from opaque preprocessing and ambiguous definitions of characters and words in existing text alignment tools. To resolve this, the authors propose Stringalign, a lightweight Python library that introduces FAIR principles to transcription evaluation for the first time. Stringalign enables reproducible, fine-grained error analysis through Unicode-aware transparent normalization, flexible tokenization strategies, and character- and word-level alignment algorithms. Coupled with interactive visualizations, Stringalign significantly enhances evaluation consistency and interpretability across OCR, handwritten text recognition (HTR), and automatic speech recognition (ASR) tasks, thereby facilitating effective model diagnosis and selection.

automatic transcription evaluationcharacter error rateevaluation transparency

Effect-driven interpretation: Functors for natural language composition

Apr 01, 2025
DB
Dylan Bumford
🏛️ University of California, Los Angeles | Yale University

This paper addresses the challenge of jointly modeling pure semantic values and contextual effects—such as reference, tense, interrogation, and negation—in compositional natural language semantics. We propose an effect-driven semantic framework inspired by denotational semantics in programming languages. Methodologically, we systematically introduce functorial structures from category theory to formally characterize the hierarchical interaction between value propagation and side-effectful processes; we integrate type-logical syntax with functional semantic composition to build an extensible interpretation system. Our contributions include a unified treatment of tense, questions, negation, and discourse coherence, significantly enhancing compositional productivity, interpretability, and theoretical unity in semantic parsing. The framework establishes a new paradigm for computational semantics that balances formal rigor with broad linguistic coverage.

Applying denotational techniques to natural language analysisExploring functors for linguistic composition interpretationModeling human language with pure and impure components

Pluto: Authoring Semantically Aligned Text and Charts for Data-Driven Communication

Feb 11, 2025
AS
Arjun Srinivasan
🏛️ Tableau Research | Massachusetts Institute of Technology

Existing visualization tools struggle to achieve semantic-level coordination between text and charts, hindering high-quality data storytelling. This paper introduces a hybrid active chart–text co-creation system featuring a novel bidirectional guidance mechanism that jointly models chart construction features and textual drafts. On one side, it parses visual encodings and aligns them with textual semantics to generate real-time suggestions—including annotations, title summarization, text completion, and data transformation recommendations. On the other, it supports brush-based interactions to trigger text generation and enables text-driven reverse inference for visualization refinement. The system integrates rule-based heuristic reasoning, visual encoding analysis, semantic similarity matching, and lightweight text generation—without relying on large language models—to ensure interpretability and author control. A user study demonstrates that our approach significantly reduces initial authoring time and achieves high recommendation adoption rates, empirically validating that semantic alignment enhances data storytelling quality.

Enhances text-chart semantic alignmentFacilitates interactive text-chart integrationSupports joint text-visual authoring

This work addresses the challenge of accurately and efficiently digitizing complex documents containing handwritten content, irregular tables, and heterogeneous layouts—tasks that remain difficult for conventional OCR systems and current large language models. The authors propose an interactive document digitization system that integrates layout-aware parsing, OCR, and a large language model, enhanced by a user-in-the-loop correction propagation mechanism. Leveraging layout-aware inference, the system automatically generalizes user edits or natural language instructions applied to a local region to structurally similar regions across the document. In a user study (n=12), this approach significantly improved correction efficiency, reduced repetitive manual operations, and enabled more controllable and effective reconstruction of document structure and content.

document digitizationhandwritten contentheterogeneous layouts

Existing element-attribute grid representations for graphic design completion tasks struggle to model variable-length, type-heterogeneous, and multimodal (text-image) structures. Method: We propose a unified interleaved multimodal tokenized document model that jointly encodes syntactic and semantic structures of markup languages (e.g., SVG/HTML) alongside variable-size, alpha-channel-aware local image generation. We introduce a specialized image quantizer for efficient transparent-image tokenization and integrate an enhanced code-large language model with an interleaved multimodal sequence architecture. Contribution/Results: Our model achieves significant improvements over baselines on three design completion tasks—missing template attributes, image synthesis, and text generation—demonstrating its effectiveness in jointly modeling structural logic and visual semantics in design documents.

Completes missing parts in graphic design documentsGenerates images and text for design automationUnifies various design tasks through multimodal representation

Latest Papers

What's happening recently
View more

This study addresses the frequent failure of large language models (LLMs) in software engineering tasks due to structurally invalid outputs—such as syntactic or formatting errors—that prevent correct parsing by downstream toolchains, even when the semantic content is accurate. The authors systematically evaluate the reliability of structured output generation across four representative tasks, categorizing errors into syntactic, structural, and semantic types. They propose a template-driven token-matching generation (TTMG) method that enforces structural consistency during autoregressive decoding. Experimental results demonstrate that TTMG nearly eliminates syntactic errors; however, structural and semantic errors remain prevalent. This work reveals, for the first time, that the fundamental bottleneck in structured output generation lies not merely in syntax but in the insufficient coordination between structural and semantic correctness, indicating that existing structural control mechanisms, while necessary, are insufficient without joint guarantees of both dimensions.

format compliancelarge language modelssoftware engineering

This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.

ambiguitycontent creationedit instruction

Existing evaluation methods for audio captioning struggle to accurately assess the fidelity of multimodal semantics and acoustic attributes in structured audio descriptions. This work proposes the first multi-axis evaluation framework tailored for structured audio captioning, integrating large language model (LLM)-based semantic judgments with deterministic acoustic metrics across five orthogonal dimensions: label sets, descriptive content, logical reasoning, numerical measurements, and spectral contours. The framework incorporates a controlled perturbation protocol to validate its ability to distinguish between semantic preservation and acoustic distortion. Experiments on the AudioCards dataset demonstrate that the proposed approach effectively differentiates semantically consistent paraphrases from genuine errors, significantly outperforming existing methods in both reliability and sensitivity.

audio descriptioncaption metricsevaluation framework

This work addresses the limited controllability of existing symbolic music generation methods, which typically rely on from-scratch synthesis and struggle to support explicit local editing. The paper introduces the first explicit editing framework for symbolic music, reframing generation as a draft-editing process. Built upon a BEAT-based rhythmic grid anchoring scheme, the framework unifies three editing mechanisms—token-wise sequence labeling, iterative accompaniment refinement, and post-hoc token infilling—within a single pre-trained backbone model to enable efficient inference. Experimental results demonstrate that the proposed approach outperforms both autoregressive and diffusion models across three distinct editing tasks, achieving inference latency under 100 milliseconds while significantly improving both generation accuracy and perceptual audio quality. These findings highlight a strong correlation between music representation design and editing efficacy.

beat-grid encodingexplicit editingmusic representation