🤖 AI Summary
This work addresses the high cost of speaker diarization annotation and the lack of quantifiable evaluation metrics by introducing an open-source annotation tool that guides human annotators through automatically generated initial hypotheses. For the first time, annotation cost—measured in edit operations and time—is treated as a primary output metric. The system features a React-based frontend integrated with the pyannote ecosystem and a stride-accelerated logging engine, supporting automatic initialization, uncertainty-aware highlighting, and a novel “phantom” attention-check mechanism to ensure annotation quality. Experiments on the AMI dataset demonstrate that automatic initialization substantially reduces annotation effort while improving accuracy, with the uncertainty-highlighting strategy yielding the best performance among the evaluated approaches.
📝 Abstract
Labeling speaker diarization data is costly, yet annotation tools rarely measure that cost. We present voxmap-studio, an open-source, React-based diarization annotation tool integrated with the pyannote-based diarization ecosystem. Its canvas is initialized by a fast stride-accelerated diarization engine so that the annotator corrects a hypothesis rather than drawing every speaker turn by hand, and the tool records annotation cost - typed edit-operation counts and time - as a first-class output, enabling quantitative comparison of how much different forms of assistance actually help. Export is gated on per-segment human confirmation and guarded by injected "phantom" attention checks, which prevent unverified automatic output from being released as ground truth. In a preliminary study on nine AMI audio files, unassisted manual annotation was the costliest and least accurate, and automatic initialization shifted the work from creating turns to correcting them; highlighting uncertain segments gave the lowest cost in our small sample. The tool and its instrumentation are open source.