Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the joint impact of latent dimensionality and frame rate in continuous audio encoders on downstream task performance. By employing matched training protocols, frozen-model PCA interventions, and automatic speech recognition (ASR) probing techniques, it systematically evaluates representational disparities across varying width and frame rate configurations. The research reveals the interaction mechanisms between these factors regarding downstream utility, demonstrating that representational organization is more critical than mere reconstruction fidelity. Experimental findings indicate that a moderate latent width paired with a high frame rate optimally benefits ASR performance, while high-dimensional models exhibit performance bottlenecks under specific compression ratios. These results challenge the conventional assumption that higher dimensionality is inherently superior, thereby establishing a new paradigm for encoder design.
📝 Abstract
Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.
Problem

Research questions and friction points this paper is trying to address.

continuous audio encoders
latent dimensionality
frame rate
downstream performance
audio compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

continuous audio encoders
latent dimensionality
frame rate
PCA intervention
width-rate interaction
🔎 Similar Papers
No similar papers found.