What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inflated accuracy often reported by data-driven models in predicting room acoustic parameters, which stems largely from biases in evaluation protocols—particularly the overestimation of performance when test locations lack actual measurements. To rectify this, the work proposes a consistent evaluation framework that explicitly distinguishes between “location interpolation” and “prediction at truly unknown locations.” Using multi-condition measured data, it systematically evaluates three approaches: random forests, hybrid CNNs, and inverse distance weighting. Results show that high predictive performance (R² = 0.80–0.88) is achievable only when measured impulse responses at test locations are available as positional fingerprints. Under realistic generalization conditions—without any test-point data—performance drops substantially (R² = 0.09–0.57), though learning-based models still demonstrate practical advantages in predicting sound strength and reverberation time. This work underscores the dominant influence of evaluation protocols on reported metrics and establishes a more reliable benchmark for acoustic modeling.
📝 Abstract
Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than by the model. Using a multi-condition measurement campaign in a 264-seat conference hall and a 180-seat concert hall, three model families were evaluated under a factorial protocol ablation: validation splits either row-based or grouped by receiver position, and input features either including measured-at-test quantities or restricted to source-receiver geometry and environmental state. Row-based splits with measured-at-test inputs reproduce the high reported accuracies (mean $R^2$ 0.81 for the core parameters); grouping the splits by position and restricting inputs to information available at an unmeasured position reduces these to 0.09-0.57, reordering the apparent difficulty of parameter classes. A hybrid CNN evaluated with the target's own impulse response as input is shown to exploit it as a position fingerprint rather than as transferable acoustic information; training-only signal access yields no gain for any parameter tested, including reverberation time. Under the deployment-consistent protocol, the spread between Random Forest, the hybrid CNN, and inverse-distance weighting is an order of magnitude smaller than the spread between protocols for a fixed model; the learnt models retain a genuine advantage for sound strength and reverberation time, and the high accuracy of the original pipelines re-emerges as condition interpolation at measured positions (band means 0.80-0.88), a distinct and operationally useful task.
Problem

Research questions and friction points this paper is trying to address.

room acoustic prediction
evaluation protocol
input availability
generalization
machine learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation protocol
room acoustic prediction
deployment-consistent validation
input availability
position fingerprinting
🔎 Similar Papers
No similar papers found.
A
Akın Oktav
Vibration and Acoustics Laboratory (VAL), Alanya Alaaddin Keykubat University, 07425, Antalya, Türkiye; Department of Mechanical Engineering, Alanya Alaaddin Keykubat University, 07425, Antalya, Türkiye