🤖 AI Summary
This study addresses the long-standing challenge of evaluating dependency parsing performance on non-human species’ sequences in the absence of gold-standard annotations. By integrating unsupervised dependency parsing with network science approaches, the work reveals for the first time that vocal or gestural sequences produced by non-human primates possess intrinsic evaluability due to their rapidly decaying length distributions. Leveraging this property, the authors propose a theoretical framework capable of effectively estimating the proportion of correctly recovered dependency edges without requiring gold-standard data, and empirically validate its efficacy. This research overturns the conventional assumption that evaluation is impossible without gold standards and demonstrates that such assessment is, in fact, more feasible for non-human communicative sequences than for human languages.
📝 Abstract
Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, dependency parsing is unfeasible in other species. However, here we apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce due to the fast decay of the sequence length distribution. In contrast, human language sequences lack that property. Therefore, evaluation without a gold standard is feasible in non-human primates but a hard problem in humans.