🤖 AI Summary
This study addresses the limitations of traditional macro-areal linguistic classifications, which rely on expert intuition and are constrained to continental or fixed geographic units, lacking a systematic approach applicable across arbitrary spatial scales. To overcome this, the authors propose an automated framework based on geographic clustering that leverages geospatial language data to identify linguistic contact zones at both global and local scales. This method represents the first systematic technique capable of delineating language areas at any desired resolution, thereby transcending the spatial constraints of conventional approaches. Empirical validation demonstrates that the resulting global macro-areas exhibit strong alignment with established authoritative classifications, while the framework also successfully reconstructs the fine-grained structure of well-known linguistic areas, confirming its effectiveness and generalizability.
📝 Abstract
Macroareas are geographical areas used in typological research for grouping variables of interest. In linguistic typology, languages in a given macroarea are considered to have potential for contact, in contrast to those outside the area, where contact is less likely. Along with language family membership, macroareas are used as controls for models in linguistic typology, in an attempt to address the problem of autocorrelation - the observation that historical developments or typological patterns may be due to contact between neighboring languages and/or inheritance from a common ancestral language. Macroareas are therefore a central aspect of research that seeks to separate universal properties of language from local (or language-specific) properties. Existing macroareas largely depend on expert determinations of what constitutes a geographical area of potential contact, and to date have mainly aligned with continents or landmasses (Hammarström and Donohue 2014; Nichols, Witzlack-Makarevich, and Bickel 2013). While there are various historical and theoretical reasons for these groupings, there as of yet has been no systematic approach to identifying such areas for a given region. This paper attempts to address such a gap and move beyond macroarea to identification of language areas of relatively arbitrary size, presenting a simple geographical clustering method for identifying groupings over any area. The method produces a set of worldwide macroareas that largely align with existing groupings, as well as local groupings for a well-known sprachbund.