🤖 AI Summary
This study investigates whether four morphosyntactic variants associated with pronouns in Brazilian Portuguese can reliably indicate speakers’ regional dialect origins. Integrating sociolinguistic theory with computational methods, the research develops a model of morphosyntactic covariation and systematically evaluates the effectiveness of correlation analysis versus clustering algorithms in uncovering regional dialect patterns. Findings demonstrate that clustering methods more effectively capture geographically coherent groupings, significantly outperforming traditional correlation-based approaches. The work not only confirms the diagnostic value of morphosyntactic covariation for inferring regional provenance but also offers a scalable, interdisciplinary analytical framework for dialect identification, thereby advancing the integration of computational linguistics and sociolinguistics.
📝 Abstract
This paper investigates morphosyntactic covariation in Brazilian Portuguese (BP) to assess whether dialectal origin can be inferred from the combined behavior of linguistic variables. Focusing on four grammatical phenomena related to pronouns, correlation and clustering methods are applied to model covariation and dialectal distribution. The results indicate that correlation captures only limited pairwise associations, whereas clustering reveals speaker groupings that reflect regional dialectal patterns. Despite the methodological constraints imposed by differences in sample size requirements between sociolinguistics and computational approaches, the study highlights the importance of interdisciplinary research. Developing fair and inclusive language technologies that respect dialectal diversity outweighs the challenges of integrating these fields.