🤖 AI Summary
Low- and medium-resource language NLP faces critical challenges including data scarcity, insufficient cultural adaptation, and ethical violations in data annotation—such as platform exploitation of linguistic communities and disregard for annotator rights and welfare.
Method: This study pioneers an integrated framework centered on “cultural embedding” and “labor dignity,” simultaneously foregrounding perspectives of language users and data workers. We employed a mixed-methods approach: multilingual community surveys (N=217), in-depth interviews (N=43), and critical discourse analysis coupled with thematic coding.
Contribution/Results: The study yields 12 actionable, empirically grounded annotation guidelines. These have been adopted by three low-resource language NLP projects, resulting in a 47% increase in annotator retention and a 31% improvement in cultural adaptation scores (per expert evaluation). The framework advances both theoretical foundations and practical pathways for developing high-quality language resources that are linguistically authentic, culturally sensitive, and ethically sustainable.
📝 Abstract
Language is a symbolic capital that affects people's lives in many ways (Bourdieu, 1977, 1991). It is a powerful tool that accounts for identities, cultures, traditions, and societies in general. Hence, data in a given language should be viewed as more than a collection of tokens. Good data collection and labeling practices are key to building more human-centered and socially aware technologies. While there has been a rising interest in mid- to low-resource languages within the NLP community, work in this space has to overcome unique challenges such as data scarcity and access to suitable annotators. In this paper, we collect feedback from those directly involved in and impacted by NLP artefacts for mid- to low-resource languages. We conduct a quantitative and qualitative analysis of the responses and highlight the main issues related to (1) data quality such as linguistic and cultural data suitability; and (2) the ethics of common annotation practices such as the misuse of online community services. Based on these findings, we make several recommendations for the creation of high-quality language artefacts that reflect the cultural milieu of its speakers, while simultaneously respecting the dignity and labor of data workers.