π€ AI Summary
This study addresses the critical challenges in Lombard, a low-resource language, where existing NLP corpora suffer from widespread mislabeling, templated content, noise, and severe dialectal imbalance. For the first time, the authors conduct a systematic manual audit, integrating language identification validation, orthographic analysis, and dialect classification to assess linguistic authenticity and regional representativeness. Their findings reveal that mainstream datasets contain an extremely low proportion of genuine Lombard texts, exhibit inconsistent orthography, and display a pronounced representational bias favoring Western dialects while marginalizing Eastern varieties. The work underscores the necessity of moving beyond quantity-driven data collection toward community-informed, dialect-sensitive strategies for building high-quality linguistic resources, offering crucial methodological guidance for corpus development in other low-resource languages.
π Abstract
Several of the world's languages are still under-resourced in terms of Natural Language Processing (NLP) tools. This is mostly due to the lack of high-quality datasets to train, develop, and evaluate systems and models for several tasks, such as Machine Translation (MT). We conduct a manual audit of the parallel and monolingual corpora available for Lombard, an under-resourced language continuum from Italy. Our analysis reveals that the perceived abundance of web-scraped data is an illusion, with massive datasets plagued by severe language misidentification, boilerplate text, and non-linguistic noise. Furthermore, we analyze the orthographic composition of the valid Lombard portions across web-scraped datasets, curated corpora, and benchmarks. Our findings show conflicting orthographical systems and severe representational bias across all corpora: high-quality data is heavily skewed towards Western Lombard varieties, with Eastern ones left on the margins. This underscores the need for variety-aware, community-driven data curation rather than purely quantity-driven scraping.