Score
Design and train sequence-tagging systems and toolchains that assign part-of-speech labels to individual tokens, producing token-level POS annotations compatible with annotation schemes such as Universal Dependencies and variants for child-directed text. Integrate taggers with parsers and downstream syntactic-analysis pipelines and evaluate their output using tagging accuracy and related token-level metrics.
The taggedPBC dataset—a massive multilingual POS-tagged corpus covering over 1,500 languages—lacks dependency annotations, hindering cross-linguistic syntactic and typological research. Method: We propose the first dependency parsing transfer framework for taggedPBC, integrating cross-lingual POS alignment, dependency structure projection, and word-order statistical modeling. We further conduct empirical correlation analyses with authoritative typological databases (WALS, Grambank, Autotyp). Contribution/Results: Experiments demonstrate high agreement between automatically inferred argument–predicate word-order types in transitive clauses and expert annotations, validating corpus-based typological inference under noisy conditions. We publicly release the complete set of CoNLLU-formatted dependency annotations, establishing the first large-scale, broadly covered, and fully reproducible benchmark for cross-linguistic dependency syntax and linguistic typology research.
To address challenges in part-of-speech (POS) tagging for low-resource languages—including difficult cross-lingual transfer, high adaptation cost, and label imbalance—this paper proposes a language-agnostic, modular, lightweight Transformer framework. The architecture decouples language-specific preprocessing, label-space mapping, and the model backbone, enabling rapid adaptation to new languages with minimal code changes. It employs a unified cross-lingual tagset and targeted data augmentation to enhance robustness under data scarcity and partial language overlap. Evaluated on Bangla and Hindi, the framework achieves 96.85% and 97.00% token-level accuracy, respectively, with stable F1 scores. This work significantly lowers the barrier to model design and hyperparameter tuning for low-resource POS tagging, providing a reusable, extensible infrastructure for cross-lingual NLP.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
This study systematically investigates how dependency annotation schemes affect the performance of transition-based parsers. Method: Addressing language-specific non-canonical structures in Universal Dependencies (UD) treebanks, we design standardization transformation rules and comparatively evaluate parser performance—measured by LAS and UAS—under both original and standardized annotations within a unified, multilingual evaluation framework. Contribution/Results: We empirically demonstrate, for the first time, that annotation standardization does not universally improve parsing accuracy. Crucially, we reveal that linguistic typological features significantly moderate the effectiveness of annotation schemes: for certain languages, the original non-standard annotations yield higher accuracy than standardized ones. This finding challenges the implicit assumption that standardization is inherently optimal and underscores the necessity of considering language-specific syntactic properties when selecting or designing syntactic representations.
Uzbek, a low-resource, morphologically rich language, lacks publicly available universal part-of-speech (UPOS) annotation resources and benchmark datasets for POS tagging. Method: We construct the first open UPOS-annotated benchmark dataset for Uzbek following Universal Dependencies guidelines, fine-tune two monolingual Uzbek BERT models on this data, and systematically evaluate their performance against multilingual BERT and a rule-based tagger. Contribution/Results: Fine-tuned monolingual Uzbek BERT achieves an average accuracy of 91%, substantially outperforming all baselines. This work provides the first empirical validation that monolingual pretraining effectively captures suffix-driven POS variation and context-sensitive morphology—capabilities beyond the reach of traditional rule-based systems. It establishes the first publicly available UPOS benchmark for Uzbek, fills a critical gap in Uzbek NLP infrastructure, and offers a reproducible evaluation framework and effective methodology for POS tagging in low-resource, morphologically complex languages.
This work proposes an unsupervised cross-lingual part-of-speech (POS) tagging framework that operates without parallel corpora, addressing the challenge of scarce labeled data in low-resource languages. By leveraging unsupervised neural machine translation (UNMT), the method generates pseudo-parallel sentence pairs from monolingual corpora alone. It then integrates word alignment with a multi-source cross-lingual label projection mechanism to accurately transfer POS tags to the target language. Notably, this approach achieves effective cross-lingual POS tagging for the first time without relying on genuine parallel data. Evaluated across 28 language pairs, the framework matches or surpasses the performance of baseline methods that depend on authentic parallel corpora, yielding an average accuracy improvement of 1.3%.
Existing linearization methods for dependency graph parsing suffer from excessively large label spaces and fail to preserve complex structural phenomena—including re-entrant edges, cycles, and null nodes. To address this, we propose a novel graph-to-sequence paradigm based on hierarchical bracket encoding, which recursively represents dependency graphs as nested bracket sequences. This approach requires only a constant-size label set (O(1)), enables linear-time parsing, and fully retains both topological and hierarchical graph structure. Unlike conventional action-sequence or flattened encodings, our method leverages hierarchical syntactic grammar to guide structure-aware representation learning. We conduct systematic evaluation across multilingual and multi-formalism benchmarks (UD, DM, PSD), demonstrating significantly higher exact-match accuracy than state-of-the-art linearization methods. Our framework offers a more concise, robust, and scalable modeling paradigm for graph-structured parsing.
This study addresses the widespread assumption in large language models that tokens serve as stable units of measurement, despite significant variation in token length across different tokenizers and text domains—a discrepancy that introduces bias in model evaluation and billing. For the first time, this work systematically quantifies the sequence compression behavior of mainstream tokenizers under diverse text distributions through large-scale empirical analysis, revealing the high variability of token lengths. By challenging the common simplifying assumption that token length is approximately constant, the research elucidates the limitations of treating tokens as a universal metric. These findings provide a more accurate theoretical foundation for assessing model performance, estimating computational resources, and designing fairer usage-based billing mechanisms.
This work addresses a key limitation of conventional large language model training, which relies on token-level next-token prediction and consequently struggles to distinguish between semantically equivalent expressions that differ in surface form, leading to a bias toward superficial patterns rather than deep semantic understanding. To overcome this, the authors propose elevating the training objective to the conceptual level by introducing a systematic concept-level supervision signal through a concept mapping framework—e.g., unifying surface variants like “mom” and “mother” under a shared concept such as MOTHER. Their approach integrates a concept alignment loss with a multi-surface aggregation strategy, encouraging the model to prioritize semantic correctness. Experiments demonstrate that the resulting concept-aware models achieve lower perplexity, superior performance across multiple NLP benchmarks, and enhanced robustness in domain transfer scenarios.