Concept Direction Reliability Across Languages with Different Tokenizer Fertility

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the poor cross-lingual reproducibility of sentiment vector directions and their disconnection from classification accuracy by proposing a direction reproducibility metric independent of classification performance. Through representation analysis of large language models, combined with split-half consistency testing and multilingual comparison techniques, this work systematically evaluates the stability of sentiment directions across multiple models in English, Hausa, and Yoruba. The findings reveal that high predictive accuracy does not entail directional consistency; notably, English exhibits significantly superior directional consistency compared to other languages, an advantage partially attributable to sentence length effects rather than solely tokenizer discrepancies. This research establishes a new paradigm for assessing the reliability of cross-lingual sentiment representations.
📝 Abstract
Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducibility, we measure split-half agreement in English, Hausa, and Yoruba representations across four language models using both native and translated texts. We identify layers selected for agreement using ten topics and evaluate direction agreement across separate groups of fifteen topics. Using the final token, split-half agreement ranges from 0.737 to 0.870 for English, 0.589 to 0.762 for Hausa, and 0.101 to 0.399 for Yoruba, maintaining this language rank order across all 77 complete model comparisons. Classifiers trained on these same layers consistently predict sentiment above chance, demonstrating that predictive accuracy does not imply directional consistency. Furthermore, averaging token representations yields less consistent agreement, and high agreement can partially reflect sentence length. Ultimately, our findings highlight the need to measure vector direction reproducibility independently of classification performance, though they do not establish that tokenizer fertility which is the average number of tokens per whitespace separated word causes cross-lingual differences.
Problem

Research questions and friction points this paper is trying to address.

concept direction reliability
tokenizer fertility
cross-lingual representations
direction reproducibility
sentiment classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

direction reproducibility
tokenizer fertility
cross-lingual representations
split-half agreement
concept extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muhammad Abdullahi Said
African Institute for Mathematical Sciences; University of Cape Town
A
Abass Oguntade
African Institute for Mathematical Sciences
E
Elisha Komolafe
African Institute for Mathematical Sciences
Babangida Sani
Babangida Sani
Kalinga University India
Artificial IntelligenceMachine LearningNatural Language Processing
F
Fatima Muhammad Adam
Federal University Dutse
M
Muhammad Sammani Sani
University of Vienna