🤖 AI Summary
This work addresses the scarcity of high-quality datasets for mathematical word problem (MWP) understanding and reasoning in low-resource languages by introducing PatiGonit22K, the largest Bengali MWP dataset to date, comprising 22,441 problems spanning both single-step and multi-step operations. The dataset was meticulously constructed through human translation, structured annotation, cultural localization, and multiple rounds of validation to ensure linguistic naturalness and mathematical correctness. PatiGonit22K substantially expands the scale and complexity of available Bengali MWPs, offering the first comprehensive and balanced benchmark resource for advancing research on mathematical reasoning in low-resource languages.
📝 Abstract
Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets. In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with a substantially larger collection of complex mathematical problems. The dataset includes both simple and multi operation equations, providing a balanced benchmark for evaluating mathematical reasoning across different difficulty levels. Each problem is carefully translated, annotated, culturally adapted, and verified to ensure linguistic consistency and mathematical correctness. By increasing both the scale and complexity of Bengali MWPs, PatiGonit22K provides a more comprehensive resource for future research on mathematical reasoning and educational NLP applications in low resource languages.