Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between theoretical memorization and case-based reasoning exhibited by large language models in Bazi (Chinese Four Pillars of Destiny) fortune-telling. We introduce the first Chinese metaphysics evaluation framework that explicitly decouples theoretical knowledge from case reasoning dimensions, constructing a 3,000-item multi-category multiple-choice benchmark. Employing model-assisted refinement and paired comparative analysis techniques, we systematically evaluate six prominent models. Our findings reveal that all models achieve significantly higher accuracy on theoretical questions than on case analyses, performing particularly poorly in career- and family-related tasks. This work demonstrates that reasoning deficiencies within specific cultural domains do not stem from inadequate general capabilities. Consequently, it underscores that evaluating domain knowledge application necessitates task-specific assessments rather than reliance on single aggregate metrics.
📝 Abstract
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Bazi
Domain Knowledge Application
Evaluation Benchmark
Theory-Case Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bazi benchmark
theory-case gap
domain-specific evaluation
reasoning configuration
cultural-domain NLP
🔎 Similar Papers
No similar papers found.