SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

๐Ÿ“… 2026-10-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the reliance of existing character-level reasoning evaluations for large language models on isolated probes and aggregated metrics, which lack fine-grained diagnostics. We construct a statistical diagnostic benchmark comprising six task categories. Through controlled input design and a paired statistical testing framework integrating McNemarโ€™s test, bootstrap confidence intervals, and tokenization visualization, we systematically evaluate modelsโ€™ character-processing capabilities under zero- to four-shot settings. Our findings reveal differential effects of tokenization mechanisms and reasoning paradigms on character-level accuracy: random strings elicit stronger character awareness than natural English text, chain-of-thought prompting does not necessarily improve performance, and substring extraction tasks yield notably low accuracy. These insights offer a novel perspective for understanding the underlying character-processing mechanisms of large language models.
๐Ÿ“ Abstract
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts. We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%. Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.
Problem

Research questions and friction points this paper is trying to address.

character-level reasoning
large language models
evaluation benchmark
tokenization
reasoning mode
Innovation

Methods, ideas, or system contributions that make the work stand out.

Character-level reasoning
Diagnostic benchmark
Statistical evaluation framework
Tokenization analysis
Reasoning mode
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mohsen Larni
Department of Computer Science, University of Nevada, Las Vegas
S
Sobhan Ebrahimi Azar
Department of Computer Science, University of Nevada, Las Vegas
P
Pouyan Nahed
Department of Computer Science, University of Nevada, Las Vegas
K
Kazem Taghva
Department of Computer Science, University of Nevada, Las Vegas