๐ค AI Summary
This study addresses the reliance of existing character-level reasoning evaluations for large language models on isolated probes and aggregated metrics, which lack fine-grained diagnostics. We construct a statistical diagnostic benchmark comprising six task categories. Through controlled input design and a paired statistical testing framework integrating McNemarโs test, bootstrap confidence intervals, and tokenization visualization, we systematically evaluate modelsโ character-processing capabilities under zero- to four-shot settings. Our findings reveal differential effects of tokenization mechanisms and reasoning paradigms on character-level accuracy: random strings elicit stronger character awareness than natural English text, chain-of-thought prompting does not necessarily improve performance, and substring extraction tasks yield notably low accuracy. These insights offer a novel perspective for understanding the underlying character-processing mechanisms of large language models.
๐ Abstract
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning. It contains five core tasks, character counting, letter containment, palindrome detection, edit distance, and longest-string selection, plus index_to_span, a harder substring-extraction stress test. The five core tasks use paired English and character-length-matched random-string inputs. index_to_span documents share a 200-500 word band and are not character-length matched. All six tasks use zero-, one-, and four-shot prompts.
We evaluate eight open-weight models from 2B to 32B parameters across 11 reasoning-mode configurations. The framework reports exact-match and relaxed accuracy, Cohen's kappa, paired McNemar tests with odds ratios, bootstrap confidence intervals, Kendall's tau, class-conditional metrics, tokenization analysis, and multiple-comparison-corrected tests. Three findings stand out. First, tokenization shapes accuracy: random strings are more character-visible than English strings (1.892 vs. 3.169 characters per token), and character-counting accuracy falls as English words occupy more tokens. Second, reasoning mode is not uniformly helpful: Gemma4-31B is nearly unchanged across modes on the near-saturated tasks, while Qwen3.6-27B is worse with thinking on palindrome detection (0.952 non-thinking vs. 0.886 thinking at four-shot). Third, index_to_span remains largely unsolved; the best four-shot exact-match accuracy is 6.75%.
Character-level evaluation needs controlled inputs, paired tests, and analyses of tokenization and reasoning mode rather than aggregate accuracy alone.