🤖 AI Summary
This work addresses the challenge of repairing syntactically erroneous markup in scientific and technical documents, a task for which existing evaluation lacks benchmarks grounded in real-world error patterns. To bridge this gap, we introduce TeXFix-Bench, the first multi-format (LaTeX/Typst/Markdown) repair benchmark built upon an empirically derived fault taxonomy. This taxonomy is constructed via grounded theory analysis of community-reported failures, and the benchmark incorporates AST-aware DocMut mutation operators to generate more challenging repair instances. Evaluating seven large language models under a unified zero-shot protocol across 48,651 attempts, we find that DocMut-induced errors are significantly harder to repair than those from conventional mutators, that Typst presents notably higher difficulty, and that 13.6–18.5% of syntactically correct repairs substantially alter the original document content—revealing that compilation success alone substantially overestimates true repair quality.
📝 Abstract
Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc edits that lack an empirical fault model. We present TeXFix-Bench, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy. A Grounded-Theory study of localized hard-crash LaTeX faults from TeX Stack Exchange, GitHub commits, and package documentation (168 verified faults, dual open coding at $κ$=0.34) yields an 18-category taxonomy instantiated as DocMut: 48 AST-aware operators across three formats. A three-model cross-benchmark shows DocMut faults are 5.6-9.2 pp harder to repair than pattern-based mutations on the same seeds, and a real-error case study (88 mined human crashes, 67.0% repair success) brackets both synthetic sets from below. We construct 10,437 instances from 743 openly licensed seeds and evaluate seven LLMs under a fixed zero-shot protocol with provider-pinned routing, collecting 48,651 attempts at about USD 200 total inference cost. A complete 6,613-instance x 7-model balanced matrix confirms all rankings. A pinned engine gate yields a 27.5-point intention-to-treat compile spread (56.7-84.2%). Typst is markedly harder than LaTeX and Markdown. A restoration oracle over 28,129 compiling repairs shows that 13.6-18.5% of compiling repairs materially alter document text, and restoration rank diverges from compile rank: the model with the lowest compile rate restores content best among its successes. Compile success alone overstates repair quality. We release the taxonomy, DocMut, and all campaign artifacts.