TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

📅 2026-08-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of repairing syntactically erroneous markup in scientific and technical documents, a task for which existing evaluation lacks benchmarks grounded in real-world error patterns. To bridge this gap, we introduce TeXFix-Bench, the first multi-format (LaTeX/Typst/Markdown) repair benchmark built upon an empirically derived fault taxonomy. This taxonomy is constructed via grounded theory analysis of community-reported failures, and the benchmark incorporates AST-aware DocMut mutation operators to generate more challenging repair instances. Evaluating seven large language models under a unified zero-shot protocol across 48,651 attempts, we find that DocMut-induced errors are significantly harder to repair than those from conventional mutators, that Typst presents notably higher difficulty, and that 13.6–18.5% of syntactically correct repairs substantially alter the original document content—revealing that compilation success alone substantially overestimates true repair quality.
📝 Abstract
Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts. Existing document-repair evaluations inject faults with ad-hoc edits that lack an empirical fault model. We present TeXFix-Bench, a multi-format benchmark for LLM-based full-source document repair grounded in a mined fault taxonomy. A Grounded-Theory study of localized hard-crash LaTeX faults from TeX Stack Exchange, GitHub commits, and package documentation (168 verified faults, dual open coding at $κ$=0.34) yields an 18-category taxonomy instantiated as DocMut: 48 AST-aware operators across three formats. A three-model cross-benchmark shows DocMut faults are 5.6-9.2 pp harder to repair than pattern-based mutations on the same seeds, and a real-error case study (88 mined human crashes, 67.0% repair success) brackets both synthetic sets from below. We construct 10,437 instances from 743 openly licensed seeds and evaluate seven LLMs under a fixed zero-shot protocol with provider-pinned routing, collecting 48,651 attempts at about USD 200 total inference cost. A complete 6,613-instance x 7-model balanced matrix confirms all rankings. A pinned engine gate yields a 27.5-point intention-to-treat compile spread (56.7-84.2%). Typst is markedly harder than LaTeX and Markdown. A restoration oracle over 28,129 compiling repairs shows that 13.6-18.5% of compiling repairs materially alter document text, and restoration rank diverges from compile rank: the model with the lowest compile rate restores content best among its successes. Compile success alone overstates repair quality. We release the taxonomy, DocMut, and all campaign artifacts.
Problem

Research questions and friction points this paper is trying to address.

document repair
LaTeX
fault taxonomy
markup languages
LLM evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

document repair
fault taxonomy
AST-aware mutation
multi-format benchmark
LLM evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Prajwal S. Venkateshmurthy
Independent Researcher