AI-Friendly LaTeX: Using LaTeX Code as a Knowledge Source for Retrieval-Augmented Generation

📅 2026-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of leveraging LaTeX source code—rich in structural and semantic information yet hindered by cross-references, custom macros, and unmarked content—for retrieval-augmented generation (RAG). The authors propose a systematic preprocessing pipeline that transforms raw LaTeX documents and their auxiliary files into structured Markdown and JSONL fragments through LaTeX parsing, macro expansion, reference resolution, and semantic annotation. This pipeline enables efficient indexing in vector databases and constitutes the first end-to-end method to convert native LaTeX into a knowledge format readily consumable by large language models. By preserving document structure, semantic labels, and authorial intent, the approach significantly enhances the accuracy and reliability of language models on mathematical and technical question-answering tasks.
📝 Abstract
Large language models can answer questions about textbooks, lecture notes, and programming exercises more reliably when their answers are grounded in an explicit knowledge source. Retrieval-augmented generation (RAG) is a common approach: relevant fragments of a document are retrieved and inserted into the model context before answering. For mathematical and technical material, the original LaTeX source can be a better starting point than a PDF, because it contains structural information, labels, sectioning commands, macros, and authorial intent that are often lost or distorted in PDF extraction. However, LaTeX source is not automatically AI-friendly. Cross-references must be resolved, custom macros must be interpreted, exercises and examples must be identified, and author-supplied semantic metadata may be needed. This article describes a focused preprocessing approach for turning LaTeX source, together with its compiled auxiliary files and optional author annotations, into Markdown and JSONL chunks suitable for indexing in a vector database.
Problem

Research questions and friction points this paper is trying to address.

LaTeX
Retrieval-Augmented Generation
AI-friendly
Knowledge Source
Technical Documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

LaTeX preprocessing
Retrieval-Augmented Generation
AI-friendly documents
structured knowledge extraction
technical content indexing
🔎 Similar Papers
No similar papers found.