Automated Textbook Auditing with Multi-Agent LLM Systems

πŸ“… 2026-07-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge that existing proofreading tools struggle to simultaneously ensure factual accuracy, domain-specific technical correctness, and linguistic quality in educational textbooks. To this end, we propose AI Textbook Auditorβ€”the first modular multi-agent system designed for comprehensive textbook auditing. The system operates dual parallel pipelines: one for fact and technical verification, and another for native PDF-based grammatical analysis, while a referee agent filters false positives to produce structured review reports. Integrating domain-customized prompt engineering, vision-aware PDF parsing (via PyMuPDF), web-augmented fact-checking, and rule-based false-positive suppression, the framework supports cross-disciplinary error categorization and native PDF processing. Evaluated on Romanian high school textbooks in computer science and history/social sciences, it identified 56 and 72 issues respectively, achieving an expert-validated precision of 62.5% and significantly enhancing manual review efficiency.
πŸ“ Abstract
Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, domain-specific technical correctness, and linguistic quality simultaneously -- a task that general-purpose grammar checkers cannot address. We present \textbf{AI Textbook Auditor}, a modular multi-agent pipeline for automated quality assurance of educational materials across subject domains. The system accepts a textbook PDF and produces a structured, human-reviewable report via two analysis tracks: a \textbf{Factual and Technical Track} in which an ensemble of specialized LLM agents detects factual inaccuracies, code errors, incorrect definitions, and conceptual inconsistencies, augmented with web search for humanities domains; and a \textbf{Grammar Track} operating PDF-natively to preserve diacritical encoding. A \textbf{Judge Agent} filters false positives using domain-specific rules before presenting findings to a human reviewer. The pipeline supports two ingestion modes -- vision-native page rendering and PyMuPDF text extraction -- and is domain-adaptable via custom prompts encoding subject-specific error taxonomies. We demonstrate the system on two Romanian upper-secondary textbooks: a CS textbook (56 technical findings across seven categories, with an expert-validated precision of 62.5\%) and a history and social sciences textbook (72 findings spanning factual errors, ideological bias, and grammar). The system is designed as a triage tool that reduces the manual effort of locating candidate issues, with human expert validation required before any editorial action.
Problem

Research questions and friction points this paper is trying to address.

textbook auditing
factual accuracy
technical correctness
linguistic quality
educational materials
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent LLM
automated textbook auditing
domain-specific error detection
PDF-native grammar checking
judge agent
C
Ciprian Cristescu
University of Bucharest, Bucharest, Romania
A
Adrian-Marius Dumitran
University of Bucharest, Bucharest, Romania
A
Angela-Liliana Dumitran
Dimitrie Cantemir Christian University, Bucharest, Romania
G
Gabriel Stefan
University of Bucharest, Bucharest, Romania