Evaluating Multimodal Large Language Models on Educational Textbook Question Answering

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing evaluations of multimodal large language models (MLLMs) inadequately assess reasoning capabilities in educational textbook question answering, particularly under integrated text-and-diagram comprehension scenarios. Method: We design a multimodal retrieval-augmented generation (RAG) pipeline that jointly processes textbook passages and accompanying figures to emulate authentic learning contexts. Contribution/Results: We identify and name the “catastrophic context interference” phenomenon—where incorporating semantically relevant retrieved context into diagram-based tasks triggers severe performance degradation (e.g., a 48.14-percentage-point accuracy drop for LLaMA 3.2-Vision). We further demonstrate that architectural differences critically influence multimodal generalization. Leveraging the CK12-QA dataset, we introduce a fine-grained, education-oriented multimodal benchmark. After supervised fine-tuning, LLaMA 3.2-Vision achieves 71.16% accuracy on the test set, validating the efficacy of our framework and benchmark.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Multimodal large language models (MLLMs) have shown success in vision-language tasks, but their ability to reason over complex educational materials remains largely untested. This work presents the first evaluation of state-of-the-art MLLMs, including LLaVA-1.5 and LLaMA 3.2-Vision, on the textbook question answering (TQA) task using the CK12-QA dataset. We introduce a multimodal retrieval-augmented generation (RAG) pipeline to simulate real-world learning by providing relevant lesson paragraphs and diagrams as context. Our zero-shot experiments reveal a critical trade-off: while retrieved context improves LLaVA's performance on text-based questions, it significantly degrades the accuracy of the more powerful LLaMA 3.2-Vision on diagram-based tasks, dropping its validation accuracy from 74.07% to 25.93%. We term this statistically significant phenomenon "catastrophic context interference." Furthermore, fine-tuning highlights architectural differences: LLaMA 3.2-Vision's performance improves to 71.16% on the test set, demonstrating its capacity to learn multimodal integration, whereas LLaVA's performance declines, indicating challenges with generalization. Our results underscore the challenges MLLMs face in modality prioritization and context integration, providing a benchmark and pointing to key directions for developing more robust AI-driven educational tools.
Problem

Research questions and friction points this paper is trying to address.

Evaluating MLLMs on complex educational textbook QA tasks
Assessing multimodal RAG pipeline impact on model performance
Identifying catastrophic context interference in MLLM reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal retrieval-augmented generation pipeline
Evaluating MLLMs on textbook question answering
Analyzing catastrophic context interference phenomenon
H
Hessa A. Alawwad
Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia
A
Anas Zafar
FAST School of Computing, National University of Computer and Emerging Sciences, Karachi, Pakistan
Areej Alhothali
Areej Alhothali
Associate Professor of Computer Science, King Abulaziz University
Machine learningNatural language processingAffective ComputingSentiment analysis
Usman Naseem
Usman Naseem
Lecturer (Asst. Prof.) @Macquarie University
Natural Language ProcessingLLM AlignmentNLP for Social GoodTrust and Safety
A
Ali Alkhathlan
Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia
A
Amani Jamal
Faculty of Computing and Information Technology & Center of Research Excellence in AI and Data Science, King Abdulaziz University, Jeddah, Saudi Arabia