🤖 AI Summary
This study addresses the widespread lack of accessibility in AI-generated educational materials and the inadequacy of existing evaluation methods, which often overlook issues of informational and visual overload. To bridge this gap, the authors propose the first reproducible evaluation protocol that explicitly incorporates content overload considerations, integrating WCAG standards with multimodal human-in-the-loop heuristic validation. The framework introduces controllable prompt engineering and a persistent accessibility configuration mechanism. Preliminary experiments conducted within the newly developed open evaluation framework demonstrate that, on a single model, explicitly prompting for WCAG compliance increases average accessibility conformance from 24.2% to 96.7%, underscoring the critical role of prompt design in enhancing accessibility outcomes.
📝 Abstract
Generative AI tools increasingly produce educational materials: documents, slides, images, audio, and video, yet little is known about whether this content meets accessibility requirements. This paper presents a protocol for evaluating the accessibility of AI-generated educational materials against the Web Content Accessibility Guidelines (WCAG), across five content types and multiple tools. The protocol compares three conditions applied to the same tool: a generic instruction with no accessibility language; a single prompt explicitly configured with WCAG criteria; and a persistent, reusable accessibility profile loaded once rather than re-specified each time. Evaluation combines a WCAG rubric per content type with heuristic validation by accessibility experts, addressing a known limitation of automated scanners. Prior evidence shows generative AI tools reproduce inaccessible practices by default, and that explicit configuration measurably improves compliance. This paper extends that discussion to a dimension WCAG checklists miss: visual and informational overload common in AI-synthesized content. The main contribution is methodological: a reproducible protocol and open evaluation instrument, with cases documenting the barriers non-configured AI content creates for people with disabilities. A first exploratory application is also reported: a rubric-based compliance score rose from a pooled mean of 24.2% under the generic condition to 96.7% under the WCAG-configured condition, using one model, one artifact per condition and type, and single-evaluator scoring; the persistent-profile condition was not exercised. The protocol is a reusable resource for the community to apply and extend at full benchmark scale. Beyond the protocol, this work aims to raise awareness of accessibility obstacles generative AI can introduce, and encourage creators to consider accessibility when using these tools.