🤖 AI Summary
Existing methods struggle to generate structured, personalized summaries for multimedia educational videos targeting middle-school students, particularly when integrating heterogeneous modalities (text, image, audio) while preserving rhetorical coherence and user-specific constraints.
Method: This paper proposes a multimodal video-document personalization summarization framework. It explicitly models rhetorical structure as a heterogeneous graph, jointly leveraging rhetorical relation parsing and graph propagation to weight segment importance. A dynamic pruning strategy, guided by user preferences and temporal constraints, ensures both global structural consistency and local semantic integrity.
Contribution/Results: Evaluated on multimodal benchmarks, the method achieves a 12.7% ROUGE-L improvement over state-of-the-art baselines. It significantly enhances summary structure, information coverage, and user satisfaction, while supporting real-time, personalized generation—demonstrating strong applicability in pedagogical multimedia learning scenarios.
📝 Abstract
In the rapidly evolving landscape of digital content, the task of summarizing multimedia documents, which encompass textual, visual, and auditory elements, presents intricate challenges. These challenges include extracting pertinent information from diverse formats, maintaining the structural integrity and semantic coherence of the original content, and generating concise yet informative summaries. This paper introduces a novel framework for multimedia document summarization that capitalizes on the inherent structure of the document to craft coherent and succinct summaries. Central to this framework is the incorporation of a rhetorical structure for structural analysis, augmented by a graph-based representation to facilitate the extraction of pivotal information. Weighting algorithms are employed to assign significance values to document units, thereby enabling effective ranking and selection of relevant content. Furthermore, the framework is designed to accommodate user preferences and time constraints, ensuring the production of personalized and contextually relevant summaries. The summarization process is elaborately delineated, encompassing document specification, graph construction, unit weighting, and summary extraction, supported by illustrative examples and algorithmic elucidation. This proposed framework represents a significant advancement in automatic summarization, with broad potential applications across multimedia document processing, promising transformative impacts in the field.