multimodal retrospective summarization

Design and build systems that ingest, align, and fuse heterogeneous temporal modalities (e.g., sensor logs, video, audio, and text) to extract objective statistics and descriptive events about activities and routines. Analyze and generate concise retrospective timeline summaries and routine-level abstractions that cope with varying data availability, temporal granularity, and uncertainty.

multimodalretrospectivesummarization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of generating structured event logs from multimodal data such as videos to support business process mining. The authors propose an end-to-end approach that first maps video frames into feature vectors using image embeddings, then performs temporal segmentation via an inter-frame similarity matrix. Subsequently, a generalized few-shot classification method automatically assigns semantic labels to the resulting segments, yielding a timestamped, structured event sequence. This work represents the first integration of image embeddings with few-shot learning for the automatic transformation of raw video into process-mining-ready event logs, thereby overcoming the traditional reliance on pre-structured input data. The method’s effectiveness and practicality are validated through experiments in real-world scenarios.

data multi-modalityevent data extractionevent discovery

This work addresses the challenge of modeling the dynamic and highly abstract evolution of information narratives during crisis events, a task where existing approaches are largely confined to static snapshots. We propose the first framework that integrates situated cognition theory with unsupervised temporal modeling, enabling adaptive representation of narrative entity trajectories within a shared semantic space. By combining semantic embeddings, density-based clustering, and rolling time-window linkage, our method requires no predefined labels and captures fine-grained narrative lifecycles, revealing heterogeneous evolution patterns characterized by coexisting transient fragments and stable anchors. Experiments on real-world crisis data demonstrate high clustering consistency and the ability to effectively identify diverse narrative evolution pathways, offering interpretable temporal representations for dynamic information monitoring and decision-making.

Crisis EventsDynamic Information EnvironmentsInformation Environment

This work addresses the challenges of fusing heterogeneous modalities—such as time-series metrics and textual logs—and the inherent difficulty large language models face in processing continuous temporal data for root cause analysis in cloud infrastructure failures. To this end, the authors propose a multimodal diagnostic framework that aligns time-series performance indicators with the embedding space of pretrained language models through temporal semantic compression, a gated cross-attention alignment encoder, and a retrieval-augmented generation mechanism. This integration enables automated root cause localization informed by historical knowledge. Experimental evaluation across six cloud system benchmarks demonstrates that the proposed method achieves a diagnosis accuracy of 48.75%, significantly outperforming existing approaches, particularly in complex, multi-fault scenarios.

cloud failurelarge language modelsmultimodal

Hybrid/remote meetings commonly suffer from prolonged duration and declining engagement, while conventional fixed-length summaries fail to satisfy heterogeneous user needs—such as rapid skimming versus deep retrospective review. To address this, we propose Recap, an LLM-driven dual-track meeting summarization system. Grounded in cognitive science and discourse theory, Recap introduces the first complementary summarization paradigm comprising “key highlights (for overview)” and “structured, hierarchical minutes (for retrospective navigation).” It integrates organizational context (e.g., slide links) with personalized adaptation mechanisms, advancing AI-generated summaries from generic outputs toward seamless workflow integration. Through a high-fidelity prototype and qualitative studies in authentic Microsoft meeting contexts (N=7), we empirically validate the synergistic value of both summary types in collaborative discussion and consensus building. Furthermore, analysis of user editing behaviors (additions, deletions, modifications) reveals critical human-AI alignment gaps, providing empirical grounding for explainable and editable AI meeting summaries.

Addressing diverse post-meeting recap needs with AI-generated summariesDesigning hierarchical and highlights-based recaps using cognitive science principlesEvaluating recap effectiveness in real-world organizational meeting contexts

This study introduces the clinical event relative timeline extraction task for PubMed case reports, aiming to convert unstructured text into temporally ordered event sequences annotated with relative temporal relations—enabling patient trajectory modeling, causal reasoning, and process prediction. Methodologically, we establish the first medical-domain relative timeline annotation guideline and design a multi-LLM consistency evaluation framework to create a new benchmark; zero-shot event identification and relative ordering are performed using large language models (e.g., O1-preview), followed by human verification and cross-model consistency analysis. Experiments achieve 0.80 event recall and 0.95 temporal ordering accuracy on real-world case reports, demonstrating high-fidelity temporal structuring. Our core contributions include: (1) formal definition of a novel NLP task in clinical text understanding; (2) construction of the first domain-specific annotation schema and evaluation framework for relative timelines; and (3) empirical validation of LLMs’ effectiveness in medical relative temporal relation extraction.

Assessing LLM performance in temporal event annotationExtracting relative timelines from clinical case reportsTransforming unstructured clinical events into time series data

Latest Papers

What's happening recently
View more

This work addresses the lack of native support for structured time series in general-purpose AI agents, which hinders end-to-end temporal reasoning within rich contextual environments. To overcome this limitation, the authors propose TimeClaw, a novel framework that integrates executable time-series tools, an experience-driven subroutine evolution mechanism, and multimodal episodic memory into large language model agents, thereby endowing them with native temporal reasoning capabilities. TimeClaw enables traceable, reusable, and open-ended time-series analysis workflows. Extensive evaluations across multiple domains—including energy, finance, meteorology, and transportation—demonstrate that TimeClaw significantly outperforms existing approaches, validating its effectiveness and generality in real-world scenarios.

contextualized reasoninggeneralist AI agentsstructured temporal signals

This work addresses the incompleteness of patient timelines caused by the absence of precise timestamps in clinical notes and the omission of numerous events in structured electronic health records (EHRs). It proposes the first multimodal alignment framework incorporating retrieval-augmented mechanisms to jointly leverage the semantic richness of clinical text and the temporal precision of EHR tabular data. Anchoring on key events, the method constructs absolute clinical timelines by inferring relative temporal offsets from text and calibrating them with structured EHR entries. The approach employs graph-based multi-stage modeling, instruction-tuned large language models, and cross-modal retrieval alignment, evaluated on MIMIC-III/IV using the AULTC metric. Experiments demonstrate significant improvements in absolute timestamp accuracy and temporal consistency on the i2m4 benchmark, with 34.8% of text-derived events absent from tabular records, underscoring the efficacy of multimodal fusion for building more complete and precise patient trajectories.

clinical narrativesclinical timeline reconstructionelectronic health records

This study addresses the limitations of conventional timeline tools, which rely on linear chronology and struggle to capture the irregularity and context-dependence of lived experiences. To overcome this, the authors propose an “orchestrated” approach to timeline construction that treats time as a spatially organized experience, integrating relative positioning, rhythm, and contextual cues among events. They developed a tablet-based augmented reality (AR) prototype enabling users to freely arrange and immersively explore temporal narratives within physical space. Through an in-situ qualitative study with twelve participants, the research uncovers diverse strategies for constructing nonlinear timelines, demonstrates the feasibility of contextualized temporal storytelling in AR environments, and distills key design principles for immersive time-based narratives.

augmented realitychronologysituatedness

This work addresses the limitation of existing time series anomaly detection methods, which predominantly focus on numerical signals and overlook the semantic complementarity of multimodal information such as text, thereby struggling to achieve cross-modal semantic alignment and efficient interaction. To this end, we propose MindTS, a novel model that introduces, for the first time, fine-grained temporal-textual semantic alignment and a content condensation–reconstruction mechanism. By integrating cross-view text fusion, multimodal alignment, and cross-modal reconstruction, MindTS effectively resolves semantic inconsistency and redundancy in heterogeneous data. Extensive experiments on six real-world multimodal datasets demonstrate that MindTS significantly outperforms current state-of-the-art methods, achieving leading or highly competitive performance in anomaly detection.

cross-modal interactionmultimodalredundant information

This study addresses the challenge of transforming multimodal passive sensing data from older adults into comprehensible and meaningful retrospective summaries for remote family members to support caregiving decisions and emotional connection. To this end, the authors propose a semantic transition framework that progresses from “What” to “How” and “Why,” leveraging a large language model (LLM)-based, multi-layered multi-agent narrative summarization mechanism that integrates objective behavioral data with contextual awareness to generate explanatory insights. Through iterative refinement via technology probes and user-centered design, the approach was evaluated by 11 remote family members and demonstrated significantly higher ratings than baseline methods in satisfaction, perceived usefulness, trustworthiness, and willingness to adopt, thereby enhancing understanding and acceptance of remote caregiving practices.

older adultspassive tracking dataremote family members

Hot Scholars

DL

Dingzeyu Li

Research Scientist @ Adobe Research
GenAIVideoCreative tools
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
HW

Huaxiaoyue Wang

PhD Student, Cornell University
RoboticsRobot Learning
RF

Rogerio Feris

Research Manager, MIT-IBM Watson AI Lab
Computer VisionMachine LearningArtificial Intelligence
AG

Anhong Guo

Assistant Professor, University of Michigan
human-computer interactionaccessibilityhuman-AI interactionaugmented reality