Score
Systematically labeling and categorizing media or interaction content to identify roles, presentation patterns, and deceptive behaviors, producing structured annotations and taxonomies for analysis and evaluation.
Traditional screen usage metrics fail to capture the complexity of real-world media experiences. To address this, we propose Media Content Atlas (MCA), the first large-scale, fine-grained, multimodal content analysis framework designed for authentic screen behavior. MCA integrates multimodal large language models (MLLMs), unsupervised vision–semantics clustering, dynamic topic modeling, interpretable image retrieval, and interactive spatiotemporal visualization to enable frame-level content understanding, semantic organization, and open-ended exploration of millions of smartphone screenshots. Evaluated on 1.12 million screenshots collected over 30 days from 112 adults, MCA achieves 96% clustering relevance and 83% accuracy in semantic content description. By unifying inductive exploration with deductive validation, MCA bridges a critical methodological gap in media research and establishes a new paradigm for hypothesis-driven analysis and intervention design in digital behavioral science.
Existing approaches to media narrative analysis often struggle to balance fine-grained detail with scalability, either sacrificing nuance through coarse-grained modeling or relying on domain-specific annotations that limit generalizability. This work proposes an unsupervised method that jointly models events and characters and incorporates structured clustering to automatically induce interpretable narrative schemata from large-scale news corpora. By avoiding manual annotation, the approach yields narrative frameworks that align with theoretical expectations while demonstrating strong generalization capabilities. The resulting schemata preserve fine-grained semantic structure and enable efficient, interpretable discovery of media narratives across diverse datasets.
This paper addresses the low accuracy and poor interpretability of automated identification of latent concepts—such as frames, narratives, and themes—in textual data. To this end, we propose a human-in-the-loop concept discovery framework. Methodologically, we introduce the first open-source large language model (LLM)-driven concept navigation paradigm, integrating iterative prompt-based sampling, cross-domain text embedding and clustering, and an expert validation feedback loop—thereby tightly coupling automated summarization with human-in-the-loop verification. Experiments on AI policy debates, cryptocurrency news, and the 20 Newsgroups dataset demonstrate substantial improvements in political discourse analysis, media frame detection, and fine-grained topic classification. Our approach achieves a superior trade-off between accuracy and interpretability, offering a robust, transparent, and reproducible pathway for concept modeling in computational social science.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.
This study addresses the challenges of automated semantic annotation in broadcast television content, which arise from complex audiovisual structures, distinctive editing patterns, and stringent operational constraints. While general-purpose multimodal large language models have shown promise in various domains, their effectiveness in this specific context remains underexplored. To bridge this gap, the authors develop a multimodal annotation framework tailored to Italian television news, integrating visual features, automatic speech recognition, speaker diarization, and metadata. They systematically evaluate state-of-the-art models—including Gemini 3.0 Pro, LLaMA 4 Maverick, Qwen-VL, and Gemma 3—across four semantic dimensions. The work establishes the first fine-grained multimodal annotation benchmark for broadcast media, reveals trade-offs between model scale and input length, achieves minute-level annotations across 14 full episodes, and demonstrates the feasibility of content-driven audience analysis by linking topical semantics to viewership behavior.
This study addresses the limitations of existing team role modeling approaches, which often lack grounding in educational theory and suffer from poor interpretability, thereby hindering their ability to effectively predict collaborative outcomes. To bridge this gap, the authors develop a theoretically informed framework comprising eight communication roles derived from educational principles. For the first time, this theory-driven role taxonomy is applied to real-world team chat data, integrating expert annotations with large language models for role identification and explicitly modeling the dynamic evolution of roles over time. The research uncovers systematic patterns in how roles shift throughout project progression and demonstrates the cross-contextual validity of these roles in predicting peer recognition and team performance. Experimental results show that the proposed approach significantly outperforms baseline methods—including lexical, conversational, and prompt-engineering strategies—on both prediction tasks.
Short-form videos pose significant challenges for standardized modeling of user engagement due to their multimodal content and platform-specific algorithms. This study addresses this gap by computationally operationalizing classical interpretive theories from narratology, rhetoric, communication studies, and semiotics at scale. Leveraging a multimodal large language model, we automatically annotated 77 theory-driven structural variables across approximately 10,000 TikTok videos from Estonian brands and institutions, supplemented by human validation to assess reliability. Controlling for account size and video age, our model yielded a stable, albeit modest, improvement in predicting user engagement. Results indicate that variables related to perception and communication were reliably annotated, whereas deeper semiotic and archetypal structures proved more challenging to capture. This work establishes a systematic computational framework for analyzing the cultural structures embedded in short-form video content.
This study addresses the urgent need for systematic strategies to identify malicious actors in online social networks by proposing a structured taxonomy that characterizes their behavioral patterns and roles in disinformation campaigns. Integrating domain expert knowledge with insights from academic literature, the framework delineates the mechanisms through which such actors operate. Developed through qualitative analysis, expert collaboration, and case studies, the taxonomy was validated using social media data centered on anti-immigration discourse. The resulting classification system effectively supports researchers and platform operators in detecting and mitigating disinformation, offering both a theoretical foundation and a practical framework for governance interventions.
This study addresses the identification and characterization of multimodal gender bias in internet memes and short videos. It proposes a late-fusion framework based on gradient-boosted regression, augmented with a hierarchical post-processing strategy that integrates visual, textual, demographic, biometric, and high-level semantic features extracted by large language models (LLMs). The findings reveal that LLM-derived semantic features substantially enhance detection performance in static meme tasks, whereas temporal modeling proves critical for dynamic video tasks. Notably, using the full, unfiltered feature set during testing yields superior generalization in video analysis, underscoring a fundamental divergence in optimal processing strategies between static and dynamic modalities.
This study addresses the limitations of existing automatic multi-label classification approaches for scholarly papers, which predominantly rely on titles and abstracts—often insufficient for accurate labeling—while full texts are lengthy and exhibit uneven information distribution. To overcome this, the authors propose segmenting full-text articles by physical position and systematically evaluating the discriminative power of individual sections and their combinations. Their analysis reveals that middle-to-late and concluding sections carry higher informational value. By integrating these informative segments with bibliographic metadata, they construct an enhanced multi-label classification model. Experiments on a corpus of 1,954 library and information science journal articles demonstrate that the proposed cross-segment combination and metadata fusion strategy significantly improves classification accuracy, offering a novel and effective pathway for fine-grained method identification in academic texts.