Multimodal Political Bias Identification and Neutralization

📅 2025-06-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal bias in political news—arising from synergistic textual and visual cues—exacerbates echo chambers; yet existing debiasing methods focus predominantly on text, neglecting image-level bias modeling. Method: We propose the first end-to-end multimodal political debiasing framework: (i) CLIP-based cross-modal semantic alignment ensures image–text coherence; (ii) ViT quantifies political bias in images; (iii) an interpretable, BERT-driven sequence-level neutralization module rewrites text; and (iv) a joint neutralization replacement mechanism enforces cross-modal consistency. Contribution/Results: This work is the first to unify political bias modeling across both modalities, enabling joint detection and co-debiasing. Experiments show high accuracy in identifying diverse subjective expressions in text, stable convergence of ViT-based image classification, and efficient CLIP alignment. Human-in-the-loop evaluation confirms strong semantic fidelity and effective neutralization, validating the framework’s efficacy in mitigating multimodal political bias.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
Due to the presence of political echo chambers, it becomes imperative to detect and remove subjective bias and emotionally charged language from both the text and images of political articles. However, prior work has focused on solely the text portion of the bias rather than both the text and image portions. This is a problem because the images are just as powerful of a medium to communicate information as text is. To that end, we present a model that leverages both text and image bias which consists of four different steps. Image Text Alignment focuses on semantically aligning images based on their bias through CLIP models. Image Bias Scoring determines the appropriate bias score of images via a ViT classifier. Text De-Biasing focuses on detecting biased words and phrases and neutralizing them through BERT models. These three steps all culminate to the final step of debiasing, which replaces the text and the image with neutralized or reduced counterparts, which for images is done by comparing the bias scores. The results so far indicate that this approach is promising, with the text debiasing strategy being able to identify many potential biased words and phrases, and the ViT model showcasing effective training. The semantic alignment model also is efficient. However, more time, particularly in training, and resources are needed to obtain better results. A human evaluation portion was also proposed to ensure semantic consistency of the newly generated text and images.
Problem

Research questions and friction points this paper is trying to address.

Detect and remove political bias in text and images
Align biased images and text using CLIP models
Neutralize biased words and score image bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses CLIP for image-text bias alignment
Employs ViT classifier for image bias scoring
Applies BERT to detect and neutralize text bias
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cedric Bernard
X
Xavier Pleimling
Amun Kharel
Amun Kharel
PhD Student, Virginia Tech
AI
C
Chase Vickery