Generating Synthetic Oracle Datasets to Analyze Noise Impact: A Study on Building Function Classification Using Tweets

📅 2025-03-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Weak supervision from geotagged tweets introduces sentence-level feature noise (e.g., irrelevant or uninformative text) in building functional classification, yet lacks controllable evaluation of its impact. Method: We propose the first synthetic oracle tweet dataset construction paradigm for geo-textual weak supervision—leveraging LLMs to generate high-quality, semantically coherent tweets with accurate labels, thereby isolating feature noise effects. Results: Experiments reveal feature noise dominates performance degradation far more than model complexity: mBERT significantly outperforms Naive Bayes on synthetic data but collapses to keyword-matching accuracy on real noisy data. Cross-domain generalization and attribution analysis further confirm its decisive influence. This open-source dataset establishes the first controllable benchmark for feature noise in geo-textual weak supervision, enabling rigorous evaluation and advancing methodology for weakly supervised geographic text modeling.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Semi-Supervised LearningKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
Tweets provides valuable semantic context for earth observation tasks and serves as a complementary modality to remote sensing imagery. In building function classification (BFC), tweets are often collected using geographic heuristics and labeled via external databases, an inherently weakly supervised process that introduces both label noise and sentence level feature noise (e.g., irrelevant or uninformative tweets). While label noise has been widely studied, the impact of sentence level feature noise remains underexplored, largely due to the lack of clean benchmark datasets for controlled analysis. In this work, we propose a method for generating a synthetic oracle dataset using LLM, designed to contain only tweets that are both correctly labeled and semantically relevant to their associated buildings. This oracle dataset enables systematic investigation of noise impacts that are otherwise difficult to isolate in real-world data. To assess its utility, we compare model performance using Naive Bayes and mBERT classifiers under three configurations: real vs. synthetic training data, and cross-domain generalization. Results show that noise in real tweets significantly degrades the contextual learning capacity of mBERT, reducing its performance to that of a simple keyword-based model. In contrast, the clean synthetic dataset allows mBERT to learn effectively, outperforming Naive Bayes Bayes by a large margin. These findings highlight that addressing feature noise is more critical than model complexity in this task. Our synthetic dataset offers a novel experimental environment for future noise injection studies and is publicly available on GitHub.
Problem

Research questions and friction points this paper is trying to address.

Analyzing impact of sentence-level noise in tweet-based building classification
Generating synthetic oracle datasets to isolate and study feature noise effects
Evaluating model performance degradation due to noise in real-world tweets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates synthetic oracle datasets using LLM
Analyzes noise impact with controlled datasets
Compares model performance across configurations
S
Shanshan Bai
Technical University of Munich, Munich Center for Machine Learning
Anna Kruspe
Anna Kruspe
Munich University of Applied Sciences
Natural Language ProcessingMachine LearningMusic Information RetrievalSpeech Recognition
X
Xiaoxiang Zhu
Technical University of Munich, Munich Center for Machine Learning