Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical gap in machine learning education: the overreliance on pre-labeled datasets, which often obscures the subjectivity and ambiguity inherent in data annotation, leading students to place undue trust in model outputs. To counter this, the authors introduce an innovative pedagogical intervention that transforms manual annotation into an active learning tool. Students annotated hair coverage in skin lesion images using a three-point scale, followed by structured reflections via questionnaires. A cross-institutional experiment involving 43 participants from Fontys University of Applied Sciences (Netherlands) and the IT University of Copenhagen (Denmark) demonstrated that this approach significantly enhanced learners’ awareness of annotation ambiguity, dataset biases, and model limitations. Most participants acknowledged the influence of personal interpretation on labeling decisions and reported higher engagement compared to traditional instruction. This work provides the first empirical evidence supporting subjective annotation as an effective strategy for cultivating critical thinking about AI systems.
📝 Abstract
Machine learning courses often use pre-labeled datasets, hiding the subjectivity of human annotation. This creates students with an overly trusting view of AI data and models, undervaluing interpretive diversity. We investigated whether manual data annotation tasks teach students about subjective labeling. Study Design: An annotation activity was implemented at two universities: Fontys (Netherlands) and IT University Copenhagen (Denmark). Students annotated skin lesion images for hair coverage on a 3-point scale. Surveys were collected from 43 participants measuring their understanding of annotation ambiguity, data quality, bias, fairness, implementation barriers, and pedagogical effectiveness. Key Findings: Self-reported familiarity with course content increased substantially across all concepts. Most students recognised that personal interpretation affects annotations. Students rated the activity as more effective than traditional lectures for understanding bias. Participants were motivated to learn more. Main Drawbacks: Emotional unease from viewing medical images was the primary issue. Many students still requested clearer guidelines to reduce disagreement, suggesting they hadn't internalised that disagreement from different perspectives is a learning feature, not a bug. Recommendations for Future Iterations: Ensure sufficient interpretive ambiguity in materials. Reduce repetitive annotation workload. Mitigate emotional unease from sensitive content. Explicitly frame disagreement as a learning opportunity rather than a problem to solve. Manual data annotations effectively teach students that human judgment shapes model behavior and that disagreement reflects domain complexity, not just noise.
Problem

Research questions and friction points this paper is trying to address.

data annotation
subjectivity
machine learning education
interpretive diversity
bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

data annotation
subjectivity
critical thinking
bias awareness
pedagogical intervention
🔎 Similar Papers
No similar papers found.
R
Ralf Raumanns
Fontys University of Applied Science, Eindhoven, The Netherlands; Eindhoven University of Technology, Eindhoven, The Netherlands
T
Theresa Elstner
Kassel University, Kassel, Germany
L
Louis Ferger-Andrews
Fontys University of Applied Science, Eindhoven, The Netherlands
L
Louise M. Carlsen
IT University of Copenhagen, Denmark
Martin Potthast
Martin Potthast
University of Kassel, hessian.AI, and ScaDS.AI
Information RetrievalNatural Language Processing
G
Gerard Schouten
Fontys University of Applied Science, Eindhoven, The Netherlands
J
Josien P. W. Pluim
Eindhoven University of Technology, Eindhoven, The Netherlands
Veronika Cheplygina
Veronika Cheplygina
IT University Copenhagen
meta-researchpattern recognitionmachine learningmedical imagingopen science