🤖 AI Summary
This study addresses the lack of mechanisms in static analysis tools for mapping personal data types in code to standardized privacy vocabularies. We construct an RDF knowledge graph and introduce a novel hybrid mapping strategy that combines deterministic rules with retrieval-augmented generation (RAG) via large language models, enabling the automated alignment of Bearer CLI scan outputs with the Data Privacy Vocabulary (DPV) ontology and GDPR provisions. To ensure semantic accuracy, the proposed method incorporates a five-fold verification gating mechanism alongside SHACL constraints, effectively revealing labeling ambiguities and coverage gaps. Experimental evaluation generated 118 ODRL resources, achieving a Gwet’s AC2 inter-rater reliability score of 0.88 in human assessment. These results validate the high reproducibility and soundness of the proposed framework.
📝 Abstract
Static-analysis scanners can identify personal-data types in source code, but they lack mechanisms to connect these findings to standardized privacy vocabularies. PrivDev maps 122 Bearer CLI data types to Data Privacy Vocabulary Personal Data (DPV-PD) categories and links them to potentially relevant GDPR provisions. Our approach combines deterministic mapping for 43 exact-label matches with a retrieval-grounded Large Language Model (LLM) to resolve the remaining 79 non-trivial mappings. The resulting RDF knowledge graph contains 118 ODRL policy resources that were structurally validated using SHACL. The artifact passed five complementary validation gates that cover structural correctness, query consistency, retrieval quality, LLM-based assessment, and human evaluation. In human evaluation, nine annotators produced 711 judgments, yielding a raw agreement of 0.72, Gwet's AC1 of 0.68, and Gwet's AC2 of 0.88. Our results indicate that the proposed mappings are plausible and reproducible, while also revealing ambiguities in scanner-defined data-type labels and coverage gaps in DPV-PD.