🤖 AI Summary
This study addresses the challenge of automatically structuring latent, unknown attributes within thematic corpora by proposing a model-driven iterative schema induction framework that operates without predefined ontologies. The framework leverages large language models to automatically discover and extract domain-specific attributes, achieving efficient induction through semantic merging and structured type assignment. Furthermore, it supports attribute extraction using smaller models, effectively balancing precision with computational cost. Experimental results demonstrate that the discovered attributes achieve 61% agreement with human annotators, while value extraction attains an F1 score of 0.8. Overall, this work presents an economical and scalable approach to structured data mining from unstructured text.
📝 Abstract
Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.