🤖 AI Summary
This study addresses the scarcity of high-quality annotated data and the prohibitive costs of fine-tuning large language models for ontology learning by reformulating the task as prompt optimization. Methodologically, a multi-agent system is employed to automatically generate training data, while prompts are iteratively refined using the GEPA greedy evolutionary algorithm within the DSPy framework, offering a lightweight alternative to conventional fine-tuning. The research demonstrates that autoregressive architectures outperform greedy decoding strategies in capturing ontological structures. Experimental evaluations on biomedical and plant ontology construction tasks reveal substantial performance improvements, validating the effectiveness and generalizability of the proposed approach in low-resource scenarios.
📝 Abstract
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimization for Ontology Learning, by casting OL as an explicit prompt optimization problem over LLM modules. To obtain training data, we employ a multi-agent system that generates text-ontology pairs from existing expert-curated ontologies. We then propose two ontology learner architectures: a greedy and an autoregressive learner, and optimize both using GEPA, a greedy evolutionary prompt optimizer built on DSPy. Experiments on two ontologies - a biomedical (DOID) and a plant ontology (PO) show consistent improvements after optimization across nearly all model and mode combinations, with autoregressive learners achieving the largest gains. Our results demonstrate that prompt optimization is a viable and lightweight alternative to fine-tuning for OL, and that the autoregressive formulation better captures ontological structure than the greedy approach.