π€ AI Summary
This study addresses the challenges of fragmented literature in theory construction and the tendency of existing large language model (LLM) approaches to overlook researcher intuition by proposing a human-AI collaborative framework for theoretical synthesis. Methodologically, the framework leverages LLMs to extract conceptual relation triplets from the literature and aligns them into a hierarchical ontology. It further develops an interactive evidence graph visualization tool that supports multi-granularity aggregation analysis. The core innovation lies in deeply integrating automated scalability with human intuition, thereby preserving researchersβ preferential control over key concepts. Through field deployments in immunology and agriculture case studies, the framework successfully uncovered latent mechanisms elusive to conventional analyses and generated novel hypotheses amenable to experimental validation.
π Abstract
A theory draws many independent observations into one framework with novel hypotheses. A researcher building such a theory must synthesize observations scattered across many papers, each describing related concepts but often in different terms. Which concepts matter most also depends on their preferences and research questions. Recent approaches scale theory synthesis with LLMs, but automate away choices and intuitions from researchers. We present Asterism, which extracts observations from hundreds of papers as concept-relation triples, with concepts unified in a hierarchical ontology. Researchers curate an evidence graph using the ontology and aggregate observations at different levels of granularity to focus theory formation on specific phenomena of interest. In a field deployment (n=10), researchers worked from observations to theories, and kept concepts and hypotheses fitting their preferences. In two case studies, teams of immunology and agriculture researchers discovered mechanisms outside their standard analyses and constructed hypotheses worth follow-up experiments.