Initialization Improves LLM-Driven Discovery

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of diversity and fragility of success in large language model (LLM)-driven discovery caused by mode collapse. We propose a general initialization intervention strategy that generates diverse initial solutions through parallel exploration to guide subsequent iterative optimization. Furthermore, this work reveals the predictive role of early-stage discoveries for ultimate success, effectively substituting conventional diversity-inducing methods. Built upon a modular Harness framework, experiments conducted across five task categories demonstrate that the proposed strategy significantly enhances LLM discovery performance. These findings validate the critical importance of high-quality initialization in LLM-driven discovery processes.
📝 Abstract
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
iterative optimization
mode collapse
discovery harnesses
diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Iterative Optimization
Mode Collapse
Parallel Exploration
Harness Design
🔎 Similar Papers
No similar papers found.