🤖 AI Summary
This study addresses the challenge of distinguishing genuinely disruptive open-source large language models (LLMs) that drive technological evolution from “consolidating” models that merely extend existing paradigms. Leveraging metadata from over 2.5 million models on Hugging Face, the authors construct a large-scale model lineage network and propose the Model Disruption Index (MDI)—a novel metric that quantifies a model’s disruptive impact by integrating graph-theoretic analysis with fine-tuning strategy patterns. Their findings reveal that the vast majority of open-source LLMs are consolidating in nature, while truly disruptive models are concentrated among large-scale foundational models and their fine-tuned derivatives. This highlights a highly centralized ecosystem characterized by pronounced path dependence in its evolutionary trajectory.
📝 Abstract
The rapid growth of open-source large language models (LLMs) has created a complex ecosystem of model inheritance and reuse. However, existing research has focused mainly on descriptive analyses of lineage evolution, with limited attention to identifying which models play a disruptive role in shaping subsequent development. Using metadata from 2,556,240 models on Hugging Face, this study reconstructs a large-scale lineage network and introduces the Model Disruption Index (MDI) to distinguish between models that reinforce existing technological trajectories and those that become new bases for later development. The results show that most models in the open-source LLM community are consolidative rather than disruptive, reflecting a highly concentrated and path-dependent evolutionary structure. Further analyses suggest that disruptive positions are more likely to emerge among large-scale models and through finetuning strategies. Overall, this study provides a new perspective for identifying disruptive models and understanding uneven technological development in open-source LLM ecosystems.