🤖 AI Summary
The absence of systematic guidance for selecting open-source large language models (LLMs) hinders efficient and compliant deployment decisions.
Method: This paper introduces and maintains an open, dynamically updated checklist assessing the deployment feasibility of LLMs. It systematically catalogs key attributes—publication year, license type (e.g., Apache 2.0, MIT, non-commercial), minimum GPU memory/compute requirements, architectural features (e.g., MoE, quantization support), fine-tuning methods, and ecosystem compatibility—for mainstream foundation and domain-specific models released between 2022 and 2024.
Contribution/Results: The work pioneers a unified evaluation framework integrating legal compliance (license constraints) with engineering feasibility (hardware limitations), publishing the structured, machine-readable dataset publicly on GitLab. This resource significantly accelerates model selection for researchers and practitioners, addressing a critical gap in open-source LLM deployment decision support tools.
📝 Abstract
Large Language Models (LLMs), such as Generative Pre-trained Transformers (GPTs) are revolutionizing the generation of human-like text, producing contextually relevant and syntactically correct content. Despite challenges like biases and hallucinations, these Artificial Intelligence (AI) models excel in tasks, such as content creation, translation, and code generation. Fine-tuning and novel architectures, such as Mixture of Experts (MoE), address these issues. Over the past two years, numerous open-source foundational and fine-tuned models have been introduced, complicating the selection of the optimal LLM for researchers and companies regarding licensing and hardware requirements. To navigate the rapidly evolving LLM landscape and facilitate LLM selection, we present a comparative list of foundational and domain-specific models, focusing on features, such as release year, licensing, and hardware requirements. This list is published on GitLab and will be continuously updated.