Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition

📅 2025-02-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of automated monolith-to-microservices decomposition—namely, low automation, weak semantic awareness, and poor generalizability—this paper proposes MonoEmbed. It leverages large language models (LLMs) to generate fine-grained, semantically rich code embeddings and introduces, for the first time in monolith decomposition, a synergistic integration of contrastive learning and LoRA-based fine-tuning to enhance embedding discriminability and robustness. Subsequently, hierarchical clustering is applied to achieve high-cohesion, low-coupling service partitioning. MonoEmbed overcomes the limitations of rule-based or static-analysis approaches by enabling end-to-end, semantics-driven automated decomposition. Extensive experiments across multi-scale monolithic systems demonstrate that MonoEmbed improves key metrics—including module cohesion, balance, and independence—by an average of 23.6% over state-of-the-art baselines.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageComputer Vision: Segmentation

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Architectures and workflows that use LLMs for crowd work
📝 Abstract
As Monolithic applications evolve, they become increasingly difficult to maintain and improve, leading to scaling and organizational issues. The Microservices architecture, known for its modularity, flexibility and scalability, offers a solution for large-scale applications allowing them to adapt and meet the demand on an ever increasing user base. Despite its advantages, migrating from a monolithic to a microservices architecture is often costly and complex, with the decomposition step being a significant challenge. This research addresses this issue by introducing MonoEmbed, a Language Model based approach for automating the decomposition process. MonoEmbed leverages state-of-the-art Large Language Models (LLMs) and representation learning techniques to generate representation vectors for monolithic components, which are then clustered to form microservices. By evaluating various pre-trained models and applying fine-tuning techniques such as Contrastive Learning and Low Rank Adaptation (LoRA), MonoEmbed aims to optimize these representations for microservice partitioning. The evaluation of the fine-tuned models showcases that they were able to significantly improve the quality of the representation vectors when compared with pre-trained models and traditional representations. The proposed approach was benchmarked against existing decomposition methods, demonstrating superior performance in generating cohesive and balanced microservices for monolithic applications with varying scales.
Problem

Research questions and friction points this paper is trying to address.

Automates monolith-to-microservice decomposition
Enhances representation vectors using LLMs
Improves microservice cohesion and balance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Learning enhances model accuracy
Low Rank Adaptation optimizes model performance
MonoEmbed automates microservice decomposition process
🔎 Similar Papers
No similar papers found.