genai architecture design

Designs and specifies end-to-end generative AI system architectures, including component selection, data flows, deployment patterns, scalability and security requirements. Builds and integrates Amazon AgentCore (and similar agent orchestration components) into AWS GenAI stacks, defining interfaces, operational controls, and end-to-end integration between agents, models, and cloud infrastructure.

genaiarchitecturedesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Toward Agentic Environments: GenAI and the Convergence of AI, Sustainability, and Human-Centric Spaces

Dec 15, 2025
PP
Przemek Pospieszny
🏛️ EPAM Systems | Warsaw School of Economics

Cloud-centric AI deployment incurs substantial carbon emissions and privacy risks due to high computational demands. Method: This paper proposes the “Proactive Environment” framework—a sustainable, distributed AI paradigm integrating generative AI, multi-agent systems (MAS), and edge computing. It intrinsically embeds carbon efficiency, data sovereignty, and human well-being into system architecture, overcoming cloud dependency. Three-tiered conceptual validations—individual, professional, and urban—are conducted using LLM-powered lightweight edge inference, MAS-coordinated resource scheduling, and human-centered empirical studies (focus groups and in-depth interviews). Contribution/Results: The framework significantly reduces energy consumption and cross-network data transmission overhead, enhances local resource utilization efficiency, and strengthens privacy preservation. It establishes a novel pathway toward low-carbon, trustworthy, and human-centric AI infrastructure.

Enhancing data privacy via decentralized, edge-driven AI solutions.Optimizing resource utilization across personal, professional, and urban domains.Reducing environmental impact of cloud-centric AI through sustainable frameworks.

Experience Deploying Containerized GenAI Services at an HPC Center

Sep 24, 2025
AM
Angel M. Beltre
🏛️ Sandia National Laboratories

High-performance computing (HPC) centers face challenges in supporting containerized generative AI (GenAI) services and ensuring cross-platform reproducibility. Method: This paper proposes a unified architecture integrating HPC and cloud-native technologies—orchestrated via Kubernetes and incorporating the vLLM inference server, multi-container runtimes (e.g., Singularity/CRI-O), object storage, and vector databases—to enable seamless deployment and coordinated execution of GenAI components across heterogeneous HPC environments. Contribution/Results: The approach breaks down traditional isolation between HPC and cloud-native ecosystems, enabling high-fidelity, cross-platform reproducibility of containerized AI workloads. Evaluated on Llama-series models, the system demonstrates superior stability, inference throughput, and deployment consistency compared to pure-HPC or pure-cloud alternatives. It establishes a reusable deployment paradigm for GenAI services, significantly enhancing HPC centers’ capability to support large-model inference workloads.

Deploying containerized GenAI services in HPC environmentsIntegrating HPC and Kubernetes platforms for AI workloadsRunning containerized inference servers across different platforms

AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services

Sep 09, 2025
SS
Shiva Sai Krishna Anand Tokal
🏛️ Indian Institute of Science | IBM India Research Lab

Agentic AI systems suffer from insufficient robustness, hallucination, and difficulties in maintaining contextual consistency—particularly when coordinating multiple tools, planning long-horizon tasks, and preserving state across extended interactions. To address these challenges, this paper introduces AgentX, a novel agent workflow paradigm featuring a three-tier collaborative architecture: “Phase Design → Stepwise Planning → Precise Execution,” integrating dedicated Phase Designer, Planner, and Executor roles. Furthermore, we propose two lightweight, highly elastic deployment strategies for Model Context Protocol (MCP) services—leveraging both MCP-native interfaces and Function-as-a-Service (FaaS) cloud functions. Extensive experiments across three real-world application scenarios demonstrate that AgentX significantly improves task success rates while reducing end-to-end latency and API invocation costs. It outperforms representative baselines—including ReAct and Magentic One—establishing a scalable, reproducible, and production-ready engineering framework for complex agentic AI systems.

Addressing complex multi-step tasks in Agentic AI systemsManaging long-context history tracking to prevent hallucinationsOrchestrating numerous tools effectively in agentic workflows

This study addresses the tension between knowledge acquisition and verification in interdisciplinary research, alongside the unclear mechanisms underlying generative AI (GenAI) adoption. Through a first-of-its-kind longitudinal investigation combining semi-structured interviews and qualitative analysis, it examines how researchers orchestrate GenAI to bridge knowledge gaps while preserving cognitive agency. The findings reveal an “expertise paradox” and elucidate the specific operational patterns through which GenAI facilitates interdisciplinary inquiry. Furthermore, this work proposes design strategies centered on calibrated verification, cross-domain synthesis, and disciplinary norm adaptation. Collectively, these insights provide empirical foundations for developing adaptive GenAI systems that prioritize and augment expert capabilities.

Epistemic AgencyExpertise ParadoxGenerative AI

Generative AI for Software Architecture. Applications, Trends, Challenges, and Future Directions

Mar 17, 2025
ME
Matteo Esposito
🏛️ University of Oulu | Tampere University | University of Arizona | IIIT Hyderabad

This study addresses the research gap concerning generative AI (GenAI) in software architecture. Through the first multi-source literature review (MLR) in this domain—integrating open coding and multilingual retrieval—the authors systematically synthesize GenAI applications across early SDLC phases, including requirements-to-architecture and architecture-to-code translation. Key findings reveal that retrieval-augmented generation (RAG) and few-shot prompting demonstrate broad applicability in architectural decision-making and refactoring; microservices and monolithic architectures are the predominant targets for GenAI adaptation; yet six critical challenges persist: hallucination, insufficient precision, ethical and privacy concerns, and the absence of architecture-specific datasets and evaluation frameworks. The primary contributions include establishing a coherent application taxonomy for GenAI in software architecture, identifying core technical bottlenecks, and providing a foundational theoretical and practical roadmap for future research.

Explores GenAI applications in software architecture.Identifies challenges like precision, ethics, and datasets.Proposes future directions for GenAI in SDLC.

Latest Papers

What's happening recently
View more

The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project

Nov 03, 2025
RG
Robin Gröpler
🏛️ ifak e.V. | Siemens AG | BT Group | Akkodis Germany Solutions GmbH | University of Hildesheim | University of Innsbruck | Bilkent University | King’s College London | Fraunhofer FOKUS | Beko | GoCodeGreen

The integration of generative AI (GenAI) into the software development lifecycle (SDLC) faces critical uncertainties regarding reliability, accountability, security, and data privacy. Method: This study—conducted collaboratively by over 30 European industry, academic, and research institutions—establishes a comprehensive GenAI adoption vision and a co-innovation framework spanning requirements, design, coding, testing, and operations. It introduces a five-year technology roadmap and develops a verifiable toolchain integrating code generation, quality assurance, security analysis, and trustworthy AI. Contribution/Results: The project delivers cross-sector consensus-based practice guidelines and functional prototypes, enabling the reliable, scalable deployment of AI-driven software engineering. It further supports the transformation and reskilling of software engineers to meet evolving role demands in GenAI-augmented development environments.

Addressing reliability, accountability, security and data privacy uncertaintiesDeveloping practical GenAI tools for industry-ready software engineering solutionsExploring GenAI application across entire Software Development Life Cycle

This study addresses the lack of a systematic understanding of generative artificial intelligence’s role across the full software development lifecycle. Through a systematic literature review complemented by structured surveys of 65 developers, this work integrates empirical data with existing research to comprehensively evaluate the real-world impact and adoption patterns of large language models (LLMs) in each development phase. Findings indicate that over 70% of developers save more than 50% of their time on boilerplate code generation and documentation tasks, and 79% use browser-based LLMs daily. While nascent governance mechanisms are emerging, benefits in early-stage activities—such as requirements elicitation and architectural design—remain limited. The results suggest that generative AI is shifting the locus of development value from coding toward upstream design activities.

AI GovernanceDeveloper SurveyGenerative AI

This work addresses the limitations of large language models in fine-grained web interaction despite their strong performance in high-level semantic planning. To bridge this gap, the authors propose CI4A, a mechanism that abstracts complex UI component interactions into unified tool primitives through semantic encapsulation, thereby constructing an agent-optimized interaction interface that transcends traditional human-centric UI constraints. Implemented on Ant Design, CI4A covers 23 common UI components and features a hybrid agent architecture with a dynamically updated action space conditioned on page state. Evaluated on a reconstructed WebArena benchmark, the CI4A agent achieves a task success rate of 86.3%, substantially outperforming existing methods while significantly improving execution efficiency.

Agent InteractionLarge Language ModelsSemantic Planning

Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap

Dec 04, 2025
JL
Jialong Li
🏛️ Waseda University | Southwest University | Zhongguancun Laboratory | KU Leuven | Peking University | Tokyo Institute of Technology

Despite growing interest in integrating generative AI (GenAI) into self-adaptive systems (SASs), its advantages and challenges remain poorly understood. To address this gap, this study conducts the first cross-domain systematic literature review spanning software engineering, human-computer interaction, autonomous systems, and AI—augmented by large language model–assisted data analysis and logical reasoning—to assess GenAI’s technical fit within each component of the MAPE-K feedback loop. We propose a novel dual-dimensional framework—“autonomy enhancement” and “human-AI collaboration”—to systematically characterize GenAI’s core strengths (e.g., dynamic modeling, intent understanding, policy generation) and critical limitations (e.g., explainability, real-time responsiveness, trustworthiness assurance). Finally, we derive a forward-looking research roadmap covering technical challenges, validation methodologies, and practical implementation pathways—providing a cohesive foundation for both theoretical advancement and industrial deployment of GenAI-powered SASs.

Explores GenAI's potential to enhance self-adaptive systems' autonomy.Identifies benefits and challenges of integrating GenAI into SAS.Proposes a research roadmap for applying GenAI in SAS.

This work addresses the lack of a systematic methodology for constructing customized AI agents, a process currently fragmented across informal resources. The authors propose a framework-agnostic, end-to-end construction methodology grounded in two foundational prerequisites—base design and building blocks—and iteratively refined through three core practices: prototyping, CLI encapsulation, and agent-driven testing. Key innovations include the introduction of the “Turtle mode,” the conceptual insight that multi-agent orchestration fundamentally amounts to CLI composition, and the novel practice of “agent-testing-agent.” To validate the approach, a single developer leveraged AI pair programming to implement AAC, a production-grade open-source agent, within ten days, demonstrating both the feasibility and transferability of the proposed methodology.

agent methodologycustom AI agentsend-to-end development