Building an Internal Coding Agent at Zup: Lessons and Open Questions

📅 2026-04-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of deploying in-house code generation agents, which often fail to transition from prototype to production due to misalignment between model capabilities and real-world engineering requirements. The authors propose CodeGen, an internal coding agent developed at Zup, that systematically enhances reliability and team adoption through string-replacement-based editing, multi-layered safety guards, explicit state management, and progressive human oversight. Their findings demonstrate that engineering design choices—such as editing strategies, safety mechanisms, and trust calibration—play a more decisive role in production effectiveness than the underlying model alone, thereby bridging a critical gap between technical prototypes and practical deployment of code generation agents.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Agent ArchitecturesMultiagent Systems: Teamwork

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Enterprise teams building internal coding agents face a gap between prototype performance and production readiness. The root cause is that technical model quality alone is insufficient -- tool design, safety enforcement, state management, and human trust calibration are equally decisive, yet underreported in the literature. We present CodeGen, an internal coding agent at Zup, and show that targeted tool design (e.g., string-replacement edits over full-file rewrites) and layered safety guardrails improved agent reliability more than prompt engineering, while progressive human oversight modes drove organic adoption without mandating trust. These findings suggest that the engineering decisions surrounding the model -- not the model itself -- determine whether a coding agent delivers real value in practice.
Problem

Research questions and friction points this paper is trying to address.

coding agent
production readiness
tool design
safety enforcement
human trust
Innovation

Methods, ideas, or system contributions that make the work stand out.

tool design
safety guardrails
human oversight
code generation agent
production readiness
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Gustavo Pinto
Gustavo Pinto
UFPA & Zup Innovation
Software EngineeringRefactoringSoftware RepositoriesML4SE
P
Pedro Eduardo de Paula Naves
Zup Innovation, Brazil
A
Ana Paula Camargo
Zup Innovation, Brazil
M
Marselle Silva
Zup Innovation, Brazil