RedShell: A Generative AI-Based Approach to Ethical Hacking

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The scarcity of high-quality malicious code datasets has hindered the application of large language models (LLMs) in penetration testing. This work proposes RedShell, a generative AI–based red-teaming tool designed to automatically produce offensive PowerShell scripts, and introduces the first task-specific, real-world malicious code dataset tailored for this purpose. By fine-tuning LLMs with natural language processing techniques and evaluating outputs using semantic similarity metrics—including Edit Distance and METEOR—the generated scripts achieve a syntax error rate below 10%, with semantic similarity exceeding 50% (Edit Distance) and 40% (METEOR). These results demonstrate a significant improvement in both correctness and practical utility of the synthesized code.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationWeb Mining and Content Analysis: Web data generation and simulation
📝 Abstract
The application of Machine Learning techniques in code generation is now a common practice for most developers. Tools such as ChatGPT from OpenAI leverage the natural language processing capabilities of Large Language Models to generate machine code from natural language descriptions. In the cybersecurity field, red teams can also take advantage of generative models to build malicious code generators, providing more automation to Pentest audits. However, the application of Large Language Models in malicious code generation remains challenging due to the lack of data to train and evaluate offensive code generators. In this work, we propose RedShell, a tool that allows ethical hackers to generate malicious PowerShell code. We also introduce a ground truth dataset, combining publicly available code samples to fine-tune models in malicious PowerShell generation. Our experiments demonstrate the strong capabilities of RedShell in generating syntactically valid PowerShell, with fewer than 10% of the generated samples resulting in parse errors. Furthermore, our specialized model was able to produce samples that were semantically consistent with reference snippets, achieving a competitive performance on standard output similarity metrics such as Edit Distance and METEOR, with their mean similarity scores exceeding 50% and 40%, respectively. This work sheds light on the state-of-the-art research in the field of Generative AI applied to Pentesting, and also serves as a steppingstone for future advancements, highlighting the potential benefits these models hold within such controlled environments.
Problem

Research questions and friction points this paper is trying to address.

Generative AI
Ethical Hacking
Malicious Code Generation
PowerShell
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative AI
Ethical Hacking
Malicious Code Generation
PowerShell
Red Teaming
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ricardo Bessa
NOVA University Lisbon — FCT & NOVA LINCS, Portugal
Rui Claro
Rui Claro
INESC-ID, Instituto Superior Técnico, Universidade de Lisboa, Portugal
J
João Trindade
Layer8 - Shield Domain SA, Portugal
J
João Lourenço
NOVA University Lisbon — FCT & NOVA LINCS, Portugal