SURE: Framework for Safety to Construct Trustworthy AI

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmentation of safety standards for large language models arising from diverse cultural policies and the consequent absence of a unified, customizable safety alignment framework. To bridge these discrepancies, this work proposes SURE, a data-driven framework that enables controllable AI safety customization through the construction of an adversarial prompt taxonomy, the design of response templates, and the introduction of an absolute safety scoring mechanism. By systematically integrating these components, SURE facilitates precise and adaptable safety alignment across varying cultural contexts. Extensive experiments conducted on multiple foundation models demonstrate that the proposed framework significantly enhances both model safety and alignment efficacy, exhibiting strong generalization capabilities. Ultimately, this research effectively reconciles divergent safety standards across multicultural settings, offering a robust and scalable solution for customizable AI alignment.
📝 Abstract
Warning: This paper contains harmful and offensive text. Recently, large language models such as GPT-4, and Claude have revolutionized tasks in various domains. As the use of these large language models increases, people are increasingly concerned about AI safety and demand that large language models behave responsibly and safely. As a result, there has been growing global interest in developing methods to ensure AI safety. However, the detailed criteria for AI safety may vary depending on the country, culture, and policies of the company you serve. In this study, we propose SURE (A Safe and Unified AI Framework foR Everyone), which is designed as a framework for customizing the attributes of AI safety and ensuring the defined AI safety. Within SURE, we establish taxonomies for adversarial prompts that could threaten AI safety and construct prompts based on the taxonomies. We then define templates for desirable AI responses to these prompts and design an absolute safety scoring scheme. Finally, we conduct AI alignment using the datasets to gradually ensure AI safety. The effectiveness of SURE is demonstrated through experiments with various base models.
Problem

Research questions and friction points this paper is trying to address.

AI safety
large language models
trustworthy AI
safety customization
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Safety
Adversarial Prompts
AI Alignment
Safety Scoring
Customizable Framework
💼 Related Jobs
No related jobs found.
S
Soeun Han
Korea Telecom(KT), Seoul, Korea
Jisoo Lee
Jisoo Lee
Indiana University
Human-AI collaborationCybersecurity
J
Jeongyong Shim
Korea Telecom(KT), Seoul, Korea
E
Eunkyeong Lee
Korea Telecom(KT), Seoul, Korea
E
Eunmi Kim
Korea Telecom(KT), Korea, Korea