Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications

📅 2025-07-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current LLM security evaluations predominantly focus on base models, neglecting the critical impact of application-layer components—such as system prompts, retrieval processes, and safety mitigations—on end-to-end security. Method: We propose the first security assessment framework tailored to real-world LLM applications, integrating a domain-customized risk taxonomy, system prompt analysis, retrieval pipeline auditing, and mitigation mechanism evaluation to enable holistic, multi-component risk identification. The framework supports cross-use-case scalability and organization-level customization. Contribution/Results: Validated across multiple internal production scenarios, it significantly improves security test coverage and high-severity risk detection rates. This work bridges the gap between AI safety theory and engineering practice, delivering a reusable methodology and actionable guidelines for secure, large-scale LLM deployment.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Application Domains: Security

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSecurity and Privacy: Large-scale security measurementsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Most safety testing efforts for large language models (LLMs) today focus on evaluating foundation models. However, there is a growing need to evaluate safety at the application level, as components such as system prompts, retrieval pipelines, and guardrails introduce additional factors that significantly influence the overall safety of LLM applications. In this paper, we introduce a practical framework for evaluating application-level safety in LLM systems, validated through real-world deployment across multiple use cases within our organization. The framework consists of two parts: (1) principles for developing customized safety risk taxonomies, and (2) practices for evaluating safety risks in LLM applications. We illustrate how the proposed framework was applied in our internal pilot, providing a reference point for organizations seeking to scale their safety testing efforts. This work aims to bridge the gap between theoretical concepts in AI safety and the operational realities of safeguarding LLM applications in practice, offering actionable guidance for safe and scalable deployment.
Problem

Research questions and friction points this paper is trying to address.

Evaluating safety risks in real-world LLM applications
Developing customized safety risk taxonomies for LLMs
Bridging AI safety theory with practical deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Framework for evaluating application-level LLM safety
Customized safety risk taxonomies development principles
Practices for assessing safety risks in LLM apps
Jia Yi Goh
Jia Yi Goh
GovTech Singapore
safetyfairnessrobustness
S
Shaun Khoo
GovTech Singapore, Singapore
N
Nyx Iskandar
University of California, Berkeley, USA
Gabriel Chua
Gabriel Chua
Data Scientist
LLM
L
Leanne Tan
GovTech Singapore, Singapore
J
Jessica Foo
GovTech Singapore, Singapore