AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of transparency and oversight in system prompts deployed by commercial large language models, which undermines user trust and accountability. The authors propose AISPA, a novel framework for systematically auditing system prompts from the user perspective. They develop an eight-dimensional evaluation schema and conduct a large-scale empirical analysis of 3,249 instructions across 88 commercial products, combining qualitative categorization with quantitative metrics. Findings reveal substantial heterogeneity in prompt design across products: while 98.9% include protective instructions, only 24% comprehensively cover all eight dimensions, and 40% still contain content potentially detrimental to user interests. These results underscore the urgent need for transparent prompt disclosure mechanisms and independent regulatory oversight.
📝 Abstract
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.
Problem

Research questions and friction points this paper is trying to address.

system prompts
trust gap
accountability
AI transparency
commercial AI products
Innovation

Methods, ideas, or system contributions that make the work stand out.

system prompt auditing
user-centric AI
AISPA
foundation model governance
AI transparency
🔎 Similar Papers
X
Xiangning Lin
CMU
Shenzhe Zhu
Shenzhe Zhu
University of Toronto
Trustworthy AIAI Agent
Shu Yang
Shu Yang
King Abdullah University of Science and Technology
Human-Centered AIAlignmentReasoning
Z
Zhenyu Zhang
Stanford University
H
Haoqian Zhang
University of Toronto
Y
Yipeng Zhao
University of Toronto
C
Chengxuan Qian
UCSB
Tianwei Wang
Tianwei Wang
South China University of Technology
Computer VisionOCRDocument Analysis
Z
Ziheng Zhang
OSU
Z
Zhenlong Yuan
UCSC
D
Dingcheng Wang
Northwestern University
Juncheng Wu
Juncheng Wu
University of California, Santa Cruz
Foundation ModelsReasoning LLMs
Y
Yuan Si
Northwestern University
J
Jiaxin Liu
UIUC
Baolong Bi
Baolong Bi
University of Chinese Academy of Sciences
trustworthy large language models
Robert Mahari
Robert Mahari
Associate Director, Stanford CodeX Center
Computational Law
Tobin South
Tobin South
Massachusetts Institute of Technology
D
Dazza Greenwood
MIT
Zexue He
Zexue He
University of California, San Diego
Trustworthy NLPLLM
Rishi Bommasani
Rishi Bommasani
CS PhD, Stanford University
Societal Impact of AIAI PolicyAI GovernanceFoundation Models
Sophia Kazinnik
Sophia Kazinnik
Stanford University
Applied Artificial IntelligenceMachine LearningBehavioral FinanceAlternative Data
Andreas Haupt
Andreas Haupt
Stanford University
EconomicsArtificial IntelligencePersonalisationMarket Design
Samuele Marro
Samuele Marro
University of Oxford
Machine LearningRobustnessOptimizationGenerative Modelling
Erik Brynjolfsson
Erik Brynjolfsson
Professor at Stanford; NBER; Stanford Digital Economy Lab
EconomicsInformation EconomicsEconomics of AIProductivityIntangible Assets
A
Alex Pentland
Stanford University