SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决LLM在安全、安保和隐私评估中的静态基准问题,提出SSP-Bench框架,通过动态生成数据实例并保持领域一致性来提高评估准确性。
📝 Abstract
Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets often fail under semantically equivalent rephrasings. We introduce SSP-Bench, a dynamic benchmarking framework that generates evaluation instances on demand while preserving domain consistency. The framework ensures label validity through externally grounded sources, enforces scope via service-specific validation, and calibrates difficulty using a multi-model steering panel. Benchmark construction is formulated as a multi-objective optimization problem over difficulty, separability, novelty, and diversity. Across 24 models and four SSP services, SSP-Bench reveals systematic failures of static evaluation, including near-zero correlation in safety rankings due to construct mixing, strong safety--over-refusal coupling, and hidden within-family regressions. These results show that static benchmarks can misrepresent model behavior, motivating dynamic, deployment-relevant evaluation.
Problem

Research questions and friction points this paper is trying to address.

safety
security
privacy
large language models
static benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamic benchmarking
on-demand instance generation
multi-objective optimization
external validation
🔎 Similar Papers
No similar papers found.