🤖 AI Summary
This paper addresses the fundamental tension in large language models (LLMs) between preventing misuse and preserving scientific expressiveness in safety-critical responses. To this end, we introduce the first controllable, dual-use AI benchmark specifically designed for scientific refusal evaluation in the chemical domain. Methodologically, we propose a structured query set coupled with a prompt mutation testing paradigm, integrating multi-model comparison, consistency analysis, and chain-of-thought attribution diagnostics. Our key contributions are: (1) the first systematic demonstration of LLM inconsistency in scientific safety responses—average single-prompt consistency is 85%, dropping to 65% across five prompt variants; (2) empirical evidence of divergent safety strategies across models (e.g., Claude-3.5 exhibits a 73% refusal rate versus 0% for Mistral); and (3) an open-source, reproducible benchmark and evaluation framework that provides quantitative grounding for balancing safety enforcement with scientific freedom.
📝 Abstract
The development of robust safety benchmarks for large language models requires open, reproducible datasets that can measure both appropriate refusal of harmful content and potential over-restriction of legitimate scientific discourse. We present an open-source dataset and testing framework for evaluating LLM safety mechanisms across mainly controlled substance queries, analyzing four major models' responses to systematically varied prompts. Our results reveal distinct safety profiles: Claude-3.5-sonnet demonstrated the most conservative approach with 73% refusals and 27% allowances, while Mistral attempted to answer 100% of queries. GPT-3.5-turbo showed moderate restriction with 10% refusals and 90% allowances, and Grok-2 registered 20% refusals and 80% allowances. Testing prompt variation strategies revealed decreasing response consistency, from 85% with single prompts to 65% with five variations. This publicly available benchmark enables systematic evaluation of the critical balance between necessary safety restrictions and potential over-censorship of legitimate scientific inquiry, while providing a foundation for measuring progress in AI safety implementation. Chain-of-thought analysis reveals potential vulnerabilities in safety mechanisms, highlighting the complexity of implementing robust safeguards without unduly restricting desirable and valid scientific discourse.