SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

πŸ“… 2026-07-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Current safety evaluations of large language models (LLMs) lack grounding in realistic scientific risk scenarios, relying instead on templated queries and LLM-as-a-Judge methodologies that inadequately capture the risks of scientific knowledge misuse. This work proposes SciHazard, a benchmark comprising 2,400 real-world hazardous queries and 600 overly safe queries across 12 scientific disciplines, alongside DeHarm-Scoreβ€”a decomposable, interpretable evaluation framework that jointly assesses query-level hazard, refusal behavior, and response-level risk. For the first time, it integrates real regulatory entities and incident scenarios, introduces dynamic weighting for quantifying exploitability, and employs retrieval-augmented assessment of net incremental risk, all validated by domain experts. Experiments show that DeHarm-Score achieves a 90.17% improvement in expert alignment over the strongest baseline, and evaluation of 31 state-of-the-art models reveals that advanced research agents exhibit 32.3% higher average risk, exposing critical blind spots in current safety mechanisms.
πŸ“ Abstract
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute \textsc{DeHarm-Score} , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into \textsc{Executability}, quantified via dynamic checklists with importance weighting, and \textsc{Net-new risk}, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that \textsc{DeHarm-Score} improves agreement with expert annotations by 90.17\% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, deep research agents yield 32.3\% higher mean \textsc{DeHarm-Score} than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.
Problem

Research questions and friction points this paper is trying to address.

scientific safety risks
hazardous knowledge
large language models
misuse guidance
safety benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

SciHazard
DeHarm-Score
scientific safety
harm decomposition
autonomous agents
πŸ”Ž Similar Papers