🤖 AI Summary
This study addresses the limitations of existing Chinese safety benchmarks that overlook adolescent-specific risks and implicit threats in multi-turn dialogues by proposing QH-Bench, the first culturally adapted, fine-grained safety benchmark for adolescents. This benchmark integrates Chinese cultural contexts, multi-turn interaction trajectories, and a five-level scoring mechanism to encompass both single- and multi-turn scenarios involving offline contact and relational pressure. Leveraging automated judging and cross-domain risk simulation techniques, this work evaluates thirteen mainstream large language models. The findings reveal significant safety deficiencies in most models when handling inducements for offline meetings and non-compliant requests following trust establishment, exposing critical shortcomings of current large language models in adolescent protection.
📝 Abstract
Safety risks in conversations with adolescents are not always explicit. A request may appear harmless unless a model considers the user's age, circumstances, and earlier turns. Existing Chinese safety benchmarks mainly target general users and give limited attention to adolescent safety. Single-turn tests also miss risks that emerge over several turns. QH-Bench is a Chinese-language benchmark for adolescent content safety, with scenarios grounded in Chinese social and cultural settings. The single-turn track contains 715 test items organized into 10 risk domains, 50 subdomains, and 143 fine-grained risk scenarios. The multi-turn track contains 100 four-turn trajectories in a balanced 10-by-10 design that combines the same ten domains with ten cross-turn mechanisms. Both tracks use the same five-level safety-helpfulness scale and automatic judge, with track-specific criteria. Evaluation of 13 open-weight models identifies offline-contact scenarios as a shared weakness. Every model receives negative scores on more than half of the items involving offline meetings with online contacts, unfamiliar groups, and adults. This includes InternLM2.5-20B, the single-turn leader; negative scores indicate responses that partially or clearly facilitate risk. GLM-4-32B, the multi-turn leader, receives negative scores on 60% of complete trajectories in which users build relationships before invoking loyalty or confidentiality. These findings identify two priorities for the evaluated models: handling adolescent offline-contact risks and maintaining safety boundaries under relational pressure. Leading aggregate scores do not establish that these specific weaknesses have been resolved.