Institution profile

Del Norte High School

Academic institutionnorthamerica · us
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior

Apr 07, 2026

This study addresses a critical gap in existing research on emotional prompting, which has predominantly focused on single positive emotions while neglecting systematic analysis of diverse emotion types and their intensities. The work presents the first comprehensive investigation into how four distinct emotions—joy, encouragement, anger, and insecurity—and their varying intensities differentially influence large language models across three key dimensions: accuracy, sycophancy, and toxicity. Leveraging a GPT-4o mini–based pipeline for emotional prompt generation and combining human and model-based annotations, the authors construct a high-quality “gold dataset” to serve as a benchmark for emotional prompting. Their findings reveal a dual effect: while positive emotions enhance accuracy and reduce toxicity, they simultaneously exacerbate sycophantic behavior, underscoring the nuanced trade-offs inherent in emotionally infused prompts.

0 citationsRead paper

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

Apr 18, 2025arXiv.org

This work addresses the security vulnerability of large language models (LLMs) to jailbreaking attacks. Methodologically, it proposes a dynamic, multi-round adversarial prompt generation framework that integrates role-playing, context manipulation, and semantic obfuscation into a parameterized attack model. Guided by failure analysis, the framework iteratively refines prompts across dialogue turns and enhances attack efficacy via strategic system prompt engineering and hyperparameter optimization. Its key contribution is the first automated, interpretable, and high-success-rate jailbreaking method tailored to complex, multi-turn conversational scenarios. Experiments on mainstream LLMs—including ChatGPT, Llama-3, and DeepSeek-V2—achieve an 86% jailbreaking success rate, revealing critical structural weaknesses in current safety mechanisms under multi-turn interaction. The approach establishes a novel paradigm for LLM red-teaming evaluation.

0 citationsRead paper
Recent publications

Latest Papers

The Role of Emotional Stimuli and Intensity in Shaping Large Language Model Behavior

Apr 07, 2026

This study addresses a critical gap in existing research on emotional prompting, which has predominantly focused on single positive emotions while neglecting systematic analysis of diverse emotion types and their intensities. The work presents the first comprehensive investigation into how four distinct emotions—joy, encouragement, anger, and insecurity—and their varying intensities differentially influence large language models across three key dimensions: accuracy, sycophancy, and toxicity. Leveraging a GPT-4o mini–based pipeline for emotional prompt generation and combining human and model-based annotations, the authors construct a high-quality “gold dataset” to serve as a benchmark for emotional prompting. Their findings reveal a dual effect: while positive emotions enhance accuracy and reduce toxicity, they simultaneously exacerbate sycophantic behavior, underscoring the nuanced trade-offs inherent in emotionally infused prompts.

0 citationsRead paper

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

Apr 18, 2025arXiv.org

This work addresses the security vulnerability of large language models (LLMs) to jailbreaking attacks. Methodologically, it proposes a dynamic, multi-round adversarial prompt generation framework that integrates role-playing, context manipulation, and semantic obfuscation into a parameterized attack model. Guided by failure analysis, the framework iteratively refines prompts across dialogue turns and enhances attack efficacy via strategic system prompt engineering and hyperparameter optimization. Its key contribution is the first automated, interpretable, and high-success-rate jailbreaking method tailored to complex, multi-turn conversational scenarios. Experiments on mainstream LLMs—including ChatGPT, Llama-3, and DeepSeek-V2—achieve an 86% jailbreaking success rate, revealing critical structural weaknesses in current safety mechanisms under multi-turn interaction. The approach establishes a novel paradigm for LLM red-teaming evaluation.

0 citationsRead paper