Scholar
Jianshuo Dong
Google Scholar ID: CY23PzAAAAAJ
Tsinghua University
Trustworthy AI
Explainable AI
Agent Security
Follow
Homepage
↗
Google Scholar
↗
Citations & Impact
All-time
Citations
100
H-index
5
i10-index
4
Publications
10
Co-authors
6
list available
Contact
CV
Open ↗
GitHub
Open ↗
Publications
10 items
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
2026
Cited
0
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
2026
Cited
0
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
2026
Cited
0
LeakDojo: Decoding the Leakage Threats of RAG Systems
2026
Cited
0
Towards Understanding the Cognitive Habits of Large Reasoning Models
Cited
0
Revisiting the Reliability of Language Models in Instruction-Following
2025
Cited
0
SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
2025
Cited
0
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
2025
Cited
0
Load more
Resume
Academic Achievements
Paper '“I’ve Decided to Leak”: Probing Internals Behind Prompt Leakage Intents' accepted to EMNLP 2025 main conference (oral)
Paper 'An Engorgio Prompt Makes Large Language Model Babble on' accepted to ICLR 2025
Paper 'One-bit Flip is All You Need: When Bit-flip Attack Meets Model Training' accepted to ICCV 2024
Preprint 'SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents' (arXiv:2509.23694)
Preprint 'Towards Understanding the Cognitive Habits of Large Reasoning Models' (arXiv:2506.21571)
Preprint 'Can Large Language Models Automate the Refinement of Cellular Network Specifications?' (arXiv:2507.04214)
Reviewer for ICLR'26; Top Reviewer for NeurIPS'25; Notable Reviewer for ICLR'25
Co-authors
6 total
Han Qiu (邱寒)
Tsinghua University
Tianwei Zhang
Nanyang Technological University
Yiming Li
Nanyang Technological University
Qingjie ZHANG
Tsinghua University
Hao Wang
Tsinghua University
Liu Yan
Researcher and Director of Ant Group