Score
Design and evaluate methods, models, and pipelines that extract stable signatures from prompts, system messages, or user interactions and use those signatures to link or fingerprint prompts and users across conversations. This includes building similarity/scoring models and aggregation techniques to detect identical or unseen system prompts from observed responses and to infer reading and selection signatures for attributing or linking prompts and users.
Frequent jailbreaking attacks against large language models (LLMs) and fragmented evaluation criteria hinder progress in prompt security research. Method: This paper introduces the first systematic framework for prompt security, featuring a multi-level taxonomy of attacks and defenses, formalized threat models and cost assumptions, machine-readable safety evaluation profiles, and an open-source benchmarking toolchain. Contributions: (1) We release JAILBREAKDB—the largest human-verified dataset of jailbreaking and benign prompts to date, containing over 120,000 samples; (2) we establish the first open, reproducible, and auditable standardized evaluation benchmark for prompt security; and (3) we conduct a unified, cross-method assessment and ranking of 56 state-of-the-art attack and defense techniques. The framework significantly enhances comparability and reproducibility across studies, providing foundational infrastructure for rigorous, scalable prompt security research.
研究通过提出黑盒行为指纹识别方法解决系统提示被克隆后无法验证的问题,该方法基于模型输出相似性进行检测。
To address risks of large language model (LLM) theft and misuse, this paper proposes a verifiable, tamper-resistant model fingerprinting technique. Methodologically, it constructs a cryptographic hash chain using question-answer pairs and SHA-256, enforcing fine-grained response constraints and hash binding to ensure strong integrity verification. It is the first work to formally define and satisfy five core fingerprint properties: transparency, efficiency, persistence, robustness, and unforgeability. Extensive experiments across multiple LLMs demonstrate that the fingerprint withstands benign modifications—including fine-tuning and pruning—as well as adversarial erasure attacks, while preserving near-original model performance post-embedding. This work delivers the first complete solution for LLM copyright protection and provenance tracking that simultaneously achieves theoretical rigor and engineering practicality.
Large language models (LLMs) exhibit poor reproducibility in text annotation tasks due to sensitivity to minor prompt perturbations, yet no standardized metric exists for quantifying prompt stability. To address this, we systematically adapt inter-annotator agreement principles from coding reliability research to prompt engineering, introducing the Prompt Stability Score (PSS)—a unified, computationally tractable metric for stability assessment. Our method integrates multi-prompt sampling, batched LLM inference, consistency analysis via Cohen’s and Fleiss’ Kappa, and an automated Python evaluation framework (open-sourced as PromptStability). Empirical validation across six benchmark datasets and twelve annotation task types—encompassing over 150,000 samples—demonstrates PSS’s effectiveness in precisely identifying low-stability prompting configurations. This work establishes the first standardized diagnostic paradigm for evaluating prompt robustness, thereby enabling reproducible, interpretable, and empirically grounded prompt engineering practices.
Current LLM-native software engineering lacks a systematic practical framework—particularly in verification and falsification—necessitating unified task taxonomies and prompt-pattern conceptualizations. Method: We conduct a systematic literature review of over 100 papers, employing bibliometric analysis and conceptual clustering to map, classify, and abstract LLM-based downstream tasks in software engineering (SE). Contribution/Results: We propose the first fine-grained SE-specific taxonomy for LLM downstream tasks, encompassing six core clusters: testing, fuzzing, bug localization, vulnerability detection, static analysis, and program verification. Our taxonomy uniquely balances cross-task abstraction with task-specific variation modeling, uncovering generalizable prompt-engineering principles. It provides a foundational framework for targeted LLM adaptation, benchmark construction, and empirically grounded engineering practice in SE.
This work addresses the challenge of identifying underlying large language model (LLM) versions in LLM-integrated applications, where model provenance is often opaque. We propose LLMmap, the first active fingerprinting technique tailored to this setting. Methodologically, it leverages domain-informed query generation, response pattern modeling, and statistical significance testing to capture fine-grained behavioral signatures of LLM outputs. Its design ensures robustness across vendors, architectures (e.g., RAG, chain-of-thought), system prompts, and sampling perturbations. Evaluated on 42 open- and closed-source LLMs—including real-world deployments with unknown prompts and generation frameworks—LLMmap achieves >95% identification accuracy using only eight queries. To our knowledge, this is the first approach enabling high-accuracy, low-overhead, and generalizable LLM version fingerprinting at the application layer.
This work addresses the lack of transparency in current large language model (LLM) services, where users cannot easily verify whether the deployed model matches its claimed identity. Existing identification methods typically require long inputs, fine-grained outputs, or cooperation from model providers. In contrast, this study proposes a lightweight authentication approach that constructs behavioral fingerprints using the output distribution of just a single token elicited by minimal prompts—such as “say a random number between 1 and 100.” The method demonstrates for the first time that single-token distributions alone suffice to reliably distinguish among LLMs without complex inputs or internal access. Evaluated across 165 commercial models using multilingual single-token queries, empirical distribution modeling, and Jensen–Shannon divergence, it achieves intra-model distances an order of magnitude smaller than inter-model distances, 59.5% accuracy in model family identification, and an equal error rate as low as 7.3%, requiring only approximately 100 queries.
This study investigates whether brief, task-oriented user prompts in interactions with large language models contain stable and identifiable identity signals. Grounded in the lexical stability hypothesis, the research demonstrates that identity cues are primarily encoded in surface-level word choices rather than abstract communicative intent, and uncovers a “uniqueness–consistency paradox” in stylistic features. By systematically comparing lexical and semantic representations, extracting stylistic metrics, and evaluating robustness through adversarial perturbations, the authors achieve high-accuracy identity recognition across 20,680 real-world prompts from 1,034 users. This work provides the first empirical validation that user prompts can serve as reliable behavioral biometrics.
This study addresses the gap in existing research, which predominantly relies on expert-crafted hypothetical questions, by offering the first systematic analysis of real user inquiries in the domain of digital security and privacy (S&P). Leveraging 14,727 authentic user dialogues from the WildChat dataset, the authors employ topic modeling, multi-turn repeated prompting, and human evaluation to characterize users’ actual S&P concerns and rigorously assess the quality and consistency of responses from leading large language models. Findings reveal that commercial models—such as GPT-5.5—deliver “sufficiently good” answers on 98% of advisory prompts, substantially outperforming open-source counterparts like Llama-4 (47%), yet exhibit notable inconsistencies across conversational turns. This work provides an empirical foundation for understanding genuine user needs and evaluating model reliability in real-world S&P contexts.
This work demonstrates that subtle numerical discrepancies introduced by different components—such as inference engines, attention backends, and hardware platforms—in large language model (LLM) inference systems can serve as distinctive fingerprints, inadvertently revealing system configurations and posing security risks. The paper presents the first fingerprinting methodology based on prompt-response behavior and numerical deviation analysis, achieving high-accuracy identification of these components through empirical evaluation. The study shows that this approach reliably infers system configurations even under non-zero temperature sampling, highlighting inherent limitations in existing defense mechanisms. Furthermore, it explores potential mitigation strategies and analyzes their practical implications, underscoring the tension between deployability and robustness in securing LLM inference pipelines.
Existing approaches struggle to systematically compare the outputs of large language models under varying generation conditions. This work proposes a “visual fingerprint” framework that models model outputs as distributions over multidimensional linguistic choices—encompassing content, expression, and structure—and enables cross-condition comparison of generative behavior through an integrated natural language processing pipeline and distribution visualization techniques. For the first time, this method facilitates intuitive, distribution-level insights into stable behavioral patterns that persist across diverse settings yet remain undetectable via single-sample inspection or conventional aggregate metrics. The efficacy of the framework is demonstrated across four distinct application scenarios.