ClaimMirage: When Self-Claims in Domain Names Change LLM Threat Judgments

๐Ÿ“… 2026-09-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study reveals that large language models (LLMs) are susceptible to semantic interference from self-asserting claims (e.g., โ€œsecure,โ€ โ€œofficialโ€) embedded within domain names, resulting in biased threat assessments. We identify a novel attack surface capable of manipulating model decisions without explicit prompt injection. Through large-scale empirical analysis encompassing 620,000 cases and controlled experiments, we systematically evaluate the vulnerabilities of multiple mainstream LLMs. Results demonstrate that such self-declarations significantly alter security alert rates, with reductions reaching up to 45.3%. These findings highlight the risks associated with LLMsโ€™ overreliance on contextual semantics in security-critical scenarios and underscore the necessity of incorporating independent verification mechanisms to mitigate such vulnerabilities.
๐Ÿ“ Abstract
Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Domain Name Security
Threat Judgment Manipulation
Self-Claims
Phishing Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

ClaimMirage
Self-Claims Manipulation
LLM Threat Judgment
Domain Name Security
Prompt Robustness
Daiki Chiba
Daiki Chiba
NTT
Cyber SecurityNetwork SecurityInternet Measurement
H
Hiroki Nakano
NTT Security Holdings Corporation & NTT, Inc., Tokyo, Japan
T
Takashi Koide
NTT Security Holdings Corporation & NTT, Inc., Tokyo, Japan