๐ค AI Summary
This study reveals that large language models (LLMs) are susceptible to semantic interference from self-asserting claims (e.g., โsecure,โ โofficialโ) embedded within domain names, resulting in biased threat assessments. We identify a novel attack surface capable of manipulating model decisions without explicit prompt injection. Through large-scale empirical analysis encompassing 620,000 cases and controlled experiments, we systematically evaluate the vulnerabilities of multiple mainstream LLMs. Results demonstrate that such self-declarations significantly alter security alert rates, with reductions reaching up to 45.3%. These findings highlight the risks associated with LLMsโ overreliance on contextual semantics in security-critical scenarios and underscore the necessity of incorporating independent verification mechanisms to mitigate such vulnerabilities.
๐ Abstract
Short claims such as not-phishing or official can change how a large language model (LLM) judges a domain name, without explicit prompt-injection commands. We study this manipulation as ClaimMirage: a name under inspection claims its own safety or approval. We analyze 622,080 judgments across 64 brands and five LLMs, comparing ten claims with length- and hyphen-matched controls in constructed brand-like names. Self-claims can substantially reduce or increase alerts, depending on the LLM and input setting. In one setting, risk-denial terms inside the registrable name reduce alerts by 45.3 percentage points even with a basic safeguard: the prompt supplies the potentially impersonated brand and its official domain for comparison. Without these references, endorsement terms at that position increase alerts by 65.6 points in the same LLM. References and component annotation remove some alert reductions but leave others or make them larger. These findings motivate testing resistance to self-claims and seeking independent evidence before treating a domain name under inspection as safe or authorized.