🤖 AI Summary
This study addresses the tendency of large language models (LLMs) to erroneously absorb user instructions following pasted text as part of the content to be edited, a vulnerability that has lacked systematic evaluation in the absence of explicit delimiters. To this end, this work proposes SEAM, a benchmark comprising 300 editing instances across six boundary conditions, enabling the first controlled assessment of instruction absorption under unmarked prompt boundaries for twenty mainstream LLMs. The findings reveal that simple newline or blank-line separators fail to effectively isolate instructions from data, resulting in absorption rates ranging from 7.7% to 66.7%. While introducing explicit boundary markers significantly reduces error rates for nineteen of the twenty models, such vulnerabilities cannot be entirely eliminated. This research bridges the evaluation gap regarding prompt boundary separation and provides empirical evidence for mitigating instruction absorption in LLMs.
📝 Abstract
Large language models (LLMs) receive each user message as plain text, even when it combines text from different sources. For example, a user may paste text into a prompt and keep typing a comment directly below it. We study absorption: a phenomenon where the model treats a trailing user comment as part of the pasted text, returning it inside the edited text. This happens even though the user did not intend the comment to become part of that text. Existing instruction-data separation benchmarks tell the model which text is instruction and which is data, then test whether it obeys that separation. They do not test harmless user speech following an unmarked paste. We introduce SEAM, a controlled benchmark of 300 editing examples. Each example is tested under six matched conditions that vary how the boundary between pasted text and later user speech is expressed. Across 20 models, absorption at a bare newline ranges from 7.7% to 66.7%. Adding a blank line does not significantly reduce absorption in any model, while boundary markers reduce it in 19 of 20 models. Comments that fit the pasted text, such as a code comment typed after code, are absorbed significantly more often in 17 of 20 models. Models often fail to separate pasted material from later user speech, and explicit boundaries reduce but do not remove this failure.